Menlo Security is the leader in Browser Security for human and agentic workforces. Our mission is to enable humans and agents to connect, communicate, and collaborate securely, without compromise. The Menlo Browser Security Platform protects organizations from cyberattacks by stopping threats across the web, documents, and email before they reach the user. With Menlo Agent Runtime Security (MARS), that protection now extends to the AI agents working alongside every employee. Menlo Security is trusted by major global businesses, including Fortune 500 companies and government agencies, to protect their most valuable asset, their data, and is backed by top-tier investors.SummaryPlatform Infrastructure Engineering builds and operates Menlo Security's Infrastructure Platform, enabling our customers to connect to the Internet without compromise. As a Platform Infrastructure Engineer, you'll join a globally distributed team of experienced engineers building and managing the company's core infrastructure services on a cloud-native platform built on Google Kubernetes Engine and VMs spanning multiple regions and environments. The team manages infrastructure as code with Terraform and Spacelift, deploys with Helm, and emphasizes security-first design, comprehensive observability, and multi-region resilience. The team also uses AI-assisted development and code-review tools, including Gemini Code Assist, as part of the standard engineering workflow, and this role is expected to use LLM-based tooling to build and troubleshoot infrastructure code efficiently.Outcomes & KPIsKey Outcome(s) Owned:Reliable, secure, and scalable infrastructure across GCP and AWS supporting Menlo's platform globally.Reduced operational toil and incident recurrence through automation and Infrastructure as Code practices.Comprehensive, end-to-end observability framework providing deep platform visibility, proactive health monitoring, and accelerated incident detection and resolution.Success Metrics / KPIs:Infrastructure uptime/availability across regions (e.g., 99.9%+)Mean time to detect (MTTD) and mean time to resolve (MTTR) for incidentsPercentage of infrastructure changes deployed via IaC (Terraform) vs. manual changesOn-call incident volume and reduction in repeat/preventable incidentsLead time for provisioning new infrastructureWhat You'll DoImplement, deploy, and maintain VM and Kubernetes infrastructure on GCP and AWS across dozens of clusters spanning development, staging, and production environments in multiple regionsBuild and maintain Infrastructure as Code using Terraform modules and Spacelift (or equivalent TACOS), provisioning networking, compute, storage, and security components, and implementing multi-layer configuration management workflowsImplement and maintain observability solutions using Grafana Cloud, Prometheus/Mimir, and OTel collectors, designing dashboards and alerting rules across all platform componentsManage certificate lifecycle, DNS automation, ingress controllers, and service mesh networking with CiliumPartner with peers and across Engineering, Product, Compliance, and Security teams to align on requirements and consult on capacity planning, disaster recovery, and architectural decisionsIdentify and eliminate toil through automation — writing scripts, building CI/CD pipelines, and using AI-assisted coding tools to move fasterParticipate in a 24x7 on-call rotation as part of a globally distributed team, responding to incidents and driving post-incident reviewsFunctional CompetenciesRequired:Bachelor's degree in Computer Science, a related technical field, or equivalent practical experienceProficiency in common programming and scripting languages, particularly Python, Bash, and GoUnderstanding of network topologies, communication protocols (e.g., TCP/IP, HTTP/S, UDP, TLS), and enterprise-grade connectivity solutionsKubernetes expertise, including cluster administration, RBAC, networking, workload management, and troubleshooting in production environmentsProven experience with Terraform for infrastructure provisioning and managementKnowledge of Google Cloud Platform services including GKE, VPC networking, Cloud DNS, Artifact Registry, Secret Manager, IAM, Gemini Code Assist, and Workload IdentityClear understanding of how to use LLM-based code-assist tools to effectively build and troubleshoot softwarePreferred / Nice to Have:Experience with GitOps methodologies and toolsOur Compensation and BenefitsAt Menlo Security, Base Salary is one part of our competitive total compensation and benefits package and is determined using a salary range. The base salary range for this role is 112,000 CAD - 168,000 CAD.In accordance with Canadian law, the range provided is Menlo Security’s reasonable estimate of the base compensation for this role. The actual amount may be higher or lower, based on non-discriminatory factors such as experience, knowledge, skills, abilities, and location. All employees may be eligible to become Menlo Security shareholders through eligibility for stock-based compensation grants, which are awarded to employees based on company and individual performance.Why Menlo?At Menlo, we don't settle for the status quo — in our technology or our culture. How we think and act is just as important as what we build. Our culture is defined by five core mindsets: Proactive Leadership, Straight Talk, United Impact, Elevated Talent, and Customer-Compelled. We take ownership and drive outcomes without waiting to be told. We communicate directly and seek hard truths. We break down silos and win together. We hold a high bar for ourselves and the people around us. And we treat every customer interaction as mission-critical. If you're someone who sees it, owns it, solves it, and does it — you'll thrive here.