SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Salary: USD 105,000 - 185,000 / annual
Menlo Security is the leader in Browser Security for human and agentic workforces, protecting organizations from cyberattacks across the web, documents, and email. The Platform Infrastructure Engineering team builds and operates Menlo's core infrastructure services on a cloud-native platform, enabling customers to connect to the Internet securely.
As a Platform Infrastructure Engineer, you will join a globally distributed team of experienced engineers responsible for building and managing infrastructure services on Google Kubernetes Engine and VMs spanning multiple regions and environments. You'll work with infrastructure as code using Terraform and Spacelift, deploy with Helm, and emphasize security-first design, comprehensive observability, and multi-region resilience. The team uses AI-assisted development tools, including Gemini Code Assist, as part of standard workflow, and this role is expected to leverage LLM-based tooling to build and troubleshoot infrastructure code efficiently.
Key responsibilities include: implementing, deploying, and maintaining VM and Kubernetes infrastructure on GCP and AWS across dozens of clusters in development, staging, and production environments across multiple regions; building and maintaining Infrastructure as Code using Terraform modules and Spacelift, provisioning networking, compute, storage, and security components; implementing and maintaining observability solutions using Grafana Cloud, Prometheus/Mimir, and OTel collectors; managing certificate lifecycle, DNS automation, ingress controllers, and service mesh networking with Cilium; partnering with Engineering, Product, Compliance, and Security teams on capacity planning, disaster recovery, and architectural decisions; identifying and eliminating toil through automation and CI/CD pipelines; and participating in a 24x7 on-call rotation as part of a globally distributed team.
Key outcomes owned include reliable, secure, and scalable infrastructure across GCP and AWS supporting Menlo's platform globally; reduced operational toil and incident recurrence through automation and IaC practices; and comprehensive, end-to-end observability providing deep platform visibility, proactive health monitoring, and accelerated incident detection and resolution. Success metrics include infrastructure uptime/availability (99.9%+), mean time to detect and resolve incidents, percentage of infrastructure changes deployed via IaC, on-call incident volume reduction, and lead time for provisioning new infrastructure.
REQUIREMENTS:
- Bachelor's degree in Computer Science, related technical field, or equivalent practical experience
- Proficiency in Python, Bash, and Go
- Understanding of network topologies, communication protocols (TCP/IP, HTTP/S, UDP, TLS), and enterprise-grade connectivity solutions
- Kubernetes expertise including cluster administration, RBAC, networking, workload management, and production troubleshooting
- Proven experience with Terraform for infrastructure provisioning and management
- Knowledge of Google Cloud Platform services including GKE, VPC networking, Cloud DNS, Artifact Registry, Secret Manager, IAM, Gemini Code Assist, and Workload Identity
- Clear understanding of how to use LLM-based code-assist tools to effectively build and troubleshoot software
PREFERRED:
- Experience with GitOps methodologies and tools