SlipstreamJobsFresh Startup & VC-Backed Jobs

Senior Cloud Infrastructure Engineer

Scale - San Francisco, CA, United States - In-office - posted 2026-09-24

Apply on the company site

SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.

Salary: USD 143,000 - 178,000 / annual

Scale is building the data and infrastructure powering the generative AI wave. The Platform team is seeking an Operations-focused AWS Engineer to manage and optimize large-scale cloud infrastructure serving major AI companies including OpenAI, Microsoft, Adept, and Stability AI. You will own the end-to-end lifecycle of AWS infrastructure, including routine patching, version upgrades, and system maintenance to ensure high availability. You'll work closely with the Security team to identify, prioritize, and remediate vulnerabilities across the cloud footprint. Your responsibilities include continuously monitoring and optimizing AWS resource utilization for performance, reliability, and cost-efficiency; automating operational tasks such as fleet-wide patching and configuration audits; providing technical expertise for infrastructure-level troubleshooting and incident response; and building systems that maintain strict security compliance standards. This is a hands-on technical role focused on operational excellence, infrastructure hardening, and automation at scale. You'll be working with a team responsible for the core abstractions and infrastructure that enable rapid product iteration. REQUIREMENTS: - 6+ years of experience with core AWS services (EC2, VPC, IAM, S3, RDS, EKS) and multi-account environment management - Proven track record managing large-scale patching programs and performing major version upgrades for OS and middleware with minimal downtime - Proficiency with AWS Systems Manager (SSM), Terraform, Atlantis, and Kubernetes for configuration management and automation - Familiarity with scripting languages (Bash, Python, etc.) to automate operational workflows, with focus on system management rather than application development - Experience with vulnerability scanning tools and strong understanding of cloud infrastructure hardening against common threats - Experience using Datadog, CloudWatch, or similar monitoring tools to monitor system health and drive optimization - Bonus: Experience with Azure or Google Cloud Platform

Similar roles