SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Upstart is an AI-powered lending marketplace that partners with banks and credit unions to expand access to affordable credit. The Cloud Platform team, part of the Reliability organization, is responsible for building and operating the shared cloud infrastructure that powers all product and machine learning workloads across the company.
As a DevOps Engineer II, you will help evolve the platform to support increasing scale and complexity while partnering closely with SRE, Delivery, InfoSec, and Product/ML teams. Your responsibilities include designing and operating a fleet of Kubernetes (EKS) clusters across production, staging, and ephemeral environments to ensure reliability and high availability. You'll evolve AWS infrastructure and network architecture (VPCs, subnets, IAM, account structure) to support scalable, multi-team workloads.
You will build and maintain infrastructure-as-code and GitOps workflows using tools such as Terraform, CDK, and ArgoCD. You'll improve platform reliability and performance by defining and driving SLOs, analyzing incidents, and implementing systemic fixes. Participation in the on-call rotation is expected, where you'll lead incident response and post-incident reviews to drive systemic platform improvements.
Key focus areas include driving improvements in developer experience by simplifying platform usage and reducing toil, enabling faster product and ML development, and contributing to cost efficiency initiatives by optimizing resource utilization across Kubernetes and cloud infrastructure. You'll influence technical decisions across teams and drive adoption of platform standards.
Minimum qualifications include a Bachelor's degree in Computer Science, Engineering, Mathematics, or related field (or equivalent practical experience) and 3+ years of professional experience. You need 3+ years operating Kubernetes in production environments, proficiency with AWS infrastructure, proven expertise implementing infrastructure-as-code, and experience with GitOps workflows. Preferred qualifications include knowledge of service mesh technologies (Istio, Envoy), multi-cluster Kubernetes architectures, cloud networking at scale, and cloud security/identity frameworks.