SlipstreamJobsFresh Startup & VC-Backed Jobs

Staff Software Engineer, Production Engineering

Harvey - New York, NY, United States - In-office - posted 2026-07-30

Apply on the company site

SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.

Harvey is an AI platform for legal and professional services, combining frontier agentic AI with enterprise-grade infrastructure. The company is scaling rapidly with strong product-market fit and world-class investor backing. As a Staff Software Engineer in Production Engineering, you'll be responsible for designing, building, and operating Harvey's core compute and networking infrastructure, Kubernetes platform, workflow orchestration, and production foundations. This is a technical leadership role focused on enabling engineering teams to move quickly while maintaining reliability and scale. Key responsibilities include: **Infrastructure Engineering & Technical Leadership**: Design and operate production infrastructure powering Harvey's products and AI workloads. Drive technical direction across compute, networking, Kubernetes, workflow orchestration, and production operations. Lead complex cross-functional initiatives improving reliability, scalability, security, and efficiency. Partner with Product Engineering, Security, AI Infrastructure, and Platform teams to translate requirements into resilient solutions. Establish reusable patterns, tooling, and paved paths for safe service deployment. Raise engineering standards through design reviews, documentation, operational rigor, and mentorship. **Infrastructure Foundation & Production Operations**: Build and operate global compute and network infrastructure ensuring high availability and performance. Improve utilization and service availability for rapidly growing AI workloads. Develop capacity models, demand forecasts, and fleet automation. Operate and enhance the Kubernetes platform including provisioning, upgrades, networking, monitoring, and automation. Drive cost efficiency through capacity management and resource optimization. Build secure foundations with IAM, network isolation, secrets management, and compliance controls. Develop Infrastructure-as-Code using Terraform and Pulumi. Enhance observability, monitoring, alerting, and incident response. Participate in on-call rotation and drive continuous improvement from production learnings. Required qualifications: 10+ years in software, infrastructure, SRE, or production engineering. Deep experience with large-scale cloud infrastructure on AWS, Azure, or GCP. Strong hands-on Kubernetes production experience including cluster lifecycle and networking. Experience building distributed systems with reliability and performance focus. Infrastructure automation and IaC expertise with Terraform or Pulumi. Strong understanding of compute, networking, capacity planning, and fleet management. Experience designing observability systems. Infrastructure security knowledge including IAM and compliance. Track record driving complex cross-functional initiatives and influencing without formal authority. Excellent communication skills and systems-thinking mindset.

About Harvey

Legal / Compliance / Risk; AI / Data / Infrastructure — AI platform for legal and professional services work.

Similar roles