SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Varda Space Industries is building commercial space infrastructure, including in-orbit pharmaceutical processing systems, reentry capsules, and satellite buses. The company develops products and infrastructure to enable manufacturing in space that benefits life on Earth.
As a Site Reliability Engineer II, you will be critical in building, scaling, and maintaining infrastructure powering Varda's systems on Earth and in orbit. This is a hands-on role requiring deep expertise in Kubernetes, containerized technologies, and modern DevOps practices applied to mission-critical environments.
Key responsibilities include:
- Deploy, maintain, and operate mission-critical applications supporting spacecraft and company-wide systems
- Build and evolve Infrastructure as Code (IaC) frameworks using Terraform
- Implement observability systems (metrics, logging, tracing) and alerting using tools like Prometheus and Grafana
- Build and maintain CI/CD pipelines for safe, repeatable deployments
- Partner with software and hardware engineers to deliver reliable, scalable systems
- Identify and resolve system bottlenecks, perform performance tuning, and implement stability improvements
- Respond to production incidents, perform root cause analysis, and drive corrective actions
- Rotate through on-call schedule
- Occasionally travel to customer sites and other Varda locations for troubleshooting and deployment
Required qualifications: Bachelor's degree in computer science, engineering, or related STEM field with 3+ years of SRE experience (or 5+ years of progressive DevOps/Systems Engineering without degree). Must have hands-on experience with Infrastructure as Code (Terraform), Kubernetes in production, observability tools (Prometheus/Grafana/InfluxDB), software-defined networking, and scripting (Python/Bash/PowerShell). Strong communication skills required.
Preferred experience includes Azure cloud infrastructure provisioning, configuration management tools (Ansible, Salt), and GitOps practices.