SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Salary: USD 180,000 - 290,000 / annual
Gray Swan is building AI security infrastructure to help frontier labs and organizations safely deploy AI models at scale. The company evaluates AI models, detects real-time threats, and runs adaptive adversarial red teaming. With ~50 employees and strong funding, Gray Swan is growing rapidly and seeking an Infrastructure Engineer to design and scale the backend systems powering their platform.
In this role, you will own the design, build, and maintenance of highly available backend services and distributed systems. You'll architect cloud infrastructure across Kubernetes, AWS, networking, storage, and compute to ensure reliable production environments. You'll build scalable APIs, internal platform services, and infrastructure tooling that improve developer productivity. You'll enhance system observability through logging, metrics, tracing, dashboards, and automated alerting, and optimize performance, latency, and costs while maintaining reliability and security.
You'll partner closely with machine learning engineers, security researchers, and product teams to deliver production-ready infrastructure for AI workloads. This is a high-ownership role in a fast-moving startup where you'll influence technical direction and have significant impact on foundational systems.
REQUIREMENTS:
- 5+ years of experience building backend infrastructure or distributed systems in production environments
- Strong programming skills in C/C++, Go, Python, Rust, or Java
- Experience operating services on Kubernetes and modern cloud platforms (AWS, GCP, or Azure)
- Deep understanding of networking, distributed systems, containers, service orchestration, and scalable architectures
- Experience designing APIs, microservices, asynchronous systems, and event-driven architectures
- Comfortable debugging complex production issues and improving reliability through automation and operational excellence
- Passion for writing clean, maintainable code and building infrastructure that other engineers love using
- Excitement to work in a fast-moving startup with significant ownership and ambiguity
BONUS EXPERIENCE:
- Supporting machine learning or LLM infrastructure
- Infrastructure-as-code tools
- Kafka, Redis, PostgreSQL, ClickHouse, or similar distributed data systems
- Building internal developer platforms or platform engineering tooling
- Cloud security, infrastructure hardening, or zero-trust architectures
- High-growth startup or zero-to-one product building experience
- Interest in AI safety, cybersecurity, or adversarial machine learning