SlipstreamJobsFresh Startup & VC-Backed Jobs

Senior Infrastructure Engineer

Voxel51 - Remote - Remote

Apply on the company site

SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.

Salary: USD 210,000 - 240,000 / annual

Voxel51 is building FiftyOne, a mission-critical platform for managing unstructured data, model development, and AI systems at the world's largest enterprises. The company is at the center of the data-centric AI revolution, with 4 million open-source downloads and impact across self-driving cars, medical imaging, agriculture, and more. The company is fully remote and distributed, with a human-first culture. As a Senior Infrastructure Engineer, you will shape the architecture and strategy of systems powering Voxel51's platform—from individual researchers to Fortune 500 enterprise deployments. You'll lead the design of containerized systems, CI/CD pipelines, and deployment solutions across cloud (GCP, AWS, Azure) and on-premises environments, solving the unique challenges of serving unstructured data (images and video) at scale. Key responsibilities include: - Shape the architecture and evolution of Voxel51's infrastructure to support deployments ranging from researchers to Fortune 500 enterprises - Design, build, and scale deployment systems across cloud and on-premises environments, ensuring reliability, security, and repeatability - Partner with enterprise customers and Customer Success ML Engineers to deliver and support production-grade deployments, guiding installation, troubleshooting, and scaling - Lead infrastructure initiatives across engineering teams, enabling peers to develop, test, and ship features faster with robust internal tooling and automation - Drive best practices in CI/CD, evolving pipelines (currently GitHub Actions + Google Cloud Build) - Develop and maintain deployment solutions for Voxel51-hosted environments (GKE) and customer on-prem installations (Kubernetes or Docker Compose) - Champion developer productivity and improve workflows for development and automated cloud deployments - Troubleshoot and resolve complex infrastructure issues spanning build failures, runtime failures, and customer deployment challenges - Design monitoring, alerting, and predictive solutions for both internal and customer environments - Mentor engineers and set technical direction, ensuring infrastructure remains ahead of customer needs and industry trends Requirements: - Deep experience with containerized environments: building, packaging, and debugging container images - Kubernetes and Docker Compose for orchestration; building, maintaining, and deploying Helm charts - Infrastructure as Code expertise (Terraform, Ansible, or equivalent) - Scripting and automation skills (Bash or similar) - Python expertise, including build and environment management, packaging/distribution, release management, and dependency debugging - CI/CD systems experience, ideally GitHub Actions - Cloud infrastructure knowledge, especially GCP (IAM, VPC, load balancing, ingress/egress routing, proxies, firewall rules) - Database fundamentals, ideally MongoDB or similar NoSQL systems - Observability skills: designing meaningful monitors, logging, tracing, and alerting - Security best practices: certificates, service accounts, least privilege, and role assumptions - Troubleshooting ability across complex, distributed systems (including with customers in the loop) - Testing mindset: comfortable designing and applying different types of tests to validate functionality - Strong communication skills, able to work directly with enterprise customers and collaborate across teams in a remote-first environment - Adaptability and curiosity, with ability to ramp quickly on unfamiliar concepts and technologies The company is fully remote, hiring for people based in the United States, and requires willingness to travel to at least 2 in-person retreats per year.

Similar roles