SlipstreamJobsFresh Startup & VC-Backed Jobs

Staff DevOps Engineer

Nexxa.ai - San Francisco, CA, USA - In-office - posted 2026-08-27

Apply on the company site

SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.

Nexxa.ai is building AI systems for heavy industries, enabling autonomous decision-making and operations across manufacturing, infrastructure, and logistics. The company focuses on translating deep technical breakthroughs into operational reality for some of the hardest systems-level problems in industry. As a Staff DevOps Engineer, you will own and evolve Nexxa's core infrastructure end-to-end, including compute, networking, storage, and deployment systems. You'll design and operate CI/CD pipelines supporting fast, safe iteration across AI, data, and product engineering teams. Your responsibilities include building and maintaining infrastructure-as-code (Terraform, Pulumi) for reproducible environments across cloud and on-prem/edge deployments, and architecting Kubernetes-based platforms for training, inference, and application workloads with GPU scheduling and autoscaling. You'll partner closely with data and AI teams to support data warehouses and lakehouse architectures (Snowflake, BigQuery, Redshift, Databricks), feature stores, embedding indices, retrieval pipelines, and model training/serving infrastructure. You'll define and drive observability practices—metrics, logging, tracing, alerting—across distributed systems, and establish reliability practices including SLOs/SLIs, incident response, postmortems, and on-call rotations. Additional responsibilities include designing for security and compliance across cloud infrastructure, secrets management, and access control (particularly relevant to industrial and legacy-environment integrations), making pragmatic tradeoffs across cost, latency, reliability, and developer velocity, collaborating with engineering leadership on infrastructure roadmap and platform strategy, and mentoring engineers on infrastructure best practices. Success means owning ambiguous, high-stakes infrastructure problems end-to-end; building systems that stay reliable as usage and scale grow; bringing strong technical judgment on reliability/cost/speed tradeoffs; raising the bar for operational rigor across the team; and helping define the platform's future direction. REQUIRED QUALIFICATIONS: - 6+ years of experience in DevOps, Site Reliability Engineering, Platform Engineering, or infrastructure-focused software engineering - Deep hands-on experience with cloud platforms (AWS, GCP, or Azure) at production scale - Kubernetes in production, including GPU workload scheduling - Infrastructure-as-code tooling (Terraform, Pulumi, or equivalent) - CI/CD systems (GitHub Actions, GitLab CI, CircleCI, Jenkins, ArgoCD) - Strong track record designing and operating observability stacks (Prometheus, Grafana, Datadog, OpenTelemetry) - Experience supporting ML/AI infrastructure (training clusters, model serving, data pipelines) strongly preferred - Excellent scripting/programming skills (Python, Go, or Bash) for automation and tooling - Proven ability to independently scope and lead infrastructure projects from design through production rollout - Strong incident management instincts and ability to lead through outages calmly PREFERRED QUALIFICATIONS: - Experience operating infrastructure bridging cloud and edge/on-prem environments, especially in industrial or manufacturing contexts - Familiarity with data warehouse/lakehouse platforms (Snowflake, BigQuery, Redshift, Databricks) - Experience with service mesh, zero-trust networking, or compliance frameworks for industrial/critical infrastructure (SOC 2, IEC 62443) - History of building internal developer platforms or self-service infrastructure tooling - Experience scaling infrastructure teams or setting technical direction at Staff level

Similar roles