SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Nexxa.ai is building AI systems for heavy industries, enabling autonomous decision-making and operations across manufacturing, infrastructure, and logistics. This Staff DevOps Engineer role offers deep infrastructure ownership at a company where uptime and reliability directly impact physical operations.
You will own and evolve Nexxa's core infrastructure end-to-end, including compute, networking, storage, and deployment systems. Key responsibilities include:
- Design and operate CI/CD pipelines supporting fast, safe iteration across AI, data, and product teams
- Build and maintain infrastructure-as-code (Terraform, Pulumi) for reproducible environments across cloud and on-prem/edge deployments
- Architect and manage Kubernetes-based platforms for training, inference, and application workloads, including GPU scheduling and autoscaling
- Partner with data and AI teams to support data warehouses, lakehouse architectures (Snowflake, BigQuery, Redshift, Databricks), feature stores, embedding indices, and model serving infrastructure
- Define and drive observability practices—metrics, logging, tracing, and alerting—across distributed systems
- Establish reliability practices: SLOs/SLIs, incident response, postmortems, and on-call rotations
- Design for security and compliance across cloud infrastructure, secrets management, and access control, particularly for industrial and legacy-environment integrations
- Make pragmatic tradeoffs across cost, latency, reliability, and developer velocity
- Collaborate with engineering leadership on infrastructure roadmap and platform strategy
- Mentor engineers on infrastructure best practices and raise the bar for operational excellence
Success means owning ambiguous, high-stakes infrastructure problems end-to-end; building systems that stay reliable as scale grows; bringing strong technical judgment on reliability/cost/speed tradeoffs; raising operational rigor across the team; and helping define the platform's future.
REQUIREMENTS:
- 6+ years in DevOps, Site Reliability Engineering, Platform Engineering, or infrastructure-focused software engineering
- Deep hands-on experience with cloud platforms (AWS, GCP, or Azure) at production scale
- Kubernetes in production, including GPU workload scheduling
- Infrastructure-as-code tooling (Terraform, Pulumi, or equivalent)
- CI/CD systems (GitHub Actions, GitLab CI, CircleCI, Jenkins, ArgoCD)
- Strong track record designing and operating observability stacks (Prometheus, Grafana, Datadog, OpenTelemetry)
- Excellent scripting/programming skills (Python, Go, or Bash) for automation and tooling
- Proven ability to independently scope and lead infrastructure projects from design through production rollout
- Strong incident management instincts and ability to lead through outages
PREFERRED:
- Experience operating infrastructure bridging cloud and edge/on-prem environments, especially in industrial or manufacturing contexts
- Familiarity with data warehouse/lakehouse platforms (Snowflake, BigQuery, Redshift, Databricks)
- Experience with service mesh, zero-trust networking, or compliance frameworks for industrial/critical infrastructure (SOC 2, IEC 62443)
- History of building internal developer platforms or self-service infrastructure tooling
- Experience scaling infrastructure teams or setting technical direction at Staff level