SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Augury is a leader in Industrial AI, helping manufacturers leverage real-time production insights to drive efficiency and minimize downtime. As an MLOps Engineer, you will design and scale the production platform behind Augury's Industrial AI Workforce, enabling teams across the company to develop, evaluate, deploy, and operate ML and AI systems consistently and safely.
Key responsibilities include:
- Design and evolve production MLOps capabilities across the full ML lifecycle: datasets, features, models, evaluations, deployments, monitoring, retraining, and feedback signals.
- Build systems for experiment tracking, artifact management, reproducibility, versioning, lineage, promotion workflows, and production readiness.
- Develop reusable platform tooling, golden paths, and engineering standards that improve consistency and delivery velocity across teams.
- Build operational infrastructure for LLM and agentic systems including prompts, tools, traces, evaluations, observability, safety boundaries, and production monitoring.
- Design evaluation and monitoring frameworks for AI systems including answer quality, latency, grounding, reliability, and operational regressions.
- Build and optimize large-scale training pipelines supporting heterogeneous data sources and scalable compute patterns.
- Write clean, modular, production-grade Python services and platform libraries.
- Drive engineering quality through automated testing, CI/CD, observability, deployment standards, and operational best practices.
Required qualifications:
- 5+ years of professional software engineering, MLOps, or ML platform engineering experience in production environments.
- Significant experience building or owning production ML infrastructure and lifecycle systems.
- Strong Python engineering skills with production-grade architecture, modular design, testing, packaging, and robust error handling.
- Strong understanding of the end-to-end ML lifecycle including training, deployment, monitoring, retraining, reproducibility, and lineage.
- Experience with large-scale data platforms such as Databricks, Spark, Delta Lake, or equivalent ecosystems.
- Experience with ML platform and MLOps frameworks such as MLflow, Metaflow, Kubeflow, or equivalent ML lifecycle-management systems.
- Proven ability to design reusable workflow orchestration using Airflow, Metaflow, or Databricks.
- Familiarity with operational patterns for LLMOps, AgentOps, and production AI systems.
- Strong written and verbal communication skills in English.
Nice-to-have experience includes industrial, IoT, or manufacturing platforms; feature stores, model registries, dataset versioning, and lineage systems; and AI agents, RAG systems, production GenAI applications, or evaluation frameworks.