SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Abacus Insights is transforming how data works for health plans by making healthcare data usable so decision-makers can act faster with confidence. The company helps health plans break down data silos to create a single, trusted data foundation that powers better decisions, improves outcomes, reduces waste, and delivers better member and provider experiences. Backed by $100M in funding, Abacus is leveraging GenAI use cases through clean, connected, and reliable healthcare data.
The Senior AI/ML Operations Engineer is a senior individual contributor responsible for the infrastructure, pipelines, and operational reliability powering both classical machine learning and generative AI/agentic systems on Databricks and Snowflake. This role blends platform and pipeline engineering, classical ML operations, and GenAI/agentic systems infrastructure with hands-on depth in AI/ML.
Key responsibilities include:
**Platform & Pipeline Engineering:** Deploy and promote ML and GenAI models, pipelines, and code across environments using CI/CD infrastructure. Develop reusable deployment patterns and tooling to reduce effort for new AI/ML use cases. Build and maintain data pipelines supporting classical ML and GenAI workloads from ingestion through feature engineering to serving. Operate within Databricks and Snowflake governance frameworks (Unity Catalog, access controls, environment boundaries) to ensure secure, compliant promotion of code, data, and models. Independently diagnose and resolve production issues across pipelines, infrastructure, and model-serving systems.
**Classical ML Operations:** Automate and monitor production ML inference and feature engineering workflows with alerting and incident response. Own model lifecycle management using MLflow and Unity Catalog for experiment tracking, model registration, versioning, and controlled promotion across environments.
**GenAI & Agentic Infrastructure:** Build and maintain infrastructure for retrieval-augmented generation (RAG) systems, including vector search indexing and retrieval pipelines. Deploy, host, and maintain MCP servers and tool integrations for agentic applications. Build and maintain evaluation infrastructure for AI systems and contribute to evaluation methodology in partnership with AI engineering. Support agent observability through logging, tracing, and monitoring for agent and model behavior in production.
**Cross-Functional Work:** Partner with business stakeholders to scope data and feature requirements. Coordinate with Software Engineering, Data Engineering, Data Science, security, and DevOps on infrastructure changes and shared platform needs. Mentor junior engineers on platform practices and operational standards. Occasionally contribute to customer-specific implementation work such as semantic layer configuration and domain-specific analytics builds.
The role offers unlimited paid time off, work-from-anywhere flexibility, comprehensive health coverage, equity for every employee, a growth-focused environment, home office setup allowance, and monthly cell phone allowance.
**Requirements:**
- 5+ years of experience in AI/ML engineering, MLOps, or closely related discipline
- Extensive hands-on experience in AI/ML with meaningful depth in at least one of: (1) AI/GenAI including agentic frameworks (e.g., LangChain), RAG systems, vector search, MCP or comparable tool-integration protocols, model serving/gateway layers, evaluation design; or (2) Classical ML including model development across common algorithm families (XGBoost, gradient boosting, random forest, neural networks), feature engineering, training pipelines, production deployment. Some working exposure to the other area required.
- Deep, hands-on experience with Databricks and/or Snowflake, including ML/AI pipeline development, working within governance/access-control frameworks (Unity Catalog), and integrating with existing CI/CD infrastructure
- Strong hands-on experience with a model lifecycle/registry tool such as MLflow: experiment tracking, model registration, versioning, and promotion across environments
- Strong proficiency in Python and SQL
- Demonstrated ability to independently diagnose and resolve production infrastructure issues
- Excellent communication skills and comfort working directly with non-technical stakeholders
- Track record of proactively learning new tools and frameworks and applying them quickly to real work
**Nice to have:**
- Experience in healthcare, health insurance, or regulated data environments
- Experience building or operating multi-agent systems
- Experience with Mosaic AI Gateway or comparable model-serving/gateway platforms