SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Salary: USD 211,000 - 249,000 / annual
Horizon3.ai is a fast-growing cybersecurity company building NodeZero, an autonomous pentesting platform that helps organizations proactively find and fix exploitable attack vectors. The platform is used by ITOps/SecOps teams, consulting pentesters, and MSSPs across organizations of all sizes, from educational institutions to Global 100 enterprises.
As a Senior Machine Learning Engineer on the Defensive Agent team, you will own the end-to-end path from research models to production systems serving thousands of customer tenants daily. This role sits at the critical intersection where most AI products fail—bridging the gap between research notebooks and reliable, scalable production capabilities.
You will build and maintain training pipelines (data preparation, fine-tuning, preference optimization, experiment tracking, artifact management), design inference and serving infrastructure (model gateways with provider routing, regional pinning, batching, caching), and own the release machinery for model artifacts including versioning, canary deployments, and rollback procedures. You'll implement comprehensive monitoring for behavioral drift, regression detection, output quality, latency, and per-tenant cost accounting with budget enforcement.
Additionally, you will build data and context pipelines feeding inference (retrieval, embeddings over attack path and remediation data with tenant isolation), optimize cost and latency across the inference path with visibility into tradeoffs, develop core product features in ETL and GraphQL, and partner with researchers to move prototypes into production while feeding production constraints back into research direction.
The team includes former U.S. Special Operations cyber operators, startup engineers, and cybersecurity practitioners committed to solving ineffective security tools, alert fatigue, blind spots, and the cybersecurity skills shortage.
**Requirements:**
- Bachelor's degree in Computer Science, Computer Engineering, or related field (or equivalent practical experience)
- 5+ years professional software engineering experience with strong production Python
- Demonstrated experience taking ML or LLM-backed systems from prototype to production and operating them
- Hands-on experience with ML pipelines and tooling: training/fine-tuning workflows, experiment tracking, artifact and model registries, reproducible data preparation
- Experience building and operating inference or model-serving infrastructure in production, including latency and cost optimization
- Experience with cloud computing platforms (AWS, Azure, GCP) and container technologies (Docker, Kubernetes)
- Solid proficiency in SQL and experience with production data pipelines
**Preferred qualifications:**
- Experience operationalizing LLM or agentic systems
- Hands-on post-training experience: supervised fine-tuning, distillation, preference optimization, or RL with surrounding infrastructure
- Experience with model gateways, multi-provider routing, self-hosted or customer-hosted inference (vLLM, TGI, Bedrock, etc.)
- Experience with GPU infrastructure, quantization, or inference optimization
- Experience with relational (PostgreSQL) and graph (Neo4j) databases, and GraphQL backends
- Experience with observability tooling (Datadog, Prometheus, Grafana) and distributed tracing
- Experience shipping ML into regulated, air-gapped, or customer-controlled environments, or under compliance regimes like FedRAMP