SlipstreamJobsFresh Startup & VC-Backed Jobs

Member of Technical Staff - ML Operations

Veeda AI - Toronto, ON, Canada - In-office - posted 2026-09-04

Apply on the company site

SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.

Veeda AI is building multimodal foundation world models for Physical AI. We're a small, fast-moving team of engineers and researchers from leading AI labs, tackling challenging problems at the intersection of AI, robotics, and embodied intelligence. In this role, you will own critical infrastructure and tooling across the ML operations stack: **Experiment Lifecycle Tracking & Tooling**: Design, build, and deploy systems that manage how runs are defined, launched, resumed, and terminated. Build and operate experiment databases for code version tracking, data version tracking, reproducibility, and checkpoint ancestry. **Inference Fleet Orchestration**: Design, build, and operate the serving control plane that handles large volumes of concurrent client requests and distributes them across inference clusters. Develop cache-aware admission, routing, batching, and scheduling policies that improve cache locality, balance workload, protect tail latency, and maintain fleet reliability and utilization. Partner with ML Performance teams on model runtime and per-worker throughput optimization. **Model Evaluation in CI**: Design, build, and operate automatic model checkpoint evaluation systems that run per-change and nightly on seeded rollout and policy-success suites. **Data Pipeline Operations**: Design, build, and operate high-performance, fault-tolerant, distributed backend services and event-driven systems for large-scale data processing. **Visualization Platforms**: Design, build, and deploy experiment observability and dataset visualization platforms with interactive data visualization, progress tracking, search, and comparison capabilities. **End-to-End Ownership**: Lead projects through the complete software lifecycle, including technical specs, implementation, CI/CD, on-call support, and production observability. **Requirements:** - Bachelor's degree or equivalent hands-on experience in Computer Science, Engineering, or related technical field - Proficiency in at least one scripting language (Python or Bash) and one compiled language (Java, Rust, or Go) - Experience shipping production-quality developer tools - Proficiency in CI/CD automations (pipelines, runners, deployment) - Experience building reproducible pipelines end-to-end, with clear understanding of which parts are bit-reproducible and why - At least one of: (1) proven understanding of event-driven architecture, concurrency models, fault tolerance, and data consistency patterns; (2) experience building or operating large-scale inference control planes or distributed serving infrastructure with hands-on work in traffic management, admission control, request scheduling, routing, batching, or cache-aware load balancing; (3) experience with multi-node workloads and SLURM-based scheduling systems; (4) experience working on in-production model evaluation frameworks, particularly regression testing of large multimodal models **Nice to Have:** - Run experiment tracking at scale with video, 3D, and trajectory artifacts - Built evaluation harnesses for generative or embodied models - Full stack development with web-based front-end - Orchestrated ML workflows with Argo Workflows, Flyte, or Ray - Built GPU-hour attribution systems - Open-source ML tooling contributions or publications on evaluation/reproducibility

Similar roles