SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Ambient.ai is the category leader in Agentic Physical Security, powered by Ambient Pulsar, a reasoning Vision-Language Model purpose-built for physical security. The platform integrates with existing security cameras and access control systems to provide unified monitoring, threat assessment, and response capabilities. The company has achieved strong momentum, doubling new ARR in FY26, and serves world-class customers including Cisco, ServiceNow, SentinelOne, TikTok, Bayer, and MoMA. Founded in 2017 and backed by Andreessen Horowitz, Y Combinator, and Allegion Ventures, Ambient.ai is on a mission to prevent every security incident possible.
In this role, you will design, build, and optimize the AI infrastructure powering Ambient.ai's real-time intelligence platform. You will work on systems required to run state-of-the-art deep learning models across terabytes of video data in real time, building and scaling infrastructure for inference, evaluation, and continuous model improvement across computer vision models, large language models, large vision models, and multimodal AI systems.
Key responsibilities include:
- Design, build, and maintain cutting-edge AI infrastructure for real-time computer vision, LLM, LVM, and multimodal inference workloads
- Build scalable systems for running state-of-the-art models across large volumes of video and sensor data
- Optimize inference performance across latency, throughput, GPU utilization, reliability, and cost
- Develop robust evaluation harnesses and benchmarking systems to measure model quality and production readiness
- Build infrastructure for continuous model evaluation, experimentation, and deployment
- Partner with research scientists to productionize advances in computer vision, LLMs, LVMs, RAG, and multimodal AI
- Improve model-serving architecture including batching, caching, routing, quantization, and model parallelism
- Develop data engines and feedback loops for collecting training data and continuously improving AI performance
- Create reliable observability, monitoring, and debugging tools for production AI systems
- Help define best practices for deploying and operating AI systems in enterprise environments
You will report to Raghu Nallamothu and work closely with research scientists and product engineering teams to bring the latest AI advancements into production.
The company operates on a hybrid model with three days per week in the Redwood City office, plus Friday collaboration days for Bay Area employees.
REQUIREMENTS:
- 2+ years of industry experience building infrastructure, distributed systems, machine learning platforms, or production AI systems
- BS/MS in Computer Science or related technical field, or equivalent practical experience
- Strong programming background, especially in Python, with solid software engineering fundamentals
- Experience designing and building scalable machine learning infrastructure for training, inference, evaluation, and deployment
- Hands-on experience running deep learning models in production, ideally including LLMs, LVMs, vision-language models, or multimodal models
- Strong understanding of inference optimization techniques including batching, caching, quantization, parallelism, memory optimization, and GPU utilization
- Experience with model-serving frameworks such as vLLM, Triton Inference Server, or similar technologies
- Experience building evaluation frameworks, test harnesses, benchmarks, regression tests, or model-quality measurement systems
- Strong background in machine learning and deep learning; computer vision experience is a strong plus
- Experience designing data engines or pipelines for collecting and curating training and evaluation data
- Familiarity with integrating advanced AI systems such as LLMs, LVMs, RAG pipelines, embedding models, or multimodal models into production applications
- Experience with cloud infrastructure, containers, orchestration, distributed systems, and GPU-based workloads
- Strong collaboration and communication skills with ability to work effectively with research scientists, product teams, and infrastructure teams
- Proactive problem-solving ability, strong ownership mindset, and adaptability to new AI technologies
NICE TO HAVE:
- Experience operating large-scale GPU infrastructure or distributed inference systems
- Experience with CUDA, NCCL, PyTorch, TensorRT, ONNX, or similar ML systems technologies
- Experience with video understanding, real-time computer vision, multimodal AI, or physical-world AI systems
- Experience with model compression, speculative decoding, distillation, pruning, or low-latency serving techniques
- Experience with prompt evaluation, model regression testing, human-in-the-loop evaluation, or automated quality gates
- Familiarity with retrieval-augmented generation, vector databases, embedding models, or search infrastructure
- Experience building internal ML platforms or tools used by researchers and applied ML teams