SlipstreamJobsFresh Startup & VC-Backed Jobs

Sr. Software Engineer, AI / ML Inference Platform

Dialpad - Buenos Aires, Argentina - In-office - posted 2026-09-07

Apply on the company site

SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.

Dialpad is an AI platform for customer experience that deploys AI agents to resolve customer problems in real time across voice and digital channels. The company is backed by Andreessen Horowitz, GV, ICONIQ Capital, and T-Mobile, and serves market-leading brands including Netflix, Motorola Solutions, and Randstad. The AI/ML Platform team builds and operates the shared infrastructure that takes Dialpad's model-backed capabilities from training through production inference. This includes GPU training infrastructure, model evaluation and lifecycle tooling, and production inference systems running on NVIDIA GPUs in GCP. As a Senior Software Engineer on this team, you will focus on inference as a center of gravity—turning trained models into reliable, observable, efficient production services. The role is intentionally end-to-end: you'll work across training clusters, model artifacts, evaluation and release workflows, serving runtimes, production operations, and feedback loops. Key responsibilities include: - Design, build, and improve shared platform capabilities spanning model training, evaluation, artifact management, release, production inference, and operational feedback - Build and operate shared GPU training infrastructure providing scientists with reliable, reproducible, and efficient environments - Improve training-cluster scheduling, workload isolation, capacity management, storage, networking, observability, and accelerator utilization - Develop production-serving pathways for low-latency, high-throughput, and highly available inference workloads - Integrate and adapt model-training frameworks and inference runtimes to meet Dialpad's requirements for automation, observability, security, and operational control - Improve GPU workload performance and efficiency by reasoning across compute, memory, storage, networking, batching, concurrency, and scheduling - Partner with ASR and NLP scientists to translate evolving model capabilities into scalable production designs - Counsel scientific teams on production concerns including reproducibility, evaluation coverage, artifact design, resource requirements, and serving feasibility You will serve as a senior engineering partner to ASR and NLP scientists, helping teams reason about reproducibility, evaluation, scalability, hardware and runtime constraints, latency, reliability, cost, and release safety. You are not expected to conduct original ML research, but must understand training, data, evaluation, and model behavior well enough to translate scientific work into dependable enterprise ML systems. This is an implementation-heavy engineering role focused on building systems directly and leading substantial technical work.

Similar roles