SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Desty is building an internal AI platform that product teams across the company depend on—a critical layer between data, AI, and engineering teams that connects to external model providers and self-hosted GPUs, ultimately serving thousands of businesses and millions of users.
You will own full-stack work end-to-end: designing and building secure REST APIs and real-time streaming endpoints, backend services that handle concurrency and partial failure with 99.9% availability targets, web interfaces that make complex systems legible to operators, and careful data modeling in PostgreSQL. You will build and operate a unified API gateway over external model providers and self-hosted models on your own GPUs, with low-latency routing, load balancing, and failover. You will implement multi-tenant boundaries (OAuth2, role-based access control, quotas, rate limits), measure usage and cost accurately for billing, instrument systems for incident debugging, and participate in delivery through code review, CI/CD, progressive rollout, and post-incident analysis.
You will work closely with Product Managers and internal teams consuming the platform to turn their needs into adopted solutions, collaborate with the Technical Program Manager on the development lifecycle (concept, design, test, release, support), and partner with DevOps on platform operations and maintenance. You will document technical requirements, API contracts, deployment notes, and post-mortems, and mentor other engineers while raising the bar on design and code quality.
This role is ideal for a senior engineer who has shipped LLM-backed features to real users and understands the practical differences between model providers, streaming responses, tool calling, embeddings/RAG, multimodal input, and batch APIs. You should have deployed and operated self-hosted models in production (LLMs, embeddings, speech-to-text, text-to-speech, multimodal) using inference servers like vLLM, SGLang, TGI, or Triton, and built agent workflows with appropriate guardrails. You will balance cost and latency as product constraints and use AI-assisted development tools effectively.
REQUIREMENTS:
- 4+ years of software engineering experience in a team setting, building and running production web applications
- Strong JavaScript and TypeScript with production Node.js experience
- Production experience with at least one modern frontend framework (Vue preferred; React or Next.js acceptable)
- Working knowledge of Go language
- PostgreSQL in production: schema design, indexing, transactions, query diagnosis
- REST API design with practical experience in streaming responses and long-lived connections
- Testing as part of change (TDD or close to it) and comfort with code review
- Good understanding of microservices design patterns
- Demonstrated ownership of scalability and reliability in high-traffic systems, including API gateways, load balancing, and availability/error-rate targets
- Security fundamentals: OAuth2, role-based access control, credential handling, tenant isolation, input validation, least privilege
- Docker and container orchestration with Kubernetes and Helm
- CI/CD pipelines and Git-based workflows
- Production experience with a public cloud (Alibaba Cloud, AWS, GCP, or Azure)
- Shipped LLM-backed features to real users (not prototypes)
- Familiarity with multiple model providers (OpenAI, Anthropic, others) and aggregators
- Practical grasp of streaming responses, tool/function calling, embeddings/RAG, multimodal input, batch and file APIs
- Experience deploying and operating self-hosted models in production with inference servers (vLLM, SGLang, TGI, Triton)
- Experience building agent workflows with guardrails
- Ability to evaluate model output quality (evaluation sets, human review, or production signals)
- Awareness of cost and latency as product constraints
- Working proficiency with AI-assisted development tools (Claude Code, Codex)
- Experience collaborating with product teams and AI engineers
- Familiarity with Scrum and Kanban
- Strong written and verbal communication
- Ability to build and deploy solutions independently
NICE TO HAVE:
- Experience deploying models across multiple GPUs or nodes, or fine-tuning models
- Microfrontend architecture experience (Module Federation, single-spa)
- Ruby framework experience (Rails, Sinatra)
- Experience building internal developer platforms or APIs consumed by other engineering teams