SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Telnyx operates a proprietary B300 GPU fleet across global facilities and is building a production inference platform to serve both internal AI-agent traffic and external customers. This role is a founding member of the China team, architecting and operating the complete inference stack from bare metal to OpenAI-compatible endpoints.
You own two core mandates: (1) operate and expand the fleet efficiently—maximizing useful inference throughput per GPU-dollar while meeting latency and reliability SLOs and continuously reducing cost per token; (2) build the product layer on top—serverless inference for open-weight models and dedicated, tuned deployments for enterprises.
Key responsibilities include designing and operating serverless serving pools using vLLM/SGLang with continuous batching, prefix caching, low-precision serving (FP8, FP4, INT4), and MoE expert parallelism; building the fleet routing layer with KV-cache-aware routing, prefill/decode disaggregation, KV tiering, and multi-LoRA serving; operating Kubernetes on bare metal with GPU Operator, topology-aware scheduling, LeaderWorkerSet, and Kueue; implementing weight logistics and elasticity via P2P model distribution, warm pools, and inference-driven autoscaling; operating the dedicated tenant tier with per-tenant pools, GPU-hour metering, latency SLOs, and model-tuning integration; and building observability and capacity-planning infrastructure grounded in roofline analysis.
The stack is open-source from bare metal to endpoint. You will evaluate, benchmark, and integrate components including vLLM, SGLang, llm-d, NVIDIA Dynamo, Kubernetes Gateway API, Envoy, Ray Serve, Mooncake, Dragonfly, Kata Containers, OpenStack Ironic, Prometheus, and DCGM. You work upstream in the core serving stack and make architectural decisions about which components earn their place in production.
As a founding member of the China team, you also have input into hiring and team structure. The role is fully remote, based in mainland China, with no relocation required. Telnyx is a profitable, financially stable company with a global async-friendly culture. Open-source contribution is part of the job, and conference travel is supported.
REQUIREMENTS:
- Owned production LLM serving under meaningful traffic and latency constraints; ability to explain architecture, diagnosed bottlenecks, interventions made, and measured improvements in latency, reliability, or cost. Scale of thousands of GPUs or millions of requests per day is a strong signal.
- Kubernetes on GPU fleets end-to-end: GPU Operator, device plugins, node pools, topology-aware placement, gang scheduling (LeaderWorkerSet, Kueue), GitOps rollouts, and debugging pod placement.
- Deep operational command of vLLM or SGLang: deploying, tuning, upgrading (parallelism with TP/EP, quantization, batch and KV-cache settings, prefix caching, disaggregation), and knowing which knob moves which metric.
- Performance engineering at system level: reading engine and DCGM metrics, reasoning from roofline, sizing deployments with numbers.
- Python and Go for automation; real Linux, networking, and storage depth. No requirement to write CUDA; must know when a problem is one and route it upstream.
- Work with community in English and Chinese; write runbooks and design docs.
EXPERIENCE ESPECIALLY VALUED:
- Operated inference platforms at scale of Ant Group, Alibaba Cloud/PAI, Tencent, ByteDance, Baidu, Huawei, DaoCloud, Moonshot, DeepSeek, or comparable teams.
- Contributor to or heavy production user of vLLM, SGLang, llm-d, NVIDIA Dynamo, Ray/KubeRay, Kata Containers, Volcano, Kueue, HAMi, Dragonfly, Mooncake, or Envoy.
- Presence in CNCF/OpenInfra community: talks at KubeCon China, OpenInfra Days, vLLM or SGLang meetups.
NICE TO HAVE:
- Fine-tuning/RL infrastructure: LoRA pipelines, evaluation harnesses, champion/challenger rollout.
- Real-time voice latency work: sub-second time-to-first-token budgets on live calls.
- Multi-region deployments and data-residency requirements.