SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Cloudflare is seeking a Senior Systems Engineer to design and build core infrastructure powering AI inference across its global network. You'll work on real-time voice, frontier open LLMs, and customer-deployed models running on heterogeneous GPU and accelerator fleets across hundreds of cities worldwide.
Key responsibilities include developing and maintaining serverless inference platform components for high availability and scalability; optimizing model scheduling systems to increase efficiency and resource utilization; implementing improvements to request routing logic to reduce latency; driving measurable improvements in platform reliability and resilience; expanding observability stacks (metrics, logging, tracing) and tuning alerts for proactive issue resolution; leading complex cross-functional technical projects from concept through deployment; and mentoring junior engineers while contributing to collaborative engineering culture.
You'll tackle foundational distributed systems and high-performance computing problems: sub-second model cold starts, multi-accelerator workload scheduling, efficient KV cache management, and building a model deployment platform serving both Cloudflare and customers bringing their own models. This role puts you at the center of building an AI inference platform embedded in the internet fabric—something that doesn't exist yet.
Desired qualifications include proven systems engineering experience with distributed, high-performance systems; expert proficiency in Rust programming, particularly in asynchronous environments; deep understanding of networking and application protocols (TCP, HTTP, WebSocket); and solid experience with scaling and performance optimization techniques including load balancing and caching in distributed environments.