SlipstreamJobsFresh Startup & VC-Backed Jobs

Lead Software Engineer, Model Serving Platform

Sciforium - San Francisco, CA, USA - In-office - posted 2026-08-17

Apply on the company site

SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.

Sciforium is an AI infrastructure company building next-generation multimodal AI models and a proprietary high-efficiency serving platform, backed by multi-million-dollar funding and direct AMD sponsorship. The company is scaling rapidly to develop the full stack powering frontier AI models and real-time applications. In this role, you will architect and lead development of Sciforium's model serving platform—the high-performance engine delivering multimodal, efficient foundation models to market. As a senior technical leader, you'll build core components yourself while guiding and mentoring other engineers, shaping engineering direction, standards, and execution quality. You will own the full AI inference stack: GPU kernels and quantized execution paths, distributed serving, scheduling, and APIs powering real-time AI applications. Key responsibilities include leading technical direction and architecture decisions, building core serving components (execution runtimes, batching, scheduling, distributed inference), developing high-performance C++ and CUDA/HIP modules with custom GPU kernels, collaborating with ML researchers to productionize multimodal models, building Python APIs and services, mentoring engineers through code reviews and design discussions, driving performance profiling and observability, ensuring reliability through testing and best practices, and troubleshooting complex issues across GPU, runtime, and service layers. Ideal candidates have a Bachelor's in Computer Science or equivalent, 5+ years designing scalable backend systems or distributed infrastructure, strong understanding of LLM inference mechanics (prefill vs decode, batching, KV cache), experience with Kubernetes/Ray and containerization, strong C++ and Python proficiency, advanced debugging and performance optimization skills, and ability to collaborate with ML researchers and lead technical discussions. Nice-to-have qualifications include ML systems engineering experience, distributed GPU scheduling, familiarity with open-source inference engines (vLLM, Sglang, TRT-LLM), large-scale ML/MLOps infrastructure, CUDA/ROCm proficiency, experience at AI/ML startups or Big Tech infrastructure teams, knowledge of multimodal architectures and efficient inference techniques, and open-source contributions.

Similar roles