SlipstreamJobsFresh Startup & VC-Backed Jobs

Software Engineer, Inference Runtime

LM Studio - New York, NY, United States - Hybrid - posted 2026-08-07

Apply on the company site

SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.

LM Studio is seeking a Software Engineer to advance its inference runtime stack for both on-device and cloud execution. The role focuses on integrating new inference engines, optimizing model execution across diverse CPU and GPU targets, and bringing up new open-weight models and modalities. You'll work on performance-critical infrastructure that powers LM Studio's platform, used by millions globally. Key responsibilities include maintaining and improving the inference stack across multiple hardware backends (CPU, CUDA, Metal, Vulkan, ROCm), bringing up new model architectures and multimodal models, and optimizing for latency, throughput, memory efficiency, and reliability. You'll build runtime capabilities for model loading, batching, scheduling, caching, and distributed execution. The role also involves benchmarking and diagnosing correctness and performance issues, and contributing upstream to open-source projects like llama.cpp and MLX. Required qualifications include significant production experience with ML systems, inference runtimes, or performance-sensitive infrastructure. You need strong programming ability in Python and C++, deep understanding of transformer architectures and model inference mechanics, and hands-on experience profiling CPU/GPU workloads. Familiarity with PyTorch and inference systems such as llama.cpp, MLX, ExecuTorch, vLLM, SGLang, or TensorRT-LLM is essential. Strong debugging skills across model code, runtime internals, operating systems, and CPU/GPU execution are required. Bonus qualifications include past contributions to open-source inference runtime projects. The team values high technical intensity, personal responsibility, curiosity, and self-motivation. The company prioritizes putting humans at the center and building tools they'd recommend to friends and family. The office is located in SoHo, NYC, with flexible work-from-home options.

Similar roles