SlipstreamJobsFresh Startup & VC-Backed Jobs

Staff HPC Systems Architect

Lambda - San Jose, CA, USA - Hybrid - posted 2026-09-14

Apply on the company site

SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.

Lambda is a leader in AI cloud infrastructure serving tens of thousands of customers, from AI researchers to enterprises and hyperscalers. The company's mission is to make compute as ubiquitous as electricity and give everyone access to superintelligence. In this Staff HPC Systems Architect role, you will architect and define scalable compute platforms optimized for AI/ML, simulation, and high-throughput workloads. You'll develop compute system standards and design patterns to ensure consistency, performance, and maintainability across infrastructure. Key responsibilities include evaluating emerging CPU, GPU, and accelerator technologies, owning architectural tradeoff decisions that impact compute density, power, cooling, and total cost of ownership. You will collaborate closely with product and engineering teams to map workload requirements to compute platform capabilities across bare metal and cloud deployments. You'll convert ambiguous business or customer needs into measurable platform requirements, technical specifications, and acceptance criteria. You'll define compute platform roadmaps and architectural reference designs that guide hardware selection, firmware baselines, rack-level, and cluster design. As a technical lead during new platform introductions, you'll guide validation and performance characterization efforts. You'll mentor systems engineers and cross-functional stakeholders on compute performance tuning, sizing, and architectural decisions, influencing technical strategy across teams. The role requires presence in the San Jose, San Francisco, or Bellevue office 4 days per week, with Tuesday designated as the work-from-home day. REQUIREMENTS: - 7+ years of proven experience architecting large-scale 10k-100k+ GPU HPC or cloud compute platforms - Deep knowledge of CPU/GPU architectures, memory hierarchies, and accelerator topologies - Experience designing systems around high-bandwidth, low-latency fabrics (NVLink, InfiniBand, RoCE) - Strong understanding of system performance tuning, resource scheduling, thermal and power optimization, and compute lifecycle management - Comfortable working across hardware and software boundaries, especially at the intersection of compute architecture, OS behavior, and orchestration layers - Skilled at balancing architectural tradeoffs for density, power efficiency, cooling, and performance - Strong analytical and communication skills with a track record of influencing technical strategy - Strong ownership and can-do attitude; self-starter comfortable working in ambiguity NICE TO HAVE: - Hands-on experience with AI/ML workloads and their compute performance characteristics - Familiarity with orchestration tools used in HPC (Slurm, Kubernetes, etc.) - Experience with virtualization technologies, specifically GPU virtualization - Exposure to hardware validation, vendor collaboration, and long-term OEM roadmap alignment - Background in compute telemetry, real-time performance profiling, or large-scale A/B infrastructure testing

Similar roles