SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Lambda is a leader in AI cloud infrastructure serving tens of thousands of customers from AI researchers to enterprises and hyperscalers. The GTM team powers growth by enabling customers to realize business goals with AI infrastructure, partnering with leading AI researchers and enterprise engineering teams to design, scale, and optimize high-performance GPU cloud solutions.
In this role, you will drive technical sales and executive influence by partnering with Account Executives to lead complex deals with large enterprises and digital native businesses. You'll build trusted relationships with technical leaders (CTOs, Heads of AI/ML, Platform Leads), evaluate customer architectural needs, uncover bottlenecks, and design end-to-end GPU cloud solutions. You'll author comprehensive proposals, architecture diagrams, and collaborate on Bills of Materials and rack elevations for multi-node GPU clusters.
You'll lead hands-on proof-of-concept activities and benchmarking for customers, designing and executing technical PoCs and custom prototypes to demonstrate Lambda's performance, reliability, and value. You'll run benchmark evaluations across training and inference workloads to show tangible performance and cost advantages.
Architecturally, you'll guide enterprise engineering teams on structuring their AI lifecycle—from data ingestion and distributed training (SLURM, Kubernetes) to inference optimization (vLLM, TensorRT-LLM) and observability. You'll provide guidance on high-performance networking (InfiniBand, RoCE), distributed storage, and cluster topologies to ensure maximum GPU utilization.
You'll champion customer feedback and product advocacy, serving as the technical voice of the customer internally, funneling field insights and feature requests to Product and Engineering teams. You'll create field enablement assets, technical whitepapers, architectural blueprints, and lead technical workshops for prospective clients and partners. You'll represent Lambda as a subject matter expert at industry conferences, webinars, and technical community events.
Required: 8+ years designing, deploying, and scaling enterprise cloud infrastructure; 4+ years in Solution Architect/Engineer or technical customer-facing roles supporting complex cloud environments; 3+ years architecting and deploying cloud-based AI/ML workloads; proven track record deploying, benchmarking, and optimizing workloads on NVIDIA GPU architectures (HGX platforms, NVLink) using deep learning frameworks (PyTorch, NeMo) and inference engines (vLLM, TensorRT-LLM); strong experience with Kubernetes, Docker, SLURM, Terraform, Ansible; deep knowledge of cloud networking (InfiniBand, RoCE, distributed file systems); coding experience in Python, Go, C/C++ (CUDA); experience partnering with Account Executives on complex cloud deals and presenting to C-level stakeholders; demonstrated impact mentoring junior SEs or architects; thrives in dynamic settings with radical ownership.
Nice to have: Direct experience with end-to-end LLM fine-tuning, algorithm selection, pipeline design, or distributed training (3D parallelism, Megatron-LM); prior experience with product launches, GTM initiatives, or publishing technical whitepapers/benchmarks; experience integrating RESTful APIs, gRPC, and service-oriented cloud architectures.
Position requires presence in San Francisco, San Jose, or Bellevue office 4 days per week; Tuesday is designated work-from-home day.