SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Lambda is an AI cloud infrastructure company serving tens of thousands of customers from AI researchers to enterprises and hyperscalers. The Operations team ensures seamless end-to-end execution of Lambda's AI-IaaS infrastructure and hardware, managing everything from procurement to deployment and operational efficiency.
As a Hardware Quality Engineer, you will be the quality backbone of Lambda's data center operations. You'll track, log, and manage all quality issues arising during deployment and production, performing root cause analysis (RCA) for every failure—hardware, software, or process-related. Your responsibilities include analyzing production system metrics and quality data to detect trends and anomalies, improving turnaround times for Return Merchandise Authorization (RMA) processes, and designing corrective and preventive actions (CAPA).
You'll implement and verify containment actions to keep systems operational while permanent fixes are applied, collaborate across operations, hardware, engineering, supply chain, and vendors to resolve issues, and capture failure analysis reports in the Quality Management System (QMS). You'll verify the quality of incoming and outgoing spares, define and track quality KPIs/SLAs, oversee Material Review Board (MRB) inventory decisions, and ensure the QMS stays current with proper training rollout.
The role requires strong data analysis and statistics skills to turn raw data into actionable insights, expertise in root cause analysis methods (5 Whys, fishbone, 8D, A3), and experience with quality tools or QMS software. You'll manage cross-team communication, stakeholder expectations, and conflict resolution while maintaining a detail-oriented, process-driven mindset. Experience with hardware, data center, or infrastructure systems is essential. Nice-to-have skills include ML/AI infrastructure background, data center standards knowledge, vendor quality experience, firmware/embedded systems understanding, scripting (Python, SQL), and cloud/hyperscaler operations exposure.
The position requires presence in the San Jose office 4 days per week (Tuesday is work-from-home). Up to 30% travel may be required.