SlipstreamJobsFresh Startup & VC-Backed Jobs

Lead Engineer, Issue Management & Triage

Diligent Robotics - Austin, TX, United States - Hybrid - posted 2026-09-11

Apply on the company site

SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.

Diligent Robotics builds autonomous robots that work safely in real-world environments. This role leads the systems, tooling, and team at the intersection of customer operations, the Remote Operations Center (ROC), and Engineering—focusing on how issues are detected, triaged, diagnosed, and resolved across a deployed robotic fleet. You will own end-to-end issue management and triage systems, designing severity frameworks, SLAs, and ensuring issues are consistently structured for engineering prioritization. You'll build automation pipelines and internal tools (primarily in Python and backend services) to ingest, process, and classify operational data, reducing manual triage effort. This is a hands-on role where you contribute directly to codebases and partner with Engineering on system integrations including logs, telemetry, and alerts. You'll act as the primary technical interface between the ROC and Engineering teams, translating real-world operational issues into prioritized, categorized technical problems. You'll develop performance measurement systems and taxonomies to systematically classify robot failure modes and degradation across the fleet, building dashboards and reporting systems to track trends, severity, and impact. A key responsibility is establishing best practices for Root Cause Analysis (RCA), identifying systemic issues, and driving long-term fixes while creating feedback loops to influence improvements in hardware, software, and autonomy. The role combines remote data analysis with hands-on bench/lab failure analysis at the Austin HQ, investigating how and why robots fail in the field across mobile base, charging/docking, motion, power, connectivity, and sensor hardware. You'll work with engineering, operations, manufacturing, and vendors to turn field reports and fleet data into validated root causes and corrective actions. Location: Austin preferred; remote possible within the US with up to ~50% travel to Austin, TX (especially in the first 90 days). REQUIREMENTS: - 7+ years in relevant technical or program management roles (e.g., engineering, incident management) - 3+ years of people management experience - Experience with complex, real-world systems (robotics, autonomous/distributed systems, or hardware-software products) - Proven track record building operational tools, systems, or infrastructure for workflows - Strong programming experience (Python preferred; backend or data systems experience a plus) - Experience with data pipelines and telemetry systems, monitoring/alerting/logging infrastructure, and internal tools/automation systems - Ability to design scalable systems for classification, prioritization, and workflow automation - Familiarity with platforms like Jira, Zendesk, SQL, Looker, Foxglove, or similar - Strong systems thinking; ability to translate ambiguous operational problems into structured technical solutions - Experience defining metrics, taxonomies, and performance frameworks - Data-driven approach to prioritization and decision-making - Hands-on mindset with willingness to dive into technical problems; strong ownership and bias toward action - Comfortable in fast-paced, scaling environments with passion for improving real-world system performance and reliability

Similar roles