SlipstreamJobsFresh Startup & VC-Backed Jobs

Director Site Reliability Engineer, AI Infrastructure

d-Matrix - Santa Clara, CA, United States - In-office

Apply on the company site

SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.

d-Matrix designs and manufactures purpose-built AI inference silicon. This Director-level role builds and leads the Site Reliability Engineering function from the ground up, owning the infrastructure that development, validation, and customer-facing deployments run on across colocation facilities, on-premises lab clusters, cloud environments (AWS, Azure, GCP), and customer-facing platform services. You will hire and grow the SRE team (1–3 engineers initially), set technical direction, own SLOs for critical systems, and serve as the senior escalation point for incidents while remaining hands-on. You will also direct a dedicated Data Center & Lab Technician team, setting work priorities and operational standards across on-premises and colocation facilities. Key responsibilities include: - Building the SRE function from scratch: define its charter, establish SRE as a discipline, and hire/develop/retain the team - Owning 24×7 reliability across colocation, on-premises, cloud, and customer-facing services; designing for failure domains and progressive delivery - Establishing SRE processes: define SLIs/SLOs and error budgets, design on-call rotations (including follow-the-sun model), and build incident/RCA management framework - Owning the full observability stack (Prometheus, Grafana, Datadog, Splunk or equivalent), instrumenting metrics, traces, and logs - Driving Infrastructure as Code automation (Terraform, Ansible) and building self-healing infrastructure - Owning FinOps and capacity planning across cloud, colocation, and on-premises, including spend attribution and TCO modeling - Leading migration from ad-hoc JBOD storage to enterprise-grade shared storage platform, including architecture, vendor selection, and DR design - Partnering with Software Tools, QA DevOps, Engineering, Networking Operations, and IT Operations teams Requirements: - Bachelor's or Master's in Computer Science, Electrical Engineering, or related field - 15+ years in SRE, infrastructure engineering, or production engineering - 5+ years leading SRE or infrastructure engineering teams, including experience building or significantly rebuilding a function (not just managing steady-state) - Demonstrated track record of establishing SRE as a discipline in an organization that lacked it: defining SLOs, creating on-call frameworks, standing up observability, driving cultural change with reactive ops teams - Deep Linux systems expertise: bare-metal operations, enterprise shared storage platforms (NAS/SAN, NFS/SMB at scale, snapshot and replication architectures), hybrid-cloud storage integration - Proven experience operating colocation and on-premises hardware at scale: server lifecycle, power/cooling awareness, rack-level networking - Hands-on Infrastructure as Code fluency with Terraform and Ansible at production scale: module design, remote state, environment isolation, change governance - Kubernetes cluster operations: lifecycle management, workload reliability, storage, RBAC at scale - Full observability stack ownership (Prometheus, Grafana, Datadog, Splunk or equivalent) for SLO definition, alert design, signal quality - Strong Python and/or Go scripting for production services and infrastructure automation - Executive communication skills, translating infrastructure health and operational risk for senior leadership and non-technical stakeholders - Ability to operate in high-ambiguity, low-process environment, building structure rather than inheriting it Preferred qualifications: - Experience operating customer-facing infrastructure or platform services with reliability expectations beyond internal tooling - Knowledge of high-speed interconnect fabrics (InfiniBand, RoCE, NVLink): setup, troubleshooting, performance tuning - HPC job scheduler experience (Slurm, LSF or equivalent): setup, tuning, infrastructure automation integration - Multi-cloud hybrid operations across AWS, Azure, GCP alongside on-prem/colo with unified observability and IaC - FinOps expertise: cloud spend attribution, TCO modeling, workload placement recommendations - ITIL knowledge or equivalent structured incident/problem/change management framework - Published technical writing, conference talks, or open-source contributions in reliability, observability, or HPC infrastructure

Similar roles