SlipstreamJobsFresh Startup & VC-Backed Jobs

Director Site Reliability Engineer, AI Infrastructure

d-Matrix Corporation - Santa Clara, CA, USA - In-office - posted 2026-09-26

Apply on the company site

SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.

d-Matrix designs and manufactures purpose-built AI inference silicon. This Director-level role builds and leads the Site Reliability Engineering function from the ground up, owning infrastructure spanning colocation facilities, on-premises lab clusters, cloud environments (AWS, Azure, GCP), and customer-facing platform services. You will hire and grow an SRE team of 1–3 engineers, establish SRE as a discipline across the organization, and direct a dedicated Data Center & Lab Technician team. You own 24×7 reliability across all infrastructure tiers, designing for failure domains, progressive delivery, and strict change control. You will define SLIs/SLOs and error budgets, design on-call rotations (including follow-the-sun models as the company expands globally), and build the incident and RCA management framework. Key responsibilities include: - Building the SRE function from scratch: charter definition, team hiring and development, and establishing SRE discipline across hardware and software engineering. - Owning full observability stack (Prometheus, Grafana, Datadog, Splunk or equivalent), instrumenting metrics, traces, and logs for SLO visibility and alert design. - Driving Infrastructure as Code automation (Terraform, Ansible) and building self-healing infrastructure with host lifecycle automation, fleet auto-remediation, and AIOps-driven alerting. - Owning FinOps and capacity planning across cloud, colocation, and on-premises tiers, including spend attribution, TCO modeling, and workload placement decisions. - Leading migration from ad-hoc JBOD storage to enterprise-grade shared storage platform spanning on-prem, colocation, and cloud, including architecture, vendor selection, and DR design. - Partnering with Software Tools, QA DevOps, Engineering, Networking Operations, and IT Operations teams to drive reliability, cost efficiency, and technical debt removal. - Translating infrastructure health and operational risk into clear narratives for senior leadership and non-technical stakeholders. The culture emphasizes respect, collaboration, humility, and direct communication. You will operate in a high-ambiguity, low-process environment, building structure rather than inheriting it. REQUIREMENTS: - Bachelor's or Master's in Computer Science, Electrical Engineering, or related field; 15+ years in SRE, infrastructure engineering, or production engineering. - 5+ years leading SRE or infrastructure engineering teams, including experience building or significantly rebuilding a function (not managing steady-state teams). - Demonstrated track record of establishing SRE as a discipline in organizations that lacked it: defining SLOs, creating on-call frameworks, standing up observability, and driving cultural change with teams from reactive ops backgrounds. - Deep Linux systems expertise: bare-metal operations, enterprise shared storage platforms (NAS/SAN, NFS/SMB at scale, snapshot and replication architectures), and hybrid-cloud storage integration. - Proven experience operating colocation and on-premises hardware at scale: server lifecycle, power and cooling awareness, rack-level networking. - Hands-on Infrastructure as Code fluency with Terraform and Ansible at production scale: module design, remote state, environment isolation, and change governance. - Kubernetes cluster operations: lifecycle management, workload reliability, storage, and RBAC at scale. - Full observability stack ownership (Prometheus, Grafana, Datadog and/or Splunk or equivalent) for SLO definition, alert design, and end-to-end signal quality. - Strong Python and/or Go scripting for production services and automation that safely touches real infrastructure. - Executive communication skills, translating infrastructure health and operational risk into clear narratives for senior leadership, including non-technical stakeholders. - Ability to operate in high-ambiguity, low-process environments, building structure rather than inheriting it. PREFERRED QUALIFICATIONS: - Experience operating customer-facing infrastructure or platform services with reliability expectations beyond internal tooling. - Knowledge of high-speed interconnect fabrics (InfiniBand, RoCE, or NVLink): setup, troubleshooting, and performance tuning. - HPC job scheduler experience (Slurm, LSF, or equivalent): setup, tuning, and infrastructure automation integration. - Multi-cloud hybrid operations across AWS, Azure, and GCP alongside on-prem/colo with unified observability and IaC. - FinOps expertise: cloud spend attribution, TCO modeling across cloud vs. on-prem vs. colo, and translating cost data into workload placement recommendations. - ITIL knowledge or equivalent structured incident/problem/change management framework. - Published technical writing, conference talks, or open-source contributions in reliability, observability, or HPC infrastructure.

Similar roles