SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
d-Matrix designs and manufactures purpose-built AI inference silicon. This Director-level role builds and leads the Site Reliability Engineering function from the ground up, owning infrastructure spanning colocation facilities, on-premises lab clusters, cloud environments (AWS, Azure, GCP), and customer-facing platform services.
You will hire and grow an SRE team of 1–3 engineers, establish SRE as a discipline across the organization, and direct a dedicated Data Center & Lab Technician team. You own 24×7 reliability across all infrastructure tiers, designing for failure domains, progressive delivery, and strict change control. You will define SLIs/SLOs and error budgets, design on-call rotations (including follow-the-sun models as the company expands globally), and build the incident and RCA management framework.
Key responsibilities include:
- Building the SRE function from scratch: charter definition, team hiring and development, and establishing SRE discipline across hardware and software engineering.
- Owning full observability stack (Prometheus, Grafana, Datadog, Splunk or equivalent), instrumenting metrics, traces, and logs for SLO visibility and alert design.
- Driving Infrastructure as Code automation (Terraform, Ansible) and building self-healing infrastructure with host lifecycle automation, fleet auto-remediation, and AIOps-driven alerting.
- Owning FinOps and capacity planning across cloud, colocation, and on-premises tiers, including spend attribution, TCO modeling, and workload placement decisions.
- Leading migration from ad-hoc JBOD storage to enterprise-grade shared storage platform spanning on-prem, colocation, and cloud, including architecture, vendor selection, and DR design.
- Partnering with Software Tools, QA DevOps, Engineering, Networking Operations, and IT Operations teams to drive reliability, cost efficiency, and technical debt removal.
- Translating infrastructure health and operational risk into clear narratives for senior leadership and non-technical stakeholders.
The culture emphasizes respect, collaboration, humility, and direct communication. You will operate in a high-ambiguity, low-process environment, building structure rather than inheriting it.
REQUIREMENTS:
- Bachelor's or Master's in Computer Science, Electrical Engineering, or related field; 15+ years in SRE, infrastructure engineering, or production engineering.
- 5+ years leading SRE or infrastructure engineering teams, including experience building or significantly rebuilding a function (not managing steady-state teams).
- Demonstrated track record of establishing SRE as a discipline in organizations that lacked it: defining SLOs, creating on-call frameworks, standing up observability, and driving cultural change with teams from reactive ops backgrounds.
- Deep Linux systems expertise: bare-metal operations, enterprise shared storage platforms (NAS/SAN, NFS/SMB at scale, snapshot and replication architectures), and hybrid-cloud storage integration.
- Proven experience operating colocation and on-premises hardware at scale: server lifecycle, power and cooling awareness, rack-level networking.
- Hands-on Infrastructure as Code fluency with Terraform and Ansible at production scale: module design, remote state, environment isolation, and change governance.
- Kubernetes cluster operations: lifecycle management, workload reliability, storage, and RBAC at scale.
- Full observability stack ownership (Prometheus, Grafana, Datadog and/or Splunk or equivalent) for SLO definition, alert design, and end-to-end signal quality.
- Strong Python and/or Go scripting for production services and automation that safely touches real infrastructure.
- Executive communication skills, translating infrastructure health and operational risk into clear narratives for senior leadership, including non-technical stakeholders.
- Ability to operate in high-ambiguity, low-process environments, building structure rather than inheriting it.
PREFERRED QUALIFICATIONS:
- Experience operating customer-facing infrastructure or platform services with reliability expectations beyond internal tooling.
- Knowledge of high-speed interconnect fabrics (InfiniBand, RoCE, or NVLink): setup, troubleshooting, and performance tuning.
- HPC job scheduler experience (Slurm, LSF, or equivalent): setup, tuning, and infrastructure automation integration.
- Multi-cloud hybrid operations across AWS, Azure, and GCP alongside on-prem/colo with unified observability and IaC.
- FinOps expertise: cloud spend attribution, TCO modeling across cloud vs. on-prem vs. colo, and translating cost data into workload placement recommendations.
- ITIL knowledge or equivalent structured incident/problem/change management framework.
- Published technical writing, conference talks, or open-source contributions in reliability, observability, or HPC infrastructure.