SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
d-Matrix is seeking a Senior Staff Site Reliability Engineer to own the infrastructure layer that powers the company's generative AI hardware and software platform. This is a hands-on, high-ownership role responsible for reliability, automation, and observability across colocation facilities, on-premises GPU clusters, cloud environments (AWS, Azure, GCP), and customer-facing platform services.
You will own systems end-to-end, from provisioning through live incident response. Key responsibilities include:
- Owning reliability and availability across colo server fleets, on-premises lab clusters, cloud environments, and customer-facing platform services.
- Performing hands-on infrastructure work: server provisioning, OS configuration, networking, storage, and hardware troubleshooting from bare metal through auto-scaling Kubernetes environments.
- Leading capacity planning and hardware lifecycle management, tracking cloud spend to support FinOps and workload placement decisions.
- Driving all provisioning, deployment, and operational changes through Terraform and/or Ansible, contributing to shared IaC modules across global SRE and data center services teams.
- Building and documenting automation that eliminates toil: host lifecycle management, fleet health checks, auto-remediation, self-service tooling, and networking automation for cluster interconnects and lab configurations.
- Designing and maintaining monitoring, alerting, and SLIs using Prometheus/Grafana, DataDog, Splunk, or equivalent, contributing to AIOps-driven detection workflows.
- Participating in on-call rotation, triaging and resolving incidents from bare metal to application layer, and producing high-quality RCAs for P0/P1 incidents.
- Supporting platform services used by internal teams and external customers, ensuring QoS and uptime commitments and documenting operational runbooks.
This role partners closely with hardware and software teams on CI/CD, QA, and HPC workloads for silicon development, as well as supporting customer-facing environments where d-Matrix partners collaborate on deployments.
REQUIREMENTS:
- Bachelor's or Master's in Computer Science, Electrical Engineering, or related field (or equivalent experience)
- 7+ years in SRE, infrastructure engineering, or systems administration
- Strong Linux systems knowledge with hands-on colocation or on-premises server infrastructure experience (networking, storage, systemd, kernel parameters, performance diagnostics, physical hardware, rack networking, bare-metal provisioning)
- Production IaC experience with Terraform and/or Ansible (writing and maintaining configurations, not just running existing playbooks)
- Kubernetes operational experience: cluster troubleshooting, workload management, storage, and networking
- Experience with observability tooling (Prometheus/Grafana, DataDog, Splunk, or equivalent), including building dashboards and writing alert rules
- Production-quality Python and/or Bash scripting, paired with incident response experience (structured triage, RCA production, follow-through on action items)
PREFERRED QUALIFICATIONS:
- Experience operating customer-facing infrastructure or platform services with external reliability expectations
- Cloud infrastructure operations across AWS, Azure, or GCP, including hybrid environments spanning cloud and on-prem
- Experience deploying and operating AI-driven infrastructure tools (AIOps platforms, intelligent alerting, anomaly detection, LLM-assisted diagnostics) in production
- HPC job scheduler experience (Slurm, LSF, or equivalent)
- Knowledge of high-speed interconnect fabrics (InfiniBand, RoCE, NVLink)
- Experience with large-scale infrastructure automation (host lifecycle management, fleet auto-healing, AIOps-driven operations)