SlipstreamJobsFresh Startup & VC-Backed Jobs

Sr Staff Site Reliability Engineer, AI Infrastructure

d-Matrix Corporation - Santa Clara, CA, United States - In-office - posted 2026-09-29

Apply on the company site

SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.

d-Matrix is seeking a Senior Staff Site Reliability Engineer to own the infrastructure layer that powers the company's generative AI hardware and software platform. This is a hands-on, high-ownership role responsible for reliability, automation, and observability across colocation facilities, on-premises GPU clusters, cloud environments (AWS, Azure, GCP), and customer-facing platform services. You will own systems end-to-end, from provisioning through live incident response. Key responsibilities include: - Owning reliability and availability across colo server fleets, on-premises lab clusters, cloud environments, and customer-facing platform services. - Performing hands-on infrastructure work: server provisioning, OS configuration, networking, storage, and hardware troubleshooting from bare metal through auto-scaling Kubernetes environments. - Leading capacity planning and hardware lifecycle management, tracking cloud spend to support FinOps and workload placement decisions. - Driving all provisioning, deployment, and operational changes through Terraform and/or Ansible, contributing to shared IaC modules across global SRE and data center services teams. - Building and documenting automation that eliminates toil: host lifecycle management, fleet health checks, auto-remediation, self-service tooling, and networking automation for cluster interconnects and lab configurations. - Designing and maintaining monitoring, alerting, and SLIs using Prometheus/Grafana, DataDog, Splunk, or equivalent, contributing to AIOps-driven detection workflows. - Participating in on-call rotation, triaging and resolving incidents from bare metal to application layer, and producing high-quality RCAs for P0/P1 incidents. - Supporting platform services used by internal teams and external customers, ensuring QoS and uptime commitments and documenting operational runbooks. This role partners closely with hardware and software teams on CI/CD, QA, and HPC workloads for silicon development, as well as supporting customer-facing environments where d-Matrix partners collaborate on deployments. REQUIREMENTS: - Bachelor's or Master's in Computer Science, Electrical Engineering, or related field (or equivalent experience) - 7+ years in SRE, infrastructure engineering, or systems administration - Strong Linux systems knowledge with hands-on colocation or on-premises server infrastructure experience (networking, storage, systemd, kernel parameters, performance diagnostics, physical hardware, rack networking, bare-metal provisioning) - Production IaC experience with Terraform and/or Ansible (writing and maintaining configurations, not just running existing playbooks) - Kubernetes operational experience: cluster troubleshooting, workload management, storage, and networking - Experience with observability tooling (Prometheus/Grafana, DataDog, Splunk, or equivalent), including building dashboards and writing alert rules - Production-quality Python and/or Bash scripting, paired with incident response experience (structured triage, RCA production, follow-through on action items) PREFERRED QUALIFICATIONS: - Experience operating customer-facing infrastructure or platform services with external reliability expectations - Cloud infrastructure operations across AWS, Azure, or GCP, including hybrid environments spanning cloud and on-prem - Experience deploying and operating AI-driven infrastructure tools (AIOps platforms, intelligent alerting, anomaly detection, LLM-assisted diagnostics) in production - HPC job scheduler experience (Slurm, LSF, or equivalent) - Knowledge of high-speed interconnect fabrics (InfiniBand, RoCE, NVLink) - Experience with large-scale infrastructure automation (host lifecycle management, fleet auto-healing, AIOps-driven operations)

Similar roles