SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
xAI is building AI systems to understand the universe and advance human knowledge. The Data Center Engineering team develops internal software platforms and systems that enable data centers to operate at scale and reliability for frontier AI training and inference.
In this role, you'll build and operate software stacks that make site operations scalable, auditable, and fast. You'll design and own multi-service production systems including repair trackers, vendor turnback workflows, operational dashboards, and their integrations. Your users are technicians, managers, NOC operators, and leadership running the fleet.
Key responsibilities include: designing and building multi-service production systems (UIs, APIs, data pipelines, auth, on-call); owning correctness of operational state including node state accuracy, queue ownership, and audit trails; building and maintaining integrations with ticketing, inventory/rack systems, telemetry stores, and vendor portals; ensuring tool reliability through uptime, data integrity, reconciliation, access control, and safe deploys; and embedding with SiteOps and NOC users to measure workflow adoption.
You'll need a Bachelor's in Computer Science or related field and 3+ years building and operating production software (backend and/or full-stack). Strong fundamentals in data structures, algorithms, operating systems, and networking are essential. You should be proficient in at least one programming language (Rust, Python, JavaScript, Java, C++, etc.), experienced with RESTful APIs, databases (Postgres, MongoDB, MySQL, DynamoDB), cloud platforms (GCP, Azure, AWS, OCI), observability tools (New Relic, Splunk, Grafana), and testing practices.
Preferred experience includes full-stack or backend + data work with workflow and state-machine systems, CI/CD automation, building high-correctness operational UIs, on-call discipline, multi-service system design, delivering internal tools in rapidly changing environments, ownership of correctness-sensitive systems, external API integrations, and collaboration with operations and infrastructure teams.
The role is fully onsite in Memphis, TN or Southaven, MS. Candidates should be located in the area or open to relocation.