SlipstreamJobsFresh Startup & VC-Backed Jobs

Infrastructure Software Engineering – Platform & Build 1

Fractile - London, United Kingdom - In-office - posted 2026-09-10

Apply on the company site

SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.

Fractile is an AI hardware company founded in 2022, focused on building chips and systems to accelerate inference speed for frontier AI models. The company recently raised $220M from Founders Fund and Accel. As an Infrastructure Platform Engineer, you will be a core member of the technical team working alongside hardware, software, and research teams to solve cutting-edge problems in AI hardware verification, simulation, and systems. You'll work within the Software organisation, which develops the full software stack for Fractile's AI inference systems—from ML compilers and device drivers to systems firmware, runtime, ML libraries, developer tooling, and datacenter-scale deployment solutions. The Infrastructure team is responsible for core platform and systems that power the entire engineering organization. You will own and scale mission-critical systems including the multi-language Bazel monorepo and CI pipelines that support hardware design, verification, kernel development, and ML compiler work. Key responsibilities: - Own CI infrastructure end-to-end: create, maintain, and debug reproducible multi-language CI pipelines; optimize CI performance across large compute clusters - Build and maintain infrastructure observability, alerting, runbooks, and incident response workflows - Scale and maintain Fractile's Bazel monorepo as the company grows across Python, C++, Rust, SystemVerilog, and ML workloads - Shape the infrastructure and platform experience for every engineer at Fractile - Apply an SRE mindset: reliability, build performance, systematic debugging, and clear visibility into infrastructure performance You bring deep infrastructure and build systems expertise with an SRE mindset. You are systematic in fault isolation and root cause analysis, apply metrics-driven thinking to reliability and developer productivity, and thrive on variety across ML, compilers, kernel drivers, simulators, and hardware verification. You understand that your work becomes the foundation every engineer builds on and take that responsibility seriously. Requirements: - 5+ years in software, platform, or infrastructure engineering - 5+ years hands-on experience with CI/CD for large-scale products, including monitoring, debugging performance, and resolving pipeline failures at scale - 3+ years working with build systems, preferably modern cross-language systems (Bazel, Buck, Pants, Please, etc.) with emphasis on correctness, performance, and extensibility - Experience with infrastructure as code (Terraform, OpenTofu, or Pulumi) - Experience with monitoring and observability tooling (Prometheus, Grafana, or similar) - Working knowledge and practice of DevOps/SRE principles: SLOs, alerting design, incident management, and on-call practice - Strong proficiency in at least one systems or scripting language for tooling (Python preferred; Go or Rust also relevant), with a track record of shipping well-tested, maintainable developer tooling that other engineers rely on: CLIs, CI integrations, build and release automation Note: This role involves technologies subject to UK, US, and other international export control regulations. Certain roles may require additional eligibility checks to ensure compliance with applicable law.

Similar roles