SlipstreamJobsFresh Startup & VC-Backed Jobs

Infrastructure Engineer (Linux)

Fractile - London, England, United Kingdom - In-office - posted 2026-08-25

Apply on the company site

SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.

Fractile is an AI infrastructure company founded in 2022, building specialized hardware and systems to accelerate inference speed for frontier AI models. The company has raised $220M from top-tier investors including Founders Fund and Accel. We are seeking an Infrastructure Engineer (Linux) to build and maintain the compute, storage, and networking foundations supporting our silicon development workloads. You will work with an on-premise compute environment used for computationally-intensive engineering, developing expertise across Linux systems administration, HPC-style cluster management, networking, storage, and performance tuning. Key Responsibilities: - Maintain and troubleshoot a fleet of on-premise Linux (Rocky/RHEL-family) servers - Support deployment and upkeep of on-premise compute infrastructure using infrastructure-as-code tooling (e.g., Ansible) - Set up and maintain monitoring and observability tooling for resource utilization, service health, and machine failures (e.g., Prometheus, node_exporter, Grafana, Zabbix) - Diagnose and resolve networking issues (DNS, VLANs, bonding, routing) and storage issues (network filesystems, capacity management, performance troubleshooting) - Support day-to-day operation of cluster compute/job scheduling environments (e.g., Slurm) - Assist with user and identity management tasks (e.g., FreeIPA/LDAP), including onboarding and access provisioning - Document infrastructure changes and contribute to runbooks and playbooks - Work cross-functionally with engineers to identify and resolve infrastructure-related bottlenecks The role offers full agency to drive work forward, rapid iteration with top leadership, and close collaboration across hardware, software, silicon, and modeling teams. You will have ownership and execution responsibility in a mission-driven environment. Requirements: - 2-3+ years of solid, hands-on production Linux system administration experience - Good working knowledge of monitoring solutions (e.g., Prometheus, Grafana, Zabbix, or similar) - Solid networking fundamentals: TCP/IP, DNS, VLANs, routing, basic troubleshooting - Good understanding of storage concepts: network filesystems, parallel/distributed filesystems, disk/volume management, basic performance troubleshooting - Comfortable working from the command line and with scripting (Bash, Python, or similar) to automate routine tasks - Methodical, curious approach to troubleshooting with ability to pick up new tools and technical areas quickly - Comfortable working independently and taking ownership of problems through to resolution Desirable: - Exposure to or knowledge of HPC (High Performance Computing) environments or cluster compute concepts - Familiarity with infrastructure-as-code tools (e.g., Ansible, Terraform) - Linux certifications (e.g., RHCSA, LFCS, or similar) - Experience working in a data center environment - Knowledge of server hardware (e.g., installation, diagnostics, component replacement) - Silicon industry background

Similar roles