SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
PhysicsX is a physics AI company for industrials, building simulation software to accelerate hardware innovation across aerospace & defence, automotive, semiconductors, materials, and energy sectors. The company is headquartered in the UK with offices in London, New York, Singapore, and the Bay Area.
You will join the Platform SRE Core Infrastructure team as a Senior Software Engineer responsible for designing, provisioning, and operating the shared infrastructure underpinning the PhysicsX platform. This role combines deep infrastructure expertise with a reliability engineering mindset and focus on developer experience for internal platform consumers.
Key responsibilities include:
- Designing and delivering core infrastructure across multiple cloud providers (AWS, GCP, Azure) and on-premises environments using Terraform and Crossplane
- Architecting and operating Kubernetes clusters supporting single-tenant and multi-tenant workloads with emphasis on isolation, performance, and reliability
- Defining infrastructure provisioning patterns using Crossplane compositions and Terraform modules for reproducibility and auditability
- Designing and operating secrets management solutions with dynamic provisioning, rotation, and fine-grained access control
- Managing GPU driver configurations and accelerated compute node pools for AI and simulation workloads
- Owning cluster networking design including CNI selection, Istio service mesh integration, ingress strategy, and cross-cluster connectivity
- Implementing vCluster-based multi-tenancy for strong workload isolation
- Developing lightweight Kubernetes Operators or controllers for infrastructure automation
- Establishing SLOs and reliability targets for core infrastructure; leading production incident response
- Partnering with security and platform teams on infrastructure governance, network policies, and compliance
- Contributing to engineering standards across the platform organization
The role offers hybrid flexibility with a base in the Shoreditch office, blending in-person collaboration with work-from-home flexibility. Benefits include equity options, 10% employer pension contribution, free office lunches, enhanced parental leave (3 months paternity/6 months maternity at full pay), YellowNest nursery scheme, 25 days annual leave plus public holidays, 100% private medical insurance, Wellhub subscription, eye tests, personal development support, Employee Assistance Programme, Bike2Work scheme, season ticket loan, and Octopus EV salary sacrifice.
REQUIREMENTS:
- Kubernetes depth: 5+ years professional experience operating Kubernetes in production with thorough understanding of cluster architecture, scheduler, networking, storage, and API lifecycle. Kubernetes certifications (CKAD, CKA, CKS) highly desirable.
- Crossplane expertise: Significant hands-on experience designing and operating Crossplane compositions, providers, and managed resources in production.
- Terraform proficiency: Strong experience authoring, structuring, and operating Terraform at scale including state management, module design, and CI integration.
- Multi-cloud and on-premises: Practical experience operating infrastructure across multiple cloud providers and on-premises environments.
- Multi-tenancy architecture: Experience designing and implementing single-tenant and multi-tenant Kubernetes architectures with strong views on isolation and resource governance.
- Secrets management: Experience with Vault, External Secrets Operator, or cloud-native secret stores including dynamic provisioning and rotation.
- Networking: Solid knowledge of Kubernetes networking, CNI plugins, Istio service mesh, and ingress patterns. Cross-cluster or hybrid connectivity experience valuable.
- vCluster and virtual clusters: Experience using vCluster or similar tooling for lightweight, isolated Kubernetes environments.
- GPU and accelerated compute: Familiarity with GPU driver management, device plugins, and operational considerations of accelerated workloads in Kubernetes.
- Kubernetes Operators: At least lightweight experience writing or extending Kubernetes Operators or controllers, ideally in Golang or Python.
- Software engineering capability: Comfortable writing code to automate and extend infrastructure. Python and Golang are primary languages; exposure to functional programming languages (Erlang, Elixir, OCaml) appreciated.
- Platform engineering mindset: Think of infrastructure as a product consumed by engineering teams; prioritize usability, documentation, and long-term maintainability.
- Distributed systems experience: Solid grounding in distributed systems concepts including failure modes, consistency, and operational challenges at scale.
IDEAL:
- Experience with GitOps workflows using Argo CD or Flux
- Contributions to open-source infrastructure or Kubernetes ecosystem projects