SlipstreamJobsFresh Startup & VC-Backed Jobs

Senior Machine Learning Infrastructure Engineer, Research

PhysicsX - Singapore, Singapore - In-office - posted 2026-08-06

Apply on the company site

SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.

PhysicsX is a deep-tech company building AI-driven simulation software for engineering and manufacturing across aerospace, defense, materials, energy, semiconductors, and automotive. The company accelerates hardware innovation by enabling high-fidelity, multi-physics simulation through AI inference across the engineering lifecycle. As Senior ML Infrastructure Engineer embedded in the Research function, you will design and operate the infrastructure powering Large Physics Model training, fine-tuning, and serving pipelines. You'll work directly with research scientists, ML engineers, and simulation data engineers to ensure efficient and reliable model training at scale. Key responsibilities include: **Training Infrastructure**: Design and operate distributed training infrastructure for neural operator architectures on NVIDIA DGX B200 platforms. Optimize training pipelines for throughput, fault tolerance, and cost efficiency through checkpointing strategies, gradient accumulation, and multi-node synchronization. Build experiment tracking and observability systems giving researchers clear visibility into training runs and model performance. **Data I/O and Performance**: Solve data loading bottlenecks for large-scale mesh datasets. Optimize data pipelines for efficient I/O from cloud storage with prefetching, caching, and format optimization. Work with heterogeneous data sources of varying formats and resolutions. **Model Serving and Deployment**: Build serving infrastructure for pre-trained LPMs supporting zero-shot inference and uncertainty quantification. Design model packaging pipelines for customer deployment with fine-tuning capabilities. Ensure reproducibility so any model checkpoint deploys consistently. **Platform and Tooling**: Improve developer experience for the Research team with fast iteration cycles, reliable CI/CD, and clear debugging tools. Collaborate with the broader Infrastructure team on shared patterns and standards. You bring 5+ years building and operating ML infrastructure at scale, with deep expertise in distributed training (NCCL, FSDP, DDP, pipeline parallelism), strong systems fundamentals (Linux, networking, storage I/O, profiling), production Kubernetes and SLURM experience, Python and PyTorch proficiency, and cloud GPU infrastructure experience. Ideally you have geometric deep learning or neural operator experience, HPC simulation engineering background, and model serving infrastructure expertise.

Similar roles