SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Vinci4D.ai is building AI-powered physics simulation tools to transform how engineers design physical products. The company's technology is used by nearly half of the top 20 semiconductor and electronics companies globally, and is backed by Khosla Ventures and Eclipse Ventures.
This is a founding-stage Staff Engineer role owning the compute and data infrastructure layer for training foundation models in physics. You will be responsible for the entire stack: GPU cluster architecture, workload scheduling, petabyte-scale data infrastructure, training/validation platforms, and cost optimization. This is a hands-on technical role, not a management position—you will write production code, lead architecture reviews, debug critical issues, and mentor a small growing team.
Key responsibilities:
- Design GPU compute architecture: provisioning, scaling, and organization as the fleet grows
- Own workload scheduling: queueing, priority, preemption, gang scheduling, and quota allocation between training, validation, and data prep
- Build petabyte-scale data infrastructure: storage, ingestion, versioning, and access patterns that feed training at full throughput
- Develop training and validation platforms: job launch, checkpointing, resumption, monitoring, and resource isolation
- Optimize for reliability, utilization, and cost: set targets for cluster uptime, job success rate, and time-to-first-batch
- Provide technical direction on compute and data requirements for upcoming models; capacity plan and cost forecast 12 months ahead
- Mentor engineers as the team grows; review designs and teach infrastructure concepts across the organization
The role spans the full stack: from GPU scheduling algorithms to data pipeline throughput, from failure recovery to cost accounting. You will work closely with the AI team to understand their needs before they become critical, and with leadership on investment and capacity planning.
REQUIREMENTS:
- 10+ years building large-scale distributed systems, including several years operating GPU infrastructure for large model training
- Direct experience running multi-node distributed training on modern accelerators (dozens to hundreds of GPUs for days), including scheduling, interconnect behavior, checkpointing, and failure recovery
- Demonstrated ability to diagnose and fix performance bottlenecks in GPU training runs
- Direct experience serving training data at petabyte scale, where both throughput and storage cost were constraints
- Experience setting scheduling and quota policy on shared GPU capacity, balancing fleet utilization against queue wait times
- Track record building infrastructure for researchers/engineers, measured by what users could accomplish, not the system itself
- Infrastructure you built that survived substantial scale or workload changes, with clear understanding of what held up and what required replacement
- History of mentoring engineers and teaching cross-functional teams
- Ability to set technical direction, make the case to non-infrastructure stakeholders, and implement it
- Advanced degree in computer science or related field, or equivalent industry experience
NICE TO HAVE:
- Production experience with Kubernetes, Ray, Slurm, Airflow, or comparable ML orchestration
- Hybrid or multi-cloud GPU capacity planning and cost negotiation
- Fluency in modern training stacks: PyTorch distributed, FSDP, DeepSpeed, mixed precision, GPU profiling
- Applied optimization for scheduling/allocation (linear programming, graph algorithms, heuristic solvers)
- Experiment tracking and model artifact management
- Work with large numerical/geometric datasets (simulation, EDA, CAD, imaging, autonomous systems)
IMPORTANT NOTE: This role is for hands-on engineers who own a cluster end-to-end. It is not a fit if your recent work has been directing an infrastructure organization, setting standards across teams, or managing vendor relationships. Lab-scale or single-node model work does not transfer to these problems.