SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Salary: USD 174,720 - 295,680 / annual
XPeng is a leading smart technology company integrating advanced AI and autonomous driving technologies into electric vehicles, eVTOL aircraft, and robotics. The company is dedicated to reshaping transportation through cutting-edge R&D in AI, machine learning, and smart connectivity.
We are seeking a Senior Machine Learning Engineer to establish state-of-the-art ML infrastructure for training very large foundation models and accelerating model training and inference. You will work with talented software engineers, machine learning engineers, and research scientists to push the boundaries of machine learning models enabling next-generation end-to-end autonomous driving solutions.
Key Responsibilities:
- Design and implement training data pipelines that stream data from hundreds of petabytes of labeled and unlabeled data from a fleet of over one million vehicles
- Implement training frameworks for all physical AI foundation models at XPeng, including VLA 2.0, XWorld, and Robotics
- Accelerate training using state-of-the-art parallelism techniques such as FSDP, Expert Parallel, Context Parallel, and advanced data types
- Accelerate model inference on the cloud for closed-loop simulation, reinforcement learning, and enterprise LLM/VLM applications
Requirements:
- Master's degree in Computer Science, Computer Engineering, or Electrical Engineering, or equivalent industry experience
- Deep knowledge of PyTorch
- Knowledge of model inference frameworks (e.g., vLLM, SGLang)
- In-depth knowledge of transformer architecture and methods to accelerate training and inference of transformer models
- Experience performing large-scale distributed training of models
- Track record of profiling models and improving model training and inference speed through performance optimization
Preferred Qualifications:
- Previous experience in the autonomous driving industry
- Experience with CUDA language for writing custom operations
- Experience with edge computing systems
- Knowledge of distributed computing frameworks such as Ray
- Track record of efficiently solving complex problems collaboratively on larger teams