SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Mecka AI is building the data infrastructure layer for robotics and embodied AI, designing and operating global systems for data capture, data labeling, and hardware-enabled workflows used by leading AI labs and robotics companies to train and validate humanoid and embodied AI systems.
You will architect and train proprietary foundation models from scratch focused on 3D human body tracking and articulated pose estimation. Your core mandate is twofold: building in-house equivalents to cutting-edge 3D human body and mesh recovery architectures, and developing highly robust interaction models tailored for complex, real-world environments characterized by severe occlusions and dynamic motion. You will serve as a lead problem-solver for emergent perception challenges as hardware and downstream robotics needs evolve.
Key responsibilities include:
• Architecting proprietary articulation models: Design, implement, and train state-of-the-art networks for 3D human pose estimation, dense full-body mesh recovery, and kinematic tracking. Scale multi-view and temporal ML architectures across multi-GPU clusters to handle massive, multi-modal datasets. Develop novel loss functions enforcing biomechanical constraints, temporal smoothness, postural balance, and physical plausibility.
• Human-scene interaction and complex motion modeling: Build custom architectures handling extreme motion blur, severe self-occlusion, and multi-person crowding. Use models to track human bodies through complex spaces, mapping foot-to-ground contact, joint torques, and environmental affordances to provide rich regularization for downstream action-conditioned robotics models, especially humanoid robots.
• Emergent perception R&D: Rapidly prototype and deploy new models for tasks spanning fine-grained action segmentation, intent prediction, and novel hardware sensor integrations. Pivot to resolve sudden algorithmic bottlenecks in the data engine, adapting latest research to unblock new product capabilities.
• Dense contact and physics-aware tracking: Connect outputs of foundational tracking models into highly optimized pipelines reasoning about physical contact surfaces, gravity, and momentum, bridging the gap between human video data and robotic control/locomotion policies.
You will have access to a massive, continuous stream of high-quality, proprietary ground-truth human motion data captured by Mecka's infrastructure, enabling you to train networks that surpass current public baselines while owning the complete human-scene perception loop for the data engine.
REQUIREMENTS:
• Deep expertise in Deep Learning, 3D Computer Vision, and specifically Articulated Tracking / Human Body Pose Estimation
• Proven experience training large-scale vision models from scratch (not just running inference or fine-tuning existing checkpoints)
• Strong theoretical and practical understanding of parametric human body models (SMPL, SMPL-X, GHUM, MHR, SOMA-X), inverse kinematics, and dense mesh estimation
• Mastery of PyTorch and deep learning scaling frameworks
• Experience handling and curating massive, multi-terabyte image and video datasets for training
• Comfortable operating in fast-paced environments where priorities shift rapidly to capitalize on new research or hardware capabilities
STRONG SIGNALS:
• First-author publications in top-tier venues (CVPR, ICCV, ECCV, NeurIPS) focusing on 3D human pose tracking, human-scene interaction, human motion capture, or human mesh recovery
• Specific experience working with massive human motion and interaction datasets (AMASS, Human3.6M, EgoBody, PROX) and solving unique optimization challenges they present