SlipstreamJobsFresh Startup & VC-Backed Jobs

Senior Manager, AI Foundation Model

Merlin Labs - Boston, MA, United States - In-office - posted 2026-09-09

Apply on the company site

SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.

Merlin Labs (NASDAQ: MRLN) is a publicly traded aerospace and defense company building autonomous flight systems. The company has proven its Merlin Pilot autonomy platform through hundreds of autonomous flights and is expanding to accelerate development and deployment across commercial and defense aviation. You will own Merlin's foundation model and world-model strategy, leading a small team of world-model and post-training engineers. This is a technical leadership role where you drive architecture decisions, build-vs-adapt trade-offs, post-training approaches, and capability roadmaps that support autonomous flight systems. Key responsibilities include: - Own technical strategy for foundation and world-model work, including architecture selection and capability roadmap - Lead and mentor a small team of world-model and post-training engineers; set technical bar and review culture - Design model interfaces to the autonomy stack with structured, schema-constrained outputs that deterministic verifiers can accept or reject - Define evaluation criteria before training begins; build evaluation harness, capability taxonomy, and regression suite that gate every model release - Establish uncertainty quantification and out-of-distribution detection as first-class model outputs for downstream safety monitoring - Deliver honest, reproducible comparisons between learned planning and rule-based behavior planning across representative mission profiles - Partner with Systems Engineering, Certification, and Chief Architect to keep model design defensible to regulators - Track external research frontier and make disciplined decisions about adoption, building, or ignoring new approaches This role requires deep expertise in AI systems that have shipped to production with real-world consequences. You understand the difference between models that perform well on benchmarks and models whose behavior can be rigorously characterized and justified to safety and regulatory stakeholders. You are comfortable with the full lifecycle from research through deployment, and you excel at building evaluation practices that catch regressions before customers do. Requirements: - Degree in Computer Science, Artificial Intelligence, Data Science, Computer Engineering, Applied Math, or related field - 8+ years building AI systems, with 3+ years leading technical teams or owning a major model program - Proven team management experience shipping high-tech, AI-powered models into production (hiring, developing AI engineers, setting technical direction, owning delivery) - Demonstrated ownership of a learned system shipped into a physical, real-time product (robotics, autonomous vehicles, aerospace, or industrial autonomy) - Depth in at least two of: world models and learned dynamics; sequence models for planning or control; post-training (SFT, preference optimization, RL fine-tuning); structured or constrained generation - Rigorous evaluation practice with ability to build eval harnesses that catch regressions and explain model behavior degradation despite metric improvements - Strong PyTorch proficiency; comfortable reading and reasoning about C++ real-time systems - Clear technical writing (architecture decisions read by systems engineers, safety engineers, and regulators) Nice to have: - Experience with learned components in certified or regulated products (DO-178C, ISO 26262, IEC 62304) - Background in classical planning, behavior trees, MCTS, or hierarchical task networks - Familiarity with aviation domain structure (flight phases, ARINC 424 procedures, ATC phraseology) - Publications or open-source contributions in embodied AI, world models, or robot learning

Similar roles