SlipstreamJobsFresh Startup & VC-Backed Jobs

Research Engineer

Constellation Systems, Inc - San Francisco, CA, United States - In-office

Apply on the company site

SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.

Salary: USD 180,000 - 250,000 / annual

Constellation Systems is building an AI-human translation layer to address deep problems in human experience: empowering people toward their goals, augmenting cognition and emotional wellness, and fostering mutual understanding. The company is generating a richly multimodal dataset to build a new class of foundation models. As Research Engineer, you will sit between data and models, optimizing the entire training loop. You'll orchestrate and optimize distributed training runs on long-horizon multimodal sequences, build pipelines that transform messy, growing data into learnable formats, and write research infrastructure (libraries, dataloaders, evaluation harnesses) that enables rapid experimentation and trustworthy results. You'll collaborate closely with both engineering and research teams. Key Responsibilities: - Orchestrate and optimize training: manage distributed configuration (DDP/FSDP), mixed precision, checkpointing and recovery on multi-day runs, and profile to identify actual bottlenecks (kernel, dataloader, communication, or pipeline). - Build high-throughput data loading over TB-scale multimodal data: implement sharding, prefetching, caching, and format choices that keep GPUs saturated rather than I/O-bound. - Wrangle complex data: align and synchronize multi-stream time series, handle evolving collection protocols, and own dataset versioning and lineage as the corpus grows. - Build reproducible pipelines that derive rich features from raw recordings, with versioning and cheap recomputation when upstream data changes. - Write research code that enables better research: clean, well-tested, well-documented libraries for datasets, models, transforms, and metrics that the team builds on. - Ensure traceability: code version, config, dataset version, and environment recoverable from any result. - Build evaluation harnesses that run automatically on new checkpoints and surface regressions early. - Convert research prototypes into repeatable pipelines without sacrificing researcher flexibility. Requirements: - 3+ years building ML systems or research infrastructure, including distributed training runs on 100+ GPUs. - Expert-level Python and deep knowledge of PyTorch internals: DDP, FSDP, mixed precision, gradient accumulation, and profiling tools to diagnose slow models vs. starved ones. - Track record of building or maintaining research libraries others depend on. Contributions to packages like torch_geometric, torchaudio, torcheeg, torch_brain, neuralsets, or comparable internal tooling are ideal. - Experience with data pipelines over large unstructured and multimodal datasets; familiarity with columnar and streaming formats (Zarr, Parquet, Arrow, Lance, WebDataset, Vortex) and tradeoffs between random access and sequential throughput. - Hands-on experience with experiment tracking and dataset versioning tools (ClearML, Weights & Biases, MLflow, or similar). - Research fluency: ability to read papers, reimplement components, and assess whether loss curves are broken. - Debugging expertise across corrupted shards, silent collate function errors, and training divergence. - Care for the well-being of people affected by this technology; thrive in high-bandwidth, collaborative environments. Nice to Have: - Custom kernel work in CUDA or Triton, or compiler-level optimization (torch.compile, TensorRT, ONNX). - Experience with time-series or multimodal data requiring cross-stream synchronization. - Low-latency or edge inference, quantization, or distillation. - Comfort with Rust or C++ when Python is insufficient. - Interest in neurotech, mental health, or human-AI interaction.

Similar roles