SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Fireworks is a Series D AI infrastructure platform (valued at $17.5B) backed by NVIDIA, Sequoia, Benchmark, and others. The company enables enterprises to build, train, and serve specialized AI models tailored to their data and workflows, supporting hundreds of state-of-the-art open models across text, image, embedding, audio, and multimodal workloads.
As an AI Forward Deployed Engineer, you will be the technical owner of customer deployments, embedding with a small number of strategic accounts. You will write code directly in customer environments, unblocking whatever stands between them and production: capacity planning, model selection, integration architecture, latency optimization, and cost efficiency. You scope high-stakes deals pre-signature and then deliver on that scope.
Key responsibilities include: owning deployment outcomes end-to-end from architecture through production; designing and running proof-of-value pilots with clear success criteria; acting as technical owner on complex deals (dedicated deployments, BYOC, compliance-sensitive architectures); diagnosing and resolving capacity, performance, integration, and model selection issues; guiding customers to optimal models, deployment tiers, and optimization paths (quantization, speculative decoding, fine-tuning); and feeding structured product signal back to engineering.
You should have 3+ years as a software engineer with depth in backend systems and infrastructure. Strong coding in Python and at least one systems language is required. You must be fluent in AI-assisted and agentic engineering—using coding agents and AI tooling as core to your workflow. Working knowledge of modern generative AI inference, serving, and deployment tradeoffs is essential. The role suits engineers with founder/CTO mentality who genuinely enjoy customer-facing technical work and can own ambiguous problems under pressure.
Preferred: founder or founding engineer experience; GPU infrastructure, distributed serving, or performance engineering background; prior forward-deployed, solutions architecture, or professional services roles; startup experience.