SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Beacon AI is building an AI platform to make flying safer, more efficient, and more capable. The company is backed by top investors, has secured a dozen Department of Defense contracts, and partners with major airlines to deliver mission-critical systems.
As a Software Engineer specializing in AI/LLM, you will ship LLM-powered product features end-to-end. This includes designing retrieval-augmented generation (RAG) and tool-calling flows, writing production services, building evaluations and guardrails, and monitoring cost, latency, and quality in production. You'll collaborate with ML/infrastructure teammates on embeddings, indexing, and model hosting, and with product teams on user experience and outcomes.
Key responsibilities include:
- Build user-facing LLM features using frameworks like LangChain, designing RAG and tool-calling flows with robust JSON and schema-bound outputs.
- Own the service layer by shipping APIs and workers in Python or TypeScript with clear contracts, streaming, backoff, caching, and prompt templates to control latency and cost.
- Collaborate on retrieval and data preparation, including chunking, embeddings, indexing, and vector backend tuning (OpenSearch, pgvector, Pinecone).
- Create offline evaluations and golden sets for prompts, retrievers, and tools; stand up online metrics for task success, hallucination rate, retrieval precision/recall, latency, and cost.
- Implement safety, privacy, and compliance measures including content checks, PII detection, access controls, and auditing for aviation data.
- Operate what you build by adding tracing, logs, dashboards, and debugging across retrieval, prompts, tools, and providers.
Success requires shipped LLM applications, strong production coding skills, deep RAG and tools knowledge, a quality-first mindset with metrics-driven iteration, cost and latency awareness, and clear communication across teams. Nice-to-have skills include experience with Bedrock, OpenSearch Serverless, pgvector, prompt versioning, multimodal work, GPU inference, and aviation domain exposure.
The role is hybrid based in San Carlos, CA, requiring 3+ days per week onsite. The company operates without silos, with small focused teams owning what they build and shipping quickly in a safety-critical domain.