SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Beacon AI is building an AI platform to make flying safer, more efficient, and more capable. The company is backed by top investors, has secured a dozen Department of Defense contracts, and partners with major airlines to deliver mission-critical systems.
You will design, implement, and operate cloud and ML infrastructure services that are scalable, reliable, and secure. This role spans AWS foundation buildout and LLM platform development, with opportunities to specialize in LLM/ML infrastructure and IoT infrastructure.
Key responsibilities include:
- Cloud Infrastructure: Design and provision AWS infrastructure using IaC tools (AWS CDK, Terraform). Build CI/CD pipelines with GitHub Actions, CodeBuild, and CodePipeline. Operate secure networking with VPCs, PrivateLink, and manage IAM, KMS, and audit logging.
- LLM Platform: Stand up and operate model endpoints using AWS Bedrock and SageMaker. Build application services that call LLMs with streaming, batching, and backoff strategies. Implement prompt and tool execution flows with LangChain.
- RAG Data Systems: Design chunking and embedding pipelines for documents, time series, and multimedia. Operate vector search using OpenSearch Serverless, Aurora PostgreSQL with pgvector, or Pinecone. Build and maintain knowledge bases with data syncs from S3, Aurora, and DynamoDB.
- Observability and Cost: Create evaluation harnesses for prompts and chains. Instrument telemetry with CloudWatch and OpenTelemetry. Build token usage and cost dashboards with budgets and alerts.
- Safety and Compliance: Implement PII detection, redaction, access controls, and content filters. Use Bedrock Guardrails to enforce safety standards and maintain audit trails.
- Data Pipelines: Build ingestion and processing pipelines for structured and unstructured data. Optimize bulk data movement in S3 and tiered storage.
- IoT Deployment: Manage infrastructure for edge device deployment, secure messaging, identity, and over-the-air updates.
- Performance Optimization: Tune retrieval quality, context window use, and caching. Optimize inference with model selection, quantization, and autoscaling.
You should have shipped or operated LLM-powered applications in production, with strong AWS depth across VPC, IAM, KMS, CloudWatch, S3, ECS/EKS, Lambda, Step Functions, Bedrock, and SageMaker. Data engineering skills in Python, Glue, and Athena are important. A security mindset and ability to use quantitative evals and metrics to guide improvements are essential. Clear communication across product, security, and engineering teams is valued.
Bonus experience includes 4+ years with serverless or container platforms on AWS, vector databases at scale, Bedrock Guardrails, GPU workloads, big data tools, aviation domain knowledge, or DevSecOps automation.