SlipstreamJobsFresh Startup & VC-Backed Jobs

Senior DevOps Engineer, AI Platform

FloQast - San Jose, CA, USA - In-office - posted 2026-08-24

Apply on the company site

SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.

FloQast is seeking a Senior DevOps Engineer to specialize in AI infrastructure and platform reliability. The company's AI products—Transform (data transformation and analytics), AI Matching (automated reconciliation), and AutoBuilder (workflow generation)—have outgrown standard infrastructure patterns and now require dedicated DevOps expertise. You will embed with the Transform and Close AI product pods and own the AI runtime infrastructure end-to-end. This is a DevOps role with deep AI infrastructure specialization, not a research or modeling role. You will not train models or tune prompts; instead, you'll architect and operate the systems that serve them. Key responsibilities include: **Transform Product**: Own the foundation-model runtime across US, EU, and AU regions. Manage sandbox isolation for AI-executed code, autoscaling for worker fleets, and cost attribution for model token spend. The product is a monorepo of containerized services on AWS, including an agentic LLM thread runtime, natural-language-to-SQL service, and queue-driven workflow workers. **AI Matching Product**: Ensure throughput and unit economics during close-cycle peak periods. Own the sandbox for AI-generated code and maintain the evaluation harness in CI to prevent blind model or prompt changes from shipping. **AutoBuilder Product**: Manage generation-queue health and backpressure. Implement triage logic that distinguishes model failures from infrastructure failures. Scale to handle spiky load and maintain latency budgets for users waiting on generated artifacts. Core focus areas: making systems fast, observable, multi-region, cost-bounded, and auditable. You'll work with foundation models in multiple regions, handle generated code execution in sandboxes, manage token-based cost models, and implement monitoring that goes beyond standard 500-rate dashboards to catch AI-specific failure modes.

Similar roles