SlipstreamJobsFresh Startup & VC-Backed Jobs

Software Engineer, Agents (Internal Audit)

Fieldguide - San Francisco, CA, USA - In-office - posted 2026-09-17

Apply on the company site

SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.

Fieldguide is automating and streamlining assurance and audit work for global commerce and capital markets, with a focus on cybersecurity, privacy, and financial audits. Over 50 of the top 100 accounting and consulting firms use Fieldguide to power mission-critical work. The company is backed by Goldman Sachs Alternatives, Bessemer Venture Partners, 8VC, Floodgate, and Y Combinator. You'll join a 0→1 team building agent-powered internal audit capabilities. This is a product-focused engineering role where you'll own agent quality, ship agents that perform real audit work, and collaborate directly with audit practitioners and design-partner firms. Key responsibilities include: - Running error analysis on real testing data and converting findings into concrete agent improvements - Navigating tradeoffs between quality, latency, and cost across multi-phase agent runs - Building structured-output pipelines that transform model outputs into real audit artifacts - Taking ambiguous problem statements and shipping features with clear documentation of tradeoffs - Working directly with embedded subject matter experts and design-partner firms, iterating on agent changes within days - Expanding agent coverage into new audit controls and areas At the Senior level, you may own a major agent area end-to-end, set evaluation and error-analysis practices for the team, collaborate on roadmaps and architectural decisions, and mentor other engineers. At the Staff level, you may drive agent initiatives across Fieldguide, set engineering standards for agent reliability, partner with leadership on long-term strategy, and represent the company externally. You should be product-minded and full-stack, with shipped LLM-backed features in production. You're fluent in evals and error analysis, have strong opinions on model selection and orchestration tradeoffs backed by evidence, and thrive on 0→1 work. You have good instincts for human-in-the-loop design, work collaboratively across PM, design, and domain experts, and ship fast without leaving technical debt. You can internalize complex domains quickly and don't need prior audit knowledge, though you'll develop it rapidly. REQUIREMENTS Must-have: - Shipped LLM-backed product features to production with real users - Applied AI skillset: evals, error analysis, and model-selection decisions you owned and can explain - Full-stack capability with sufficient backend depth for agent orchestration - Ability to work autonomously from ambiguous specifications - Collaborative working style across PM, design, and domain experts Nice-to-have: - Python, TypeScript, React, Postgres, Hasura, GraphQL - Temporal or comparable durable-execution/workflow orchestration - Hands-on eval experience (Langfuse, Braintrust, LangSmith, Arize Phoenix, or similar) - Structured-output work including schema contracts and artifact generation from model output - Startup experience as founder or early engineer - 0→1 track record with things you started from scratch - Direct customer experience and comfort in rooms where customers use your work - Background in internal audit, SOX, accounting, or regulated domains - Document processing experience (PDF, Excel manipulation and annotation) Not a fit if prompt engineering is your only skillset, your agent work never carried production traffic, you want to own eval methodology rather than ship product, you need fully specified tickets to start, or you prefer not to work directly with customers and domain experts.

Similar roles