SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Shield AI is seeking a Staff Engineer for AI Operations & Governance to own the technical operating load for workplace AI systems. This is a deeply technical individual contributor role reporting to the Head of AI Operations & Governance (Enterprise AI).
You will be responsible for platform configuration, connector and integration management, observability wiring, secrets and access hygiene, model/prompt lifecycle mechanics, and hands-on production changes. The role bridges technical and governance layers, translating production realities into risk assessments, governance artifacts, and training materials.
Key responsibilities include:
- Implement and maintain AI platform configurations (models, routes, guardrails, tenants, policies, role mappings, prompt libraries)
- Design, configure, and maintain connectors and extensions into SaaS systems, data sources, and workflow tools
- Set up and maintain logging, metrics, and alerts for AI workflows; ensure key signals (latency, errors, usage, drift) are captured and visible
- Implement secure storage and rotation for API keys, tokens, and credentials; maintain access control configurations
- Execute model swaps, policy updates, prompt changes, version upgrades, and rollout plans in production
- Run experiments and benchmarks on models, tools, and configurations to inform governance decisions
- Translate technical signals (logs, metrics, incidents) into clear risk, reliability, and compliance narratives for Security, Legal, HR, and business sponsors
- Co-author technical sections of operational playbooks, runbooks, and training materials
- Act as technical point of contact for incidents: triage, investigate, propose mitigations, execute fixes, and document learnings
Shield AI is a venture-backed defense-tech company protecting service members and civilians with intelligent systems. Products include Hivemind autonomy software, V-BAT and X-BAT aircraft, and Aechelon simulation technologies. The company operates globally with offices across the U.S., Europe, the Middle East, and Asia-Pacific.
REQUIREMENTS:
- 4–7+ years in platform engineering, DevOps/SRE, ML/AI operations, or technical SaaS operations with hands-on responsibility for production systems
- Strong fluency in APIs, integrations, and infrastructure-as-config concepts; able to work in code/JSON/YAML configuration environments
- Hands-on experience with monitoring and observability tools (logs, metrics, alerts) and using them to diagnose issues
- Practical experience working with at least one class of AI or automation platforms (LLM providers, AI productivity tools, RPA/workflow engines, or similar)
- Comfort with secure secrets management and access control practices (roles, permissions, key rotation, least-privilege patterns)
- Ability to document technical work and decisions clearly for both technical and non-technical audiences
- Strong ownership mindset, bias to action, and comfort operating close to production in a high-stakes environment
PREFERRED:
- Experience in ML/AI ops specifically (model deployment, evaluation, drift monitoring)
- Familiarity with prompt engineering, policies/guardrails, and configuration patterns for LLM-based systems
- Exposure to governance, compliance, or risk frameworks for data-driven or AI systems
- Experience collaborating with Security, Legal, and business stakeholders on technical risk and mitigation
- Prior involvement in incident response, change management, or on-call rotations for critical systems