SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Onebrief builds collaboration and AI-powered workflow software for military planning and operational coordination. The company brings planning, collaboration, simulation, and AI into one connected environment to help teams test strategies, adapt to changing conditions, and make decisions with greater clarity in high-stakes environments.
You will lead the Site Reliability Engineering team within Infrastructure & Security, working closely with platform engineering, application engineering, security, and customer success to ensure Onebrief's mission-critical deployments are reliable, secure, and well-supported across on-premises DoD and AWS environments.
Key responsibilities include:
- Lead and develop the SRE team: hire, coach, support engineers across infrastructure, operations, and application reliability; set expectations, manage performance, and support career development.
- Own capacity and work planning: maintain visibility of incoming requests, ongoing support needs, and planned engineering work; establish ownership and sequence work.
- Coordinate delivery across teams: manage dependencies, identify blockers early, resolve competing priorities, and communicate progress and risks to stakeholders.
- Set the reliability roadmap: translate customer needs, production data, and incident patterns into a prioritized improvement plan; protect capacity for work that reduces recurring failures.
- Establish operational ownership: ensure production deployments have clear support responsibilities, escalation paths, and readiness criteria.
- Support team-led incident response: establish sustainable on-call coverage, coach incident commanders, and lead blameless postmortems and After Action Reviews.
- Guide technical priorities: evaluate approaches to infrastructure, automation, observability, and application reliability; ensure plans account for operability and security requirements in on-premises and air-gapped environments.
- Make reliability and workload visible: guide the team's use of SLIs, SLOs, and operational metrics to explain where investment is needed.
- Reduce operational toil: give engineers time and support to automate repetitive work and turn lessons into reusable improvements.
You care deeply about reliability and understand the challenges of operating software in constrained environments. You are an effective people manager who sets clear expectations, gives useful feedback, and helps engineers grow. You are comfortable managing changing workloads, turning competing requests into actionable plans, and accounting for operational interruptions. You bring calm and structure when priorities shift or incidents occur. You have the technical judgment to determine whether recurring problems need infrastructure changes, better automation, application fixes, or clearer processes.
REQUIREMENTS:
- Active Secret clearance (required)
- 5+ years in Site Reliability Engineering, Platform Engineering, DevOps, or related role with substantial infrastructure and operations experience
- Direct experience managing engineers, including coaching, performance management, career development, and hiring
- Experience planning team capacity, prioritizing competing requests, and coordinating engineering work across teams
- Track record of delivering reliability improvements while managing ongoing operational responsibilities
- Experience with incident response and post-incident review practices
- Technical judgment sufficient to evaluate engineering proposals, understand operational risks, and guide prioritization
- Clear communication skills, including ability to explain constraints and negotiate scope, timing, and commitments
- Practical experience operating production systems with breadth across: Infrastructure as Code (Terraform, Ansible, Python, Go, Bash), Kubernetes deployment and operations, CI/CD pipelines and release safety, AWS or AWS GovCloud and customer-managed infrastructure, observability tools (Grafana stack, ELK, Datadog), networking and security, application reliability
BONUS EXPERIENCE:
- Leading teams supporting mission-critical customer deployments
- Managing engineering programs or coordinating delivery across multiple customer environments
- Operating in DoD, classified, or air-gapped environments
- Familiarity with RMF, STIGs, and ICD 503
- Implementing SLIs, SLOs, and error budgets
- GitOps practices and toolchains
- On-premises virtualization (VMware, Proxmox, Nutanix, Hyper-V)
- Application development experience (TypeScript, Node.js)
- Relevant certifications (AWS DevOps Engineer, CKA/CKAD, Security+, or DoD 8570.01-approved credential)
NOTE: This role requires regularly working on-site at customer locations. If not currently within commuting distance, you must be willing to relocate. Onebrief provides relocation assistance.