SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Onebrief builds collaboration and AI-powered workflow software for military planning and operational coordination. The company brings planning, collaboration, simulation, and AI into one connected environment, helping teams test strategies, adapt to changing conditions, and make decisions with greater clarity in high-stakes operational contexts.
You will join the Infrastructure & Security team as a Senior Site Reliability Engineer, serving as the first line of support for mission-critical deployments. You'll work closely with fellow SREs, security engineers, and customer success teams, operating in both on-premise DoD environments and AWS cloud environments. Your responsibilities span reliability, scalability, and security of production applications and platforms.
Key responsibilities include:
- Designing, implementing, and managing a world-class observability platform using tools like Prometheus, Loki, Alloy, and Grafana to create actionable insights and automated alerting.
- Defining, measuring, and owning Service Level Indicators (SLIs) and Service Level Objectives (SLOs), becoming the organization's expert on system reliability measurement.
- Acting as incident responder and potential incident commander during critical incidents, leading blameless post-mortems and After Action Reviews to identify root causes and drive automated, long-term solutions.
- Partnering with platform engineers to design, build, and manage secure, resilient Kubernetes clusters and cloud/on-prem environments using Infrastructure-as-Code (Terraform, Ansible), embedding security and compliance controls (RMF, STIGs).
- Proactively identifying and eliminating operational toil through automation, sharing best practices for air-gapped environments, and supporting team readiness for production.
You care deeply about reliability as a core feature, treating infrastructure and operability as products to be automated, well-documented, and continuously improved. You're equally comfortable leading post-incident reviews or diving into kubectl shells to triage complex production issues. You mentor others, fostering a culture of blameless postmortems and proactive reliability.
Onsite requirement: This role requires regularly working on-site at customer locations in Arlington, VA. If not currently within commuting distance, you must be willing to relocate; Onebrief provides relocation assistance.
REQUIREMENTS:
- Active Top Secret clearance required; SCI eligibility is a plus.
- 5+ years in Platform, DevOps, or Site Reliability Engineering with an infrastructure and operations focus.
- Proven ability to partner with DevOps/Platform and application teams; collaborates well across functions and shares context openly.
- Deep understanding of incident response processes, with experience conducting thorough root cause analyses and driving continuous improvement.
- Infrastructure as Code: Terraform or CloudFormation, Ansible.
- Containers and orchestration: Kubernetes design, deployment, and operations.
- CI/CD: experience building and maintaining pipelines (GitLab CI/CD, Jenkins, GitHub Actions).
- Scripting: proficiency with at least one of Python, Go, or Bash.
- Cloud: Familiarity with AWS or AWS GovCloud.
- Observability: Grafana stack, ELK stack, or Datadog.
- Networking fundamentals: core protocols and secure configurations.
BONUS QUALIFICATIONS:
- Experience in DoD environments and compliance frameworks (RMF, STIGs, ICD 503).
- GitOps practices and toolchains.
- Security-minded design for sensitive environments.
- Experience designing and implementing meaningful SLIs/SLOs (including error budgets) for complex, distributed systems.
- Familiarity with on-prem virtualization (VMware, Proxmox, Nutanix, Hyper-V, etc).
- Service mesh exposure (Istio, Linkerd).
- Relevant certifications (e.g., AWS DevOps Engineer, CKA/CKAD).
- Active Security+ or another DoD 8570.01-approved security credential, or ability to obtain valid credentials within 3 months of employment.