SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Onebrief builds AI-powered collaboration and workflow software for military planning and operational coordination. This Senior Site Reliability Engineer role sits at the intersection of engineering and operations, focused on making the production application reliable, scalable, and secure—primarily by improving the software itself rather than working around problems in infrastructure.
You'll work closely with product engineers, fellow SREs, security, and customer success teams. The role is weighted heavily toward engineering: much of the reliability and performance work happens in the codebase (primarily TypeScript), so you'll fix problems at the source. You'll be a first line of support for mission-critical deployments across on-premises DoD and AWS environments, and insights from field work feed directly back into product improvements.
Key responsibilities include: improving application reliability and performance by working directly in the TypeScript codebase; designing and operating monitoring, logging, and alerting systems (Prometheus, Loki, Alloy, Grafana) that developers actually use; defining and measuring SLIs and SLOs with corresponding alerting; leading incident response and running blameless post-mortems (AARs) to identify root causes and implement lasting fixes; and automating away repetitive operational work while sharing solutions across teams, including those in air-gapped environments.
You treat reliability as a feature, not an afterthought. You understand the full software development lifecycle and know where reliability fits into design, code review, testing, and release. You're comfortable reading and writing application code and equally at home debugging production issues via kubectl. You mentor others, foster a culture of blameless postmortems, and work naturally with product and platform teams to help them move fast without breaking things.
Required: active Secret clearance, 5+ years in software engineering/SRE with real application code shipping experience, strong TypeScript skills, solid grasp of the full SDLC, incident response and root cause analysis experience, and strong cross-functional collaboration skills. Technical expertise needed: TypeScript (Node and/or modern front-end frameworks), CI/CD pipeline building (GitHub Actions, GitLab CI/CD, Jenkins), testing and quality practices, comfort with Python/Go/Bash for tooling, containers and Kubernetes, and networking fundamentals.
Bonus skills: observability tools (Grafana stack, ELK, Datadog), Infrastructure as Code (Terraform, Ansible), AWS/AWS GovCloud experience, Kubernetes cluster design, meaningful SLI/SLO design with error budgets, GitOps practices, and DoD environment experience.
The role requires regularly working on-site at customer locations in Arlington, VA. Relocation assistance is provided for candidates not currently within commuting distance.