SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Ontic provides AI-powered security software for corporate and government teams to identify threats, assess risk, and respond to incidents. The company's Connected Intelligence Platform unifies security operations and data, serving Fortune 500 companies and federal agencies.
You will architect and operate secure, scalable Kubernetes platforms across AWS and AWS GovCloud environments, including FedRAMP-authorized systems. This is a senior individual contributor role with significant technical leadership responsibilities—you'll mentor engineers, drive platform strategy decisions, and partner with Security and Engineering leadership.
Key areas of focus:
**Enterprise Platform Architecture**: Design resilient Kubernetes platforms across cloud environments and FedRAMP authorization boundaries. Define standardized blueprints for multi-tenant and dedicated deployments. Architect high-availability, multi-AZ, and disaster recovery strategies aligned with RTO/RPO objectives. Own system inventory, authorized-boundary governance, and significant-change assessment.
**DevSecOps & Zero-Trust Security**: Establish secure-by-design cloud architecture and workload identity models. Enforce RBAC, IAM governance, encryption standards, and least-privilege principles. Integrate supply-chain security and compliance validation into CI/CD pipelines. Ensure infrastructure meets FedRAMP and enterprise security-control requirements. Manage vulnerability remediation, POA&M tracking, and incident-response exercises.
**AI-Enabled Operations**: Implement AIOps capabilities for intelligent alerting, anomaly detection, and controlled remediation. Enhance observability using metrics, logs, and distributed tracing. Support AI/ML workload infrastructure. Leverage automation to reduce MTTR and operational toil.
**Infrastructure Automation & GitOps**: Drive standardization using Terraform, Helm, and GitOps workflows. Design safe deployment strategies (blue/green, canary, progressive delivery). Improve infrastructure drift detection and environment reproducibility.
**Reliability & Performance Engineering**: Define SLO/SLI frameworks and reliability benchmarks. Optimize autoscaling and workload efficiency. Improve performance of distributed stateful systems. Conduct recurring disaster-recovery testing.
**FinOps & Cost Governance**: Implement cloud cost observability and allocation strategies. Optimize compute, storage, and networking costs. Establish cost governance models aligned with business growth.
**Leadership & Governance**: Mentor senior engineers and promote DevOps culture. Drive architecture reviews and production readiness assessments. Partner with security and engineering on platform strategy alignment.
**Requirements:**
- 8–10+ years of DevOps / Platform Engineering experience in enterprise or high-scale SaaS environments
- Proven experience architecting and operating production environments in AWS and AWS GovCloud (GCP or OCI experience a plus)
- Demonstrated experience maintaining a FedRAMP Moderate/High production environment, including continuous monitoring, vulnerability remediation, incident response, configuration management, and audit evidence
- Experience owning infrastructure patterns and changes within authorized boundaries, including system inventory, data-flow implications, inherited controls, and third-party services
- Strong experience designing and operating AWS Landing Zone architectures with multi-account governance and guardrails
- Advanced Terraform experience building reusable, secure infrastructure modules and standardized environments
- Deep understanding of AWS and AWS GovCloud networking (VPC design, segmentation, routing, private connectivity, security controls)
- Deep expertise in Kubernetes-based platform design and large-scale cluster operations
- Strong experience implementing GitOps workflows (ArgoCD) and CI/CD automation (Jenkins, GitLab)
- Hands-on production experience with distributed systems (MongoDB, Elasticsearch, Kafka, Redis, ArangoDB)
- Strong understanding of IAM, encryption at rest and in transit, and zero-trust access models
- Experience implementing enterprise observability stacks
- Experience designing multi-tenant SaaS platforms and dedicated workload environments
- Exposure to service mesh and secure east-west traffic management
- Exposure to AIOps is a plus
Note: This role supports a FedRAMP environment and requires work to be performed within the United States by U.S. citizens authorized to work in the U.S. The role is based in Austin and requires three days per week in the office.