SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
GoGuardian is seeking a Senior Site Reliability Engineer to join the Tech Foundation team, which manages core cloud infrastructure, shared data services, and developer tooling. You will design, scale, and maintain the infrastructure powering GoGuardian's K-12 learning solutions, collaborating with engineering teams to drive operational excellence, optimize system performance, and ensure high availability across production environments.
Key Responsibilities:
- Architect and maintain scalable, secure cloud infrastructure to ensure high availability for core products
- Enhance observability and monitoring frameworks to deliver highly accurate alerts, minimizing noise and improving incident detection
- Participate in on-call rotations and lead incident response, ensuring comprehensive post-mortems and root cause analyses are completed to drive systemic improvements
- Optimize and modernize deployment pipelines and automation workflows to maximize engineering velocity and operational safety
- Partner with product development teams to provide infrastructure support, review architectural changes, and promote reliability best practices
- Implement and uphold robust security standards and compliance controls across all managed cloud infrastructure
GoGuardian is a remote, mission-driven learning company focused on improving K-12 education environments. The company emphasizes collaborative problem-solving, accountability, and an inclusive culture where employees bring their whole selves to work.
Requirements:
- 5+ years of professional experience in Site Reliability Engineering, Infrastructure, or DevOps roles supporting production SaaS applications
- Strong proficiency with AWS core services (EC2, VPC, S3), Serverless frameworks, and managed Kubernetes environments like EKS
- Extensive experience writing and managing Infrastructure as Code using Terraform
- Familiarity with configuring, troubleshooting, and maintaining data layers such as MongoDB, Redshift, and OpenSearch (GCP/Firestore experience is a plus)
- Experience managing or modernizing CI/CD pipelines and deployment workflows using Jenkins, AWS CodeBuild/CodePipeline, or GitHub Actions
- Deep understanding of Linux operating system fundamentals and Unix shell scripting
- Ability to read and debug code written in JavaScript/TypeScript, Python, or Go
- Strong communication and collaboration skills with a track record of driving technical decisions and establishing team-wide operational standards
- Eagerness to take initiative in a fast-paced, dynamic environment