SlipstreamJobsFresh Startup & VC-Backed Jobs

Site Reliability Engineer

GiveCampus - Remote - Remote - posted 2026-09-21

Apply on the company site

SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.

GiveCampus is the world's leading fundraising platform for non-profit educational institutions, trusted by 1,300+ colleges, universities, and K-12 schools. The company is backed by Y Combinator and has achieved profitability while maintaining rapid growth, making the Inc. 5000 list for five consecutive years. With 130+ employees distributed across 30+ US states, GiveCampus operates a remote-first culture with optional office space in Washington, DC and regular team gatherings. This is a hands-on Site Reliability Engineer role focused on improving the reliability, performance, and operational maturity of GiveCampus's platform. Following the company's migration to AWS, the position will center on operating and strengthening the production environment through improved observability, infrastructure automation, incident response, and collaboration with product engineers to build resilient systems. Key responsibilities include: - Operating, maintaining, and improving production infrastructure in AWS - Building and maintaining infrastructure as code using Terraform - Supporting workloads on Kubernetes and Amazon EKS - Improving dashboards, alerts, and service-level indicators using observability platforms like New Relic - Investigating production issues, identifying root causes, and implementing durable fixes - Participating in a shared 24/7 on-call rotation and contributing to effective incident response - Partnering with product engineers to troubleshoot performance and reliability issues - Improving application resilience using patterns such as timeouts, retries, queuing, and idempotency - Maintaining and improving CI/CD pipelines using GitHub Actions and CircleCI - Automating repetitive operational tasks and reducing engineering toil - Creating and maintaining runbooks, system diagrams, and production documentation - Contributing to capacity planning, performance testing, and production-readiness reviews - Owning small-to-medium reliability improvements from design through delivery The role requires independent work on straightforward problems, seeking guidance when needed, explaining technical tradeoffs, and keeping teammates informed at key milestones. REQUIREMENTS: - Approximately 5+ years of related experience in software engineering, infrastructure, systems engineering, SRE, Platform Engineering, DevOps, or equivalent practical experience - Hands-on experience operating production workloads in AWS - Experience building or maintaining infrastructure using Terraform or similar infrastructure-as-code tools - Experience with New Relic, Datadog, or another modern observability platform - Experience troubleshooting production incidents and participating in an on-call rotation - Experience building or maintaining CI/CD pipelines - Software development or scripting experience with ability to read, debug, and make targeted changes to application or automation code - Working knowledge of Linux, networking, distributed systems, and relational databases - Ability to articulate root causes, explain technical tradeoffs, and translate findings into practical solutions - Ability to manage well-scoped projects with general direction and provide timely updates at key milestones - Strong written and verbal communication skills and collaborative approach to working across engineering disciplines - Habit of automating repetitive work and improving system reliability BONUS EXPERIENCE: - Ruby or Ruby on Rails - PostgreSQL administration or performance-tuning - Kubernetes and Amazon EKS - Redis, OpenSearch, or Amazon RDS - Operating enterprise SaaS products at scale - SLOs, SLIs, error budgets, capacity modeling, or load testing - Payments, fintech, or regulated systems - SOC 2 or similar security and compliance programs

Similar roles