SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Salary: USD 145,000 - 170,000 / annual
Banyan Software is seeking a Senior Site Reliability Engineer to own operational excellence for modernized SaaS applications delivered by the Banyan AI Factory to its Operating Companies. This is a hands-on, individual contributor role focused on running the reliability of distributed applications rather than building the factory infrastructure itself.
You will join a team providing 24x7 coverage with rotating on-call responsibilities, serving as Tier 1 SRE for OpCos' containerized applications across AWS and Microsoft Azure. Your day-to-day work will include automated deployments, cloud service integration, application performance and availability monitoring, observability implementation, and security incident response.
Key responsibilities include:
- Operating as part of a 24x7 on-call rotation to ensure continuous availability and rapid incident response
- Managing day-to-day cloud integrations across AWS and Azure to keep production systems healthy, performant, and secure
- Implementing and maintaining robust application observability tooling (monitoring, logging, tracing) to proactively detect degradation and drive down mean-time-to-detect and mean-time-to-resolve
- Developing, maintaining, testing, and executing disaster recovery and business continuity procedures
- Responding to security incidents and operational events, executing runbooks and coordinating remediation
- Using Infrastructure-as-Code (Terraform) and CI/CD pipelines (GitHub Actions, GitLab CI) to manage and automate operational environments
- Designing and building AI agents to automate SRE tasks and incident response, reducing toil and accelerating detection and remediation
- Serving as a technical escalation point for operational challenges in distributed, multi-tenant SaaS environments
This role is ideal for an engineer with a track record of keeping secure, highly available production systems running at scale, who is comfortable with hands-on problem-solving and technical ambiguity.
REQUIREMENTS:
- 5–7 years of progressive experience in Software Engineering and/or Site Reliability Engineering, with focus on operating distributed systems
- Deep expertise in Python, Javascript, or Go for automation coding and tool integrations (authentication, parallelization, triggering, APIs, data transformation)
- Deep expertise in container technologies (Docker/Kubernetes) supporting highly scalable and resilient distributed systems
- Hands-on experience with Infrastructure-as-Code using Terraform at scale (required)
- Hands-on experience operating production workloads on AWS (EC2, Lambda, EKS, S3, RDS) and/or Microsoft Azure (Container Apps, AKS, Container Storage)
- Deep hands-on experience with CI/CD platforms (GitHub Actions, GitLab CI) and embedding DevSecOps practices into operational workflows
- Experience with application-level logging, troubleshooting, and tracing tools; proven track record operating highly available production systems
- Experience with AI-assisted engineering tools such as Claude Code or similar
- Familiarity with APM tooling and practices (Datadog, New Relic, Dynatrace, or similar) to instrument, profile, and optimize application performance
- Demonstrated experience participating in on-call rotations, responding to production and security incidents, and executing disaster recovery procedures
- Exceptional communication, presentation, and collaboration skills with proven ability to coordinate across teams
- Bachelor's degree in Computer Science or related technical field
Preferred: Familiarity with advanced cloud security tools like Wiz, Prisma Cloud, and Checkov.