SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Salary: USD 75,450 - 169,700 / annual
Remote is a global HR platform that enables businesses to recruit, pay, and manage international teams compliantly. The company operates fully remotely across 6 continents with a strong async-first culture.
You will lead Remote's Site Reliability Engineering team as a Team Leader in a 60% individual contributor, 40% people management role. The SRE team owns Kubernetes, AWS, PostgreSQL, CI infrastructure, observability, and reliability practices across the platform. You will manage 4 direct reports, own their career development and performance, and serve as the team's spokesperson to engineering leadership.
The reliability practice at Remote is maturing—SLO frameworks are live on initial teams and need broader rollout. There is meaningful work to balance operational load against project delivery, making this an opportunity to shape foundational practices while solving open technical problems.
Key responsibilities include:
- Full career lifecycle management: onboarding, feedback, performance assessment, progression, and hiring
- Team health, dynamics, and retrospective practices
- Setting SRE goals, priorities, and on-call/support rotation models
- Owning core infrastructure: Kubernetes, AWS, PostgreSQL, DNS, TLS, CI systems
- Driving reliability practices: SLOs, error budgets, incident response, observability
- Partnering with Security on threats, patching, compliance, and audit obligations
- Managing vendor relationships and commercial conversations
You will report to the Director of Engineering, Platform.
REQUIREMENTS:
People Leadership:
- Proven experience leading an SRE, infrastructure, or platform engineering team with ownership of reports' growth, performance, and career progression
- Demonstrated ability to coach both technical craft and soft skills; can point to people who grew under your mentorship
- Direct experience handling underperformance with clarity and empathy
- Track record hiring engineers and distinguishing good interviews from good engineers
- Strong ability to read team dynamics and resolve conflict constructively
- Natural talent for fostering commitment to company goals
Technical Depth:
- Hands-on background in site reliability, DevOps, or cloud infrastructure engineering with depth to review team work, challenge designs, and be credible in incidents
- Production Kubernetes experience, including operational realities beyond happy-path scenarios
- AWS at meaningful scale
- Hands-on AI building, enablement, and AI infrastructure scaling
- Solid observability practices and principles
- Infrastructure as code with Terraform
- CI/CD systems (GitLab CI, GitHub Actions, Jenkins)
- Docker and shell scripting
- Experience running a reliability practice: incident response, on-call, SLOs, error budgets, and turning incidents into lasting changes
- Understanding and history working in regulated environments
Ways of Working:
- Exceptional prioritization when operational load and project work compete; ability to protect team focus without dropping operational commitment
- Clear written communication (critical in fully distributed, async environment)
- Strong cross-team relationship building
Nice to Have:
- Backend language knowledge (Elixir preferred; Java, Clojure, Node.js, Python acceptable)
- Modern observability depth: OpenTelemetry, distributed tracing, tools like Honeycomb
- Database operations experience, especially PostgreSQL/Aurora performance tuning and connection pool health
- Linux system administration outside cloud environments
- Security capability (defensive and offensive)
- Cloud cost management and FinOps
- Experience growing teams from small base and building hiring bars