SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Spare is a fast-growing startup building cloud infrastructure for on-demand transit systems. You will lead the Infrastructure Platform Team, responsible for designing, building, and operating the reliable, secure, and cost-efficient cloud foundation that powers all Spare products.
You will split your time 50/50 between hands-on technical contribution and people leadership. Currently, four software developers report to this role. You'll work in an autonomous environment solving complex technical challenges while growing and mentoring a high-performing team.
Key responsibilities include:
- Own design and development of core infrastructure platform capabilities from inception to launch
- Build tooling, automation, and platform services that make every engineering team faster and safer
- Architect and implement high-performance, scalable distributed systems on GCP and Kubernetes
- Drive improvements in cluster reliability, application resilience, and internal access security
- Operate and maintain Redis and PostgreSQL databases (availability, performance tuning, scaling, upgrades, backups, disaster recovery)
- Drive AI SRE practices—embed AI agents into incident detection, alert triage, and operational workflows
- Manage and continuously improve the SRE on-call rotation with healthy schedules, clear escalation paths, and blameless post-mortems
- Drive FinOps practices across the organization—own cloud spend visibility, right-sizing, and cost optimization
- Use AI agentic tooling daily to accelerate your own and your team's engineering output
- Actively mentor software developers of all levels and uplift team capacity
- Collaborate cross-functionally with product managers, designers, and other engineers
- Ensure 99.99% uptime and maintain exceptional system performance
- Participate in team agile rituals and help improve software development processes
The role includes travel: up to four customer site visits per year and participation in biannual software development hackathons in Vancouver.
REQUIREMENTS:
- 7+ years of software development experience, with at least 2+ years in a people leadership role
- Expert in backend technologies with strong distributed systems experience
- Demonstrated proficiency with AI-assisted and agentic development workflows (AI coding agents, automation of engineering and operational tasks)
- Experience operating systems at scale with a strong reliability and uptime mindset
- Experience running or managing an SRE on-call rotation, including incident response and post-mortem culture
- Deep experience with cloud infrastructure (GCP) and container orchestration (Kubernetes)
- Experience driving cloud cost optimization and FinOps initiatives—right-sizing, spend visibility, and measurable cost savings
- Experience with infrastructure-as-code and configuration management tooling (Terraform)
- Understanding of security best practices, especially around internal access control and container security
- Demonstrated success in managing a team of software developers, with a focus on team and individual performance
- Demonstrated ability to mentor other developers and provide technical leadership
- Strong problem-solving, debugging, and system design skills
- Excellent communication and collaboration skills
NICE-TO-HAVE:
- Experience in the transit industry or another real-time, safety-critical domain
- Experience building internal developer platforms and golden-path tooling
- Experience applying AI/LLM tooling to SRE or infrastructure operations (AIOps, automated incident triage)
- Experience with CI/CD systems at scale, including test sharding and build performance
- Experience operating and maintaining PostgreSQL and Redis in production—performance tuning, indexing, replication, backup and restore