SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Lightspeed Commerce is seeking a Staff Site Reliability Engineer to join the Platform team in Auckland. You will shape how the company builds and operates reliable systems at scale, working across AWS infrastructure, platform engineering, and software to enable product teams to build, ship, and operate great products independently.
Key Responsibilities:
- Shape the company's approach to infrastructure reliability, developer experience, and platform capabilities
- Design and maintain scalable, automated AWS infrastructure using Infrastructure as Code
- Build tools and platform services that help product engineers ship and operate independently
- Partner with engineering teams to design resilient, secure, and cost-effective systems
- Improve engineering practices and delivery processes across distributed, cross-functional teams
- Champion observability, high availability, incident management, and disaster recovery
- Apply systems thinking to complex problems, lead incident response, and drive lasting improvements
- Mentor engineers, influence reliability practices, and participate in on-call rotation
You will have high autonomy to tackle complex problems, influence technical direction across teams, and continuously improve platform resilience, scalability, and efficiency. The role offers flexibility with a hybrid workplace model, allowing remote work from anywhere in the world for up to 60 days per year, plus modern office space in Newmarket, Auckland.
Lightspeed is a dual-listed commerce platform company (NYSE: LSPD, TSX: LSPD) founded in 2005, serving retail, hospitality, and golf businesses in over 100 countries with cloud commerce solutions that unify online and physical operations, multichannel sales, global payments, and financial solutions.
Requirements:
- Strong AWS experience operating highly available, scalable production systems
- Experience in a SaaS or product-led environment, working cross-functionally
- Strong Infrastructure as Code experience with Terraform and configuration management tooling
- Experience with Docker, Kubernetes, ECS, and Linux systems
- Ability to code in Python, Ruby, or Go, alongside complex Shell scripting
- Experience with observability, incident management, disaster recovery, and security practices
- Experience with datastores such as MySQL, PostgreSQL, Redis, or DynamoDB, plus cloud cost optimization
- Understanding of Agile, continuous delivery, testing, with strong communication, judgment, and prioritization skills
- Note: Strong candidates who don't tick every box are encouraged to apply if they have strong AWS, infrastructure, and engineering foundations