SlipstreamJobsFresh Startup & VC-Backed Jobs

Senior Site Reliability Engineer (remote within EMEA)

Printify - Remote - Remote

Apply on the company site

SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.

Join Printify's Platform Infrastructure team as a Senior Site Reliability Engineer II, a senior individual contributor role focused on building, operating, and evolving the company's container platform and cloud foundation. You'll architect and drive large-scale automation and reliability initiatives across AWS, Kubernetes, and GCP infrastructure, setting standards that other engineers follow while mentoring mid-level SREs. Your responsibilities span multiple domains: design and operate highly available, secure, and scalable infrastructure across multi-account AWS environments using infrastructure-as-code (Terraform/Terragrunt); architect and manage Amazon EKS clusters with networking policies, persistent storage, and scaling strategies; own core platform services including cloud networking, Kubernetes, databases, and messaging systems; drive large-scale automation projects and GitOps adoption (ArgoCD) to reduce manual operational work; lead observability and incident response initiatives using Grafana, Prometheus, Loki, Tempo, and Mimir; participate in on-call rotations and lead production incident response; audit infrastructure spend and drive cost optimization through FinOps practices; mentor mid-level SREs and support team onboarding; and partner with product engineering squads to understand their needs. Required qualifications include solid Linux systems administration background with Python scripting; strong AWS expertise (EKS, IAM, VPC, RDS, S3, SQS, Well-Architected Framework); hands-on production-scale Kubernetes experience including Helm, Cilium CNI, and container security; Terraform/Terragrunt proficiency and GitOps experience with ArgoCD; production experience with Postgres, MySQL, MongoDB, or Aurora; CI/CD experience with Jenkins or GitHub Actions; expertise maintaining the Grafana observability stack; practical incident management experience; and 12-Factor App and FinOps awareness. You should demonstrate methodical, data-driven troubleshooting; strong written communication for runbooks and postmortems; ability to drive initiatives with ambiguous ownership; track record of mentoring; and several years of hands-on production infrastructure/SRE experience with demonstrated senior individual-contributor ownership.

Similar roles