SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Salary: EUR 60,000 - 85,500 / annual
Nexthink is seeking an experienced Senior Site Reliability Engineer to join its SRE team in Madrid. The role focuses on strengthening infrastructure and enhancing the ability to deploy, monitor, and scale systems effectively and reliably across a global SaaS platform serving over 1,300 customers.
You will work closely with over 50 Product Engineering teams, Technical Platform Engineering, Security, and Architecture teams to understand reliability requirements, design and implement solutions, and promote adoption across the organization.
Key responsibilities include:
- Implement and manage cloud-native systems (AWS) using best-in-class tools and automation
- Operate and enhance Kubernetes clusters, deployment pipelines, and service meshes to support rapid delivery cycles
- Design, build, and maintain infrastructure powering a multi-tenant SaaS platform with reliability, security, and scalability in mind
- Define and maintain SLOs, SLAs, and error budgets; proactively address availability and performance issues
- Develop infrastructure-as-code (Terraform or similar) for repeatable and auditable provisioning
- Build internal platform tools and automation to support provisioning, monitoring, and operational efficiency
- Monitor infrastructure and applications to ensure high-quality user experiences
- Participate in shared on-call rotation, responding to incidents, troubleshooting outages, and driving timely resolution
- Act as Incident Commander during on-call duty and coordinate cross-team responses to maintain SLA
- Drive and refine incident response processes, reducing MTTD and MTTR
- Diagnose and resolve complex issues independently with minimal external escalation
- Work closely with software engineers to embed observability, fault tolerance, and reliability principles into service design
- Automate runbooks, health checks, and alerting to support reliable operations with minimal manual intervention
- Support automated testing, canary deployments, and rollback strategies for safe, fast, reliable releases
- Contribute to security best practices, compliance automation, and cost optimization
Nexthink is a leader in digital employee experience management software, headquartered in Lausanne and Boston with 9 offices worldwide and over 1,000 employees across 5 continents. The company offers a hybrid work model with a beautiful office in Prilly (near Lake Geneva), flexible hours, unlimited vacation, fitness center access, and relocation support.
REQUIREMENTS:
- Minimum Bachelor's degree in Computer Science or equivalent practical experience
- 5+ years of experience as a Site Reliability Engineer or Platform Engineer with strong knowledge of software development best practices
- Strong hands-on experience with public cloud services (AWS, GCP, Azure) and supporting SaaS products
- Strong programming or scripting skills (Python, Go, Bash, etc.) and experience with infrastructure-as-code (Terraform)
- Proficiency with Kubernetes, container-based deployment (Docker), and related ecosystems (Helm)
- Experience supporting multi-tenant microservices architectures
- Experience with CI/CD pipelines and tools (Jenkins, GitHub Actions, GitLab CI, FluxCD, Crossplane)
- Experience managing monitoring solutions (e.g., Datadog)
- Comfortable participating in rotating on-call schedule, managing critical incidents, and leading post-incident reviews
- At ease operating and managing production systems, balancing urgency with methodology
- Strong system-level troubleshooting skills and proactive mindset toward incident prevention
- Deep understanding of Linux systems, networking, and common troubleshooting practices
- Solid understanding of network stack (TCP/IP, VPN), cloud architectures (VPC, subnets, firewalls, load balancers), service mesh (Istio), and storage (S3, EBS)
- Knowledge of zero-downtime deployment strategies, blue/green and canary releases
- Exposure to compliance standards (SOC 2, ISO 27001, HIPAA); FedRAMP experience is a plus
- Experience with chaos engineering or resilience testing practices
- Excellent problem-solving skills, collaborative mindset, and strong grasp of agile, iterative development
- Self-driven, highly organized, capable of independently managing priorities
- Curiosity to learn new technologies
- Strong communication, presentation, and team collaboration skills
- Excellent written and verbal English skills