SlipstreamJobsFresh Startup & VC-Backed Jobs

Site Reliability Engineer

Viva Benefits - Marousi, Attica, Greece - In-office - posted 2026-09-22

Apply on the company site

SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.

Viva.com is Europe's first Tech Bank for businesses, operating critical infrastructure for payments and banking services across 31 countries. The company offers omnichannel payment acceptance, deposit accounts, card issuing, and financing solutions through a single integrated platform. Viva pioneered Tap on Any Device technology, enabling payment acceptance on any Android or iOS device while connecting to major international and domestic card schemes across Europe. As a Site Reliability Engineer, you will ensure the availability, performance, and scalability of production systems. You will collaborate with developers, system administrators, and support teams to improve system architecture and operational processes, with emphasis on automation, observability, and incident response. Key responsibilities include: - Ensuring reliability and uptime of critical production services and infrastructure - Designing scalable monitoring, alerting, and observability systems - Developing tools and automation to eliminate manual and repetitive work - Leading postmortems and root cause analysis, driving remediation to completion - Maintaining runbooks, escalation paths, and on-call documentation - Defining and tracking SLIs and SLOs with engineering teams - Collaborating with software and system engineers to improve system design for reliability - Identifying and fixing system weaknesses, bottlenecks, and single points of failure The company emphasizes amplifying human potential through AI, creating a future-forward environment where technology empowers every role. REQUIREMENTS: - More than 3 years of experience as a Site Reliability Engineer, DevOps engineer, or similar role - Hands-on experience operating production systems in a live on-call capacity - Proficiency in scripting languages (Python, Bash, Go, or similar) - Experience with cloud platforms (Azure preferred) and container orchestration (Kubernetes, Docker) - Practical experience applying AI to operations, including agentic and managed AI tooling for automation, diagnostics, or incident triage, with sound judgment on where human oversight is required - Strong understanding of Linux systems, networking, and troubleshooting - Familiarity with infrastructure as code tools (Terraform, Ansible, etc.) - Familiarity with observability stacks (Datadog, Prometheus, Grafana, ELK, etc.) - Familiarity with CI/CD pipelines and automated rollback - Experience with incident management and on-call tooling such as PagerDuty or incident.io - Clear communication under pressure, including with non-technical stakeholders - Strong problem-solving skills and passion for reliability and performance

Similar roles