SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Arista Networks is seeking an experienced Site Reliability Engineer to join their organization on a permanent remote basis from Ireland. You will be instrumental in building, deploying, and operating critical production systems with a focus on scalability, reliability, observability, and security.
Key responsibilities include:
- Design, build, and deploy production systems ensuring they meet stringent security and performance standards
- Develop and maintain comprehensive automation solutions to eliminate toil and streamline operational efficiency
- Proactively monitor production systems, establish intelligent alerting strategies, and implement automated incident response mechanisms
- Create and maintain detailed incident response runbooks and conduct thorough postmortem analyses
- Collaborate with software engineering teams to identify and resolve infrastructural bottlenecks
- Manage and optimize monitoring infrastructure using industry-standard tools
- Plan and execute maintenance windows on production systems with minimal service disruption
- Triage platform and infrastructural issues with analytical rigor
- Deploy new systems and updates in a staged, risk-managed manner
- Survey and adopt best practices in infrastructure and platform management
- Study design and implementation details of open-source systems to enhance troubleshooting
- Work transparently with stakeholders on system status and infrastructure improvements
Arista Networks is an engineering-centric company with a flat management structure led by engineers. The company is headquartered in Santa Clara with development offices globally including Ireland, Australia, Canada, and India. Engineers have complete ownership of projects and access across the company. Arista is a well-established, profitable company with over $8 billion in revenue, specializing in data-driven, client-to-cloud networking solutions.
REQUIREMENTS:
Essential:
- Bachelor's degree in Computer Science, Engineering, or equivalent professional experience (5+ years in related infrastructure or systems role)
- Proficiency in one or more programming languages: Go, Python, or bash shell scripting, with ability to implement medium-complexity automation workflows
- Strong knowledge of Linux or UNIX from both administration and debugging perspectives
- Hands-on experience operating software systems, infrastructure, and complex applications at scale in production environments
- Demonstrated expertise in infrastructure-as-code principles and practices
- Strong problem-solving and software troubleshooting skills with methodical, analytical approach
- Experience with server provisioning, particularly from storage and networking perspectives
- Proven ability to work collaboratively within cross-functional teams and communicate technical concepts clearly
- Experience with incident response, postmortem analysis, and continuous improvement methodologies
Desirable:
- Experience with container orchestration platforms, particularly Kubernetes
- Hands-on experience with Docker and virtualization technologies
- Proficiency in managing monitoring stacks, including Prometheus and Grafana
- Experience with CI/CD systems such as GitLab tools or Spinnaker
- Knowledge of infrastructure-as-code frameworks, particularly Terraform
- Experience managing databases such as PostgreSQL or equivalent relational database management systems
- Experience with artifact repositories and Docker registries
- Familiarity with cloud platforms (GCP, AWS, or Microsoft Azure)
- Understanding of distributed systems architecture and principles
- Experience with performance tuning and system optimization
- Knowledge of security best practices in infrastructure and systems design
- On-call support experience and comfort with incident response responsibilities