SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
VAST Data is seeking a Senior Lab Reliability Engineer to own the operational reliability of VAST clusters in a presales lab environment that supports field engineering demonstrations, customer evaluations, and internal enablement. This is a senior individual contributor role with technical leadership responsibilities on a growing lab and platform engineering team.
You will own end-to-end reliability of VAST storage clusters in the lab, including proactive health monitoring, upgrade planning, and incident resolution. You'll serve as the primary technical escalation point for complex cluster issues, working hands-on to diagnose and resolve problems, and partnering with VAST engineering teams when deeper investigation is needed. You'll shape the automation and tooling strategy for the lab environment, establishing standards for infrastructure-as-code and configuration management practices that the rest of the team builds upon.
Key responsibilities include reproducing and isolating difficult issues in controlled environments to produce high-quality diagnostic data for engineering teams, owning broader systems infrastructure (virtualization platforms like VMware vSphere and Proxmox, compute, networking, storage), and mentoring other lab operations team members to raise the collective technical bar. You'll partner closely with pre-sales SEs, professional services, and engineering to reproduce customer-relevant scenarios and validate solutions.
Required qualifications: 4+ years in systems engineering, storage engineering, customer support engineering, or related role. Deep hands-on experience with enterprise storage systems (VAST, Pure, NetApp, Isilon, Ceph, or similar). Strong Linux systems administration skills including networking, storage, filesystems, and CLI tooling. Strong scripting/programming in Python, Bash, or similar. Operational experience with Docker and Kubernetes. Experience with infrastructure-as-code or configuration management (Ansible or equivalent). Solid networking foundations and virtualization platform experience. Methodical troubleshooting approach and excellent communication skills.
Preferred: hands-on VAST cluster experience, prior customer support or reliability engineering background at storage/infrastructure companies, familiarity with observability platforms (Grafana, Prometheus, Elasticsearch), network switch administration, or HPC/ML training workload experience.