SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
VAST Data is seeking an experienced Solutions Data Engineer to join a fast-growing AI infrastructure company. You'll partner with engineering, product, and customers to design and deliver high-impact distributed data systems that move, transform, and serve data at scale.
In this role, you'll own end-to-end platform lifecycle from ingestion through transformation, storage, and compute. You'll build distributed data pipelines using Kafka, Spark (batch & streaming), Python, Trino, Airflow, and S3-compatible data lakes—designed for scale, modularity, and seamless integration across real-time and batch workloads. You'll design, deploy, and troubleshoot hybrid cloud/on-premises environments using Terraform, Docker, Kubernetes, and CI/CD automation tools.
Key responsibilities include implementing event-driven and serverless workflows with precise control over latency, throughput, and fault tolerance; creating technical guides, architecture documentation, and demo pipelines to support onboarding and accelerate adoption; integrating data validation, observability tools, and governance into the pipeline lifecycle; benchmarking and tuning storage backends (S3/NFS/SMB) and compute layers for throughput, latency, and scalability using production datasets; and operating and debugging object store–backed data lake infrastructure for schema-on-read access, high-throughput ingestion, and advanced searching strategies.
You'll work cross-functionally with R&D to push performance limits across interactive, streaming, and ML-ready analytics workloads. This role requires balancing various project aspects—from safety to design—while researching advanced technology and identifying cost-effective solutions. You'll be comfortable switching between low-level debugging, high-level architecture, and communicating clearly with stakeholders of all technical levels.
REQUIREMENTS:
- 2–4 years in software, solutions, or infrastructure engineering
- 2–4 years focused on building/maintaining large-scale data pipelines, storage, and database solutions
- Proficiency in Trino, Spark (Structured Streaming & batch), and solid working knowledge of Apache Kafka
- Python coding (must-have); familiarity with Bash and scripting tools is a plus
- Deep understanding of data storage architectures including SQL, NoSQL, and HDFS
- Solid grasp of DevOps practices: containerization (Docker), orchestration (Kubernetes), infrastructure provisioning (Terraform)
- Experience with distributed systems, stream processing, and event-driven architecture
- Hands-on familiarity with benchmarking and performance profiling for storage systems, databases, and analytics engines
- Excellent communication skills to explain thinking clearly, guide customer conversations, and collaborate across teams