SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Cyvl is a Physical AI company building purpose-built sensors, computer vision, and AI infrastructure for municipal infrastructure management. The company's Infrastructure Intelligence Platform enables 500+ cities and towns to manage roads, bridges, and sidewalks with real-time data on physical conditions and asset locations.
As a Data Engineer, you will design, build, and operate the data pipelines and platform that transform raw sensor data (LiDAR, imagery, GPS) into client-ready geospatial deliverables. You'll work across the full data lifecycle: ingestion, processing, AI inference, quality assurance, and delivery to city engineers and planners.
Key responsibilities include:
- Designing and building production data pipelines that move sensor data from ingest through processing to client-ready outputs
- Building and maintaining the data platform used by product, AI, and analytics teams for data discovery and access
- Implementing automated data quality checks and monitoring systems to catch issues before customers encounter them
- Partnering with Data Operations and GIS teams to automate manual workflows and increase processing throughput
- Optimizing cost and performance of large-scale data processing on AWS
The role is based in Somerville, MA (on-site ~90% of the time) with an option to work remotely from San Francisco, CA. The company emphasizes ownership, end-to-end system responsibility, and direct collaboration with customers and cross-functional teams. You'll work on hard technical problems with real-world impact—your work directly influences how efficiently cities maintain public infrastructure.
REQUIREMENTS:
- 3+ years building production data pipelines or data platforms
- Strong Python and SQL
- Experience with cloud data infrastructure (AWS preferred) and workflow orchestration tools (Airflow, Dagster, Prefect, or similar)
- Demonstrated ability to make large, messy datasets reliable
- Comfort owning systems end-to-end, from design through on-call support
NICE TO HAVE:
- Geospatial data experience (PostGIS, GDAL, GeoParquet, H3, or other spatial indexing)
- Sensor data experience (LiDAR point clouds, imagery, GPS/IMU)
- Lakehouse formats and engines (Iceberg, Delta, DuckDB, Spark)
- Docker, Kubernetes, and Terraform
- Experience supporting ML training and inference pipelines