SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Treeswift is building physical AI for field workers—hardware, sensors (LiDAR, camera), and software that multiplies productivity for energy company engineers, linemen, and vegetation crews. Since launching pilots in mid-2024, the company has grown to work with three of the five largest US utilities, helping reduce wildfire risk, regulatory outage risk, and accelerate storm recovery. The team combines robotics expertise (Penn, Caltech, CMU) with enterprise software talent (Palantir, Stripe, Oracle, MongoDB), backed by leading investors including Penny Pritzker's Inspired Capital.
You will be Treeswift's first full-time DevOps Engineer, owning infrastructure strategy and execution as the platform scales. The role spans a complex, mission-critical stack: Apache Airflow on Astronomer orchestrating high-volume data pipelines across AWS and Kubernetes, machine learning training and inference, and a web application serving utility customers.
Key responsibilities include designing and implementing observability and reliability for production pipelines—monitoring, alerting, performance/cost visibility, and operational runbooks. You'll partner with data platform and engineering teams to understand how changes propagate across Airflow DAGs, containerized workers (Kubernetes), and AWS services (S3, SQS, Lambda, Step Functions, ECS). You'll own CI/CD guardrails for safe production deployments, instrument ML inference operations for reliability, and create operational tooling to reduce toil.
While there is no established on-call rotation currently (pipelines don't require real-time processing), you'll lead reliability improvements and operational readiness so the team can diagnose issues faster, alert better, and release safer.
You bring 7–10 years of software engineering with significant time in observability, systems/infrastructure, SRE, or DevOps in cloud environments. You reason about architecture end-to-end with product impact in mind, have hands-on infrastructure-as-code (Terraform), container orchestration (Kubernetes/ECS), strong Linux debugging skills, and can collaborate effectively across teams. Nice-to-haves include early-stage startup experience, Apache Airflow/Astronomer, AWS, geospatial/imagery/LiDAR domains, and MLOps.