SlipstreamJobsFresh Startup & VC-Backed Jobs

Senior Site Reliability and Infrastructure Engineer

Treeswift - New York, NY, USA - Hybrid - posted 2026-08-04

Apply on the company site

SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.

Treeswift is building physical AI for field workers in the energy sector, enabling utility companies to modernize operations and reduce wildfire risk, outages, and costs. The company combines hardware (LiDAR, cameras, sensors), AI, and software to deliver 10x productivity gains for engineers, linemen, and vegetation crews. Founded in 2024 with pilots across three of the five largest US utilities, Treeswift is backed by leading investors and headquartered in Manhattan with offices in San Francisco and Philadelphia. As the first full-time SRE/Infrastructure Engineer, you will lead the design and scaling of Treeswift's platform infrastructure to support rapid growth. You'll own reliability and observability across a complex stack: Apache Airflow on Astronomer orchestrating high-volume data pipelines, Kubernetes-based containerized workers, AWS services (S3, SQS, Lambda, Step Functions, ECS), and machine learning inference operations. Key responsibilities include: - Partner with data platform and engineering teams to understand how changes propagate across pipeline execution, containerized workers, and cloud services. - Design and implement comprehensive monitoring, alerting, and dashboards for DAG/task failures, operational workflows, and pipeline health with clear SLO/SLI definitions. - Own CI/CD guardrails for production changes, including build/deploy validation and safe rollout mechanics for Astronomer deployments. - Instrument and operationalize machine learning inference runs, adding visibility into model checkpoint resolution, S3 sync behavior, inference outcomes, and failure modes. - Create operational tooling, runbooks, incident learnings, and engineering standards to reduce toil and improve debugging at scale. - Lead reliability improvements and operational readiness to enable faster diagnosis, better alerts, and safer releases. While there is no current on-call rotation and pipelines do not require real-time processing, you'll establish the observability and reliability foundations for confident operation as customer data volume grows. You bring 7-10 years of software engineering experience with significant focus on observability, systems/infrastructure, SRE, or DevOps in cloud environments. You reason about architecture end-to-end with product impact in mind, have hands-on experience with infrastructure-as-code (Terraform), container orchestration (Kubernetes/ECS), and strong Linux debugging skills. You communicate clearly across teams and thrive in fast-moving, ownership-driven environments. Nice-to-haves include Apache Airflow/Astronomer experience, AWS expertise, geospatial/imagery/LiDAR domain knowledge, and MLOps skills.

Similar roles