SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
YipitData is a market research and analytics firm specializing in alternative data for the disruptive economy. Recently valued at over $1B after raising $475M from The Carlyle Group, the company analyzes billions of data points daily to provide insights on ridesharing, e-commerce, payments, and more to major investment funds and corporations.
You will join a small, high-leverage infrastructure team reporting to the Head of Cloud Infrastructure. This is a hands-on senior engineering role focused on modernizing and scaling the shared cloud infrastructure that powers YipitData's data-intensive products and agentic customer experiences.
Key responsibilities include:
- Design, build, and operate shared cloud infrastructure using AWS, Kubernetes, Terraform, Databricks, and Cloudflare
- Deliver SRE and DevOps initiatives improving reliability, scalability, observability, and operational readiness
- Build reusable infrastructure modules and self-service workflows that enhance developer experience
- Define and implement SLIs, SLOs, monitoring, alerting, and error-budget practices for critical systems
- Lead incident response, post-incident reviews, and corrective actions
- Strengthen disaster-recovery readiness through planning, automation, and testing
- Improve CI/CD workflows and infrastructure delivery for faster feedback and confident deployments
- Partner with application, Data Platform, and Data Engineering teams on infrastructure needs
- Optimize cloud costs through architecture, capacity planning, and Kubernetes resource optimization
- Contribute to technical standards, architecture decisions, and sustainable 24/7 operating practices
You bring 4+ years of software, infrastructure, SRE, platform engineering, or DevOps experience with hands-on production use of AWS, Kubernetes, Datadog, Databricks, and Terraform. You have strong interest in AI-native and data-intensive applications, experience operating business-critical systems, and familiarity with reliability concepts like SLIs, SLOs, and error budgets. You write reliable, maintainable code and are comfortable working across infrastructure, systems, and application boundaries.
Note: This role requires working East Coast hours despite remote flexibility.