SlipstreamJobsFresh Startup & VC-Backed Jobs

Platform / Site Reliability Engineer

Sunset - New York, NY, USA - In-office - posted 2026-08-13

Apply on the company site

SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.

Sunset is a rapidly scaling B2B SaaS platform that helps businesses unlock value through data partnerships with frontier AI labs. Founded to support startup shutdowns, the company has pivoted to become a primary source of proprietary business data for training next-generation AI models. In 2025, Sunset scaled from $0 to a multi-eight-figure run rate and raised from top-tier investors including Floodgate, Afore, Ludlow, and Hustle Fund. As Platform / Site Reliability Engineer, you will design and operate the shared infrastructure foundation that enables Sunset's product, data, and AI teams to ship reliable, secure, observable, and cost-aware systems. You will own the coherent platform layer across diverse workloads: customer-facing SaaS services, data connectors and ingestion pipelines, asynchronous workers, high-volume data processing, model-backed systems, and review tools. Key responsibilities include: - Establish Sunset's baseline for platform architecture, workload patterns, reliability metrics, ownership models, toil, recovery capabilities, costs, and technical controls - Build reusable infrastructure-as-code modules, runtime templates, deployment workflows, and operational tooling that support multiple workload shapes (services, batch jobs, pipelines, model workloads) - Create supported deployment and operational paths that enable teams to own their systems without manual infrastructure work or linear operational risk growth - Improve deploy safety, workload visibility, backup/recovery, incident response, and durable remediation across the platform - Define service and pipeline objectives, ownership models, escalation paths, and recovery procedures with engineering teams - Build self-service capabilities for infrastructure provisioning, environment management, access control, deployments, and debugging - Make cloud and vendor costs transparent by service and workload; improve efficiency within explicit reliability and security bounds - Partner with the Security Lead on cloud identity, secrets management, isolation, audit logging, vulnerability response, and automated control evidence - Leverage AI tools for platform engineering and operations while maintaining verification rigor for generated code, plans, and state changes Success metrics include: visibility into environments, runtimes, deploy paths, and infrastructure costs; materially reducing one failure or toil class in 90 days; enabling product/data/AI teams to ship faster while retaining clear ownership; establishing useful objectives and tested recovery paths for priority services; and making common platform work self-service while keeping exceptions explicit and monitored.

Similar roles