SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Snorkel AI is hiring Senior Platform Engineers to build and operate the infrastructure powering their data-centric AI platform. The Platform organization owns the foundational systems—pipelines, evaluation frameworks, access layers, event systems, governance, compute, and agent infrastructure—that every product team and customer deployment depends on.
You'll work on a small, high-impact team in the middle of a major architectural transformation: migrating from a single-database data path to a multi-source, event-driven, agent-first platform. The decisions made now will define how the platform scales for years.
Key responsibilities include:
- Design and build agent infrastructure for safe, reliable workflow acceleration
- Implement event-driven data flows using event brokers, CDC connectors, schema registries, and dead letter queues
- Build systems for data lineage tracking, governance/RBAC enforcement, and audit logging (PII handling, retention policies, enterprise compliance)
- Set strategy for build systems, testing frameworks, and CI/CD pipelines; drive transition to robust automated continuous deployment
- Instrument services with OpenTelemetry, define and monitor SLOs, build alerting; participate in on-call rotation
- Contribute to infrastructure cost visibility and optimization (query cost estimation, workload right-sizing, storage tier routing)
- Collaborate with engineers, product managers, and designers on consistency and standards
Required: 5+ years building platform infrastructure, backend services, or data systems in production. Strong Python proficiency and REST API design experience. Deep background in distributed systems and cloud platforms (AWS preferred)—hands-on with S3, RDS, EKS, EventBridge, IAM, Terraform. Familiarity with data orchestration tools (Prefect, Airflow, Dagster) and transformation frameworks (dbt). Understanding of data governance (RBAC, PII, audit logging, lineage). Track record leading complex initiatives, influencing stakeholders, delivering measurable impact. Fast-paced environment with strong technical communication.
Nice-to-have: shared libraries/SDKs, event-driven architectures (CDC, event buses, schema registries), OpenTelemetry/ClickHouse observability, regulated environments (SOC 2, FedRAMP, HIPAA), Ray or distributed compute frameworks.