SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Checkbox is a Series A AI-native SaaS company building intelligent workflow automation for in-house legal teams. The company is transforming its platform into agentic-first products, powered by AI agents and intelligent automation. This Staff Data Engineer role reports directly to the VP of Engineering and sits within a tier-one Data team, responsible for owning Checkbox's data strategy and architecture as a competitive moat.
You will own the data and AI reference architecture underpinning the agentic product direction, including base data sources, intelligence services, context and retrieval layers, API and MCP surfaces, and secure data access patterns for AI agents. This is a hands-on role requiring deep involvement in design and delivery—not a distance-design position.
Key responsibilities include:
**Data and AI Reference Architecture**: Define how data sources, intelligence services, APIs, MCP surfaces, context engines and retrieval layers fit together. Establish architectural principles for storage, retrieval, serving, tenancy, observability and compliance. Partner with the Principal Engineer on architecture spanning application, platform, eventing, data and AI systems.
**Context, Retrieval and AI Data Foundations**: Architect and build context and retrieval layers powering AI agents and GenAI experiences. Design secure, tenant-isolated data access patterns. Define how customer, matter, workflow, document, event and operational data should be modelled and served. Evaluate and select appropriate data patterns (transactional stores, warehouses, vector stores, semantic layers, knowledge graphs, event streams, APIs). Make pragmatic buy-vs-build decisions.
**Data Platform, Integrations and Eventing**: Own architecture across integrations, eventing, operational and transactional data. Define data contracts, event models, ingestion patterns and transformation approaches. Partner with product engineering teams to ensure new capabilities generate usable, reliable, well-governed data. Build shared data capabilities supporting analytics, AI agents, workflow automation and customer-facing experiences.
**Security, Tenancy and Compliance**: Treat security and data isolation as first-order architectural concerns. Ensure data access patterns are compliant, auditable and appropriate for enterprise customers. Design systems supporting tenant isolation, data segregation, least privilege access and secure retrieval. Partner with Platform, Security and Engineering teams on encryption, access control and auditability.
**Team Leadership and Technical Direction**: Set technical direction while staying close to implementation. Partner with one other Staff Engineer in the Greenfield Data Platform, sharing load across Analytics, Warehousing, AI & ML. Mentor engineers on data architecture, retrieval patterns and pragmatic systems design. Build operating rhythms, standards and documentation. Act as senior technical voice for data architecture in engineering leadership discussions.
Success means product engineering streams can build AI features on dependable shared data layers; the reference architecture is documented and actively used; storage and retrieval choices fit problems rather than forcing uniformity; data access is compliant and secure by default; AI agents access relevant context reliably with clear controls; data contracts reduce duplication; the Data team becomes a strategic enabler rather than bottleneck; and Checkbox's data capability becomes a visible competitive moat.
**Requirements**:
- Significant experience as senior data engineer, principal data engineer, data architect, staff engineer or similar technical leadership role
- Proven experience building data and context foundations powering AI products in production
- Strong experience designing data architectures across transactional, operational, analytical and AI-ready systems
- Deep understanding of modern data platform patterns: data contracts, event-driven architecture, ingestion, transformation, observability, lineage and governance
- Strong understanding of AI data patterns: retrieval systems, embeddings, semantic modelling, vector search, knowledge graphs, context engineering, agentic workflows
- Experience with multi-tenant SaaS systems where data segregation, tenancy and access control matter
- Ability to make pragmatic architecture decisions and select appropriate tools and patterns
- Strong judgement on buy-vs-build decisions across data infrastructure, retrieval, orchestration and AI platform layers
- Comfortable leading technical direction while building, reviewing and shipping
- Experience mentoring engineers or leading small technical teams
- Strong communication skills and ability to work closely with product engineering, platform teams and senior leadership
- Comfortable in fast-moving environments where systems are built, scaled and refined simultaneously
**Bonus**: AI agents, agentic workflows, GenAI platforms, AI-native SaaS; MCP or API surfaces for data access; legal tech, workflow automation, enterprise SaaS or document-heavy products; AWS-based data infrastructure; event-driven systems, queues, Pub/Sub or streaming; data security, compliance and auditability; growing small data functions into scalable teams.