SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Cortea is building intelligent systems for a high-stakes domain where accuracy matters. You will join as a Staff Platform Engineer responsible for making the company's rapidly scaling AI agent infrastructure explicit, reliable, and maintainable.
The role sits at the intersection of software engineering, DevOps, and SRE. Your customers are product engineers. This is not a DevOps role: you will spend more time in application code than infrastructure configuration. The core thesis is that you cannot design systems well without being able to run them.
You will own four interconnected areas:
**Product Platform**: Design and build the paved road—the easiest, most secure, reliable, and scalable way for product engineers to ship features. Own shared primitives, libraries, and APIs that hide complexity while carrying quality and observability standards. Define the core stack and architecture roadmap.
**Infrastructure, Observability & Reliability**: Implement infrastructure-as-code across cloud resources, dashboards, and alerts. Provision and tune the Kubernetes cluster. Define SLIs for critical workloads, build SLOs, and create high-signal alerting. Optimize AI and infrastructure spend.
**Security & Compliance**: Secure the internal foundations by default—IAM, dependency management, secrets management. Own the authentication and authorization stack, including ReBAC models for both humans and agents. Implement SOC 2 and ISO 27001 controls without creating friction for future changes.
**AI Dev Tooling**: Keep product engineers and their agents on the paved road through documentation, agent guidelines, and standardized automations. Shorten the path from design to implementation with development environments that mirror production.
Practical examples of work include: building a system of record for agent configurations with auditability and tenant isolation; creating harnesses that abstract over rapidly changing use cases; centralized progress tracking across parallel agent executions; usage-based tenant billing with quotas and rate limiting; SLIs for background job systems (execution latency, AI spend, memory, CPU); and a document pipeline handling dozens of file formats and hundreds of parallel uploads under strict reliability requirements.
The team ships fast and intends to keep shipping fast. You will use AI heavily as a tool but retain full judgment—the mental model must live in your head, not in a context window. You will design, reason, and write the decisions yourself, and be able to defend every choice you ship.
You will work face-to-face with experienced founders, learn directly from customer insights, and have real influence on strategy in a small team of excellent engineers with high autonomy. Decisions happen in hours, not weeks. The mission is building intelligent systems for a $200bn industry from Berlin, with AI at the core.
**Requirements:**
- Designed, built, and operated distributed systems end to end; understand how every part interacts
- Strong backend software engineering AND DevOps/SRE skills; do not view these as separate disciplines
- Write design docs, RFCs, ADRs, postmortems
- Prefer simple, boring solutions; distinguish essential from accidental complexity; know which corners are safe to cut
- Can reframe hard problems and find unconventional solutions
- Concrete experience: scaled Postgres or another relational transactional database under real load; run production Kubernetes; built observability from scratch (SLIs, SLOs, distributed tracing)
- If you have designed a system from scratch, run it in production, and can explain in writing why it is shaped the way it is, you are a strong candidate
**Nice-to-haves:**
- Experience with durable workflow orchestrators like Temporal, background job and queue processing
- Infrastructure for LLM-based products or agentic systems
- Background in audit, finance, compliance, or another high-accuracy domain
- Azure and/or GCP experience
**Not a fit if:**
- You view application code as someone else's responsibility (platform work here happens inside the product codebase)
- Your approach to reliability is primarily process-based; Cortea wants guardrails, not gates
- You need a well-defined system to work in; this system is still being shaped, and that shaping is the job