SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
WRITER is an enterprise generative AI platform helping leading companies like Mars, Marriott, Uber, and Vanguard build and deploy AI agents grounded in their data. The company is valued at $1.9B and backed by top-tier investors including Premji Invest, Radical Ventures, and ICONIQ Growth.
As an Infrastructure Engineer, you will be responsible for ensuring WRITER's platform remains available, performant, and reliable 24/7 for hundreds of enterprise customers. You'll work at the intersection of SRE, DevOps, infrastructure, and platform engineering, focusing on building resilient systems, automating across the stack, and championing reliability best practices.
Key responsibilities include:
- Design and maintain scalable, fault-tolerant infrastructure across AWS (preferred), GCP, and Azure, working fluently with Kubernetes, Helm, Terraform, and AI tooling
- Lead incident response, post-mortems, and root-cause analyses; trace failures to underlying problems and prevent recurrence
- Own reliability, performance, and efficiency of core services end-to-end; define and uphold SLOs and error budgets
- Balance critical operational work with 6–12-month platform direction; ship urgent fixes while shaping multi-year observability, cost, and reliability investments
- Automate operational tasks and infrastructure management using Python or Go; treat manual on-call work as a defect to be designed out
- Integrate AI agents (Claude Code, Droid, Codex, internal skills) into your daily workflow to investigate incidents, draft infrastructure changes, write runbooks, and review PRs
- Collaborate cross-functionally with product, security, and engineering peers to provide expert guidance on system design for reliability, performance, and scalability
- Challenge the status quo, proactively identify systemic weaknesses, and propose innovative solutions to complex reliability problems
You'll report to the Director of Engineering and work from either the London or New York City hub on a hybrid basis.
REQUIREMENTS:
- 5+ years of experience in infrastructure engineering, DevOps, or similar role focused on building and operating large-scale, high-availability production systems at a high-growth product company
- Experience running containerization in production (real cluster, not lab) with hands-on experience in Helm and Terraform or Pulumi on at least one major cloud provider (AWS preferred)
- Strong proficiency in Python or Go for automation and tooling
- Daily use of AI tooling (Claude Code, Droid, Codex, internal skills) in your workflow; agentic tooling must be part of how you ship, not something you've read about. Candidates without actual daily AI-assisted workflows will not be advanced.
- Demonstrated ability to challenge the status quo, proactively identify systemic weaknesses, and propose innovative solutions; reason from constraints and failure modes rather than analogy or vendor defaults
- Fluency with monitoring and logging stacks (Prometheus, Grafana, ELK or equivalent)
- Strong sense of ownership and accountability; at least one 0-to-1 infrastructure build you owned end-to-end with outcome metrics attached
- Excellent communication, collaboration, and problem-solving skills; ability to build strong relationships and connect with cross-functional teams
BONUS:
- Software-engineering background with experience designing, building, and shipping non-trivial production code (services, libraries, internal frameworks) in Python, Go, or comparable language
- Ability to read and modify codebases your infrastructure runs; comfort moving between infra automation and feature engineering
About Writer
AI / Data / Infrastructure; SaaS / Enterprise Software — enterprise generative AI platform for business teams.