SlipstreamJobsFresh Startup & VC-Backed Jobs

Infrastructure engineer

Writer - New York, NY, United States - Hybrid - posted 2026-09-22

Apply on the company site

SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.

WRITER is an enterprise generative AI platform trusted by leading companies like Mars, Marriott, Uber, and Vanguard to build and deploy AI agents grounded in company data. The company is valued at $1.9B and backed by top-tier investors including Premji Invest, Radical Ventures, and ICONIQ Growth. As an Infrastructure Engineer, you will be responsible for ensuring WRITER's platform remains available, performant, and reliable 24/7 for hundreds of enterprise customers. You'll work at the intersection of SRE, DevOps, infrastructure, and platform engineering, focusing on building resilient systems, automating across the stack, and championing reliability best practices. Key responsibilities include: - Design and operate scalable, fault-tolerant infrastructure across AWS (preferred), GCP, and Azure, working fluently with Kubernetes, Helm, Terraform, and AI tooling that backs WRITER's high-traffic platform. - Lead incident response, post-mortems, and root-cause analyses—trace failures to underlying problems and prevent recurrence through architectural improvements. - Own reliability, performance, and efficiency of core services end-to-end, defining and upholding SLOs and error budgets while carrying the on-call pager. - Balance critical tactical work with 6–12-month platform direction, shipping on-call-driving fixes while shaping multi-year observability, cost, and reliability investments. - Integrate AI agents (Claude Code, Droid, Codex, internal skills) into your daily workflow to investigate incidents, draft infrastructure changes, write runbooks, and review PRs. Build agentic setups where humans and digital teammates work as one team with shared skills and context. - Challenge the status quo, remove toil before adding features, and automate operational tasks with Python or Go. Treat manual on-call work as a defect to be designed out. - Collaborate cross-functionally with product, security, and engineering peers, providing expert guidance on system design for reliability, performance, and scalability from conception through launch. - Demonstrate breadth across disciplines—bring deep focus to one substantial initiative at a time (on-call posture, release pipeline, multi-region Terraform layout, internal platform surface) while maintaining cross-layer fluency to identify the right next initiative. This is a hybrid role based out of WRITER's New York City, San Francisco, Seattle, or London hubs, reporting to the Director of Engineering. The company is open to hiring at multiple levels with compensation scaled to experience and expertise. REQUIREMENTS: - 5+ years of experience in infrastructure engineering, DevOps, or similar role focused on building and operating large-scale, high-availability production systems at a high-growth product company. - Track record running containerization in production (real cluster, not lab) with hands-on experience in Helm and Terraform or Pulumi on at least one major cloud provider (AWS preferred). - Strong proficiency in Python or Go for automation and tooling. - AI is a hard requirement: agentic tooling (Claude Code, Droid, Codex, internal skills) must be in your daily workflow already. You've built or adopted AI-assisted workflows others now use and have strong opinions on where it's unreliable. Candidates whose actual daily workflow does not already include AI tooling will not be advanced. - Demonstrated ability to challenge the status quo, proactively identify systemic weaknesses, and propose innovative solutions to complex reliability problems. Reason from constraints and failure modes (not analogy or vendor defaults), name tradeoffs in business terms, and reject "best practices" answers when they don't fit the problem. - Make reversible calls by default—write the rollback before touching production. Work fluently with monitoring and logging stacks (Prometheus, Grafana, ELK or equivalent). - Excellent communication, collaboration, and problem-solving skills with talent for building strong relationships and connecting with cross-functional teams. - Strong sense of ownership and accountability. Own at least one 0-to-1 infrastructure build end-to-end with the outcome metric attached. BONUS: - Software-engineering background beyond config and scripting—you've designed, built, and shipped non-trivial production code (services, libraries, internal frameworks) in Python, Go, or comparable language. - Ability to read and modify codebases your infrastructure runs and move between infra automation and feature engineering seamlessly.

About Writer

AI / Data / Infrastructure; SaaS / Enterprise Software — enterprise generative AI platform for business teams.

Similar roles