SlipstreamJobsFresh Startup & VC-Backed Jobs

Senior Platform Engineer

Earnin - Mountain View, CA, United States - Hybrid - posted 2026-09-15

Apply on the company site

SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.

Salary: USD 228,000 - 279,000 / annual

EarnIn is a fintech pioneer in earned wage access, enabling workers to access their earnings in real-time without mandatory fees, interest rates, or credit checks. The Platform-as-a-Service Engineering team builds the foundational infrastructure and "Paved Road" that enables product teams to ship with velocity and safety. You will design, build, and operate agentic systems and core infrastructure powering EarnIn's platform. Your work spans Kubernetes (AWS EKS), GitOps (Argo CD), CI/CD (GitHub Actions), and the developer portal (Cortex). You'll build AI agents and automation that reduce operational toil and increase platform leverage for engineering teams across the company. As a Senior engineer, you independently design and implement solutions for well-scoped but complex problems, re-architect components, adopt new technologies, and improve existing processes with minimal guidance. You'll build agents that handle incident triage, environment provisioning, config generation, and developer self-service, partnering with Staff engineers and your manager to align work with the team's broader technical direction. Key responsibilities include: - Design, build, and operate AI agents in production following established governance patterns for model selection, evaluation, and safety. Contribute to infrastructure-as-code practices for agentic systems, ensuring prompts, tools, and evaluation criteria are versioned, reviewed, and tested. - Implement agentic patterns for cloud infrastructure in your area of ownership, applying and refining team best practices. Support and mentor less experienced engineers on agentic patterns, LLM integration, and prompt engineering. - Operate and improve high-availability distributed systems on AWS, identifying and resolving performance, scalability, and stability issues. Use AI-driven observability and anomaly detection to catch problems early. - Contribute to the evolution of the developer control plane (Cortex) as an AI-augmented self-service platform, building features that let engineers scaffold services, debug deployments, and resolve issues through natural language. - Document agentic architecture, best practices, and operational procedures. Participate in on-call rotations and use post-mortems as feedback loops to improve system reliability and agentic automation. Requirements: - Bachelor's or Master's degree in Computer Science, Engineering, or related field - 4+ years in cloud infrastructure, working on large-scale, high-availability, customer-facing distributed systems - Experience mentoring junior engineers and leading well-scoped platform initiatives - Hands-on experience building and operating AI-driven systems in production, including agentic workflows that perform operational tasks; ability to speak to specific agent designs or automation built - Experience reducing operational toil through agentic AI (e.g., LLM-powered incident diagnosis, intelligent CI/CD with test selection, self-service assistants) on at least one meaningful workflow - Strong working knowledge of AWS (EKS, Lambda, Bedrock, etc.) and experience with containerized and serverless architectures - Solid expertise in Kubernetes at scale and ability to implement complex, resilient solutions - Strong knowledge of infrastructure-as-code tools (Terraform, Ansible) and experience applying both traditional IaC and agentic automation - Strong command of Datadog and observability practices, using metrics to drive decisions - Strong adherence to security, privacy, and compliance best practices, with awareness of governance considerations for production AI systems (model safety, prompt injection prevention, data isolation) - Experience with LLM orchestration frameworks (LangChain, LlamaIndex, CrewAI, or custom agentic architectures) and production prompt engineering - Strong coding ability in Python and/or Go, with experience treating infrastructure and agentic systems as software - Ability to collaborate effectively across engineering, product, and security teams Plus (preferred): - Experience with service mesh (Linkerd, Istio) and traffic management at scale - Proficiency with GitOps (Argo CD, Flux CD) and CI/CD orchestration (GitHub Actions, Argo Workflows) - Experience with MLOps or LLMOps concepts (model versioning, evaluation frameworks, production monitoring for AI systems) - Familiarity with security frameworks relevant to AI systems (guardrails, audit logging, data governance for LLMs)

Similar roles