SlipstreamJobsFresh Startup & VC-Backed Jobs

SDE 2 Infra

Fam - Bengaluru, India - In-office - posted 2026-10-01

Apply on the company site

SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.

Fam (formerly FamPay) is India's first payments app for users aged 11+, enabling UPI and card payments at scale. The company is backed by top-tier investors including Y Combinator, Elevation Capital, and Peak XV (Sequoia Capital India). The Core Infrastructure team builds the foundational platforms and systems powering Fam's entire product ecosystem. This is a high-leverage, small team with deep ownership. The infrastructure currently handles 1 billion+ API requests per day, ~10M daily transactions, 10M+ users, and 20,000+ RPS across a 470-node Kubernetes cluster (~2,500 vCPUs, ~8 TiB RAM). The team manages 12+ TB of active transactional data, 400+ event-streaming topics with 3,800+ partitions at ~46,700 messages/sec, and dynamic multi-dimensional auto-scaling via 195+ HPAs and Karpenter. As an SDE 2 in Infrastructure, you will: - Build and operate compute, storage, and networking infrastructure at scale powering the FamApp platform - Design platform contracts for internal teams: modules, golden paths, and abstractions with clear inputs, outputs, and guarantees - Own medium-to-large infrastructure projects from design through rollout, manage cross-team dependencies, and work through ambiguity independently - Build idempotent, retry-safe automation with integrated observability, CI/CD, ticketing, compliance, and incident workflows - Develop reusable infrastructure using Infrastructure as Code and Configuration as Code with clear interfaces, validation, versioning, and safe provisioning CI/CD - Operate Kubernetes as the operating system: own cluster lifecycle (worker plane, node groups, add-ons), extend with controllers/operators/CRDs so teams consume platform capabilities as Kubernetes-native APIs - Define user-facing SLIs, SLO targets, burn-rate alerts, and error-budget actions; carry the pager, reduce alert noise, and lead blameless post-incident follow-ups - Track cloud spend and identify savings opportunities without compromising reliability or compliance - Maintain golden pipeline templates, GitOps release patterns, and self-service onboarding with security gates and least-privilege pipeline identities - Author technical design docs covering full system lifecycle: adoption, operations, maintenance, reliability (monitoring, SLIs, backup, DR, incident response) The team is moving toward a vision where internal engineering teams consume platform primitives (compute, storage, event streams) via self-serve contracts, and building AI-native platform operations with agent planes and agent skills for safe infrastructure interaction. The team follows CNCF best practices and contributes to open source. REQUIREMENTS: Must Have: - 3–6 years in DevOps, SRE, or platform engineering with ownership of production systems - Cloud compute and Linux fundamentals: instances, images, block storage on major clouds; ability to diagnose throughput/IOPS limits, memory pressure, inode and file-descriptor exhaustion, systemd issues, and correlate OS-level evidence with cloud metrics - Cloud networking: network and subnet design, CIDR planning, routing, firewall and security-group rules, DNS; ability to design cross-account/cross-project connectivity, weigh peering vs. transit hubs vs. private endpoints, debug asymmetric routing and overlapping CIDRs - Hands-on Kubernetes in production managed services (EKS, GKE, or AKS) - Terraform or equivalent IaC with reusable, versioned modules and safe change workflows - Programming ability in Python or Go to build idempotent, retry-safe, observable automation - CI/CD and GitOps experience: pipeline templates, release and rollback patterns, security gates - Reliability practice: SLIs/SLOs, actionable alerting, on-call, blameless post-mortems - Everyday use of AI tools with habit of validating output before shipping Good to Have: - Karpenter in production (consolidation, spot handling, multi-NodePool strategy) - Observability at scale: Prometheus/VictoriaMetrics, Loki, Jaeger, OpenTelemetry, high-cardinality metrics - CI/CD supply-chain security (SAST, secret detection, image scanning, artifact promotion) - ClickHouse Operator on Kubernetes or other analytical store solutions managed at scale in production - PCI, RBI, SOC 2, or similar compliance-heavy environments - Experience building, scaling, and managing infrastructure for customer-focused products

Similar roles