SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Fam (formerly FamPay) is India's first payments app for users aged 11+, enabling UPI and card payments at scale. The company is backed by top-tier investors including Y Combinator, Elevation Capital, and Peak XV (Sequoia Capital India).
The Core Infrastructure team builds the foundational platforms and systems powering Fam's entire product ecosystem. This is a high-leverage, small team with deep ownership. The infrastructure currently handles 1 billion+ API requests per day, ~10M daily transactions, 10M+ users, and 20,000+ RPS across a 470-node Kubernetes cluster (~2,500 vCPUs, ~8 TiB RAM). The team manages 12+ TB of active transactional data, 400+ event-streaming topics with 3,800+ partitions at ~46,700 messages/sec, and dynamic multi-dimensional auto-scaling via 195+ HPAs and Karpenter.
As an SDE 2 in Infrastructure, you will:
- Build and operate compute, storage, and networking infrastructure at scale powering the FamApp platform
- Design platform contracts for internal teams: modules, golden paths, and abstractions with clear inputs, outputs, and guarantees
- Own medium-to-large infrastructure projects from design through rollout, manage cross-team dependencies, and work through ambiguity independently
- Build idempotent, retry-safe automation with integrated observability, CI/CD, ticketing, compliance, and incident workflows
- Develop reusable infrastructure using Infrastructure as Code and Configuration as Code with clear interfaces, validation, versioning, and safe provisioning CI/CD
- Operate Kubernetes as the operating system: own cluster lifecycle (worker plane, node groups, add-ons), extend with controllers/operators/CRDs so teams consume platform capabilities as Kubernetes-native APIs
- Define user-facing SLIs, SLO targets, burn-rate alerts, and error-budget actions; carry the pager, reduce alert noise, and lead blameless post-incident follow-ups
- Track cloud spend and identify savings opportunities without compromising reliability or compliance
- Maintain golden pipeline templates, GitOps release patterns, and self-service onboarding with security gates and least-privilege pipeline identities
- Author technical design docs covering full system lifecycle: adoption, operations, maintenance, reliability (monitoring, SLIs, backup, DR, incident response)
The team is moving toward a vision where internal engineering teams consume platform primitives (compute, storage, event streams) via self-serve contracts, and building AI-native platform operations with agent planes and agent skills for safe infrastructure interaction. The team follows CNCF best practices and contributes to open source.
REQUIREMENTS:
Must Have:
- 3–6 years in DevOps, SRE, or platform engineering with ownership of production systems
- Cloud compute and Linux fundamentals: instances, images, block storage on major clouds; ability to diagnose throughput/IOPS limits, memory pressure, inode and file-descriptor exhaustion, systemd issues, and correlate OS-level evidence with cloud metrics
- Cloud networking: network and subnet design, CIDR planning, routing, firewall and security-group rules, DNS; ability to design cross-account/cross-project connectivity, weigh peering vs. transit hubs vs. private endpoints, debug asymmetric routing and overlapping CIDRs
- Hands-on Kubernetes in production managed services (EKS, GKE, or AKS)
- Terraform or equivalent IaC with reusable, versioned modules and safe change workflows
- Programming ability in Python or Go to build idempotent, retry-safe, observable automation
- CI/CD and GitOps experience: pipeline templates, release and rollback patterns, security gates
- Reliability practice: SLIs/SLOs, actionable alerting, on-call, blameless post-mortems
- Everyday use of AI tools with habit of validating output before shipping
Good to Have:
- Karpenter in production (consolidation, spot handling, multi-NodePool strategy)
- Observability at scale: Prometheus/VictoriaMetrics, Loki, Jaeger, OpenTelemetry, high-cardinality metrics
- CI/CD supply-chain security (SAST, secret detection, image scanning, artifact promotion)
- ClickHouse Operator on Kubernetes or other analytical store solutions managed at scale in production
- PCI, RBI, SOC 2, or similar compliance-heavy environments
- Experience building, scaling, and managing infrastructure for customer-focused products