SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
fomo is a trading app designed for accessibility and social features. The platform allows users to sign up instantly and access on-chain assets without external wallets or bridges, while following top traders and discovering tokens early. The company combines social trading features with best-in-class execution and data for experienced traders.
As a Staff Distributed Systems Engineer, you will own the reliability, scalability, and performance of fomo's multi-region backend platform. This is a hands-on role with direct ownership of production systems. You will design and operate critical shared infrastructure including datastores, caches, messaging systems, and regional application services.
Key responsibilities include designing high-throughput, multi-region services that remain predictable during traffic surges, dependency failures, infrastructure changes, and partial regional outages. You will improve datastore and cache performance across replication, failure handling, and capacity planning. You'll implement resilience patterns such as backpressure, concurrency limits, load shedding, rate limiting, circuit breakers, and bounded retries.
You will establish new failover and disaster-recovery capabilities, defining recovery objectives and implementing systems and testing required to safely recover services and data. Additional responsibilities include reducing cross-region latency, improving data locality, designing and testing service failover procedures, and helping architect new features to operate at scale from day one. You will also mentor the team on distributed systems thinking and scalability principles.
Required qualifications include 8+ years of backend, platform, or infrastructure engineering experience. You must have strong experience designing and debugging distributed, high-throughput production systems, with deep PostgreSQL expertise (query performance, indexing, connection pooling, replication, transaction contention, failure modes). Strong experience with Redis-compatible systems (Redis, Valkey, Dragonfly, KeyDB) including sharding, replication, memory management, and failure handling is essential. AWS operations experience (ECS, RDS, ElastiCache), infrastructure-as-code proficiency (Terraform), and systems-oriented language skills (Go, TypeScript/Node.js) are required. Hands-on experience designing and testing failover and disaster-recovery systems is critical.
Nice-to-have skills include NATS JetStream or Kafka experience, Datadog APM and AWS Performance Insights familiarity, live datastore/cache topology migration experience, and background with high-throughput systems in financial, trading, cryptocurrency, or gaming sectors.