SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Revolut is a global fintech company with 80+ million customers, building a financial super app that includes spending, saving, investing, exchanging, and travel products. The company has 13,000+ employees worldwide and is certified as a Great Place to Work.
You will join the AI team within the Technology organization, working on the core infrastructure that powers Revolut's machine learning and AI capabilities. This role focuses on building and optimizing the MLOps platform that enables scientists and engineers to develop and deploy AI solutions at scale.
Key Responsibilities:
- Design, build, and maintain scalable backend services and tooling supporting the AI lifecycle and GPU-based workloads
- Profile and optimize model training and inference workloads (latency, throughput, stability) using PyTorch and CUDA-enabled libraries
- Improve GPU compute utilization, memory consumption, and bandwidth across shared, multi-tenant, and multi-node environments
- Implement mixed-precision execution, quantization, efficient batching, gradient accumulation, and memory-efficient attention mechanisms
- Create repeatable benchmarking practices and automated regression tests; guide engineering teams in selecting execution frameworks and GPU types
Requirements:
- Proven track record designing and operating scalable backend systems in production (Linux, Kubernetes, containerized workloads) using Python
- Experience running, profiling, and optimizing machine learning workloads built with PyTorch or equivalent on NVIDIA GPUs
- Solid understanding of GPU execution concepts: memory hierarchy, CPU/GPU synchronization, host-to-device transfers, CUDA stream behavior
- Familiarity with multi-GPU communication (NCCL), data/tensor parallelism, and resource allocation in Kubernetes
- Degree in STEM subject (or equivalent) with strong foundation in computer science principles, testing, observability, and operational ownership