SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Revolut is a global fintech company on a mission to give people more from their money through spending, saving, investing, exchanging, and travel products. With 80+ million customers and 13,000+ employees worldwide, the company is scaling rapidly.
The Technology team builds the systems and infrastructure powering Revolut's innovative app and features. You'll join the AI team as a Python Engineer focused on MLOps infrastructure, working on the core platform that enables scientists and engineers to solve complex challenges.
In this role, you will:
- Design, build, and maintain scalable backend services and tooling supporting the AI lifecycle and GPU-based workloads
- Profile and optimize model training and inference workloads (latency, throughput, stability) using PyTorch and CUDA-enabled libraries
- Improve GPU compute utilization, memory consumption, and bandwidth across shared, multi-tenant, and multi-node environments
- Implement mixed-precision execution, quantization, efficient batching, gradient accumulation, and memory-efficient attention mechanisms
- Create repeatable benchmarking practices, automated regression tests, and guide engineering teams in selecting execution frameworks and GPU types
You'll work on heavily regulated financial systems, writing high-quality code and building automated solutions that power Revolut's infrastructure.
REQUIREMENTS:
- Proven track record designing and operating scalable backend systems in production (Linux, Kubernetes, containerized workloads) using Python
- Experience running, profiling, and optimizing machine learning workloads built with PyTorch or equivalent on NVIDIA GPUs
- Solid understanding of GPU execution concepts: memory hierarchy, CPU/GPU synchronization, host-to-device transfers, CUDA stream behavior
- Familiarity with multi-GPU communication (NCCL), data/tensor parallelism, and resource allocation in Kubernetes
- Degree in STEM subject (or equivalent) with strong foundation in computer science principles, testing, observability, and operational ownership