SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Airwallex is an AI-native financial operating system serving over 676,000 businesses globally, including McLaren Racing, Qantas, SHEIN, and TikTok. The company provides regulated financial infrastructure spanning North America, Europe, the Middle East, and Asia-Pacific, with headquarters in San Francisco and Singapore and 2,300+ employees across 27 offices.
You will design and operate distributed systems that power Airwallex's data and AI platforms. This role focuses on building highly available infrastructure for high-throughput data processing, real-time workloads, and production AI applications. Key responsibilities include evolving Kubernetes and cloud foundations, improving reliability and scalability of platforms like Kafka, Spark, and Flink, and building infrastructure for AI traffic management and model serving.
You will work across application, data, machine learning, security, and infrastructure teams to establish technical foundations for the company's next growth phase. Specific deliverables include designing and operating highly available data and AI infrastructure on Kubernetes and public cloud platforms; developing scalable platforms for streaming, batch processing, and real-time data workloads; building and evolving AI infrastructure including gateways, model-routing layers, traffic management, rate limiting, authentication, observability, and usage controls; developing self-service capabilities for data, AI, and application teams; and partnering with engineering teams to translate emerging requirements into durable platform capabilities.
This is a high-impact role for an engineer who enjoys solving complex infrastructure problems, writing production software, and providing reliable self-service platforms to other engineering teams.
REQUIREMENTS:
- 5+ years in DevOps, SRE, or platform engineering, owning production systems end to end
- Strong experience designing, operating, and troubleshooting production Kubernetes environments
- Experience building or operating distributed data infrastructure with Kafka, Spark, or Flink OR experience developing AI infrastructure such as AI gateways or model-routing platforms
- Hands-on experience with at least one major public cloud platform (AWS, Google Cloud, or Microsoft Azure)
- Strong knowledge of cloud and container networking (DNS, load balancing, ingress, service discovery, TLS, routing, network security)
- Proficiency in one or more of Go, Python, or Java, with experience writing maintainable production software
- Solid understanding of distributed-systems concepts (availability, consistency, fault tolerance, backpressure, horizontal scalability)
- Experience operating critical infrastructure using infrastructure-as-code, automated delivery, and modern observability practices
- Strong debugging skills and ability to work methodically across multiple system layers
- Clear communication skills and track record of collaborating effectively across engineering disciplines
- Ownership mindset: identifying important problems, driving them to resolution, and improving underlying systems
Especially valuable: platform engineering experience building internal developer platforms; hands-on experience with serving technologies (SGLang, vLLM, NVIDIA Triton Inference Server); knowledge of GPU scheduling, batching, model parallelism, memory management, autoscaling, and inference-performance optimization; experience improving cost efficiency of large-scale data processing or AI inference workloads; contributions to infrastructure, data-platform, Kubernetes, or AI-serving open-source projects.