SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Wealthsimple, Canada's leading financial innovator serving 4+ million Canadians with $155B+ in assets under administration, is seeking an experienced Senior Software Engineer to join the ML Platform & Infrastructure team.
The ML Platform & Infrastructure team builds foundational architecture powering AI and GenAI initiatives across Wealthsimple, sitting at the intersection of production MLOps and cutting-edge GenAI enablement. As the company's AI footprint expands rapidly, the priority is evolving robust MLOps foundations into a scalable, high-performance LLM serving and routing platform. The team builds self-serve systems enabling Data Scientists and Engineers to host open-source LLMs reliably, optimize inference latencies, manage GPU infrastructure, and benchmark model performance safely in production.
In this role, you will bridge traditional MLOps (model lifecycles, pipeline orchestration, serving infrastructure) and modern GenAI stack requirements (vLLM, GPU cluster management, intelligent model routing, automated Evals). You will take end-to-end ownership of setting technical direction for operating enterprise-grade LLM systems company-wide.
Key responsibilities include: transitioning traditional ML lifecycle and serving patterns into state-of-the-art LLM inference engines and GPU orchestration systems; designing low-latency routing frameworks (e.g., LiteLLM integration) to dynamically direct requests across managed cloud providers (AWS Bedrock) and self-hosted open-source models; architecting and managing high-performance GPU serving environments on Kubernetes using vLLM, Ray, and Triton; building automated Evals and observability frameworks to empower engineers and data scientists to validate model quality, latency, and drift; partnering with product engineering and data science teams to build framework-agnostic platform tooling; and driving cost and performance optimization across self-hosted and managed inference.
The team is hybrid with 1,500+ employees across North America. Benefits include top-tier health and life insurance, employer-matched group savings, 20 vacation days plus 4 wellness days and unlimited sick/mental health days annually, 90-day work-from-anywhere program, and employee resource groups.
REQUIREMENTS:
- 7+ years of software engineering experience in ML Infrastructure, MLOps, ML Tooling, or Data Platform engineering
- Deep experience in MLOps/ML Platform practices: proven track record building and operating self-serve ML platforms, model registry workflows, experiment tracking, or production serving infrastructure (Kubeflow, MLflow, Ray, Triton, SageMaker)
- Strong platform fundamentals: advanced proficiency in Python, container orchestration via Kubernetes, infrastructure-as-code (Terraform), and cloud provider ecosystem (AWS)
- Strong appetite to specialize in LLM serving: genuine desire to leverage existing MLOps skillset to tackle LLM-specific challenges (vLLM, model routing, prompt engineering tooling, vector databases, GPU memory optimization, LLM evaluation frameworks)
- Backend performance and observability focus: experience designing highly available, observable microservices (e.g., FastAPI) handling real-time, low-latency requests
- End-to-end technical ownership: proven capability to lead architectural roadmaps, guide multi-functional projects with high autonomy, and maintain complex platform systems
NICE-TO-HAVES:
- Direct experience serving open-source Large Language Models in production (vLLM, SGLang, TensorRT-LLM, Dynamo)
- Hands-on work with CUDA, GPU partitioning, or distributed inference frameworks (Ray Serve)
- Familiarity with vector search and retrieval engines (Elasticsearch, Qdrant, Pinecone)