SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Nearmap is seeking a Principal Cloud Platform Engineer to architect and lead the ML infrastructure platform that powers the company's AI innovation. You will be the technical owner and key architect of the platform serving Nearmap's AI Computer Vision (AICV) organization, reporting to the Director of AICV Platform Engineering.
You will lead a team of 3–4 mid-level and senior ML systems engineers and set the technical vision for the ML infrastructure ecosystem. Nearmap processes aerial imagery at massive scale—thousands of EKS nodes running batch inference, distributed GPU training across AWS and GCP, real-time model endpoints, and LLM/agentic systems moving from prototype to production. Your role is to define and evolve the platform that enables all of this.
This is a platform engineering role, not an application role. Your customers are the AICV product teams (AI Model R&D, Computer Vision, Insurance Data Science, Agentic AI, AI Map Data), and your objective is to build a robust, scalable, efficient ecosystem that acts as a force multiplier for the entire AI organization.
Day-to-day responsibilities include:
- Define and own the technical roadmap for core ML infrastructure: workflow orchestration on Kubernetes, distributed training, batch inference at thousands-of-nodes scale, real-time serving on Ray, and LLMOps
- Make critical design decisions and evaluate new technologies (orchestrators, serving frameworks, vector databases)
- Lead and mentor a team of 3–4 mid-level and senior ML systems engineers; own technical hiring for the platform team
- Spearhead the highest-risk work: multi-cloud GPU capacity strategy, foundational platform components, and the observability stack
- Write production Python for the hardest parts of the shared platform and prototype new capabilities before team commitment
- Establish and champion MLOps and AIOps best practices across the organization (automation, infrastructure as code, CI/CD, security)
- Partner with Data Science and ML Engineering teams to translate challenges into actionable platform roadmap
- Own service level objectives and GPU cost efficiency across AWS and GCP; lead major incident response; build detection for silent production ML failures
Requirements:
- 10+ years in software engineering, with at least 4 years focused on building and operating large-scale infrastructure, platform engineering, or distributed systems in production
- Proven experience leading technical projects and mentoring or managing a team of engineers
- Deep, production-level expertise designing, building, and operating systems on Kubernetes, including GPU scheduling
- Expert-level Python and distributed systems judgment to reason clearly about failure, backpressure, and cost at scale
- Hands-on experience building and managing cloud infrastructure on AWS and GCP with infrastructure as code (Terraform, Pulumi, or similar)
- Strong, practical grounding in modern software development: Git, CI/CD, monitoring, alerting, automated testing
- Bachelor's or Master's degree in computer science, engineering, or related technical field, or equivalent practical experience
Highly desirable:
- Workflow orchestration frameworks (Argo Workflows, Kubeflow Pipelines, Flyte, Ray/KubeRay)
- Designing and managing large-scale, GPU-intensive ML workloads and GPU capacity strategy across multiple clouds
- Deep familiarity with MLOps and LLMOps ecosystem (model registries, feature stores, serving frameworks, vector databases, RAG systems)
- Advanced Kubernetes concepts (custom operators, multi-cluster networking)
- Petabyte-scale data pipelines or geospatial/computer vision workloads
- Track record of contributing to open-source projects in cloud-native or MLOps space