SlipstreamJobsFresh Startup & VC-Backed Jobs

Technical Product Manager, Observability – remote in the US

Mirantis - Remote - Remote - posted 2026-09-15

Apply on the company site

SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.

Mirantis is seeking a Technical Product Manager to own the observability strategy and roadmap for k0rdent AI, a control plane for GPU infrastructure and distributed AI workloads. In this role, you will define how operators gain visibility into the health, performance, and resource utilization of GPU clusters running large-scale training and inference workloads. You will shape observability across the full stack: GPU compute, east-west fabric (InfiniBand, RoCE), high-performance storage, DPU/SmartNIC telemetry, workload schedulers, inference serving, and data services. The observability plane will be powered by the OpenTelemetry ecosystem and Prometheus-compatible metrics pipelines. Key responsibilities include: - Own the vision, roadmap, and priorities for k0rdent AI observability across GPU compute, networking fabric, storage, and AI workload layers - Translate requirements from NeoClouds, GPU clouds, telcos, sovereign clouds, and enterprise platform teams into clear product direction - Partner with engineering to define requirements, evaluate trade-offs, and manage the observability backlog - Track and shape response to emerging observability standards and technologies (OpenTelemetry, DCGM GPU metrics, fabric counters, storage telemetry APIs, AI workload profiling) - Define integration strategies for vendor telemetry sources (NVIDIA, BlueField DPUs, VAST, Weka, DDN, SLURM, inference stacks, vector/relational databases) into a unified observability plane - Partner with product marketing and field teams on positioning, technical briefs, and reference architectures - Represent Mirantis with customers, analysts, and ecosystem partners You will work directly with engineering on technical requirements, with marketing on positioning, and with customers to ensure their success. REQUIREMENTS: - 5+ years in product management or a senior technical role owning an observability product or operating large-scale monitoring infrastructure - Working knowledge of Prometheus, OpenTelemetry, distributed tracing (Jaeger, Tempo), and log aggregation (Loki, Elasticsearch/OpenSearch) - Fluency in Kubernetes observability, cloud-native monitoring, or metrics and alerting pipeline architecture - Ability to work directly with engineering on technical trade-offs and with field teams in competitive GPU cloud and NeoCloud deals STRONGLY PREFERRED: - Exposure to GPU observability, including DCGM metrics, AI workload profiling and performance analysis - Familiarity with east-west fabric telemetry (InfiniBand counters, RoCEv2 congestion metrics, switch-level fabric health) - Experience with high-performance storage telemetry from VAST Data, Weka, or DDN - Familiarity with NVIDIA BlueField DPU telemetry, SR-IOV, or offload pipeline observability - Exposure to workload-level visibility for SLURM job scheduling, inference serving stacks (vLLM, Triton, TensorRT-LLM), or data service telemetry from vector/relational databases in AI pipelines

Similar roles