SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Cognite is building industrial digitalization and AI solutions at scale. This role focuses on designing, building, and operating the core serverless execution engine and Workflows orchestration layer that power CDF's AI and automation capabilities.
You will own platform reliability, ensuring Functions execute deterministically and Workflows progress without data loss or silent failures. You'll architect for multi-tenant, multi-cloud, high-throughput workloads, designing scheduling, queueing, and retry mechanisms that degrade gracefully under pressure. Your responsibilities include defining clean, API-first architecture with versioned REST and event-driven APIs that downstream teams and external customers depend on.
Observability is critical: you'll instrument services with distributed tracing, structured logging, and alerting using the Open-telemetry / Prometheus / Grafana / Honeycomb stack so failures surface before customers notice. You'll champion test automation—unit, integration, and smoke tests—and maintain deployment pipelines that ship to production with confidence. Performance profiling and optimization of execution throughput, cold-start latencies, and cross-service call chains will drive a snappy platform experience for industrial workloads. As a bonus, you may model compute and storage costs and identify optimizations that reduce cloud spend without sacrificing reliability.
Cognite operates globally with presence in Phoenix, Houston, Oslo, Tokyo, Bengaluru, and Abu Dhabi, and is pursuing a bold Moonshot: unlock $100B in customer value by 2035 and redefine how global industry works.
**Requirements:**
- 3–6 years of engineering experience with a proven track record building and operating production backend services at scale
- Deep mastery of JVM languages (Kotlin preferred, Java acceptable) and Python (FastAPI)
- Strong expertise in distributed systems patterns and cloud-native service design (Kubernetes, Azure, GCP, AWS, Private cloud)
- Hands-on experience with workflow engines (Conductor, Apache Airflow, or equivalent) and event-driven architectures (Kafka, Pub/Sub)
- Comfortable working with relational databases (PostgreSQL), non-relational databases, object storage (Data-lakes), and caching layers (Redis) in multi-tenant environments
- Practical experience with Open-telemetry, Prometheus, and Grafana for instrumentation and operational insight
- Experience supporting ML workloads and notebooks in production (job scheduling, resource management, experiment tracking integration, or model serving infrastructure)
**Nice to have:**
- Familiarity with industrial knowledge graph construction, entity resolution, or NLP/CV pipelines as they relate to industrial asset data
- Familiarity with React or TypeScript for consuming and dogfooding platform developer tooling
- ML Workload Support: building platform primitives for compute scheduling, environment management, and secrets handling
- Contextualisation Pipelines: engineering infrastructure for entity matching, asset hierarchy inference, P&ID parsing
- Vector and embedding infrastructure experience for semantic search and RAG-based contextualisation
- Model lifecycle awareness: versioning, A/B experiment tracking, and understanding the boundary between platform and ML framework concerns