SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Salary: USD 215,000 - 260,000 / annual
Crusoe is a vertically integrated AI infrastructure company building cloud infrastructure for AI workloads. The Cloud Monitoring Services team owns observability across Crusoe Cloud, including metrics, logs, alerting, and telemetry agents running on every node in the fleet.
You will lead the Platform team as an Engineering Manager, managing 4-6 engineers directly. The Platform team owns time series and log storage systems, and the query layer serving dashboards, APIs, and investigations. This is a first-line management role reporting to the Engineering Manager for Cloud Monitoring Services.
Key responsibilities include:
- Direct people management: 1:1s, career development, performance reviews, and team health for a growing team
- Own storage and query systems: accountability for time series and log storage, query layer performance under load, and operational costs
- Delivery ownership: plan and sequence work across a roadmap mixing customer-facing features with infrastructure, making early risk calls
- Technical direction: partner with Staff engineers on technical strategy, ask hard questions, and make sound tradeoff decisions
- Operational excellence: maintain high availability, query performance, and correctness on critical-path systems; own oncall rotation
- Cost and scale management: make product decisions on retention policy, downsampling, cardinality, and storage tiering
- Cross-team coordination: sequence dependencies with collection and ingestion teams, serve internal and external customers
- Team growth: work with recruiting, run interview loops, close candidates, and onboard effectively
- Cross-functional collaboration: align with product, infrastructure teams, peer managers, and leadership
This is a people-first leadership role with real delivery stakes. Your success is measured through what your team accomplishes. Query latency, retention, and storage cost are live tradeoffs your team will make continuously, and customers feel all three directly.
Requirements:
- 5+ years of hands-on engineering experience in backend or distributed systems (databases, storage engines, query engines, streaming pipelines, or data-intensive services)
- Experience with Go, Rust, Java, or C++; hands-on familiarity with Kubernetes
- 2+ years directly managing software engineers, including performance cycles and career conversations
- Experience coaching senior and staff-level engineers; comfort partnering with strong technical leads
- Track record of shipping multi-phase projects against fixed deadlines; ability to balance infrastructure work with feature work
- Operational judgment: experience running teams that own systems customers depend on during incidents
- Strong communication and judgment; ability to align cross-functional stakeholders and explain tradeoffs
Bonus experience:
- Time series databases or columnar storage at scale (Prometheus, Thanos, Mimir, VictoriaMetrics, InfluxDB, ClickHouse)
- Query engines and query cost control (planning, pushdown, caching, concurrency limits)
- Retention, compaction, downsampling, and storage tiering for large telemetry datasets
- Log storage and search systems (Loki, Elasticsearch, OpenSearch)
- Observability domain experience (OpenTelemetry, Grafana, metrics and log pipelines, cardinality management)
- Multi-tenant systems with per-tenant isolation, quotas, and fairness
- GPU or accelerated computing environments
- Managing or scaling a team through growth