SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Demandbase is a pipeline AI platform for go-to-market teams, enabling B2B enterprises to automate growth and execute account-based strategies at scale. The company is recognized as one of the best places to work in the San Francisco Bay Area.
As a Staff DevOps Engineer on the Developer Experience (DevEx) team, you will set the technical direction for the platforms, tooling, and workflows that every engineering team at Demandbase depends on to ship reliably to production. This is a strategic, organization-wide role—not a single-team execution position.
You will define multi-quarter platform strategy across build, test, deploy, and operate stages for services and data pipelines running on AWS and GCP. You'll be the escalation point and trusted advisor when platform decisions have broad organizational impact. Your work will act as a force multiplier at scale: abstracting infrastructure complexity into paved roads, driving adoption of self-service platforms, and embedding reliability, security, cost-awareness, and observability across the engineering organization.
Key responsibilities include:
**Platform Strategy & Technical Leadership:** Set and drive the technical roadmap for the Internal Developer Platform (IDP). Author and review RFCs/design docs adopted org-wide. Serve as escalation point for highest-severity production incidents involving platform infrastructure; lead retrospectives and turn findings into structural fixes. Mentor senior and staff-level engineers across teams, raising the technical bar for platform engineering, reliability, and secure software delivery.
**Developer Platform & Workflow Enablement:** Own and evolve Demandbase's Internal Developer Platform (Backstage service/API/group catalog, golden-path scaffolder templates, CI/CD components) as a product, measured by adoption, time-to-first-deploy, and reduction in platform-support toil. Partner with software, data, security, and product engineering leadership to identify systemic friction across the SDLC and resolve it through automation and platform abstraction. Define what "good" looks like for AI-assisted development workflows across the org—evaluate, pilot, and roll out coding-agent tooling with appropriate guardrails.
**CI/CD & Release Engineering:** Own the architecture of CI/CD orchestration (GitLab CI/CD, shared CI/CD components) supporting high release velocity, strong security guardrails, and local-to-production parity. Drive standardization of build, test, and deployment patterns across application and data workloads at scale. Own GitOps-based deployment strategy and evolve it as the org's service count grows.
**Cloud Infrastructure & Kubernetes Platforms:** Own the architecture and lifecycle of multi-cluster Kubernetes platforms across environments (EKS and GKE)—upgrades, node lifecycle (Karpenter), networking, and security posture at fleet scale. Own multi-account cloud IAM, networking, and security architecture; set the guardrails other teams build against. Drive Infrastructure-as-Code strategy (Terraform, Crossplane) for consistency and security across accounts. Own service mesh (Istio) architecture across multi-cluster environments, including sidecar injection strategy and traffic policy.
**Platform Reliability, Security & Cost:** Own core platform components—GitOps tooling, secrets management, service mesh—as production systems. Own observability platform strategy (Prometheus, Datadog, Grafana/Loki)—what's collected, indexed, and actionable—balancing signal quality against cost. Own FinOps practice for platform infrastructure as a standing responsibility. Define and evolve org-wide SLIs/SLOs and incident response practices; lead blameless postmortem culture at the organizational level.
You are fluent with AI-assisted and agentic engineering workflows and are expected to help define how the broader engineering org adopts these tools safely and effectively. This role is critical to improving local-to-production parity, reducing cognitive load fleet-wide, and sustaining high release velocity across Kubernetes-based services and data platforms.
**Requirements:**
- 10+ years of overall engineering experience, including meaningful hands-on software development (Go, Python, or Java) and 7+ years building and operating cloud infrastructure at scale
- Demonstrated track record operating as a technical leader across multiple teams—driving cross-team RFCs, architecture decisions, or platform strategy
- Deep, hands-on experience designing and operating multi-account, multi-cluster Kubernetes platforms in production (EKS and/or GKE) at meaningful scale (dozens+ of services/clusters)
- Strong proficiency with Infrastructure as Code (Terraform, Crossplane) and GitOps practices (Flux or Argo CD) applied at organizational scale
- Deep experience designing and owning CI/CD systems supporting high release velocity and strong security guardrails (GitLab CI/CD preferred)
- Hands-on experience with service mesh (Istio) in multi-cluster environments, including security/traffic policy
- Strong experience building and owning observability platforms (Prometheus, Datadog, Grafana, Loki/Mimir/Thanos) with a cost-aware lens
**Nice-to-have:**
- Working knowledge of container/supply-chain security practices (image scanning, hardened base images, vulnerability remediation)
- Experience owning or meaningfully contributing to cloud cost governance/FinOps practice
- Fluency with AI-assisted/agentic development tooling (Cursor, Claude Code, Codex, or equivalent) and informed opinions on safe, effective rollout at organizational scale
- Solid understanding of SLIs/SLOs, alerting strategy, and incident response, with experience leading blameless postmortems
- Excellent written and verbal communication—staff-level influence runs through docs and reviews as much as code
- Experience supporting data pipeline/batch platforms (Airflow, EMR, Dataproc)