SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Salary: USD 268,000 - 351,750 / annual
Temporal is an open-source programming model that simplifies code, improves application reliability, and helps developers ship features faster. The company is building the reliable foundation for every developer's toolbox.
As Senior Engineering Manager of Test Systems & Tooling, you will lead a team responsible for building shared systems, environments, and tooling that enable Temporal engineers to validate the platform under production-like and adverse conditions. Your mission is to ensure system-level testing gaps don't fall through organizational boundaries by driving cross-team alignment, clear ownership, and measurable improvements in pre-production confidence.
Key responsibilities include:
• Lead and grow the Test Systems & Tooling team: hire, coach, and develop engineers building shared testing capabilities used across Engineering.
• Make production-like testing practical: improve the ability for engineers to create test environments that resemble production topology, configuration, scale, and behavior—so reproducing issues and validating changes is faster and more reliable.
• Enable production-representative workload testing: build or evolve capabilities such as workload replay, realistic workload generation, multi-tenant testing, load testing, and stress testing to surface behaviors that only emerge at scale.
• Drive failure-mode validation: establish repeatable ways to test Temporal through real failure conditions (dependency failures, latency injection, resource exhaustion, degraded-state recovery, regional failure scenarios).
• Create shared frameworks and guidance: provide opinionated testing frameworks, patterns, and tooling that make it easier for product teams to write high-quality tests consistently, with visibility into test health and systemic gaps.
• Own system-level coverage outcomes: identify critical system-level scenarios that need explicit coverage, drive them to completion (often through partner teams), and ensure they're integrated into continuous and/or release validation.
• Collaborate across org boundaries: partner closely with Cloud/Infrastructure, Reliability, Release Engineering, and product teams to turn production learnings into stronger pre-production detection.
You bring experience managing and developing high-performing engineering teams with strong coaching, feedback, and performance management skills. You have a strong technical background building or operating reliable distributed systems and adjacent platform infrastructure, with judgment to prioritize the failure modes that matter most. You have a track record delivering internal platforms and tooling adopted by many teams, balancing ergonomics, speed, and long-term maintainability. You're comfortable operating in ambiguous problem spaces where success depends on aligning stakeholders, clarifying ownership, and driving execution across multiple teams. You can define and drive measurable outcomes (adoption, confidence signals, reduced escaped issues, faster incident-to-coverage time), not just ship tooling.
Nice-to-have experience includes large-scale testing practices (chaos engineering, fault injection, workload replay/shadowing, performance testing, multi-tenant test strategy), partnering with Release Engineering or CI/CD platform teams to institutionalize validations in pipelines, and familiarity with observability and incident response practices.