SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Figma is seeking a Manager of Software Engineering for its Observability team. You will lead a team of five engineers responsible for building and operating the systems that provide deep visibility into the health, performance, and efficiency of Figma's platform.
Key responsibilities include:
- Lead and grow a team of 5 engineers responsible for the reliability, scalability, and evolution of Figma's Observability and AI Trace Observability platforms.
- Own and operate the AI Trace Observability ecosystem, including a safe telemetry pipeline that aggregates data, runs it through classification, and enables engineers to run evaluations against this data.
- Own and operate Figma's core observability stack, including vendor platforms such as Datadog, ensuring high availability, strong data quality, and effective signal-to-noise across metrics, logs, and traces.
- Define and drive the technical strategy for instrumentation standards, observability libraries, agents, and operators used to monitor internal and external-facing services.
- Explore and implement innovative, AI-driven approaches to anomaly detection, root cause analysis, signal correlation, and operational automation.
- Partner with infrastructure, product engineering, finance, and security teams to improve visibility into system health and cost efficiency at scale.
- Coach and mentor engineers through career development, performance feedback, and technical leadership, fostering a culture of ownership, collaboration, and high-quality execution.
Required qualifications:
- 4+ years of experience leading infrastructure, observability, or platform engineering teams, with a track record of delivering highly reliable production systems.
- Deep hands-on experience with modern observability platforms (e.g., Datadog, OpenTelemetry) across metrics, logs, and distributed tracing.
- Strong understanding of distributed systems, instrumentation best practices, SLO design, and incident response workflows.
- Experience driving cost transparency and accountability initiatives, including cost attribution, budgeting, forecasting, and alerting in cloud environments.
- Demonstrated ability to set technical direction, drive cross-functional alignment (Engineering, Finance, Security), and make sound architectural decisions in complex environments.
Preferred qualifications include experience building observability and telemetry systems for AI/ML products, designing company-wide observability standards, cost optimization for infrastructure tooling, applying AI/ML to anomaly detection and operational automation, and familiarity with OpenTelemetry.