SlipstreamJobsFresh Startup & VC-Backed Jobs

Manager, Software Engineering - Observability

Figma - San Francisco, CA, United States - Hybrid - posted 2026-02-27

Apply on the company site

SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.

Figma is seeking a Manager of Software Engineering for its Observability team. You will lead a team of five engineers responsible for building and operating the systems that provide deep visibility into the health, performance, and efficiency of Figma's platform. Key responsibilities include: - Lead and grow a team of 5 engineers responsible for the reliability, scalability, and evolution of Figma's Observability and AI Trace Observability platforms. - Own and operate the AI Trace Observability ecosystem, including a safe telemetry pipeline that aggregates data, runs it through classification, and enables engineers to run evaluations against this data. - Own and operate Figma's core observability stack, including vendor platforms such as Datadog, ensuring high availability, strong data quality, and effective signal-to-noise across metrics, logs, and traces. - Define and drive the technical strategy for instrumentation standards, observability libraries, agents, and operators used to monitor internal and external-facing services. - Explore and implement innovative, AI-driven approaches to anomaly detection, root cause analysis, signal correlation, and operational automation. - Partner with infrastructure, product engineering, finance, and security teams to improve visibility into system health and cost efficiency at scale. - Coach and mentor engineers through career development, performance feedback, and technical leadership, fostering a culture of ownership, collaboration, and high-quality execution. Required qualifications: - 4+ years of experience leading infrastructure, observability, or platform engineering teams, with a track record of delivering highly reliable production systems. - Deep hands-on experience with modern observability platforms (e.g., Datadog, OpenTelemetry) across metrics, logs, and distributed tracing. - Strong understanding of distributed systems, instrumentation best practices, SLO design, and incident response workflows. - Experience driving cost transparency and accountability initiatives, including cost attribution, budgeting, forecasting, and alerting in cloud environments. - Demonstrated ability to set technical direction, drive cross-functional alignment (Engineering, Finance, Security), and make sound architectural decisions in complex environments. Preferred qualifications include experience building observability and telemetry systems for AI/ML products, designing company-wide observability standards, cost optimization for infrastructure tooling, applying AI/ML to anomaly detection and operational automation, and familiarity with OpenTelemetry.

Similar roles