SlipstreamJobsFresh Startup & VC-Backed Jobs

Senior Full Stack Engineer, Observability

NetBox Labs - Remote - Remote - posted 2026-09-28

Apply on the company site

SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.

NetBox Labs is seeking a Senior Full Stack Engineer to join the Observability team, which builds products connecting network source of truth to actual running infrastructure. The team owns NetBox Discovery (finding devices and interfaces), NetBox Assurance (spotting drift between intended and actual state), Fleet Management and the Orb agent (deploying and managing data collection agents), and Diode (the data ingestion pipeline). You will take ownership of features end-to-end across the full stack: Go and Python services with gRPC APIs, data flows carrying discovery and telemetry into NetBox, and React dashboards where customers monitor network and device health. You'll work closely with product, design, and other engineering teams, and help run services in production. Key responsibilities include: - Design, build, and operate backend services in Go and Python for discovery, assurance, fleet management, and data ingestion - Define and evolve gRPC and REST APIs using Protocol Buffers and OpenAPI - Build React and TypeScript monitoring dashboards showing device, interface, and network health with time-series charts, status views, and drill-down capabilities - Design dashboard experiences that help operators spot problems quickly with sensible defaults, time-range and filter controls, thresholds, and clear links to underlying devices - Collaborate with backend engineers on query and aggregation APIs to keep dashboards fast with large fleets and high-frequency telemetry - Build type-safe frontend-backend integration using generated API clients and shared schemas - Participate in on-call rotation for owned services - Add automated tests across the stack (unit, integration, contract, end-to-end) and enforce quality gates in CI - Collaborate with product managers, designers, customer-facing teams, and other engineers to solve real network operator problems - Use AI-enabled development tools and agentic workflows to speed up design, coding, testing, code review, and incident triage - Review code, mentor teammates, and share best practices - Participate in planning and help shape the roadmap REQUIREMENTS: - 5+ years of professional software engineering with meaningful production experience on both backend and frontend - Backend: Production experience with Go and Python (strong in at least one, working proficiency in the other); hands-on experience designing and operating gRPC services with Protocol Buffers including schema evolution, backward compatibility, streaming RPCs, deadlines, interceptors/middleware, and error handling; experience building distributed, event-driven systems including message queues (RabbitMQ, Kafka), asynchronous job processing, and idempotent data ingestion - Frontend: Strong React and TypeScript skills including component composition, state management, and typing best practices; proven experience building monitoring, observability, or analytics dashboards with time-series charts, heatmaps, status and health views, and drill-down navigation; hands-on experience with data visualization libraries (D3, ECharts, Recharts, uPlot, Visx); experience rendering large or high-frequency datasets performantly using downsampling, virtualization, canvas or WebGL rendering; experience handling real-time data via WebSockets, server-sent events, or gRPC-Web streaming with caching and refresh strategies (TanStack Query); modern CSS (TailwindCSS or similar), responsive layout, and practical accessibility (WCAG) including color-blind-safe palettes and accessible charts; automated frontend testing with Jest and React Testing Library including visual regression testing - AI-enabled development: Practical, regular use of AI coding assistants and agents (Claude Code, Cursor, GitHub Copilot) across the development lifecycle; good judgment about when to trust AI output including verifying generated code and protecting sensitive data; familiarity with prompting and context techniques - Ways of working: Good communication skills and proven ability to work collaboratively in small cross-functional teams; comfortable in fast-moving environments and contributing to product and technical decisions NICE TO HAVE: - Networking protocols and device interfaces: TCP/IP, DNS, DHCP, BGP, OSPF, VLANs, LLDP/CDP; network management and telemetry interfaces such as SNMP, gNMI/gNOI, streaming telemetry - Network automation: NetBox, NAPALM, Nornir, Netmiko, or building agents for customer environments - Telemetry and observability: OpenTelemetry/OTLP, Prometheus, time-series databases (Mimir, InfluxDB, ClickHouse), PromQL - Dashboard tooling: Grafana panel/plugin development, user-configurable dashboards, alerting and incident UX design - Building with AI: LLM APIs for product features, Model Context Protocol, tool-using agents, agentic development frameworks, eval suites and guardrails

Similar roles