SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
NetBox Labs is seeking a Senior DevOps Engineer to join the rapidly expanding engineering team. This role focuses on building and maintaining reliable, scalable infrastructure through infrastructure-as-code, owning the reliability and performance of deployed systems. The ideal candidate is infrastructure-minded, treats operations like software through automation and instrumentation, and is comfortable navigating and contributing to software systems while collaborating closely with developers to bridge the gap between code and production.
You will design, build, and operate infrastructure systems supporting NetBox Labs' SaaS and on-premise engineering needs. Key responsibilities include operating and optimizing AWS and other cloud infrastructure with focus on cost efficiency, security, and performance; contributing to internal platform tooling, CI/CD automation (GitHub Actions), and developer self-service capabilities; enhancing observability and incident response systems including monitoring, alerting, and SLOs; collaborating with product teams to understand their needs and improve internal platforms; helping enforce and improve security and compliance standards including SOC 2 controls; contributing to documentation and onboarding materials; and participating in on-call rotation.
NetBox Labs helps companies build and manage complex networks through open, composable products. The company is the commercial steward of open-source NetBox (the world's most popular network source of truth) and Orb (next-generation network observability platform). Products include NetBox Enterprise and NetBox Cloud. NetBox powers thousands of companies and is backed by Notable Capital, Grafana Labs CEO Raj Dutt, Flybridge, IBM, Salesforce Ventures, and Mango Capital.
REQUIREMENTS:
- 5+ years of experience in DevOps, SRE, or platform engineering roles
- 2+ years of experience at a B2B software startup
- Strong experience with AWS (EC2, VPC, IAM, RDS, etc.), especially EKS/Kubernetes and infrastructure-as-code (Terraform, Helm)
- Experience with CI/CD pipelines and automation tooling, ideally GitHub Actions
- Familiarity with observability tools like Prometheus, Grafana, Mimir, Loki, or similar
- Proficiency in Python, Go, or shell scripting
- Comfortable operating in a fast-paced, ambiguous startup environment
- Strong communication and documentation skills
- Experience with Change Data Capture (CDC) and event streaming systems OR experience with scaling large, multi-tenant observability systems including ingest, analysis, and alerting
NICE TO HAVE:
- Experience in SOC 2 compliant or security-focused environment
- Experience with MQTT, AMQP, and other messaging technologies
- Familiarity with NetBox ecosystem or network automation tooling
- Open-source experience or contributions
- Experience with AI tools (e.g., Copilot, ChatGPT, Cursor)