SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Salary: USD 130,000 - 0 / annual
Dragos is the global leader in xOT (extended operational technology) cybersecurity, protecting critical infrastructure systems including water, power, and healthcare. The company combines technology, threat intelligence, and expert services to defend the systems that power civilization against daily cyberattacks.
As a Senior Engineering Support Engineer, you will serve as a primary escalation point for complex customer issues, bridging the gap between customers and engineering teams. You'll own Tier 3 technical escalations, application-layer observability, and the feedback loop that turns production signals into better software.
Key responsibilities include:
- Lead escalation triage and incident response for complex, multi-component customer support cases, owning investigation and resolution coordination from intake to close and running technical incident command.
- Validate, reproduce, and document confirmed defects emerging from customer escalations, producing clear bug reports with supporting evidence and reproduction steps.
- Build and maintain application observability by developing and refining Datadog monitors, dashboards, and alerts that provide visibility into customer-facing SLOs and platform health.
- Own customer communication through the lifecycle of escalated incidents, ensuring accurate and timely updates.
- Translate field experience into documentation by authoring troubleshooting guides, runbooks, and playbooks for Tier 1/2 support staff.
- Drive pattern recognition across the customer base to identify recurring failure modes and collaborate with product teams on permanent fixes.
- Participate in on-call rotation, including occasional weekend coverage, triaging application-layer alerts and executing documented remediations.
- Mentor Tier 1/2 Customer Experience support staff and maintain the knowledge base.
Requirements:
- 3+ years of experience in technical support engineering, site reliability engineering, or closely related customer-facing engineering role with demonstrated ownership of complex issue resolution.
- Demonstrated experience troubleshooting complex distributed systems in production environments; comfortable reading logs, tracing component interactions, and isolating root cause under pressure.
- Strong Linux system administration skills including processes, filesystem, networking, and diagnosing application behavior from first principles.
- Experience with containerized application environments (Kubernetes and/or Docker) sufficient to investigate pod health, examine logs, and understand deployment state.
- Familiarity with observability platforms (Datadog or equivalent) and experience building and maintaining monitors, dashboards, and alerts.
- Experience supporting Elasticsearch, PostgreSQL, or similar database platforms in production.
- Strong written communication skills with demonstrated ability to produce clear, accurate technical documentation (RCAs, runbooks, postmortems) for internal and customer-facing audiences.
- Comfort working with AI tools and assistants as part of day-to-day engineering workflow, including prompt engineering and AI-assisted triage.
- Experience in customer-facing or customer-adjacent engineering roles with judgment to balance urgency and thoroughness under SLA pressure.
- Experience supporting or securing software in a cybersecurity context.
- Ability to read and navigate application source code to trace bugs, understand service behavior, or follow data flows.
- Preferred: Direct experience with OT/ICS cybersecurity environments or familiarity with industrial control systems.
- Preferred: Experience with AWS (EKS, RDS, EC2) or comparable cloud infrastructure at the application layer.
- Preferred: Background in customer success, customer engineering, or professional services roles requiring deep technical credibility alongside relationship management.