SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Twilio is seeking a Principal Software Engineer to lead the technical depth and evolution of systems that continuously test, measure, and assure the quality of Twilio's global carrier connections and operational platforms. Every call and message Twilio makes crosses carrier networks outside their control; this team ensures those connections remain healthy through continuous monitoring and testing worldwide for both voice and messaging reliability.
In this role, you will design and deliver major components of Twilio's carrier test and observability platform, covering active voice and messaging testing, quality scoring, and global test coverage. You'll build high-throughput data and alerting systems that transform raw data into actionable signals for operations teams without overwhelming them with noise. A key focus is extending and hardening production LLM systems for automated carrier troubleshooting—expanding the classes of issues they resolve autonomously while owning evaluation, observability, and guardrails that make autonomous action safe.
You'll diagnose cross-boundary failures where test results, carrier behavior, and platform telemetry disagree, establishing the observability needed to distinguish these cases. You'll contribute to technical design and review across the team, mentor engineers, and help shape written standards that outlast individual projects. This is a remote-first role based in Ireland, though occasional travel may be required for team gatherings, functional off-sites, or customer meetings.
Twilio is a remote-first company delivering innovative communications solutions to hundreds of thousands of businesses and empowering millions of developers worldwide. The company emphasizes connection, global inclusion, and rewarding work.
**Requirements:**
Required:
- 12+ years of software engineering experience, including significant time at staff or principal level working on systems at large scale
- Deep expertise designing, operating, and debugging distributed systems that ingest and process high-volume time-series or event data in production, including alerting, anomaly detection, and failure isolation
- Hands-on experience taking LLM-based systems to production and maintaining them: context and tool design, evaluation against real outcomes, cost and latency management, and guardrails required when model output triggers real actions
- Strong proficiency in at least one backend language used at scale (Java, Go, Scala, or similar) and fluency with event streaming and cloud infrastructure (Kafka, AWS, Kubernetes)
- Ability to communicate complex technical trade-offs clearly in writing and influence peers through design review rather than authority
Desired:
- Experience with telecommunications, messaging, or voice systems
- Experience with active network testing, monitoring, or quality-of-service measurement
- Experience with agentic systems that take action on live infrastructure and evaluation practices that make that trustworthy
About Twilio
SaaS / Enterprise Software — authentication and identity infrastructure for developers.