SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Union.ai is hiring a Systems Development Engineer to improve the reliability, operability, and customer experience of their production platform. This is a 50/50 operations and engineering role focused on investigating customer-impacting production issues and building tools, automation, and system design to prevent recurrence.
You will work from real operational signals—customer issues, incidents, on-call pages, support patterns, and observability gaps—to drive durable platform improvements. This is not a traditional support role but a systems engineering position for someone who can debug deeply, communicate clearly, and write software that reduces operational load. The full engineering team backs you on on-call.
Key responsibilities include investigating and resolving production issues across cloud infrastructure, workflow execution, access control, storage, networking, deployment systems, and observability. You'll identify patterns in customer issues and convert them into automation, product improvements, runbooks, tests, or design changes. You'll build internal tools and diagnostics to make production issues easier to detect, understand, and resolve, improve platform observability (logs, metrics, dashboards, alerts), and participate in design and development to ensure systems are easier to operate and debug.
You'll define and uphold operational engineering practices including production readiness, alert quality, runbook discipline, observability standards, regression prevention, and code quality. Success means driving measurable reductions in on-call pages, recurring customer issues, manual operational work, and time-to-resolution.
Required skills include strong software engineering in Python, Go, Java, or similar languages; experience debugging production systems across multiple stack layers; practical knowledge of Kubernetes, Linux, cloud infrastructure, distributed systems, networking, storage, and IAM; experience with infrastructure-as-code, deployment systems, CI/CD, observability, and operational automation; and ability to move from ambiguous customer symptoms to clear technical diagnosis and durable remediation.
Preferred experience includes operating customer-facing SaaS, cloud infrastructure, self-hosted or on-prem deployments, or workflow orchestration systems; batch workloads, autoscaling, capacity management, identity and access systems, storage systems, or platform observability; improving on-call health or building production diagnostics; and working across support, customer success, product, and engineering teams.