SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Precisely is a global leader in data quality, data enrichment, and location intelligence, serving thousands of trusted brands. The company is AI-first, embedding artificial intelligence throughout product development and operations.
As a Senior Site Reliability Engineer, you will be responsible for the reliability, performance, and scalability of Precisely's infrastructure platforms across three environments: CEDAR (CCX)—an on-premises, private cloud managed services environment; RapidCX (RCX)—an AWS cloud environment for SaaS-delivered customer communications management; and Hosted Managed Services (HMS)—an AWS cloud environment supporting managed client deployments.
This role bridges software engineering and systems operations. You will build automation, observability tooling, and reliability standards to ensure platform availability and operational excellence. As a Senior SRE, you are regarded as a platform expert across all three environments, taking on complex reliability engineering work, mentoring less experienced engineers, and contributing to release triage and coordination.
Key responsibilities include:
- Define and maintain reliability standards (SLOs, SLIs, error budgets) across all platforms
- Establish standards for alerting, logging, and diagnostic tracing; partner with engineering teams to ensure compliance
- Lead complex reliability engineering work across on-premises, private cloud, and AWS environments
- Build and maintain infrastructure-as-code, deployment automation, monitoring, and observability tooling using Terraform, Ansible, Datadog, and scripting languages
- Partner with engineering teams to embed reliability, operability, scalability, backup, and recovery considerations into service design
- Lead Operational Readiness Reviews and validate production and disaster recovery readiness
- Contribute to release triage, deployment coordination, and change reviews
- Lead response to complex P1/P2 incidents, serve as incident commander when required
- Author root cause analyses and post-incident reports; identify patterns and implement preventive automation
- Maintain operational runbooks, reliability backlogs, and team knowledge resources
- Use Precisely-provided AI tools (GitHub Copilot, Claude) for automation, incident analysis, troubleshooting, and documentation
- Ensure infrastructure meets security, compliance, and data-protection requirements
- Mentor Associate SREs and SREs; coach engineering teams on production diagnostics and operational practices
- Lead project-based work and coordinate on-call coverage
- Participate in rotating on-call schedule covering after-hours, weekends, and holidays
This is a 100% remote position within the U.S., with a preference for candidates in Mountain or Pacific Time zones. No travel required.
REQUIREMENTS:
Required:
- Bachelor's degree in Computer Science, Information Systems, Engineering, or equivalent practical experience
- 5+ years of systems or infrastructure engineering experience in an enterprise production environment
- High-level, developing subject-matter-expert competency in one or more complex infrastructure domains (cloud platform engineering, IaC automation, observability, or on-premises virtualization)
- Advanced proficiency with Linux (RHEL/Oracle Linux) across multiple environments (on-premises and cloud)
- Proficient with Terraform for infrastructure-as-code; experienced with Ansible role development beyond basic use
- Intermediate-level AWS experience: EC2, ECS, S3, VPC, IAM, CloudWatch, Auto Scaling
- Proficiency with at least one scripting language (Python, Bash) for automation development
- Demonstrated experience designing monitoring system architecture and alerting strategy (Datadog preferred)
- Solid understanding of TCP/IP networking, DNS, load balancing, and distributed systems
- Experience designing and managing CI/CD pipelines and deployment automation standards
- Strong analytical skills; demonstrated experience authoring root cause analyses for complex incidents
- Ability to define SLOs and lead Operational Readiness Reviews; comfortable partnering with engineering teams on production readiness
- Demonstrated ability to work cross-functionally with engineering teams on reliability standards and observability
- Experience participating in change advisory processes and contributing to capacity and reliability planning
- Active, proficient use of AI tools (GitHub Copilot, Claude, or equivalent) for complex automation, solution testing, and architecture documentation; ability to mentor others on AI-assisted engineering practices
Preferred:
- Experience with containerization and orchestration (Docker, ECS, Kubernetes)
- Familiarity with GitOps workflows and source control best practices (Git, GitLab)
- Knowledge of enterprise virtualization platforms in hybrid cloud context
- Understanding of change management and ITIL operational practices
- Experience with enterprise security tooling (Qualys, CrowdStrike, Rapid7)
- AWS Solutions Architect, SysOps Administrator, or DevOps Engineer certification
- Wireshark and protocol analysis experience
- Prior experience mentoring engineers or leading on-call rotations