SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
PDI Technologies is seeking a CloudOps Engineer III to manage and support enterprise customers' cloud infrastructure on AWS, ensuring availability, performance, and security of production environments.
Key Responsibilities:
- Manage and support enterprise customers' cloud infrastructure on AWS, ensuring availability, performance, and security of production environments
- Monitor and respond to alerts via observability tooling, triaging and resolving incidents in line with agreed SLAs and SLOs
- Lead incident response and root cause analysis (RCA) for complex or high-severity issues, producing clear write-ups and remediation actions
- Carry out change management activities including deployments, security patching, and infrastructure updates following established change control and CR governance processes
- Administer and optimize AWS database infrastructure (e.g., RDS), including performance tuning, troubleshooting, and exposure/risk remediation
- Support compliance and audit activity (e.g., SOC audits), ensuring operational practices meet security and regulatory requirements
- Identify and help deliver cloud cost optimization opportunities
- Maintain accurate tickets, documentation, and workaround registers, and help close down accumulated technical debt
Communication & Stakeholder Management:
- Communicate clearly and confidently with customers, Professional Services, and internal stakeholders, both verbally and in writing
- Plan, schedule, and run operational and governance meetings, setting clear agendas, driving decisions, and tracking follow-up actions to completion
- Provide regular, well-structured status updates and reporting back to the wider CloudOps team, keeping colleagues informed of progress, risks, and changes
- Represent CloudOps professionally in customer-facing forums, translating technical detail into clear, actionable communication for a range of audiences
Requirements:
- Significant experience in a Cloud Operations, DevOps, or Site Reliability Engineering role, ideally within AWS environments
- Strong working knowledge of core AWS services (e.g., EC2, RDS) and cloud infrastructure operations
- Experience with monitoring/observability tooling, incident management, and change management processes
- Understanding of security patching and compliance/audit processes (e.g., SOC 2)
- Excellent written and verbal communication skills, with the confidence and credibility to lead meetings and communicate outcomes back to the team and to customers
- Strong organizational skills and the proven ability to manage competing priorities under your own initiative
- A collaborative team player who is equally comfortable working independently and knowing when to seek support from the team
- Relevant AWS certification(s) desirable