SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
PDI Technologies is seeking a CloudOps Engineer IV to lead cloud infrastructure and operations at scale. This is a senior individual contributor role focused on building, operating, and troubleshooting production infrastructure across AWS and Azure, with significant mentorship responsibilities for junior engineers.
Key Responsibilities:
Cloud Infrastructure & Operations: Build, operate, and troubleshoot infrastructure across AWS and Azure supporting production workloads. Operate and maintain Kubernetes clusters, including deploying and maintaining Helm charts. Participate in on-call rotation, respond to incidents, and drive them to resolution. Contribute to capacity planning, cost optimization, and resilience improvements.
Automation & Continuous Delivery: Build and maintain GitOps-based deployment pipelines using Argo CD/Argo Workflows with rollout and promotion configuration across environments. Write and maintain Infrastructure-as-Code using Terraform and OpenTofu following team module standards. Build and maintain CI/CD pipelines in Jenkins, improving build/deploy automation and reliability. Support progressive delivery practices including blue-green/canary deployments and automated rollback.
Reliability & Observability: Build and maintain Datadog dashboards, monitors, and alerts for services, tuning thresholds to reduce noise. Contribute to defining SLIs/SLOs and track them over time. Participate in postmortems for incidents and follow through on remediation items.
Collaboration & Mentorship: Partner with SRE and product engineering teams to troubleshoot issues and improve system design. Mentor less-experienced engineers (CloudOps Engineer I-III) on cloud infrastructure, Kubernetes, and CI/CD practices. Contribute to documentation, runbooks, and onboarding materials.
Required Qualifications:
- 10+ years of experience in Cloud Operations, Site Reliability Engineering, DevOps, or Infrastructure Engineering roles
- Hands-on experience with AWS: building, troubleshooting, and operating cloud infrastructure directly
- Hands-on experience with Kubernetes and Helm: deploying, operating, and troubleshooting production workloads
- Hands-on experience with Argo CD/Argo Workflows for GitOps-based continuous delivery
- Hands-on experience with Infrastructure as Code (Terraform, OpenTofu)
- Hands-on experience with Jenkins and Rancher for CI/CD pipeline development and maintenance
- Hands-on experience with Datadog or equivalent observability platform, including dashboards, monitors, and alerts
- Experience participating in on-call rotation and responding to production incidents
- Strong communication skills and ability to work effectively across teams
Preferred Qualifications:
- Experience supporting payments, fuel/retail, loyalty platforms, or systems with PCI DSS or similar compliance obligations
- Relevant certifications: CKA/CKAD, AWS Certified Solutions Architect – Associate, Microsoft Certified: Azure Administrator, or HashiCorp Terraform Associate
- Experience with messaging systems (Kafka/SQS/SNS) and multi-region/multi-AZ resilience patterns
- Prior experience mentoring junior engineers or leading small technical initiatives