SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
HealthEdge is seeking a Senior Cloud Infrastructure Engineer to own the design, resilience, and operational health of its infrastructure across AWS, Azure, GCP, and hybrid on-premises environments. This is a hands-on senior individual contributor role focused on deep infrastructure ownership across a large, multi-account, multi-platform environment actively migrating off legacy on-premises infrastructure.
Key responsibilities span multiple domains:
Infrastructure & Hybrid Cloud: Design, build, and maintain infrastructure primarily on AWS with supporting work in Azure and GCP. Own secure configuration baselines, patch management for OS images and on-premises hardware, and reusable Infrastructure as Code for cloud and on-premises deployments. Support cloud networking execution, VPC provisioning, security group standards, and connectivity work. Manage storage across cloud and legacy environments as workloads migrate, and contribute to on-premises retirement and datacenter decommissioning.
Disaster Recovery & Resilience: Own disaster recovery architecture and execution across cloud environments. Maintain DR solutions, backup strategy, and run regular DR drills. Design for resilience from the start and treat recoverability as a first-class requirement.
Security, Compliance & Vulnerability Management: Contribute to vulnerability management triage, threat detection, and infrastructure security findings review. Support PHI/PII classification scanning, penetration testing coordination, and security exception approvals. Maintain HIPAA and SOC 2 compliance controls, manage access reviews, and coordinate change freezes.
Cloud Infrastructure Operations: Design and manage roles and cross-account access controls following least-privilege principles. Own cloud execution including load balancers, DNS, VPC provisioning, and security group standards. Manage compute resources at scale with attention to right-sizing and maintainability. Administer cloud storage and database services. Own infrastructure health, cost, and performance monitoring using native and third-party tooling. Administer and harden Linux and Windows Server environments across cloud and on-premises.
CI/CD, Automation & Delivery: Build and evolve CI/CD pipelines for secure, repeatable infrastructure deployments. Write automation to reduce manual toil and enforce operational consistency. Take solutions from proof-of-concept to production with long-term maintainability in mind.
Reliability, Monitoring & Incident Response: Monitor, scale, and maintain production infrastructure with availability, performance, and security as top priorities. Participate in on-call rotation and serve as L2 escalation point for cross-team infrastructure support. Lead root cause analysis and drive incident retrospectives to closure. Author and maintain runbooks that hold up under pressure.
Cost & Tagging Governance: Contribute to FinOps efforts by identifying and remediating cost anomalies, owning tagging remediation against enterprise standards, and making pragmatic cost/performance/resilience tradeoffs.
AI-Enabled Engineering: Use AI coding assistants to accelerate Infrastructure as Code development, scripting, and troubleshooting. Use AI tooling to draft runbooks, DR documentation, and incident retrospectives.
Collaboration & Documentation: Document architecture, DR runbooks, and standard operating procedures. Provide technical guidance to product teams on infrastructure resilience, migration sequencing, and recovery design.
Required qualifications include 5+ years of hands-on cloud infrastructure engineering experience with deep AWS expertise and experience in other cloud environments, direct disaster recovery design and execution experience, strong Infrastructure as Code skills (CDK, Terraform, or CloudFormation), and experience with containerized environments and Kubernetes/EKS.