SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Salary: USD 203,000 - 275,000 / annual
Altruist is transforming the wealth management industry with an AI platform for financial advisors. We're hiring a Staff Cloud Infrastructure Engineer to join the Cloud Infrastructure & Platform (CIN) team, responsible for architecting, building, and operating AWS-based infrastructure powering our broker-dealer and clearing platform.
This is a high-impact, individual contributor role with significant technical influence. You will own critical infrastructure domains end-to-end and drive technical decisions affecting reliability, security, and scalability of systems handling real financial transactions. As the industry evolves with generative AI and agentic workflows, we need an AI-forward engineer combining deep AWS and Kubernetes expertise with vision to define how AI/ML tools can transform infrastructure operations, developer productivity, and platform resilience.
Key responsibilities include:
**Cloud Infrastructure & Platform Engineering:** Architect and operate production AWS infrastructure supporting high-availability financial services workloads (EKS, MSK, RDS/Aurora PostgreSQL, OpenSearch, ElastiCache, S3, CloudFront). Own and evolve Infrastructure as Code (IaC) strategy using Terraform, defining module standards and driving adoption of reusable patterns. Lead Kubernetes (EKS) platform strategy including cluster upgrades, node group architecture, Helm governance, and service mesh evolution. Design and drive CI/CD platform improvements (GitHub Actions, ArgoCD) enabling safe, fast deployments. Architect and validate disaster recovery strategies including cross-region failover and backup automation. Lead infrastructure design reviews ensuring solutions meet scalability, security, and compliance requirements.
**Reliability, Observability & Operational Excellence:** Define and drive observability strategy across the platform (Datadog, Prometheus, Grafana, CloudWatch, OpenSearch) including SLO/SLI frameworks and alerting standards. Serve as senior on-call escalation point; lead root-cause analysis on critical production incidents and drive systemic improvements. Own monthly resource saturation reviews and capacity planning. Drive cloud cost optimization strategy including FinOps practices and vendor spend governance.
**Security, Compliance & Networking:** Define and enforce security architecture standards across AWS environments (IAM, VPC design, encryption, secrets management). Partner with Security, Compliance, and Audit teams to ensure infrastructure meets FINRA, SEC, SOC 2 requirements. Own networking architecture decisions including VPC topology, Transit Gateway strategy, load balancer patterns, and CDN optimization.
**AI/ML Infrastructure & Developer Productivity:** Define strategy for evaluating and governing AI-powered developer tools (Cursor AI, GitHub Copilot, CodeRabbit) including usage analytics and cost optimization. Architect infrastructure for AI/ML workloads (GPU compute, SageMaker, Bedrock, vector databases). Lead adoption of AI-driven automation for infrastructure operations (intelligent alerting, anomaly detection, auto-remediation). Build internal platforms leveraging generative AI to improve developer experience and reduce operational toil.
**Technical Leadership & Collaboration:** Act as trusted technical advisor to engineering leadership on infrastructure strategy and investment priorities. Lead cross-functional initiatives spanning application engineering, data, security, and DevSecOps teams. Represent CIN team in architecture review boards and incident response leadership. Author and maintain comprehensive technical documentation, runbooks, and ADRs. Mentor senior and mid-level engineers; conduct design and code reviews; contribute to hiring and onboarding.
This role follows a hybrid schedule with three days per week onsite in San Francisco or Culver City office.
**Requirements:**
- 7+ years of hands-on experience in cloud infrastructure engineering with deep, production-proven AWS expertise
- Track record of owning and driving infrastructure initiatives end-to-end from design through operational excellence
- Expert-level proficiency with Terraform including module design, state management, and establishing IaC standards
- Extensive production experience operating Kubernetes (EKS strongly preferred) at scale including cluster lifecycle management and GitOps workflows
- Strong Linux systems engineering skills and advanced scripting proficiency (Python, Bash, or Go)
- Deep expertise in CI/CD platforms (GitHub Actions, ArgoCD, Jenkins) with experience designing deployment strategies for multi-service architectures
- Expert-level understanding of cloud networking (VPC architecture, Transit Gateway, DNS, load balancing) and security (IAM, KMS, WAF, GuardDuty, Secrets Manager)
- Proven experience designing and operating observability platforms (Datadog, Prometheus/Grafana, CloudWatch, OpenSearch) at organizational scale
- Strong experience with database infrastructure and data layer architecture (Aurora PostgreSQL, RDS, ElastiCache/Redis, OpenSearch, DynamoDB)
- Demonstrated ability to lead disaster recovery planning and design HA architecture patterns for mission-critical systems
- Excellent technical communication skills including ability to write clear ADRs and present to leadership
- Proven track record of mentoring engineers and elevating team capabilities
**Bonus qualifications:**
- 5+ years operating at senior or staff level
- Experience in financial services, fintech, broker-dealer, or heavily regulated industries with FINRA/SEC compliance requirements
- AWS certifications (Solutions Architect Professional, DevOps Engineer Professional, Security Specialty, Machine Learning Specialty)
- Experience with event streaming platforms at scale (Amazon MSK / Apache Kafka)
- Hands-on experience with API gateway management (Kong, AWS API Gateway)
- Demonstrated experience with AI/ML infrastructure (GPU compute, SageMaker/Bedrock, vector databases, MLOps pipelines, AIOps automation)
- Experience defining organizational rollout strategies for AI developer tools with governance frameworks
- Proficiency with policy-as-code frameworks (OPA/Rego, Sentinel, Kyverno)
- FinOps certification or demonstrated experience leading cloud cost optimization programs
- Experience authoring RFCs or ADRs that influenced engineering-wide decisions
- Open-source contributions, conference talks, or published technical writing