SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Salary: USD 150,000 - 170,000 / annual
DriveWealth is a global B2B fintech platform democratizing access to financial markets through an API-based infrastructure. The company enables partners to offer seamless investing and trading experiences worldwide, supporting US equities, mutual funds, ETFs, fixed income, and options.
As Manager of Site Reliability Engineering, you will lead a team of SRE Automation Engineers while maintaining hands-on technical authority for the Brokerage-as-a-Service platform. This is a dual-responsibility role: you will drive the automation agenda to eliminate manual toil while also building and developing your engineering team.
Key responsibilities include:
• Team Leadership & Development: Manage and mentor SRE Automation Engineers, set technical direction, conduct 1:1s, own performance management and career development, and manage the team's Jira board.
• Engineering & Automation: Lead design and development of internal tooling and automation using Rundeck and Airflow to eliminate repetitive work and improve developer velocity. Remain hands-on with complex, high-leverage automation work.
• SRE Practice & Governance: Adapt Google's SRE principles (SLIs, SLOs, error budgets, blameless postmortems) to the regulated brokerage environment.
• Infrastructure as Code: Set architectural standards for modular, reusable IaC using Terraform and oversee GitOps workflows via ArgoCD.
• Platform Governance: Review software architecture and Kubernetes metrics to ensure high availability, capacity planning, and cost optimization across AWS regions.
• Incident Engineering: Lead incident response for critical events, drive root-cause analysis, and champion a blameless post-mortem culture.
• Collaboration & Stakeholder Management: Partner with engineering leadership to align SRE priorities with business goals and foster adoption of new tools and reliability best practices.
The role includes participation in 24/7 on-call rotations supporting global operations, with the primary mission of building systems that make manual intervention obsolete.
REQUIREMENTS:
• Prior experience managing or leading SRE/DevOps engineers, ideally in fintech or highly regulated environments. Ability to flex between hands-on principal-level engineering and coaching/developing a team.
• Working knowledge of Google's SRE practices (SLIs/SLOs, error budgets, toil reduction, blameless postmortems) and experience adapting them to regulated environments.
• Proficiency in Linux administration with deep understanding of TCP/IP stack, OSI model, DNS, and network troubleshooting.
• Experience in highly regulated financial environments or with FIX/API connectivity.
• Hands-on experience managing production-grade Kubernetes clusters, including RBAC, autoscaling, Helm, and multi-cluster patterns.
• Strong grasp of AWS core services, security, and high-availability patterns. Proficiency with boto3 and AWS CLI for automation.
• Experience building secure, automated delivery pipelines and operating GitOps workflows (ArgoCD).
• Strong scripting and development skills in Python or Golang, along with Bash and Ansible.
• Experience with Grafana or similar tools, Prometheus, logs shipping/management, and metric-first alerting.
• Experience with secrets management, vulnerability scanning, and securing the software supply chain.
• Familiarity with using LLMs, Public MCPs, or Bedrock Agent Core to enhance SRE workflows.
• Hands-on experience with Rundeck and Airflow for job orchestration and automation, plus experience managing Kafka, MQ, or SQS.
• Must be authorized to work in the United States. No visa sponsorship available.