SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Salary: USD 200,000 - 350,000 / annual
Edison Scientific builds and deploys AI scientist agents to accelerate science and the development of new medicines. As Principal Member of Technical Staff for Core Infrastructure, you will design, scale, and operate the platform infrastructure powering autonomous scientific discovery.
Your primary focus will be orchestration of AI agents at scale—building and managing Kubernetes clusters that support thousands of persistent, stateful workloads. You will develop custom resource definitions (CRDs) and operators, ensure reliability and efficiency of the compute layer, and establish infrastructure best practices across the organization.
Key responsibilities include:
- Architect and operate Kubernetes clusters supporting thousands of concurrent persistent resources with high availability and efficient resource utilization
- Design and develop CRDs and Kubernetes operators to manage domain-specific workloads such as AI agent lifecycles, research pipelines, and long-running compute tasks
- Drive cluster scaling strategy, node pool management, autoscaling policies, and resource quota frameworks
- Build and maintain infrastructure-as-code (Terraform, Pulumi, or similar) for reproducible environment management
- Design scheduling, placement, and affinity strategies to optimize cost, performance, and fault tolerance for heterogeneous workloads (CPU, GPU, memory-intensive)
- Establish observability, monitoring, alerting, and incident response best practices (Prometheus, Grafana, Datadog, or similar)
- Own storage and networking strategy within Kubernetes—persistent volumes, CSI drivers, service mesh, network policies, and ingress architecture
- Troubleshoot complex cross-system infrastructure issues and guide others through debugging in distributed environments
- Collaborate with backend, ML, and research teams to translate workload requirements into reliable infrastructure patterns
This is a senior technical leadership role focused on technical ownership and leverage—understanding how complex systems interact, making sound architectural tradeoffs, and building foundations that enable teams and science to move faster.
REQUIREMENTS:
- 10+ years of professional infrastructure or platform engineering experience with deep hands-on Kubernetes expertise in production environments
- Experience designing and implementing CRDs and Kubernetes operators (Kubebuilder, Operator SDK, controller-runtime)
- Track record operating and scaling Kubernetes clusters supporting thousands of persistent or long-lived resources
- Deep understanding of Kubernetes internals (API server, etcd, scheduler, controller manager, kubelet) and behavior at scale
- Expertise with cloud infrastructure (AWS EKS, GCP GKE, or Azure AKS) and associated networking, storage, and IAM primitives
- Proficiency in at least one systems or backend language for operator development and infrastructure tooling
- Hands-on experience with infrastructure-as-code tools (Terraform, Pulumi, or Crossplane) and GitOps workflows
- Strong working knowledge of container networking (CNI plugins, service mesh, network policies), storage (CSI, persistent volumes, StatefulSets), and security (RBAC, Pod Security Standards, secrets management)
- Ability to operate autonomously, make sound technical judgments, and drive projects from concept through production
PREFERRED:
- Experience with data-intensive platforms, scientific computing, or ML/AI infrastructure
- Prior experience in startups or small teams with significant architectural ownership and ambiguity
- Experience scaling systems, teams, or platforms through periods of rapid growth