SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Bright Machines is an AI-enabled manufacturer producing data center infrastructure hardware for hyperscalers and OEMs. The company uses proprietary AI-based robotics and software to assemble servers at scale, addressing the surge in AI computing demand and supporting U.S. manufacturing reshoring initiatives.
As a Staff Platform Engineer, you will design, build, and operate scalable, secure infrastructure that powers the company's software-driven automation solutions. You'll drive monitoring, alerting, CI/CD pipelines, and developer experience tooling that accelerate development of complex, mission-critical industrial manufacturing applications. Working closely with Platform Engineering, R&D, and Operations teams, you'll shape and maintain the core platform supporting automation solutions deployed both on-premises in factories and across AWS and Azure cloud environments.
The technology stack spans on-prem compute and networking, AWS and Azure cloud, Terraform (Infrastructure as Code), ArgoCD, Kubernetes (self-hosted and managed), GitLab, and observability tools including Grafana, Prometheus, and OpenTelemetry. The integrated solution includes full network infrastructure (switches, routers, firewalling) and cloud-native software applications that operate and monitor robotic assembly lines.
Key responsibilities include: designing and maintaining reliable, scalable, secure infrastructure; owning and improving platform engineering tools and services; automating infrastructure provisioning and configuration management; managing CI/CD systems for efficient software delivery; implementing observability solutions for uptime and performance; enforcing security best practices and incident response readiness; and participating in on-call rotation for platform services.
Requirements: Bachelor's or Master's degree in Computer Science, Engineering, or related technical field; 5+ years of experience in Platform Engineering, DevOps, or Site Reliability Engineering (SRE); strong experience with physical infrastructure including networking equipment and servers; proficiency with CI/CD tools and GitOps workflows; hands-on coding in Python and Shell scripting; solid understanding of Linux internals and system administration; experience with public cloud providers (AWS, Azure, or GCP); proficiency with Infrastructure as Code tools (Terraform, Ansible); deep knowledge of Kubernetes (self-hosted and managed) including Helm; familiarity with observability stacks (Prometheus, Grafana, OpenTelemetry). Nice-to-have: experience with bare-metal provisioning (MaaS or similar).