SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
OpenAI's compute infrastructure team operates a large, modern GPU fleet and unified platform supporting production AI workloads and training for next-generation models. This role focuses on hands-on network activation and Layer 1 operations across OpenAI's distributed data-center footprint.
You will own the end-to-end activation and troubleshooting of WAN, fiber, carrier, and cloud-interconnect circuits. This includes reconciling complete circuit documentation (circuit IDs, LOAs, carrier demarcations, patch-panel positions, fiber pairs, optics, and device ports), investigating optical faults (no-light, low-light, wrong-port, link-flap, error-rate issues), and guiding remote hands and carrier technicians through targeted diagnostics using VFL, optical power meters, OTDR/OLTS, and DOM/DDM testing.
Key responsibilities include verifying service acceptance on both A-side and Z-side, owning tickets through resolution, executing safe change work with proper approvals, and delivering complete production handoffs with verified mappings and test results. You will also build automation workflows that translate system output and LLM-assisted diagnostics into precise, approved technician actions while maintaining human oversight for intrusive work.
The ideal candidate combines strong physical-networking judgment with practical automation skills: understanding of single-mode/multimode fiber, LC and MPO/MTP connectors, 100G/400G/800G optics, wavelength and power budgets, and experience with carriers, colocation providers, AWS Direct Connect, Azure ExpressRoute, or Google Cloud Interconnect. You should be comfortable scripting, working with APIs, and structuring operational data to drive reliable circuit bring-up and troubleshooting at scale.