SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Armada is a hyperscaler for edge AI infrastructure, backed by ~$500M in funding from Founders Fund, Lux, BlackRock, and Microsoft. The company delivers modular AI infrastructure for sovereign and edge computing, deployed across 60+ countries for energy, defense, and other critical sectors.
As a NOC / Data Center Operations Engineer, you will be responsible for monitoring, operating, and maintaining the reliability of Armada's critical physical infrastructure and distributed edge data center environments. This is a Tier 2 operations role with hands-on technical depth and incident leadership responsibilities.
Key responsibilities include:
**Infrastructure Monitoring & Incident Response:**
- Monitor critical physical infrastructure using PLC, BMS, and DCIM platforms (Distech, Radix IoT/Mango, Schneider Electric, Siemens, or similar).
- Respond to alerts involving power, cooling, network connectivity, environmental conditions, and facility systems.
- Serve as Tier 2 escalation point for L1 technicians; coordinate escalations to engineering, facilities, vendors, and other teams.
- Own incidents through their full lifecycle: triage, troubleshooting, resolution, stakeholder communication, and documentation.
- Support post-incident reviews and identify opportunities to improve operational reliability.
- Develop and refine monitoring dashboards, alerting thresholds, and operational health indicators.
**Data Center Operations:**
- Perform and coordinate routine health checks of UPS systems, PDUs, CRAC/CRAH units, backup generators, and environmental monitoring systems.
- Troubleshoot and coordinate resolution of mechanical and electrical infrastructure issues using MEP systems knowledge.
- Read and interpret technical documentation including electrical one-line diagrams, network diagrams, schematics, and equipment specs.
- Coordinate scheduled and emergency maintenance with internal teams, vendors, and remote hands personnel.
- Participate in change management and assess operational/infrastructure impact of proposed changes.
- Ensure physical and logical security policies are followed in data center and edge environments.
**Modular & Edge Infrastructure:**
- Operate and support modular, containerized, micro, and distributed edge data center environments.
- Maintain availability, continuity, and resiliency across geographically distributed and remotely operated infrastructure.
- Coordinate remote troubleshooting and hands-on support for edge deployments.
- Implement and continuously improve operational practices for remote infrastructure monitoring, fault detection, escalation, and recovery.
- Support environmental sensors and IoT-based monitoring integrations across remote infrastructure.
Ideal candidates bring hands-on experience with data center operations, PLC/BMS/DCIM platforms, mechanical and electrical infrastructure, incident management, and remote or edge infrastructure.