SlipstreamJobsFresh Startup & VC-Backed Jobs

Sr. Software Engineer - Kubernetes

Armada - Thiruvananthapuram, Kerala, India - In-office - posted 2025-04-06

Apply on the company site

SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.

Armada is a hyperscaler for edge AI infrastructure, backed by ~$500M in funding from Founders Fund, Lux, BlackRock, and Microsoft. The company delivers modular AI infrastructure for sovereign and edge computing, deployed across 60+ countries for energy, defense, and other mission-critical sectors. Strategic partnerships include Microsoft, Dell, Palantir, NVIDIA, SpaceX, and Skydio. This Senior Platform Engineer role sits on the Edge Platform team, combining systems engineering, software development, and cloud-native infrastructure. You will design, build, and operate the software and platform capabilities powering Armada's distributed edge computing infrastructure across Galleon mobile data centers and Commander cloud services. Key responsibilities include: **Platform Engineering & Software Development**: Design and develop platform services, controllers, operators, and automation frameworks in Go and Python. Build internal tools and APIs for provisioning, lifecycle management, observability, and operations. Develop software to automate infrastructure workflows, hardware lifecycle management, cluster operations, and fleet-wide orchestration. Contribute to platform architecture and engineering best practices. Design self-healing and autonomous operational capabilities. **Linux Systems Engineering**: Debug and optimize Linux systems across compute, storage, networking, and container workloads. Investigate complex performance issues involving CPU scheduling, memory management, I/O subsystems, networking, filesystems, and containers. Analyze kernel and userspace behavior using perf, eBPF, strace, tcpdump, bpftrace, and profiling utilities. Drive platform reliability through deep Linux internals knowledge. Participate in root cause analysis of production incidents. **Kubernetes & Cloud Platform**: Architect, deploy, and manage highly available Kubernetes environments across edge and cloud infrastructure. Build and maintain Kubernetes operators, controllers, CRDs, admission webhooks, and platform extensions. Implement scalable networking, storage, security, and observability solutions. Optimize cluster performance and resource utilization in resource-constrained edge environments. Drive Infrastructure-as-Code using Terraform, Ansible, Helm, and GitOps. **Reliability & Operations**: Design and maintain observability platforms using Prometheus, Grafana, OpenTelemetry, and centralized logging. Establish operational excellence through automation, monitoring, incident response, and postmortem processes. Collaborate across software, infrastructure, security, and product teams. Participate in on-call rotations and improve operational maturity.

Similar roles