SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
OXIO is a NeoTelco platform enabling anyone to launch and scale their own carrier through a browser-based Carrier-as-a-Service offering called BrandVNO. The company operates a modern telecom network that helps users stay connected across carriers and geographies while reducing mobile data costs.
As a Site Reliability Engineer, you will design and implement cloud-based infrastructure to support OXIO's backend services, ensuring mission-critical production systems maintain maximum uptime. You'll automate technical operations including deployments, scaling, and recovery procedures, and participate in an on-call rotation with a blameless postmortem culture. Your work will directly enable Engineering, Telecom, and Data Engineering teams by providing them with robust operational tools and platforms.
Key responsibilities include:
- Designing and implementing cloud platforms for OXIO backend services
- Automating deployments, scaling, recovery, and other operational tasks
- Monitoring and maintaining production infrastructure for reliability
- Participating in on-call rotations and continuous improvement processes
- Enabling internal teams with operational tooling and infrastructure
Required qualifications include strong Linux/Unix systems knowledge (process management, filesystems, memory, networking), proficiency in Python, Go, or Ruby with strong Bash/Perl scripting skills, hands-on experience with infrastructure-as-code tools (Terraform, CloudFormation, Ansible), containerization (Docker) and orchestration (Kubernetes), monitoring platforms (Prometheus, Grafana, Datadog), CI/CD pipeline setup (Jenkins, GitLab CI, CircleCI), and cloud provider experience (AWS, GCP, Azure). You should understand TCP/IP, DNS, HTTP/HTTPS, load balancing, firewalls, and have practical incident management and on-call experience.
Nice-to-have skills include deployment strategies (canary, blue-green), high availability and failover mechanisms, IAM and zero-trust principles, distributed systems (Kafka, Cassandra, Elasticsearch), custom monitoring tools, database management (SQL/NoSQL), distributed tracing (Jaeger, OpenTelemetry), log aggregation (ELK, Splunk), performance profiling, load testing, and SaltStack configuration management.