SlipstreamJobsFresh Startup & VC-Backed Jobs

Staff / Senior Software DevOps Engineer

Flux Computing - Toronto, ON, Canada - In-office

Apply on the company site

SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.

OLIX is building the DX-1 Decode Accelerator, a next-generation AI inference accelerator architected for decode workloads. The company is addressing a critical infrastructure gap in AI compute—the industry's hardware and power infrastructure cannot keep pace with AI demand, and existing designs have reached their limits. You will own the build, test, and CI infrastructure that the entire DX-1 software stack depends on: compiler, runtime, simulator, and framework integration. This is a large, complex test suite that must assert token-exact correctness against golden references while running across scarce, expensive resources spanning simulation compute and hardware-in-the-loop testing (simulator, emulator, DX-1 boards, and prototype platforms). Key responsibilities: - Design and own CI pipelines across PR, merge, and nightly lanes that gate the entire software stack, balancing speed, coverage, and cost - Scale test execution by splitting large suites into staged lanes, parallelizing with real isolation, and using content-addressed caching to maintain fast feedback as the suite and team grow - Manage a heterogeneous fleet of cloud and self-hosted machines, providing fair, monitored, fail-fast shared access to scarce hardware - Build performance and readiness signals using pinned hardware, rolling baselines, sound metric aggregation, and deterministic testing - Own software observability: stand up metrics stores, dashboards for CI health and product readiness that drive real decisions - Define standards for hermetic, reproducible builds and test runs; implement fail-closed defaults with blast radius control At the Staff/Principal level, your impact is measured by the velocity of every engineer depending on this system: how fast they get trustworthy signal, how rarely they wait on resources, and how much they can self-serve. You'll partner closely with infrastructure, compiler, runtime, simulator, and modelling teams. Required experience: - Demonstrated end-to-end ownership of large test or CI systems - Scaling large test suites through staged lanes, parallelism, content-addressed caching, and cost-effective feedback loops - Managing heterogeneous CI runner fleets across cloud and on-prem (VMs, containers, bare-metal) - Providing teams shared, monitored access to scarce resources (custom accelerators, FPGAs, prototype rigs, lab hardware) - Supporting performance analysis with trustworthy regression baselines, deterministic testing, and sound metric aggregation - Building observability and metrics platforms (CI health dashboards, high-cardinality time-series stores) - Strong scripting and systems programming (Python + systems language), containers, Linux, cloud infrastructure (AWS) - Reproducible, safe-by-default mindset: hermetic builds, least privilege, fail-closed defaults - Bachelor's degree in Computer Science, Electrical Engineering, Mathematics, or related field Nice to have: GitHub Actions at scale, AWS CI runner fleet scaling, hardware-in-the-loop/lab automation for custom silicon or FPGA bring-up, time-series and observability stacks (Prometheus/Grafana, Datadog, columnar warehouses), HPC/cluster batch scheduling, release engineering, or developer-productivity platforms. Compensation: Competitive salary commensurate with experience, skills, and location; meaningful stock options.

Similar roles