SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Axiado is an AI-enhanced security processor company building silicon-rooted security and management chips for AI data center infrastructure. The company combines platform security, BMC/firmware, and on-chip AI for real-time threat detection and dynamic power/thermal management. Founded in 2017 with 100+ employees, Axiado recently closed a $100M+ Series C+ round and is scaling rapidly.
You will work across the full stack from model to silicon, optimizing training and inference performance on GPU and AI-accelerator infrastructure while adapting model and inference-engine design to underlying chip constraints. This role offers rare end-to-end exposure to the full AI silicon cycle—from algorithm through deployment—that most ML engineers at large companies never access.
Key responsibilities include:
- Optimize training and inference performance across GPU and AI-accelerator infrastructure, including MLOps pipelines
- Design, train, and evaluate ML models (deep learning, LLM, computer vision, or recommendation systems) and take them into production
- Harden and extend NPU cores (e.g., building on open RVV/tensor cores like CoralNPU) into production silicon
- Build or optimize inference engines and serving runtimes against real hardware constraints—latency, memory, and power
- Work below the application layer where needed: BMC firmware, embedded Linux, or RTOS (e.g., Zephyr) to ensure AI features run reliably on real systems
- Build automated test/verification harnesses that close the loop for AI-assisted RTL/DV, hardware bring-up, or manufacturing test
- Apply ML to security: AI-driven log/intrusion analysis, AI-assisted penetration testing, or firmware/hardware security work
- Collaborate closely with RTL/hardware, firmware, and QA teams to ship AI features end-to-end from training through deployment and monitoring
Requirements:
- 5–7+ years of hands-on AI/ML experience; Master's degree required, PhD preferred
- Hands-on experience with AI/ML infrastructure and performance: GPU clusters, distributed training, inference-serving optimization, MLOps pipelines
- Model/algorithm development experience: designing, training, and evaluating ML models
- Experience taking models into production: feature engineering, data pipelines, deployment
- AI chip/hardware-aware ML experience: optimizing inference engines for a specific chip, or adapting model architecture/quantization to chip constraints
- Deep, hands-on expertise in at least 2 of the following 5 specialty areas:
• NPU/AI-accelerator: hardening or extending an NPU core into production silicon, mapping models onto MAC/tensor-engine constraints, or NPU-aware RTL/DV work
• Systems/sys-level software: BMC firmware, embedded Linux, RTOS (e.g., Zephyr), or other low-level system software
• Inference engine/runtime: built or materially optimized an inference engine or serving runtime against real hardware constraints
• Test/verification harness: built an automated harness that closes a loop, e.g., an agent-driven RTL/DV test runner or hardware bring-up/MFG test harness
• Cybersecurity: AI-driven log/intrusion analysis, AI-assisted penetration testing, or firmware/hardware security