SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Salary: USD 150,000 - 200,000 / annual
Quadric is a growth-stage semiconductor IP company pioneering General Purpose Neural Processing Units (GPNPUs) that enable both neural network inference and conventional C++ code execution on a single programmable architecture. Founded by technologists from MIT and Carnegie Mellon, the company serves automotive, industrial, robotics, and embedded systems markets.
As an AI Performance Modeling Engineer, you will build analytical, cycle-level performance models of AI inference workloads in Python before silicon exists. These models directly guide critical hardware decisions on lane bindings, tensor placement, and architecture trade-offs.
Key responsibilities include:
- Developing cycle-level Python models of AI inference workloads on next-generation GPNPU hardware
- Deriving from first principles which hardware lanes operations bind to (compute, memory bandwidth, interconnect) and modeling software pipelining overlaps
- Modeling tensor placement, tiling across processing elements, local memory residency, and data movement across memory hierarchies
- Incorporating architectural details for vision networks and Large Language Models, including operator mix, sparsity, routing, and quantization
- Calibrating performance models against instruction-set simulators and profiling traces to meet accuracy targets
- Writing and defending technical studies that inform architecture and product decisions
- Balancing single-stream latency against scaled throughput performance
Within 6–12 months, successful candidates will own full workload models end-to-end, calibrated against simulation and trusted by engineering teams. Models should predict workload behavior within 10–15% accuracy, and you'll publish technical studies whose conclusions directly shape product decisions.
Required qualifications: Strong Python skills with experience writing and validating quantitative models; solid grasp of memory hierarchies, bandwidth/latency trade-offs, pipelining, and execution bottlenecks; comfort with technical writing; deep understanding of neural network inference operators (Transformers, attention, MoE) or proven performance modeling experience in quantitative domains; BS/MS/PhD in Computer Science, Electrical Engineering, or equivalent practical experience.
Preferred: GPU/AI accelerator experience, CUDA/Triton kernels, roofline analysis, architecture simulators (gem5, Timeloop, MAESTRO), compiler internals, or published performance studies.