SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
d-Matrix is seeking a Principal Software Engineer to lead performance analysis and modeling efforts for AI inference accelerators. This hands-on technical contributor role focuses on emerging hardware technologies (DIMC, D2D, 3D-DRAM) and cutting-edge workloads (generative inference, multi-modal LLMs, video/audio generation). You will build and maintain analytical performance models, develop architecture simulators, and translate workload analysis into concrete modeling inputs that help the architecture team project performance across current and future d-Matrix silicon.
Key responsibilities include analyzing emerging ML workloads to identify performance-relevant properties, building and extending architecture simulators to support HW/SW feature analysis, partnering with hardware design, compiler, inference server, and kernel teams to validate assumptions and surface improvement opportunities, tracking ML architecture research, and proposing targeted HW/SW optimizations based on modeling results. You will document methodology and findings for reuse across the architecture organization.
Required qualifications: BSEE with 6+ years of industry experience or MSEE with 4+ years. You must have working knowledge of computer architecture, HW/SW co-design, performance modeling, and ML fundamentals (particularly DNNs), with programming fluency in C/C++ or Python. Experience building or working with analytical performance models or architecture simulators is essential. You should be self-motivated, collaborative, and comfortable working across hardware and software teams.
Preferred qualifications include experience optimizing AI/ML workloads on accelerator technologies and research/investigation in AI/ML architecture and microarchitecture. The role is based in Santa Clara, CA, with 3 days per week onsite and hybrid flexibility. Remote work within the United States is also considered.