We are building a high-impact team at the intersection of CPU architecture, machine learning workloads, and system-level performance optimization. This role focuses on CPU softwarehardware co-design for next-generation QMX architectures, including workload characterization, simulation, kernel optimization, and driving architectural insights for future CPU designs. The ideal candidate will work across the full stack from ML models to low-level kernels to architectural feedback enabling efficient execution of ML workloads on CPU platforms.
Key Responsibilities
1. ML Workload Identification Characterization
- Identify and prioritize critical ML use cases and models for CPU-centric execution (LLMs, vision, speech, recommender systems, etc.)
- Analyze workload characteristics including:
- Compute intensity
- Memory bandwidth and cache behavior
- Parallelism and dataflow patterns
2. Simulation Trace Generation
- Generate detailed execution traces for ML workloads using QEMU or equivalent simulators
- Develop tooling to:
- Capture instruction-level execution behavior
- Extract performance counters and bottlenecks
- Enable accurate modeling of workload behavior for architectural exploration
3. Bottleneck Analysis Performance Optimization
- Identify system bottlenecks across:
- CPU pipelines
- Memory hierarchy
- Instruction utilization
- Optimize critical hotspots through:
- Kernel-level tuning
- Algorithmic improvements
- Data layout and memory optimizations
- Drive measurable improvements in workload performance
4. SoftwareHardware Co-Design
- Collaborate with CPU architecture and design teams to:
- Provide data-driven insights from real workloads
- Identify inefficiencies and propose architectural enhancements
5. ML Kernel Library Development (QMX Focus)
- Design and implement highly optimized ML kernels and libraries for QMX architecture
- Develop kernels for:
- GEMM, convolution, attention, activation functions, etc.
- Enable integration with:
- Open-source ML frameworks (e.g., PyTorch, ONNX, XNNPACK, MLAS)
- Apply advanced optimizations:
- SIMD/vectorization
- Cache-aware execution
- Parallel execution strategies
6. Benchmarking Performance Engineering
- Optimize CPU-centric ML benchmarks such as:
- Geekbench AI
- Internal benchmarking suites
- Establish performance baselines and track improvements across hardware generations
- Perform competitive analysis and performance positioning
Required Qualifications
- Strong background in:
- Computer Architecture / Systems Programming
- Machine Learning fundamentals
- Proficiency in:
- C/C++ (mandatory)
- Experience with:
- Performance profiling, benchmarking, and optimization
-
Are you interested in this position?
Apply by clicking on the “Apply Now” button below!
#AlbionarcJobs#FintechJobs
#AsiaJobs#MiddleEastCareers
#TechTalent#FintechRecruitment
#FinanceOpportunities#
