AI Kernel Optimization Engineer

Engineer

AI Kernel Optimization Engineer

Apply Now

- $0.00

  • Date posted
    July 24, 2026
  • Expiration date
    October 24, 2026
  • Application ends
    October 24, 2026

Our Client Currently looking for AI Kernel Optimization Engineer,
You will design, implement, and optimize AI compute kernels (Gen AI Large Language Model, AI Vision, CNNs, etc) and runtime components to fully exploit the underlying hardware architecture – from vector/matrix units and memory hierarchies down to the assembly level.
Your work will directly influence how efficiently AI models run on SoCs, shaping the performance of next-generation inference accelerators. You will collaborate closely with hardware architects, compiler engineers, and AI framework developers to achieve optimal hardware–software co-design.
Your main responsibilities will include:
Leading and contributing to:

  • Develop, optimize, profile, and debug AI compute-intensive kernels (e.g., GEMM, attention, activations) targeting RISC-V architectures.
  • Identify and resolve performance bottlenecks at the ISA, compiler, and runtime levels.
  • Collaborate with hardware and architecture teams to influence design decisions and improve real-world AI performance.
  • Contribute to the development and optimization of AI runtime and graph execution engines.
  • Evaluate, benchmark, and optimize AI inference workloads on platforms.
  • Develop performance analysis tools and automation scripts for profiling, validation, and performance optimization.
  • Work with AI frameworks (e.g., vLLM, SGLang, PyTorch, TensorRT-LLM) to ensure efficient mapping to targets.
  • Stay up to date with AI kernel optimization trends, emerging hardware acceleration techniques, and open-source developments.
  • Share technical expertise and contribute to knowledge sharing and continuous improvement within the team.

Technical skills:

  • Strong background in low-level performance optimization (vectorization, memory access optimization, loop unrolling, instruction scheduling, data-tiling, etc.).
  • Proficiency in C/C++ and good understanding of assembly-level optimizations (SIMD, intrinsics, compiler flags).
  • Solid understanding of CPU/GPU/AI accelerator architecture (pipelines, caches, memory hierarchies, compute units).
  • Experience with profiling and performance analysis tools (perf, VTune, nvprof, etc.).
  • Strong knowledge of parallel programming (SIMD, multithreading, OpenMP, CUDA, or similar).
  • Solid software engineering skills (version control, CI/CD, testing).
  • Experience with RISC-V architectures or other custom ISAs is a plus.

Nice to have:

  • Experience with AI inference workloads or libraries (e.g., BLAS, cuDNN, oneDNN, TVM, or similar).
  • Familiarity with MLIR/LLVM or other compiler infrastructures.
  • Contributions to open-source AI inference engines or kernel libraries.
  • Understanding of NUMA architectures or heterogeneous computing.
  • Experience with quantization and mixed precision inference.
  • Are you interested in this position?

     

    Apply by clicking on the “Apply Now” button below!

     

    #AlbionarcJobs#FintechJobs

    #AsiaJobs#MiddleEastCareers

    #TechTalent#FintechRecruitment

    #FinanceOpportunities#

     

Apply Now

- $0.00

Select your currency