We are seeking an experienced AI Validation Engineer to lead validation, benchmarking, and quality assurance activities for AI/ML software stacks running on embedded and heterogeneous computing platforms. The ideal candidate will possess strong expertise in AI frameworks, ROCm ecosystems, Linux-based environments, performance analysis, and automation. This role will drive end-to-end AI pipeline validation while collaborating closely with architecture, compiler, runtime, driver, and hardware teams to ensure production-quality AI solutions.
Key Responsibilities :
AI/ML Validation & Quality Ownership :
– Lead validation efforts for complex AI/ML compute stacks across multiple hardware and software platforms.
– Define validation strategies, test plans, methodologies, and quality metrics for AI software and system pipelines.
– Own the complete defect lifecycle, including issue reporting, triage, root-cause analysis, tracking, and closure.
– Ensure comprehensive coverage across functional, performance, regression, stress, scalability, and reliability testing.
End-to-End AI Pipeline Validation :
– Validate complete AI workflows across training, optimization, and inference pipelines.
– Validate ROCm libraries and AI software stack functionality.
– Verify :
1. Model training, conversion, and optimization workflows (e.g., PyTorch to ONNX)
2. Inference runtimes such as ONNX Runtime, TensorRT, ROCm/HIP, and OpenVINO
3. AI compilers and toolchains including TVM, Vitis AI, XDNA, and XLA
4. Kernel execution, memory movement, inference correctness, and accuracy
– Validate AI workload stability, performance, and correctness on Ubuntu and Yocto-based Linux platforms.
AI Benchmarking, Profiling & Performance Optimization :
– Define and execute benchmarking strategies for AI training and inference workloads.
– Profile AI models to identify compute, memory, throughput, and latency bottlenecks.
– Collaborate with compiler, runtime, and hardware teams to drive system-level and model-level optimizations.
– Validate performance improvements across :
1. Model architectures
2. Batch sizes
3. Precision modes (FP32, FP16, INT8)
4. Execution paths and hardware configurations
– Ensure performance regressions are detected early and release performance targets are consistently achieved.
AI Framework & Compute Stack Validation :
– Validate functionality, integration, and performance of AI frameworks including :
1. PyTorch
2. TensorFlow
3. ONNX Runtime
– Execute and validate workloads across heterogeneous compute environments utilizing :
1. ROCm/HIP
2. CUDA
3. OpenCL
4. AI accelerators
– Analyze the impact of framework, compiler, and runtime changes on real-world AI workloads.
Automation & Tool Development :
– Design and develop Python-based validation, benchmarking, and profiling frameworks.
– Build reusable automation for :
1. Test execution
2. Benchmarking
3. Performance profiling
4. Result analysis
5. Reporting and dashboards
– Continuously improve validation efficiency, scalability, and coverage through automation.
Technical Leadership :
– Provide technical leadership and mentorship to validation engineers and junior team members.
– Partner with architecture, compiler, runtime, driver, and hardware teams to resolve functional and performance issues.
– Collaborate effectively with globally distributed cross-functional teams.
– Present validation status, benchmarking results, quality metrics, and performance risks to stakeholders.
Required Skills & Qualifications :
Technical Expertise :
– 8-12 years of experience in AI/ML validation, performance analysis, or software quality engineering.
– Strong understanding of :
1. Deep Learning
2. Large Language Models (LLMs)
3. Recommender Systems
– Strong hands-on experience with ROCm technologies and ROCm stack validation.
– Experience validating AI/ML compute stacks including :
1. HIP
2. CUDA
3. OpenCL
4. OpenVINO
5. PyTorch and TensorFlow ecosystems
– Expertise in end-to-end AI pipeline validation including :
1. Model conversion
2. Inference runtimes
3. AI compilers
4. Kernel execution
5. Accuracy validation
– Advanced Python
– Strong experience in AI benchmarking, profiling, and performance optimization.
– Deep understanding of Linux environments, particularly Ubuntu and Yocto.
Are you interested in this position?
Apply by clicking on the “Apply Now” button below!
#AlbionarcJobs#FintechJobs
#AsiaJobs#MiddleEastCareers
#TechTalent#FintechRecruitment
#FinanceOpportunities#
