Machine Learning & AI Infrastructure Engineer

Engineer

Machine Learning & AI Infrastructure Engineer

Apply Now

- 0.00

  • Date posted
    August 12, 2026
  • Expiration date
    November 12, 2026
  • Application ends
    November 12, 2026

 

 

Our Client is seeking a Machine Learning & AI Infrastructure Engineer. This is not a traditional AI Engineer or Data Scientist role. The hiring team is specifically seeking a unique blend of: HPC Administrator + Kubernetes Administrator + AI Infrastructure Operations Engineer. Candidates who have owned, operated, supported, and troubleshot production AI or HPC environments will be the strongest fit. Experience administering and maintaining systems is significantly more important than architecture-only experience.

Responsibilities:

  • Administer and support AI and HPC cluster environments
  • Manage day-to-day operations of large-scale compute infrastructure
  • Deploy, maintain, and troubleshoot Kubernetes-based platforms
  • Ensure reliability, performance, and scalability across compute, storage, and networking environments
  • Support AI model training and inference infrastructure
  • Automate operational processes through scripting and tooling
  • Partner with engineering teams and customers to optimize platform performance
  • Troubleshoot complex infrastructure, networking, storage, and containerization issues
  • Support both internal platforms and customer-facing environments

Why Consider This Opportunity?

  • 100% Remote Environment
  • Exposure to cutting-edge AI, GenAI, and HPC technologies
  • Flat organizational structure with minimal bureaucracy
  • Direct impact on strategic technology initiatives
  • Opportunity to work on platforms that support healthcare, research, drug discovery, and other meaningful AI-driven innovations
  • High visibility and collaboration with industry-leading technical teams

Compensation & Benefits:

  • Base Salary: $175,000 – $200,000+
  • Annual Bonus: Typically 10%-15%
  • Medical, Dental, and Vision Coverage
  • 401(k)
  • Additional performance-based incentives
Requirements
  • Strong experience administering High Performance Computing (HPC) environments
  • Experience with AI cluster administration and infrastructure operations
  • Hands-on Kubernetes administration experience in on-premises environments
  • Experience provisioning and managing PV/PVC storage through Kubernetes CSI drivers
  • Strong Linux administration skills, specifically Ubuntu
  • Scripting experience with Bash and/or Python
  • Proven troubleshooting and operational support experience
  • Ability to manage and maintain production infrastructure environments

Experience with one or more of the following:

  • Dell PowerScale/Isilon
  • VAST Storage
  • NetApp ONTAP
  • DDN IntelliFlash
  • DDN Exascaler
  • Lustre Parallel File Systems

Successful candidates may come from organizations focused on:

  • AI Infrastructure
  • Machine Learning Platforms
  • HPC Operations
  • Research Computing
  • Biotechnology
  • Academic Medical Centers
  • Digital Biology
  • Financial Services AI Platforms
  • Automotive AI Initiatives
  • Large-Scale Data Science Environments

Preferred:

  • NVIDIA ecosystem experience
  • NVIDIA Base Command Manager (BCM)
  • Bright Cluster Manager
  • MLOps platform exposure
  • Containerization technologies and orchestration platforms
  • High-performance networking experience
  • RDMA technologies
  • InfiniBand networking
  • NVIDIA UFM
  • Parallel file system administration
  • Storage Technologies (highly desired)
  • Are you interested in this position?

     

    Apply by clicking on the “Apply Now” button below!

     

    #AlbionarcJobs#FintechJobs

    #AsiaJobs#MiddleEastCareers

    #TechTalent#FintechRecruitment

    #FinanceOpportunities#

     

Apply Now

- 0.00

Select your currency