Our Client is immediately seeking an experienced High Performance Compute Administrator in support of our enterprise storage and compute customer .
Responsibilities:
- On-site technical support for the HPC environment (SGI 8600 compute systems and supporting infrastructure)
- Administration and operational support for HPCM cluster management
- Support for the DMF storage environment, including coordinating upgrades and system health monitoring, working with DMF team.
- CDU cooling system maintenance and troubleshooting
- Hardware diagnostics and coordination with the site admin to maintain system stability
- Supporting the site admin
- Facilitation of weekly operational meetings (Monday’s 12PM) and project coordination- PM role
Requirements
- Bachelor’s degree in Computer Science, Information Technology, Engineering, or related field; equivalent experience considered
- 5+ years of experience supporting enterprise High-Performance Computing (HPC) environments
- Hands-on experience with Cray EX, SGI 8600, or similar large-scale platforms
- Experience administering and supporting Performance Cluster Manage and Linux-based environments (RHEL/SLES)
- Knowledge of Data Management Framework (DMF), storage lifecycle management, system monitoring, and operational support
- Experience supporting storage, compute, and networking infrastructure, including Lustre, InfiniBand, and management systems
- Ability to perform hardware diagnostics, troubleshooting, root cause analysis, and coordinate repairs with vendors and data center personnel
- Experience maintaining and troubleshooting Coolant Distribution Units (CDUs) and related cooling infrastructure
- Knowledge of system administration, firmware management, health monitoring, and upgrade coordination
- Experience supporting site administrators and collaborating with cross-functional infrastructure, storage, and networking teams
- Strong project coordination skills, including meeting facilitation, action item tracking, status reporting, risk management, and stakeholder communication
- Ability to lead recurring operational meetings and coordinate technical activities across multiple teams
- Strong verbal and written communication skills with the ability to interact effectively with technical and non-technical stakeholders
Preferred:
- HPE, Linux, or HPC-related certifications
- Experience with Cray EX Supercomputing environments
- Experience with automation and scripting using Python, Bash, or Ansible
- Familiarity with Slurm workload management
- Experience supporting government, research, scientific, or mission-critical computing environments
-
Are you interested in this position?
Apply by clicking on the “Apply Now” button below!
#AlbionarcJobs#FintechJobs
#AsiaJobs#MiddleEastCareers
#TechTalent#FintechRecruitment
#FinanceOpportunities#
