HPC Fleet Reliability Engineer – GPU Clusters

Company: Dormont Manufacturing Co
Apply for the HPC Fleet Reliability Engineer – GPU Clusters
Location: London
Job Description:

Dormont Manufacturing Co is seeking a Fleet Reliability Operations team member to oversee the management and uptime of supercomputing clusters. The successful candidate will configure and troubleshoot issues in a fast-paced environment, ensuring optimal performance of our systems and infrastructure.

This role requires a bachelor’s degree and at least 2 years of experience in troubleshooting and maintaining data center infrastructure, ideally in a Linux environment. The position offers various benefits including medical and dental insurance, and pension contributions.

#J-18808-Ljbffr…

Posted: July 19th, 2026