HPC Systems Administrator
Apply
Apply for this job directly on SHORTList.
Referral
Share your custom referral link for this job with qualified candidates. Earn the referral you lead to a hire.
Helsing operates on-premises high-performance computing (HPC) infrastructure that supports electromagnetics, computational fluid dynamics, and multi-physics simulation. As an HPC Systems Administrator based in Munich, you will take ownership of this critical environment, ensuring that our team of simulation engineers remains unblocked, productive, and equipped to solve complex problems. You will play a vital role in maintaining rigorous technical standards, optimising compute resources, and scaling our infrastructure to support continuous, large-scale modelling. The role is based on-site in Munich with regular travel to our Tussenhausen site.
The day-to-day
- Own the day-to-day administration of compute nodes, workload schedulers, parallel storage, high-speed interconnects, and licence servers
- Ensure the environment remains highly available and consistently performant through proactive monitoring, patching, firmware updates, and incident response
- Administer the workload scheduler (Slurm, PBS Pro, or similar), managing queues, fair-share policies, accounting, and quotas to optimise resource utilisation
- Manage the simulation software stack and user environments using tools such as Lmod, Spack, or EasyBuild
- Collaborate with hardware and software vendors to resolve support cases, process RMAs, and ensure upgrade quality
- Automate operational workflows using Bash, Python, and Ansible to improve system efficiency and reduce manual intervention
- Maintain the strict security posture required for cleared work and support ongoing compliance reviews
- Onboard users and maintain comprehensive documentation to empower engineers to self-serve
You should apply if you
- have administered Linux systems within a production HPC or large shared compute environment
- have hands-on experience managing workload schedulers such as Slurm, PBS Pro, LSF, or similar
- possess production experience with parallel filesystems (Lustre, BeeGFS, or GPFS) and high-speed interconnects (InfiniBand or RoCE)
- are capable of scripting and automating complex workflows with Bash, Python.
Apply
Apply for this job directly on SHORTList.
Referral
Share your custom referral link for this job with qualified candidates. Earn the referral you lead to a hire.