HPC System Administrator Consultant at Louisiana State University
- Company: Louisiana State University
- Location: 0340 Fred C. Frey Computing Services Building
- Job type: full time
- Workplace: onsite
- Posted: 2026-05-26
Job description
All Job Postings will close at 12:01a.m. CST (1:01a.m. EST) on the specified Closing Date (if designated). If you close the browser or exit your application prior to submitting, the application progress will be saved as a draft. You will be able to access and complete the application through “My Draft Applications” located on your Candidate Home page. Job Posting Title: HPC System Administrator Consultant Position Type: Professional / Unclassified Department: LSUAM FA - ITS - TA - RETS - HPC - Systems (Timothy William Wright (00008005)) Work Location: 0340 Fred C. Frey Computing Services Building Pay Grade: Professional Job Description: This position is for a "hands-on" IT Consultant in the High Performance Computing group in the Information Technology Services Department at LSU. The HPC IT Consultant specializes in hardware architecture and advanced troubleshooting to support, optimize, and maintain research computing infrastructure. This role is responsible for enabling both existing and emerging high performance computing initiatives through direct technical support, training, and system hardware and software support. This role is designed for a technical expert who is equally comfortable troubleshooting physical hardware in the data center as they are writing complex automation scripts in a Red Hat Enterprise Linux (RHEL) environment. All Information Technology Services employees are expected to demonstrate a commitment to exemplary customer service in all facets of their work. Job Responsibilities: Operations: Expertise and leadership in verifying the quality of operations of Linux supercomputers, infrastructure systems, and other research computing systems. This includes, but is not limited to, performing daily system checks, analyzing system logs, troubleshooting hardware and software problems, monitoring and analyzing storage/infrastructure/job performance, helping users recognize job performance problems, writing scripts to enhance monitoring, and responding to unplanned system events such as power outages. This may require travel to various HPC sites to maintain physical installation of systems located off-site. Proactively perform hardware maintenance on the clusters, cluster infrastructure, and other systems as needed. This includes diagnosing and fixing problems which includes, but is not limited to, running diagnostics, re-seating dimms, replacing hard disks, calling vendors for RMA support, replacing mother boards, and return shipping replacement parts. Plan and perform software maintenance on both the clusters and the cluster infrastructure as needed. This includes, but is not limited to, installing operating systems, installing security patches, installing or upgrading drivers, upgrading firmware, installing or upgrading software licenses, installing or upgrading software specific to HPC cluster management. (50%) Research:Investigates, architects and implements new technology as appropriate to add new features to both the user environment and to our deployment environment. This requires the ability to work without training to take a new technology through installation to production. This also includes the ability to develop and document procedures related to that technology and to train other members of the group. (25%) Customer Support: Respond to tickets which include complaints, requests, troubleshooting, assessing storage options, etc. Provide training to groups or individuals as needed. (15%) Other duties as assigned. (10%) Minimum Qualifications: Bachelor's Degree with 3 years of experience (Ph.D. in Computation Science, Engineering or other computationally intensive disciplines substitutes for 2 year exp). Experience in IT systems administration in Linux/HPC environments. Strong knowledge of Linux/Unix operating systems. Expertise in scripting and programming in bash and other languages. Experience with HPC cluster resource managers and other management software such as Kickstart, DNF, RACADM, SSH keys,
HPC System Administrator Consultant on JobPost.