HPC Assistant Director

Apply Now

How to Apply

A cover letter and resume are important submissions for the hiring team to get a sense of your experience. In the cover letter, in one page or less, please let us know how this role aligns with your career aspirations and skills. Submit both a cover letter and resume as one file.

Salary will be commensurate with the selected candidate's qualifications, experience, and education.
 

Job Summary

Information and Technology Services (ITS) at the University of Michigan has an exciting opportunity for an HPC Assistant Director to lead the people, services, and systems that support large-scale research computing across the University, including regulated and researcher-owned environments. This is primarily a managerial and service-leadership role, accountable for team development, strategic and operational planning, and the availability, reliability, security, performance, capacity, and lifecycle of HPC services. The Assistant Director will work in close partnership with the HPC Technical Lead, who serves as the principal technical advisor and leads HPC architecture and technical direction. The successful candidate does not need to be the team's foremost technical expert but must have sufficient understanding of Linux, HPC clusters, networking, and storage to evaluate recommendations, ask informed questions, weigh risks and tradeoffs, and make sound operational and organizational decisions. This position also leads major incident response, establishes operational priorities and policies, defines requirements for significant procurements, and builds partnerships across ITS and the University. The successful candidate will bring a strong commitment to developing staff and enabling research, scholarship, and discovery.

The HPC Assistant Director will report to the Director of ITS-Advanced Research Computing.
 

Responsibilities*

You will provide strategic and operational leadership for the team responsible for designing, building, and operating the University's high-performance computing (HPC) environments. Key responsibilities include:

  • Collaborate with the broader Advanced Research Computing (ARC) leadership team to develop and execute University-wide research computing strategies, service roadmaps, and investment priorities.
  • Lead, manage, and develop a team of HPC and Linux professionals by setting clear expectations, providing coaching and feedback, supporting professional growth, and fostering an inclusive, collaborative, and accountable team culture.
  • Plan and prioritize the team's work across infrastructure projects, system maintenance, security, technical debt, and operational support; allocate resources, manage dependencies and risks, and ensure reliable delivery aligned with ARC strategy and researcher needs.
  • Develop and implement operational policies, standards, automation, metrics, and service-management practices that improve reliability, sustainability, user experience, and efficient use of University resources. Use meaningful service and capacity metrics to inform operational decisions and long-term investments.
  • Own the operational health and lifecycle of ARC's HPC services, establishing expectations for availability, reliability, security, performance, capacity, supportability, observability, continuity, and disaster recovery.
  • Lead the response to significant service disruptions, security events, and other production issues by coordinating technical teams, establishing priorities, assessing risk and service impact, communicating with stakeholders, and ensuring remediation and follow-through in accordance with the ITS major incident process.
  • Engage researchers, other ITS units (e.g. data centers and networking), unit IT organizations, and the broader research IT community to understand evolving needs, translate requirements into technical and service capabilities, and establish clear expectations about feasible solutions.
  • Develop sustainable financial and service-delivery models for HPC services, including cost projections, capacity planning, recharge rates, lifecycle investments, and funding strategies, in collaboration with the ARC director.
  • Build strong relationships with peers across the One ITS ecosystem to advance shared priorities, foster collaboration, and contribute to ITS areas of focus.
  • Monitor trends in computing technologies, research workflows, cybersecurity, data management, and artificial intelligence to guide the evolution of HPC services and architecture.
  • Represent the University in the national research computing ecosystem by engaging with other universities, national laboratories, professional communities, vendors, and software developers.
  • Apply insights from external partnerships and industry developments to identify emerging opportunities, anticipate shifts in research computing, and adopt relevant best practices.
  • Communicate strategy, priorities, risks, and resource needs effectively to technical teams, researchers, institutional leaders, and other stakeholders.
  • Establish a close and trusted partnership with the HPC Technical Lead, who serves as the principal technical advisor and leads HPC architecture and technical direction. Empower the Technical Lead to exercise appropriate technical authority while retaining managerial and service accountability for priorities, resources, risk, and outcomes.
  • In partnership with the HPC Technical Lead and appropriate security and compliance offices, provide oversight for the design, operation, and continuous improvement of research computing environments that process or store regulated, restricted, or otherwise sensitive information, including CUI, export-controlled information, NIH restricted data, and PHI. Ensure these environments comply with applicable laws, contractual obligations, security frameworks, technology control plans, and institutional policies, which may include HIPAA, CMMC, NIST SP 800-171 Revision 2, NIST SP 800-171 Revision 3, ITAR, EAR, and 10 CFR Part 810.
  • Lead HPC infrastructure lifecycle planning and major procurements, including RFP requirements development and evaluation, vendor selection, capacity forecasting, and coordination of data-center space, power, cooling, networking, and support requirements. Maintain strategic vendor relationships and awareness of technology roadmaps.
     

Required Qualifications*

  • The selected candidate must be able to obtain and maintain all access authorizations required for assigned duties under applicable U.S. export-control laws, institutional technology control plans, and U-M Standard DS-04. Eligibility will be assessed individually in coordination with the appropriate U-M compliance offices.
  • Bachelor's degree in computer science, engineering, or a related field, or an equivalent combination of education and experience.
  • At least seven years of experience in research computing or a closely related technical environment, including at least five years supporting, operating, or leading services based on production Linux systems.
  • At least one year of experience supervising staff or serving in a sustained technical or team leadership role. Experience should include mentoring staff, coordinating and prioritizing work, guiding team decisions, and assuming broader leadership responsibilities when needed.
  • Experience working in an academic research environment and demonstrated understanding of researchers' computing and data needs.
  • Sufficient technical understanding of Linux-based HPC environments-including clusters, workload scheduling, storage, networking and high-speed fabrics, automation, and monitoring - to evaluate recommendations, ask informed questions, assess risks and tradeoffs, and make sound operational decisions.
  • Demonstrated ability to lead through partnership with senior technical experts, including delegating appropriate technical authority, respecting technical ownership, and integrating technical recommendations with operational, organizational, financial, and compliance considerations.
  • Familiarity with, or demonstrated ability to learn, requirements for environments that process regulated, restricted, or sensitive information, such as PHI, CUI, human-subject research data, and export-controlled information, with support from appropriate security and compliance offices.
  • Demonstrated ability to prioritize and coordinate multiple projects and operational needs, manage dependencies and risks, and make decisions in a complex service environment.
  • Strong written, verbal, and presentation skills, including the ability to communicate technical risks, options, and recommendations to researchers, technical staff, and organizational leaders.
  • Demonstrated ability to work independently and collaboratively across the HPC team, ARC, ITS, and the broader University community.
     

Desired Qualifications*

  • Demonstrated experience applying security or compliance requirements to environments that process PHI, CUI, export-controlled information, or other regulated research data.
  • Experience with formal people-management responsibilities, such as recruiting, performance management, coaching, career development, resource planning, and addressing employee-relations matters.
  • Significant experience partnering directly with faculty and other members of the academic research community.
  • Existing relationships in national organizations such as CASC, CaRCC, Midwest Research Computing, Campus Champions, XSEDE/Access-CI,  Women in HPC, etc.
  • Experience developing technical requirements and managing procurements or RFPs for HPC systems, infrastructure, or vendor services.
  • Hands-on technical experience with Linux automation and deployment (e.g., Ansible, Salt, Puppet, etc.)
  • Deep technical expertise in one or more areas of HPC architecture and operations, such as workload scheduling, parallel file systems, high-speed networking for MPI, monitoring, AI infrastructure, or large-scale system deployment
  • Familiarity with data-center space, power, cooling, and capacity planning, including air- and liquid-cooled systems
  • Experience using scripting, programming, data-analysis, or version-control tools such as Python, SQL, Bash, and Git
  • Experience with cloud environments for research (Verge.io, OpenStack, AWS, GCP, Azure, etc.)
  • Demonstrated as AI-forward for automating work, streamlining work, and facilitating adoption where appropriate to benefit the team
     

Why Work at Michigan?

In addition to a career filled with purpose and opportunity, The University of Michigan offers a comprehensive benefits package to help you stay well, protect yourself and your family and plan for a secure future. Benefits include:

  • Generous time off
  • A retirement plan that provides two-for-one matching contributions with immediate vesting
  • Many choices for comprehensive health insurance
  • Life insurance
  • Long-term disability coverage
  • Flexible spending accounts for healthcare and dependent care expenses
  • Dental and Vision Insurance
  • Parental and Maternity Leave
     

Modes of Work

Positions that are eligible for hybrid or mobile/remote work mode are at the discretion of the hiring department. Work agreements are reviewed annually at a minimum and are subject to change at any time, and for any reason, throughout the course of employment. Learn more about the work modes.

Work Locations

This role is remote but will require periodic in-person obligations generally scheduled in advance and may span multiple days.  Candidates should reside within approximately one hour's drive of Ann Arbor.  

Application Deadline

Job openings are posted for a minimum of seven calendar days. The review and selection process may begin as early as the eighth day after posting. This opening may be removed from posting boards and filled any time after the minimum posting period has ended.

U-M EEO Statement

The University of Michigan is an Equal Opportunity Employer. We are committed to providing an environment of mutual respect where equal employment opportunities are available to all applicants, including protected veterans and individuals with disabilities.