HPC Operations Manager – Hardware Engineering

1 Month ago • 15 Years + • Research & Development • $272,000 PA - $425,500 PA

Job Summary

Job Description

NVIDIA seeks a highly motivated HPC Operations Manager to lead and mentor a multinational team in managing global HPC clusters used by hardware design teams. Responsibilities include ensuring cluster reliability, developing key metrics, identifying and resolving failures, evaluating new technologies, planning hardware deployments, collaborating with engineering leaders, managing the HPC scheduler (LSF), and communicating program status to senior management. The role requires expertise in Linux servers, NFS storage, Ethernet networks, HPC schedulers, and hardware design workflows.
Must have:
  • 15+ years experience
  • 5+ years managing IT teams
  • 10+ years running Linux servers
  • HPC schedulers (LSF preferred)
  • Hardware design workflows knowledge
  • Data center operations
Good to have:
  • HPC storage expertise
  • Infiniband knowledge
  • Software development skills
  • Relational database knowledge
  • Experience with enterprise equipment suppliers
Perks:
  • Equity
  • Benefits

Job Details

Widely considered to be one of the technology world’s most desirable employers, NVIDIA is an industry leader with groundbreaking developments in High-Performance Computing, Artificial Intelligence and Visualization. The GPU, our invention, serves as the visual cortex of modern computers and is at the heart of our products and services. Our work opens up new universes to explore, enables outstanding creativity and discovery and powers what were once science fiction inventions from artificial intelligence to autonomous cars. We are now looking for a highly motivated HPC Operations Manager to join this multifaceted and innovative infrastructure team to craft global and dynamic HPC clusters used by Nvidia’s hardware design teams. We are looking for leaders to help us grow and evolve a reliable computing environment to enable our hardware designers to build the next generation of GPUs and SOCs.

What You'll be Doing:

  • A huge part of the day-to-day job is collaborating with partners to develop programs driving around storage, networking, and compute in our growing fleet of data centers.

  • Lead, cultivate, and mentor a multi-national team of sysadmins and devops engineers, in support of the chip design teams

  • Ensure the highest reliability of HPC clusters. Develop critical metrics, program schedules to measure program health, predictability, and achievements

  • Identify failures, lead retrospective analysis, and help to develop improvement action plans. Build standard methodologies that cut through complexity and can be used across Nvidia and influence other partners for continuous improvement

  • Evaluate the latest technologies (hardware and cloud computing) and recommend future evolution of the infrastructure. Plan deployments and refresh of hardware (compute, storage, network equipment), and associated software stack (e.g. OS)

  • Work multi-functionally with hardware engineering leaders to support their future chip design needs, understand their workflow characteristics, and engineer an efficient HPC environment. Work with IT and engineering infrastructure teams on the different subsystems that comprise the computing environment.

  • Lead all aspects of the HPC scheduler (LSF), set/adjust policy, ensure delivery of forecasted compute demand to each hardware division, and drive high utilization.

  • Track software licensing servers and drive efficient license utilization

  • Develop and manage program schedules, milestones and deliverables. Adjust in the face of a highly fluid customer product roadmap.

  • Regularly communicate program status and key issues to senior management at NVIDIA’s headquarters. Accurately represent the importance of issues and call out issues appropriately. Be the evangelist of data driven project management

What We Need to See:

  • B.S. or M.S. in Computer Science, Computer Engineering, Information Science (or equivalent experience)

  • 15+ years overall

  • 5+ years managing IT infrastructure teams of 10+ people

  • 10+ years experience running Linux servers, NFS storage, and Ethernet networks

  • Knowledge of HPC schedulers (IBM LSF preferred)

  • Knowledge of hardware design workflows (EDA tools and methodology)

  • Experience using project management and capacity planning software

  • Datacenter operations (rack and stack, maintenance)

Ways to stand out from the crowd:

  • HPC storage (e.g. Netapp, Pure Storage, Lustre, ZFS, Isilon)

  • Infiniband (operations, debugging, performance tuning)

  • Software development, especially in a devops context

  • Knowledge of relational databases, data lakes, metrics/visualization/analytics platforms

  • Deploying and maintaining FlexLM-based software license servers

  • Established relationships with enterprise-level equipment suppliers

The base salary range is 272,000 USD - 425,500 USD. Your base salary will be determined based on your location, experience, and the pay of employees in similar positions.

You will also be eligible for equity and benefits. NVIDIA accepts applications on an ongoing basis.

NVIDIA is committed to fostering a diverse work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.

Similar Jobs

CloudLinux - Principal Software Engineer / Product Owner (worldwide remote, work anywhere)

CloudLinux

Valencian Community, Spain (Remote)
2 Months ago
Milestone - Senior Software Engineer - Access Control

Milestone

Sofia, Sofia City Province, Bulgaria (Hybrid)
1 Month ago
Trend Micro - Embedded Software Engineer (C/C++)

Trend Micro

Manila, Metro Manila, Philippines (On-Site)
15 Years ago
Daybreak Game Company LLC - Senior Software Engineer, Platform

Daybreak Game Company LLC

San Diego, California, United States (Remote)
3 Months ago
ION - Cyber Security Analyst, Italy

ION

Pisa, Tuscany, Italy (On-Site)
4 Months ago
ByteDance - Student Researcher (Doubao (Seed) - Foundation Model - Vision and Language) - 2025 Start (PhD)

ByteDance

San Jose, California, United States (On-Site)
3 Months ago
Krafton  - Global Strategy Manager 집중채용

Krafton

Seoul, South Korea (On-Site)
2 Months ago
NVIDIA - Senior Deep Learning Software Engineer, Inference

NVIDIA

Santa Clara, California, United States (Hybrid)
1 Month ago
Netflix - Full Stack Software Engineer, L5 - Growth Delivery and Operations

Netflix

United States (Remote)
1 Month ago
NVIDIA - Physical Design Signoff CAD Engineer

NVIDIA

Yokne'am Illit, North District, Israel (Hybrid)
1 Month ago

Get notifed when new similar jobs are uploaded

Similar Skill Jobs

Playtech - Senior Embedded Software Engineer

Playtech

Manchester, England, United Kingdom (On-Site)
4 Months ago
CloudLinux - Principal Software Engineer / Product Owner (worldwide remote, work anywhere)

CloudLinux

Warsaw, Masovian Voivodeship, Poland (Remote)
2 Months ago
N-iX - Senior Java Engineer

N-iX

Kyiv, Kyiv City, Ukraine (Hybrid)
2 Weeks ago
ION - Support Engineer

ION

Italy (Hybrid)
4 Months ago
NVIDIA - Firmware PHY Verification Engineer

NVIDIA

Tel Aviv-Yafo, Tel Aviv District, Israel (Hybrid)
1 Month ago
Tencent - Senior Cloud Network Engineer - Singapore

Tencent

Singapore (On-Site)
4 Months ago
Rackspace Technology - DEVOP Engineer (AWS Terraform)-PSDE III

Rackspace Technology

India (Remote)
3 Months ago
Nintendo - Machine Learning Operations Engineer

Nintendo

Redmond, Washington, United States (On-Site)
1 Week ago
Infoblox - Senior Software Engineer

Infoblox

Burnaby, British Columbia, Canada (Hybrid)
3 Months ago

Get notifed when new similar jobs are uploaded

Jobs in Santa Clara, California, United States

Magnopus - Director of Design

Magnopus

Los Angeles, California, United States (Hybrid)
5 Months ago
Patreon - Senior iOS Engineer

Patreon

San Francisco, California, United States (Hybrid)
2 Weeks ago
Daybreak Game Company LLC - Senior Publishing Producer

Daybreak Game Company LLC

San Diego, California, United States (Hybrid)
5 Months ago
Onward Search - Inside Sales Representative (Real Estate)

Onward Search

Roanoke, Virginia, United States (On-Site)
3 Months ago
Hasbro - Intern - Brand Marketing, Undergrad (Summer 2025)

Hasbro

Rhode Island, United States (On-Site)
1 Month ago
Azra Games - Senior Environment Artist

Azra Games

Austin, Texas, United States (On-Site)
2 Months ago
Scopely - Senior Product Manager, Economy - Monopoly GO!

Scopely

California, United States (Remote)
2 Weeks ago
The Walt Disney Company - Senior Manager, Storage Systems Engineering

The Walt Disney Company

New York, New York, United States (On-Site)
1 Month ago
Apollo - Senior Engineering Manager (EST)

Apollo

United States (Remote)
4 Months ago
Insomniac Games - Senior Character FX Artist

Insomniac Games

California, United States (On-Site)
1 Month ago

Get notifed when new similar jobs are uploaded

Research & Development Jobs

Luxoft - C/C++ Lead Software Developer with ADAS, ASPICE, Korean speaker

Luxoft

Seoul, South Korea (On-Site)
3 Months ago
Synopsys  Inc  - Mac OS Virtualization Specialist

Synopsys Inc

Bengaluru, Karnataka, India (On-Site)
3 Months ago
NVIDIA - Senior Software Technical Program Manager - Compute Software Technologies

NVIDIA

Santa Clara, California, United States (On-Site)
1 Week ago
Tesla - Senior Mechanical Design Engineer - Motors

Tesla

Athens, Greece (On-Site)
5 Days ago
Meta - Research Scientist Intern, Feed Recommendations (PhD)

Meta

Menlo Park, California, United States (On-Site)
3 Months ago
NVIDIA - Senior Physical Design Methodology Engineer, PPA Improvement Technology Scaling

NVIDIA

Santa Clara, California, United States (On-Site)
5 Days ago
Fabric - Applied Researcher, Cryptography Hardware

Fabric

France (Remote)
4 Months ago
Rivos - Silicon Power - Full-time

Rivos

Bengaluru, Karnataka, India (Hybrid)
4 Months ago
Astera Labs - Senior Digital Design Engineer - SOC

Astera Labs

Bengaluru, Karnataka, India (On-Site)
4 Months ago
Ceragon Networks - Senior Verification Engineer

Ceragon Networks

Karnataka, India (On-Site)
3 Months ago

Get notifed when new similar jobs are uploaded

About The Company

Since its founding in 1993, NVIDIA (NASDAQ: NVDA) has been a pioneer in accelerated computing. The company’s invention of the GPU in 1999 sparked the growth of the PC gaming market, redefined computer graphics, ignited the era of modern AI and is fueling the creation of the metaverse. NVIDIA is now a full-stack computing company with data-center-scale offerings that are reshaping industry.


Yokne'am Illit, North District, Israel (On-Site)

Santa Clara, California, United States (Hybrid)

Santa Clara, California, United States (Hybrid)

Santa Clara, California, United States (On-Site)

United States (Remote)

Santa Clara, California, United States (On-Site)

Santa Clara, California, United States (On-Site)

Bengaluru, Karnataka, India (Hybrid)

Bengaluru, Karnataka, India (Hybrid)

View All Jobs

Get notified when new jobs are added by NVIDIA

Level Up Your Career in Game Development!

Transform Your Passion into Profession with Our Comprehensive Courses for Aspiring Game Developers.

Job Common Plug