Distinguished Engineer, AI Resiliency Lead

2 Months ago • 15 Years + • Artificial Intelligence • $308,000 PA - $471,500 PA

Job Summary

Job Description

NVIDIA seeks a Distinguished Engineer to lead AI Resiliency, architecting and developing software resiliency features for training AI models on large-scale superclusters. This role involves designing modular, resilient software, innovating in areas like in-memory checkpointing and anomaly detection to ensure near-zero downtime. The successful candidate will collaborate with cross-functional teams, communicate progress to senior leadership, and contribute to achieving stringent uptime requirements (downtime <1%). Responsibilities include defining scalable software architecture for resilient training across hundreds of thousands of GPUs and working with frameworks like PyTorch and JAX/XLA.
Must have:
  • Master's/Ph.D. in CS/ECE
  • 15+ years software architecture experience
  • Deep understanding of AI-optimized systems
  • 5+ years HPC/AI software development
  • Strong collaboration & communication skills
Good to have:
  • Experience with large-scale AI supercomputing applications
  • 5+ years with PyTorch and JAX/XLA
  • System architecture design expertise (CPU, GPU, memory, storage, networking)
  • HPC software development best practices implementation
Perks:
  • Equity
  • Benefits

Job Details

We are seeking a Distinguished Engineer to lead AI Resiliency at NVIDIA!

Join NVIDIA and help push the boundaries of AI. In this role, you will architect, design, and develop world-class software resiliency features for training ground breaking AI models on the largest AI superclusters in the world. Leading a team of cross-functional experts, you will drive and shape our end-to-end AI software stack, ensuring seamless training of frontier models on industry-leading frameworks like PyTorch and JAX/XLA, with near-zero downtime. Your optimizations will span from algorithmic innovations to robust software architecture, with a significant impact on NVIDIA’s most critical customers. This highly visible role demands exceptional technical expertise and leadership across organizations, with direct exposure to NVIDIA's senior leadership.

What You'll Be Doing:

  • Define a scalable software architecture to enable single-job resilient training on hundreds of thousands of GPUs with minimal downtime.

  • Design and deliver modular, resilient software features to support large-scale AI training for our top customers.

  • Innovate and evolve resilient architecture designs to achieve stringent uptime requirements (downtime < 1%), through solutions like in-memory check-pointing, in-process restart, and anomaly/SDC detection.

  • Collaborate closely with internal partners, spearheading successful project execution and communicating regular progress updates to senior leadership.

What We Need to See:

  • A Master’s or Ph.D. in Computer Science, Electrical or Computer Engineering from a top-tier university, or equivalent experience.

  • 15+ years of experience in software architecture or related fields, with a deep understanding of AI-optimized systems.

  • Excellent and proven ability to collaborate and communicate effectively across multiple engineering teams.

  • At least 5 years of hands-on experience in software development on high-complexity projects involving HPC or AI.

Ways to Stand Out from the Crowd:

  • Proven experience with large-scale AI supercomputing applications, particularly in the training phase.

  • 5+ years of experience with using and contributing to modern AI frameworks like PyTorch and JAX/XLA, specifically for large-scale training workloads.

  • A strong passion for designing system architectures tailored for AI, covering CPU, GPU, memory, storage, and networking.

  • Hands-on involvement in the entire lifecycle—from design to deployment—of large-scale High-Performance Computing (HPC) systems.

  • Experience in implementing HPC software development best practices in large-scale systems.

NVIDIA continues to expand its presence in the Datacenter space, and our team plays a pivotal role in enhancing the value of our rapidly growing datacenter deployments. We also drive a data-driven approach to hardware design and system software development. You will collaborate with a wide array of teams across NVIDIA, including deep learning research, CUDA kernel and framework development, and silicon architecture. NVIDIA is widely recognized as one of the technology industry’s most desirable employers, with some of the most dedicated and innovative minds working with us. If you’re creative, driven, and autonomous, we want to hear from you!

The base salary range is 308,000 USD - 471,500 USD. Your base salary will be determined based on your location, experience, and the pay of employees in similar positions.

You will also be eligible for equity and benefits. NVIDIA accepts applications on an ongoing basis.

NVIDIA is committed to fostering a diverse work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.

Similar Jobs

NVIDIA - DFX Methodology Engineer

NVIDIA

Santa Clara, California, United States (On-Site)
1 Month ago
Arrise Solutions (India)   - Senior Data Scientist (Remote)

Arrise Solutions (India)

Hyderabad, Telangana, India (Remote)
6 Months ago
ByteDance - GPU/AI Application Platform Engineer Intern (Server Platform)

ByteDance

San Jose, California, United States (On-Site)
1 Month ago
ByteDance - Software Engineer Graduate (Applied Machine Learning - Engine) - 2025 Start (BS/MS)

ByteDance

San Jose, California, United States (On-Site)
6 Months ago
Zoox - Senior Machine Learning Engineer - Collision Avoidance System

Zoox

Foster City, California, United States (Hybrid)
6 Months ago
ByteDance - Research Engineer Intern (Doubao (Seed) - Machine Learning System) - 2025 Summer (MS)

ByteDance

Seattle, Washington, United States (On-Site)
6 Months ago
PlayStation Global - Staff Machine Learning Engineer, Enterprise Enablement

PlayStation Global

California, United States (On-Site)
3 Months ago
Canva - Senior Computer Vision Engineer - Photo AI

Canva

Vienna, Vienna, Austria (Remote)
1 Month ago
Canva - Senior Backend Engineer - AI Enablement

Canva

Surry Hills, New South Wales, Australia (Remote)
1 Month ago
Zoox - Software Engineer - Simulation Traffic & Behavior Modeling

Zoox

Seattle, Washington, United States (Hybrid)
6 Months ago

Get notifed when new similar jobs are uploaded

Similar Skill Jobs

GoTo Group - Senior Data Scientist  (Singapore)

GoTo Group

Singapore (On-Site)
6 Months ago
NVIDIA - PCB Design Layout Engineer

NVIDIA

Santa Clara, California, United States (On-Site)
2 Months ago
Hashlist - Data Scientist

Hashlist

Bengaluru, Karnataka, India (Hybrid)
5 Months ago
NVIDIA - System Memory Performance and Power Engineer

NVIDIA

Canada (Hybrid)
1 Month ago
NVIDIA - Senior Product Manager – AI Networking Orchestration

NVIDIA

Canada (On-Site)
2 Months ago
ByteDance - Software Engineer, ML System Scheduling

ByteDance

Seattle, Washington, United States (On-Site)
6 Months ago
Samsung Semiconductor - Machine Learning Engineer Intern - PEFT

Samsung Semiconductor

San Jose, California, United States (On-Site)
3 Months ago
ByteDance - Research Scientist Graduate (High-Performance Computing (Inference Optimization) - Vision AI Platform)

ByteDance

Seattle, Washington, United States (On-Site)
1 Month ago
NVIDIA - Principal Firmware Engineer - Data Center Server Management

NVIDIA

California, United States (On-Site)
2 Months ago
NVIDIA - DFX Methodology Engineer

NVIDIA

Canada (On-Site)
1 Month ago

Get notifed when new similar jobs are uploaded

Jobs in Redmond, Washington, United States

Google - Software Engineer III, Infrastructure, Google Cloud Data Management

Google

New York, New York, United States (On-Site)
5 Months ago
Globalization Partners - Director, Product Led Growth Marketing

Globalization Partners

Boston, Massachusetts, United States (Remote)
5 Months ago
ByteDance - Senior Backend Software Engineer - Customer Service Platform

ByteDance

Seattle, Washington, United States (On-Site)
1 Month ago
IGT - QA Technician III

IGT

West Greenwich, Rhode Island, United States (On-Site)
4 Months ago
ByteDance - Software Engineer

ByteDance

Seattle, Washington, United States (On-Site)
2 Months ago
Nagarro - Associate Principal Engineer, CRM Salesforce

Nagarro

New York, New York, United States (On-Site)
6 Months ago
Next Level Business Services - SAP PI/PO Consultant

Next Level Business Services

Santa Clara, California, United States (On-Site)
6 Months ago
Keywords Studios - Project Lead - AI

Keywords Studios

San Francisco, California, United States (Remote)
3 Weeks ago
Cloud Imperium Games - Accounts Payable Specialist

Cloud Imperium Games

Austin, Texas, United States (On-Site)
1 Month ago
Anthology  Inc  - Regional Sales Manager

Anthology Inc

United States (Remote)
3 Months ago

Get notifed when new similar jobs are uploaded

Artificial Intelligence Jobs

NVIDIA - Deep Learning Compiler Intern

NVIDIA

Santa Clara, California, United States (On-Site)
1 Month ago
Electronic Arts - Data Science Engineer

Electronic Arts

Hyderabad, Telangana, India (On-Site)
1 Month ago
NVIDIA - Senior Deep Learning Performance Architect

NVIDIA

Redmond, Washington, United States (On-Site)
2 Months ago
Microsoft - Platform Engineering Manager

Microsoft

Redmond, Washington, United States (Hybrid)
1 Month ago
Inworld AI - Forward Deployed Engineer - Canada

Inworld AI

Vancouver, British Columbia, Canada (Remote)
6 Months ago
Zoox - Prediction Internship/Co-op

Zoox

Foster City, California, United States (On-Site)
6 Months ago
ByteDance - Research Engineer Graduate (Vision AI Platform)

ByteDance

San Jose, California, United States (On-Site)
2 Months ago
Zoox - Software Engineer - Perception

Zoox

Foster City, California, United States (Hybrid)
6 Months ago
NVIDIA - Senior Technical Instructor - AI and Data Center Infrastructure

NVIDIA

Texas, United States (Remote)
1 Month ago
PwC - Manager_Conversational AI Developer_Advisory Corporate_Advisory_Bangalore

PwC

Bengaluru, Karnataka, India (On-Site)
7 Months ago

Get notifed when new similar jobs are uploaded

About The Company

Since its founding in 1993, NVIDIA (NASDAQ: NVDA) has been a pioneer in accelerated computing. The company’s invention of the GPU in 1999 sparked the growth of the PC gaming market, redefined computer graphics, ignited the era of modern AI and is fueling the creation of the metaverse. NVIDIA is now a full-stack computing company with data-center-scale offerings that are reshaping industry.

Santa Clara, California, United States (On-Site)

Santa Clara, California, United States (On-Site)

Massachusetts, United States (On-Site)

Santa Clara, California, United States (On-Site)

Santa Clara, California, United States (On-Site)

Santa Clara, California, United States (On-Site)

Texas, United States (On-Site)

Santa Clara, California, United States (Hybrid)

Austin, Texas, United States (Remote)

Santa Clara, California, United States (Hybrid)

View All Jobs

Get notified when new jobs are added by NVIDIA

Level Up Your Career in Game Development!

Transform Your Passion into Profession with Our Comprehensive Courses for Aspiring Game Developers.

Job Common Plug