Outscal Logooutscal logo

Distinguished Engineer, AI Resiliency Lead

1 Month ago • 15 Years + • Artificial Intelligence • $308,000 PA - $471,500 PA

Job Summary

Job Description

NVIDIA seeks a Distinguished Engineer to lead AI Resiliency, architecting and developing software resiliency features for training AI models on large-scale superclusters. This role involves designing modular, resilient software, innovating in areas like in-memory checkpointing and anomaly detection to ensure near-zero downtime. The successful candidate will collaborate with cross-functional teams, communicate progress to senior leadership, and contribute to achieving stringent uptime requirements (downtime <1%). Responsibilities include defining scalable software architecture for resilient training across hundreds of thousands of GPUs and working with frameworks like PyTorch and JAX/XLA.
Must have:
  • Master's/Ph.D. in CS/ECE
  • 15+ years software architecture experience
  • Deep understanding of AI-optimized systems
  • 5+ years HPC/AI software development
  • Strong collaboration & communication skills
Good to have:
  • Experience with large-scale AI supercomputing applications
  • 5+ years with PyTorch and JAX/XLA
  • System architecture design expertise (CPU, GPU, memory, storage, networking)
  • HPC software development best practices implementation
Perks:
  • Equity
  • Benefits

Job Details

We are seeking a Distinguished Engineer to lead AI Resiliency at NVIDIA!

Join NVIDIA and help push the boundaries of AI. In this role, you will architect, design, and develop world-class software resiliency features for training ground breaking AI models on the largest AI superclusters in the world. Leading a team of cross-functional experts, you will drive and shape our end-to-end AI software stack, ensuring seamless training of frontier models on industry-leading frameworks like PyTorch and JAX/XLA, with near-zero downtime. Your optimizations will span from algorithmic innovations to robust software architecture, with a significant impact on NVIDIA’s most critical customers. This highly visible role demands exceptional technical expertise and leadership across organizations, with direct exposure to NVIDIA's senior leadership.

What You'll Be Doing:

  • Define a scalable software architecture to enable single-job resilient training on hundreds of thousands of GPUs with minimal downtime.

  • Design and deliver modular, resilient software features to support large-scale AI training for our top customers.

  • Innovate and evolve resilient architecture designs to achieve stringent uptime requirements (downtime < 1%), through solutions like in-memory check-pointing, in-process restart, and anomaly/SDC detection.

  • Collaborate closely with internal partners, spearheading successful project execution and communicating regular progress updates to senior leadership.

What We Need to See:

  • A Master’s or Ph.D. in Computer Science, Electrical or Computer Engineering from a top-tier university, or equivalent experience.

  • 15+ years of experience in software architecture or related fields, with a deep understanding of AI-optimized systems.

  • Excellent and proven ability to collaborate and communicate effectively across multiple engineering teams.

  • At least 5 years of hands-on experience in software development on high-complexity projects involving HPC or AI.

Ways to Stand Out from the Crowd:

  • Proven experience with large-scale AI supercomputing applications, particularly in the training phase.

  • 5+ years of experience with using and contributing to modern AI frameworks like PyTorch and JAX/XLA, specifically for large-scale training workloads.

  • A strong passion for designing system architectures tailored for AI, covering CPU, GPU, memory, storage, and networking.

  • Hands-on involvement in the entire lifecycle—from design to deployment—of large-scale High-Performance Computing (HPC) systems.

  • Experience in implementing HPC software development best practices in large-scale systems.

NVIDIA continues to expand its presence in the Datacenter space, and our team plays a pivotal role in enhancing the value of our rapidly growing datacenter deployments. We also drive a data-driven approach to hardware design and system software development. You will collaborate with a wide array of teams across NVIDIA, including deep learning research, CUDA kernel and framework development, and silicon architecture. NVIDIA is widely recognized as one of the technology industry’s most desirable employers, with some of the most dedicated and innovative minds working with us. If you’re creative, driven, and autonomous, we want to hear from you!

The base salary range is 308,000 USD - 471,500 USD. Your base salary will be determined based on your location, experience, and the pay of employees in similar positions.

You will also be eligible for equity and benefits. NVIDIA accepts applications on an ongoing basis.

NVIDIA is committed to fostering a diverse work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.

Similar Jobs

NVIDIA - Senior ASIC Timing Engineer

NVIDIA

Westborough, Massachusetts, United States (On-Site)
2 Weeks ago
NVIDIA - Distinguished Engineer, AI Resiliency Lead

NVIDIA

Santa Clara, California, United States (On-Site)
2 Months ago
Google - Student Researcher, BS/MS, Winter/Summer 2025

Google

Montreal, Quebec, Canada (On-Site)
4 Months ago
ByteDance - Student Researcher (Doubao (Seed) - Foundation Model - Vision Generative AI)

ByteDance

San Jose, California, United States (On-Site)
21 Hours ago
Google - Software Engineer III, Machine Learning, Search

Google

Mountain View, California, United States (On-Site)
4 Months ago
Meta - Software Engineer, Machine Learning

Meta

New York, New York, United States (On-Site)
4 Months ago
NVIDIA - Solution Architect, Generative AI - Digital Human

NVIDIA

Canada (On-Site)
1 Month ago
Pika - Research Engineer (Applied Research)

Pika

Palo Alto, California, United States (On-Site)
11 Hours ago
Samsung Semiconductor - Intern, Machine Learning Engineer - VLMs

Samsung Semiconductor

San Jose, California, United States (Hybrid)
2 Months ago

Get notifed when new similar jobs are uploaded

Similar Skill Jobs

NVIDIA - Senior ASIC Verification Engineer

NVIDIA

Santa Clara, California, United States (On-Site)
2 Months ago
ByteDance - Machine Learning Engineer Intern (Applied Machine Learning-Algorithm) - 2025 Summer/Fall (PhD)

ByteDance

San Jose, California, United States (On-Site)
4 Months ago
NVIDIA - Senior AI Cluster Tools Developer

NVIDIA

Santa Clara, California, United States (Hybrid)
2 Months ago
ByteDance - Research Engineer- Foundation Model AI Platform- Seattle

ByteDance

Seattle, Washington, United States (On-Site)
4 Months ago
NVIDIA - Mixed Signal Design Engineer

NVIDIA

Canada (On-Site)
1 Month ago
Krafton  - Head of Deep Learning PM & Ops

Krafton

Seoul, South Korea (On-Site)
1 Week ago
Canva - Senior Machine Learning Engineer - Photo AI

Canva

Prague, Czechia (Remote)
2 Months ago
InMobiInMobi - Data Scientist II

InMobiInMobi

Bengaluru, Karnataka, India (On-Site)
2 Months ago
Zazz - Machine Learning Engineer

Zazz

(Remote)
1 Month ago
Saama Technologies,  Inc  - NLP Engineer

Saama Technologies, Inc

(Remote)
4 Weeks ago

Get notifed when new similar jobs are uploaded

Jobs in Redmond, Washington, United States

Crunchyroll - Staff Engineer, Partner Reliability

Crunchyroll

San Francisco, California, United States (On-Site)
1 Month ago
Gearbox Software - Technical Artist

Gearbox Software

Frisco, Texas, United States (On-Site)
9 Months ago
Rivos - Platform FPGA Design

Rivos

Santa Clara, California, United States (On-Site)
5 Months ago
Meta - Research Scientist Intern, Machine Perception for Input and Interaction (PhD)

Meta

Burlingame, California, United States (On-Site)
4 Months ago
Zoox - Senior Product Manager, Operations Software

Zoox

Foster City, California, United States (Hybrid)
5 Months ago
Flip Fit - Data Scientist

Flip Fit

El Segundo, California, United States (On-Site)
5 Months ago
Framestore - FREELANCE: NUKE - NEW YORK

Framestore

New York, New York, United States (On-Site)
10 Months ago
Fluence - BESS Analyst - Battery Energy Storage

Fluence

Alpharetta, Georgia, United States (Hybrid)
3 Months ago
Aristocrat Gaming - Supplier Quality Engineer

Aristocrat Gaming

Las Vegas, Nevada, United States (On-Site)
2 Weeks ago
Samsung Semiconductor - Intern, Logic Design Engineer

Samsung Semiconductor

San Jose, California, United States (Hybrid)
1 Week ago

Get notifed when new similar jobs are uploaded

Artificial Intelligence Jobs

Zoox - Staff/Senior Staff Software Engineer, ML Performance Optimization

Zoox

Foster City, California, United States (On-Site)
5 Months ago
Codeninja - Graduate Trainee - AI/ML

Codeninja

Lahore, Punjab, Pakistan (On-Site)
6 Days ago
Google - Software Engineer III, Machine Learning, Google Ads

Google

Los Angeles, California, United States (On-Site)
4 Months ago
ByteDance - Research Engineer Graduate (Vision AI Platform)

ByteDance

San Jose, California, United States (On-Site)
21 Hours ago
SiftHub - Senior NLP Engineer

SiftHub

Maharashtra, India (On-Site)
6 Months ago
Zoox - Software Engineer - Simulation Workload Orchestration

Zoox

Foster City, California, United States (Hybrid)
5 Months ago
Match Group - Sr. Software Engineer, Generative AI

Match Group

Palo Alto, California, United States (Hybrid)
5 Months ago
Keywords Studios (Player Support) - AI - Senior Research Associate (Prompts)

Keywords Studios (Player Support)

Silesian Voivodeship, Poland (On-Site)
1 Week ago

Get notifed when new similar jobs are uploaded

About The Company

Since its founding in 1993, NVIDIA (NASDAQ: NVDA) has been a pioneer in accelerated computing. The company’s invention of the GPU in 1999 sparked the growth of the PC gaming market, redefined computer graphics, ignited the era of modern AI and is fueling the creation of the metaverse. NVIDIA is now a full-stack computing company with data-center-scale offerings that are reshaping industry.


Hsinchu, Hsinchu City, Taiwan (On-Site)

Yokne'am Illit, North District, Israel (On-Site)

Seoul, South Korea (Hybrid)

Tel Aviv-Yafo, Tel Aviv District, Israel (On-Site)

Ra'anana, Center District, Israel (On-Site)

Shanghai, Shanghai, China (On-Site)

Tel Aviv-Yafo, Tel Aviv District, Israel (On-Site)

Be'er Sheva, South District, Israel (On-Site)

Santa Clara, California, United States (On-Site)

Santa Clara, California, United States (On-Site)

View All Jobs

Get notified when new jobs are added by NVIDIA

Level Up Your Career in Game Development!

Transform Your Passion into Profession with Our Comprehensive Courses for Aspiring Game Developers.

Job Common Plug