Senior HPC and AI Networking Performance Research and Analysis Engineer

1 Month ago • 5 Years + • Artificial Intelligence • Research & Development • $148,000 PA - $287,500 PA

Job Summary

Job Description

NVIDIA seeks a Senior HPC and AI Networking Performance Research and Analysis Engineer to profile and analyze AI workloads on large-scale GPU/CPU clusters for distributed deep learning. Responsibilities include benchmarking, profiling, implementing performance analysis tools, collaborating with various teams, defining performance test plans, and setting expectations for new technologies. The role focuses on high-performance networking, NCCL, and identifying performance bottlenecks. Experience with RDMA, MPI, NCCL, deep learning frameworks (TensorFlow/PyTorch), and CUDA is essential.
Must have:
  • 5+ years HPC Networking experience (RDMA, MPI, NCCL)
  • Performance analysis skills and methodologies
  • Experience with NVIDIA GPUs, CUDA, Deep Learning Frameworks
  • Python, Bash, C programming
  • Linux OS experience
Good to have:
  • In-depth knowledge of AI workloads and LLM training
  • Knowledge of CUDA and NCCL libraries
  • Understanding of Congestion Control algorithms
  • System knowledge (CPUs, GPUs, HCA, Memory, PCI)
Perks:
  • Competitive salary
  • Comprehensive benefits package
  • Diverse and supportive work environment

Job Details

Intelligent machines powered by Artificial Intelligence computers that can learn, reason and interact with people are no longer science fiction. GPU Deep Learning has provided the foundation for machines to learn, perceive, reason and solve problems. Today, visual computing is a crucial tool in helping people get along with technology, and NVIDIA has extended its technology into datacenters, mobile devices and cars. There has never been a more exciting time to join our team - if this role sounds like a fit for you, we'd love to hear from you!

NVIDIA is seeking a Senior High Performance Computing (HPC) and AI Networking Performance Research and Analysis Engineer to join our Performance group. In this exciting role, you will profile and analyze AI workloads on large GPUs and CPUs scale clusters for distributed Deep Learning LLM training focused on collectives communication and networking. You will interact with many types of hardware and platforms, such as HCAs, Switches, CPUs, GPUs, and Systems. You will develop performance analysis tools and methodologies to dive deeply into the details and understand performance expectations, limitations, and bottlenecks.

What you'll be doing:

  • Exploring and researching AI workloads and DL models specifically tailored for large-scale deep learning LLM training on NVIDIA supercomputers and distributed systems focusing on high-performance networking and Nvidia Collective Communications Library (NCCL).

  • Benchmarking, Profiling, and Analyzing the performance to find bottlenecks and identify areas of improvement and optimizations, with a strong emphasis on networking aspects.

  • Implementing performance analysis tools.

  • Collaborating with many teams from hardware to software to provide performance analysis insights.

  • Defining performance test planning , setting performance expectations for new technologies and solutions, and working to reach the performance targets limits.

What we need to see:

  • B.Sc in Computer Science or Software Engineering or equivalent experience

  • 5+ years of experience with high-performance Networking (RDMA, MPI, NCCL, Congestion Control Algorithms)

  • Demonstrated Performance Analysis skills and methodologies.

  • Experience with NVIDIA GPUs, CUDA library, deep learning frameworks like TensorFlow or PyTorch, combined with expertise in networking collective communication libraries (such as NCCL) and protocols (such as RoCE and RDMA).

  • Fast and self-learning capabilities with strong analytical and problem-solving skills.

  • Programming Languages: Python, Bash and C languages

  • Experience with Linux OS distros.

  • Great teammate with good communication and interpersonal skills

Ways to stand out from the crowd:

  • In-depth knowledge and experience with AI workloads and benchmarking for distributed LLM training.

  • Knowledge in CUDA, and NCCL libraries.

  • Knowledge in Congestion Control algorithms.

  • In-depth System knowledge and understanding (Intel / AMD / ARM CPUs, NVIDIA GPUs, HCA, Memory, PCI).

  • Strong Performance Analysis skills and methodologies using modern tools.

NVIDIA has been redefining computer graphics, PC gaming, and accelerated computing for more than 25 years. We have a unique legacy of innovation that’s fueled by great technology—and amazing people. Today, we’re tapping into the unlimited potential of AI to define the next era of computing. An era in which our GPU acts as the brains of computers, robots, and self-driving cars that can understand the world. Doing what’s never been done before takes vision, innovation, and the world’s best talent. Widely considered to be one of the technology world’s most desirable employers, NVIDIA offers highly competitive salaries and a comprehensive benefits package. As an NVIDIAN, you’ll be immersed in a diverse, supportive environment where everyone is inspired to do their best work. Come join the team and see how you can make a lasting impact on the world!

#LI-Hybrid

The base salary range is 148,000 USD - 287,500 USD. Your base salary will be determined based on your location, experience, and the pay of employees in similar positions.

You will also be eligible for equity and benefits. NVIDIA accepts applications on an ongoing basis.

NVIDIA is committed to fostering a diverse work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.

Similar Jobs

Grid Dynamics - DevOps Engineer

Grid Dynamics

Bengaluru, Karnataka, India (Hybrid)
5 Months ago
Milk Visual Effects - Systems Administrator

Milk Visual Effects

(On-Site)
2 Months ago
Axon - Senior Security Engineer

Axon

Scottsdale, Arizona, United States (Hybrid)
2 Months ago
ByteDance - Backend Software Engineer

ByteDance

San Jose, California, United States (On-Site)
5 Days ago
Blinkhealth - Senior Manager, Cloud Engineering

Blinkhealth

(Remote)
1 Week ago
Microsoft - Senior Applied Scientist

Microsoft

Cairo, Cairo Governorate, Egypt (On-Site)
1 Month ago
BigID - Sr Solutions/Presales Engineer - West

BigID

Denver, Colorado, United States (Remote)
3 Months ago
AI Fund - Machine Learning Engineer

AI Fund

(Remote)
4 Months ago
PwC - Senior AI Developer - Roma [DIG]

PwC

Rome, Lazio, Italy (On-Site)
4 Months ago
Zoox - Software Engineer - Simulation Workload Orchestration

Zoox

Seattle, Washington, United States (Hybrid)
4 Months ago

Get notifed when new similar jobs are uploaded

Similar Skill Jobs

ION - Storage Engineer, Italy

ION

Italy (Hybrid)
4 Months ago
ION - Cyber Security Analyst, Italy

ION

Turin, Piedmont, Italy (On-Site)
4 Months ago
Paytm - Data Engineer - Technical Lead

Paytm

Noida, Uttar Pradesh, India (On-Site)
2 Months ago
NVIDIA - Senior Firmware Engineer – GPU Networking

NVIDIA

Santa Clara, California, United States (Hybrid)
1 Month ago
ByteDance - Backend Software Engineer

ByteDance

San Jose, California, United States (On-Site)
5 Days ago
Scopely - Lead DevOps/SRE - Unannounced Project

Scopely

Dublin, County Dublin, Ireland (Hybrid)
1 Month ago
Interactive Brokers - Senior Systems Engineer- Microsoft M365/Active Directory

Interactive Brokers

Greenwich, Connecticut, United States (Hybrid)
4 Months ago
Funko - Cloud Systems Engineer

Funko

Washington, United States (On-Site)
2 Months ago
Respawn Entertainment - Senior Build Engineer (Apex Legends)

Respawn Entertainment

Los Angeles, California, United States (On-Site)
6 Months ago
Dolby Laboratories - AIOps Research Scientist

Dolby Laboratories

Bengaluru, Karnataka, India (Hybrid)
4 Months ago

Get notifed when new similar jobs are uploaded

Jobs in Santa Clara, California, United States

Bethesda - Senior AI Programmer

Bethesda

Rockville, Maryland, United States (On-Site)
2 Months ago
Ludeo - Marketing Director

Ludeo

Los Angeles, California, United States (On-Site)
1 Month ago
Netflix - Principal Counsel, Litigation

Netflix

Los Angeles, California, United States (On-Site)
3 Months ago
Xsolla - Anaplan Model Builder

Xsolla

Los Angeles, California, United States (Hybrid)
5 Months ago
Netflix - Operations Manager 5, Live Broadcast Technology

Netflix

Los Angeles, California, United States (On-Site)
3 Months ago
Epic Games - Principal Cloud Engineer

Epic Games

Cary, North Carolina, United States (On-Site)
3 Weeks ago
Microsoft - Software Engineer 2

Microsoft

Redmond, Washington, United States (Remote)
1 Month ago
Crunchyroll - Senior Engineering Operations

Crunchyroll

San Francisco, California, United States (On-Site)
4 Weeks ago
Warner Bros Discovery - Cybersecurity Compliance Staff (Lead)

Warner Bros Discovery

Atlanta, Georgia, United States (Hybrid)
3 Months ago
Warner Bros Discovery - Senior Manager Corporate Accounting

Warner Bros Discovery

Atlanta, Georgia, United States (On-Site)
3 Months ago

Get notifed when new similar jobs are uploaded

Artificial Intelligence Jobs

Google - Software Engineer III, AI/ML, Google Cloud

Google

Gurugram, Haryana, India (On-Site)
1 Month ago
Meta - AI Research Scientist, Language - Generative AI

Meta

New York, New York, United States (On-Site)
3 Months ago
KPIT - CTO_ML/DL Data scientist

KPIT

Pune, Maharashtra, India (On-Site)
3 Months ago
DEVOTEAM - Data Driven | MLOps Engineer

DEVOTEAM

Lisbon, Lisbon, Portugal (Remote)
4 Months ago
Inkittt - Director of AI

Inkittt

San Francisco, California, United States (On-Site)
6 Months ago
Ubisoft - Programmeur IA (W/M/NB) – Project Non Annoncé

Ubisoft

Lyon, Auvergne-Rhône-Alpes, France (On-Site)
2 Months ago
ByteDance - Research Scientist Intern (Doubao (Seed) - Machine Learning System) - 2025 Summer (PhD)

ByteDance

Seattle, Washington, United States (On-Site)
3 Months ago
Interface AI - Senior Account Manager

Interface AI

United States (Remote)
6 Days ago
ION - Senior AI Engineer, Italy

ION

Pisa, Tuscany, Italy (On-Site)
4 Months ago
Blockville Digital Assets - AI Technology Specialist for Game Development

Blockville Digital Assets

İstanbul, Türkiye (On-Site)
7 Months ago

Get notifed when new similar jobs are uploaded

About The Company

Since its founding in 1993, NVIDIA (NASDAQ: NVDA) has been a pioneer in accelerated computing. The company’s invention of the GPU in 1999 sparked the growth of the PC gaming market, redefined computer graphics, ignited the era of modern AI and is fueling the creation of the metaverse. NVIDIA is now a full-stack computing company with data-center-scale offerings that are reshaping industry.


Santa Clara, California, United States (On-Site)

Santa Clara, California, United States (On-Site)

Shenzhen, Guangdong Province, China (On-Site)

Bengaluru, Karnataka, India (On-Site)

Taipei City, Taiwan (On-Site)

Taipei City, Taiwan (On-Site)

Shanghai, Shanghai, China (On-Site)

Shanghai, Shanghai, China (On-Site)

Yokne'am Illit, North District, Israel (On-Site)

View All Jobs

Get notified when new jobs are added by NVIDIA

Level Up Your Career in Game Development!

Transform Your Passion into Profession with Our Comprehensive Courses for Aspiring Game Developers.

Job Common Plug