Home >

Jobs >

Senior HPC and AI Networking Performance Research and Analysis Engineer

NVIDIA

California, United States (Hybrid)

Senior HPC and AI Networking Performance Research and Analysis Engineer

7 Months ago • 5 Years + • Research Development • $148,000 PA - $287,500 PA

Job Summary

Job Description

NVIDIA seeks a Senior HPC and AI Networking Performance Research and Analysis Engineer to profile and analyze AI workloads on large-scale GPU/CPU clusters for distributed deep learning. Responsibilities include benchmarking, profiling, implementing performance analysis tools, collaborating with various teams, defining performance test plans, and setting expectations for new technologies. The role focuses on high-performance networking, NCCL, and identifying performance bottlenecks. Experience with RDMA, MPI, NCCL, deep learning frameworks (TensorFlow/PyTorch), and CUDA is essential.

Must have:

5+ years HPC Networking experience (RDMA, MPI, NCCL)
Performance analysis skills and methodologies
Experience with NVIDIA GPUs, CUDA, Deep Learning Frameworks
Python, Bash, C programming
Linux OS experience

Good to have:

In-depth knowledge of AI workloads and LLM training
Knowledge of CUDA and NCCL libraries
Understanding of Congestion Control algorithms
System knowledge (CPUs, GPUs, HCA, Memory, PCI)

Perks:

Competitive salary
Comprehensive benefits package
Diverse and supportive work environment

14 skills required

14 skills required for this role

Add these skills to join the top 1% applicants for this job

bash

tensorflow

algorithms

deep-learning

python

pytorch

linux

foundation

cuda

communication

innovation

test-coverage

networking

performance-analysis

Job Details

Intelligent machines powered by Artificial Intelligence computers that can learn, reason and interact with people are no longer science fiction. GPU Deep Learning has provided the foundation for machines to learn, perceive, reason and solve problems. Today, visual computing is a crucial tool in helping people get along with technology, and NVIDIA has extended its technology into datacenters, mobile devices and cars. There has never been a more exciting time to join our team - if this role sounds like a fit for you, we'd love to hear from you!

NVIDIA is seeking a Senior High Performance Computing (HPC) and AI Networking Performance Research and Analysis Engineer to join our Performance group. In this exciting role, you will profile and analyze AI workloads on large GPUs and CPUs scale clusters for distributed Deep Learning LLM training focused on collectives communication and networking. You will interact with many types of hardware and platforms, such as HCAs, Switches, CPUs, GPUs, and Systems. You will develop performance analysis tools and methodologies to dive deeply into the details and understand performance expectations, limitations, and bottlenecks.

What you'll be doing:

Exploring and researching AI workloads and DL models specifically tailored for large-scale deep learning LLM training on NVIDIA supercomputers and distributed systems focusing on high-performance networking and Nvidia Collective Communications Library (NCCL).
Benchmarking, Profiling, and Analyzing the performance to find bottlenecks and identify areas of improvement and optimizations, with a strong emphasis on networking aspects.
Implementing performance analysis tools.
Collaborating with many teams from hardware to software to provide performance analysis insights.
Defining performance test planning , setting performance expectations for new technologies and solutions, and working to reach the performance targets limits.

What we need to see:

B.Sc in Computer Science or Software Engineering or equivalent experience
5+ years of experience with high-performance Networking (RDMA, MPI, NCCL, Congestion Control Algorithms)
Demonstrated Performance Analysis skills and methodologies.
Experience with NVIDIA GPUs, CUDA library, deep learning frameworks like TensorFlow or PyTorch, combined with expertise in networking collective communication libraries (such as NCCL) and protocols (such as RoCE and RDMA).
Fast and self-learning capabilities with strong analytical and problem-solving skills.
Programming Languages: Python, Bash and C languages
Experience with Linux OS distros.
Great teammate with good communication and interpersonal skills

Ways to stand out from the crowd:

In-depth knowledge and experience with AI workloads and benchmarking for distributed LLM training.
Knowledge in CUDA, and NCCL libraries.
Knowledge in Congestion Control algorithms.
In-depth System knowledge and understanding (Intel / AMD / ARM CPUs, NVIDIA GPUs, HCA, Memory, PCI).
Strong Performance Analysis skills and methodologies using modern tools.

NVIDIA has been redefining computer graphics, PC gaming, and accelerated computing for more than 25 years. We have a unique legacy of innovation that’s fueled by great technology—and amazing people. Today, we’re tapping into the unlimited potential of AI to define the next era of computing. An era in which our GPU acts as the brains of computers, robots, and self-driving cars that can understand the world. Doing what’s never been done before takes vision, innovation, and the world’s best talent. Widely considered to be one of the technology world’s most desirable employers, NVIDIA offers highly competitive salaries and a comprehensive benefits package. As an NVIDIAN, you’ll be immersed in a diverse, supportive environment where everyone is inspired to do their best work. Come join the team and see how you can make a lasting impact on the world!

#LI-Hybrid

The base salary range is 148,000 USD - 287,500 USD. Your base salary will be determined based on your location, experience, and the pay of employees in similar positions.

You will also be eligible for equity and benefits. NVIDIA accepts applications on an ongoing basis.

NVIDIA is committed to fostering a diverse work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.

Similar Jobs

Software Integration Engineer 2 (IDN - 057)

Sagecor Solutions

Fort Meade, Maryland, United States (On-Site)

• 10 Months ago

Senior Platform Engineer - Design

Interactive Brokers

Fort Lauderdale, Florida, United States (Hybrid)

• 11 Months ago

Python (Golang) Developer

Wargaming

Warsaw, Masovian Voivodeship, Poland (On-Site)

• 6 Months ago

Solutions Architect

Keywords Studios (Player Support)

Montreal, Quebec, Canada (Remote)

• 9 Months ago

Senior HPC AI Cluster Engineer

NVIDIA

California, United States (Remote)

• 7 Months ago

AI Artist (Portrait Specialist)

Scopely

Bengaluru, Karnataka, India (On-Site)

• 6 Months ago

Research Scientist in Foundation Model (Speech & Audio Generation) - 2025 Start (PhD）

ByteDance

Seattle, Washington, United States (On-Site)

• 10 Months ago

Software Engineer, PhD, Early Career, Campus, Machine Learning, Systems and Cloud AI, 2025 start

Google

Sunnyvale, California, United States (On-Site)

• 8 Months ago

Artificial Intelligence (AI) Technical Director (TD)

Luma Pictures

Vancouver, British Columbia, Canada (On-Site)

• 8 Months ago

Research Scientist- Foundation Model, Video Generation

ByteDance

Seattle, Washington, United States (On-Site)

• 10 Months ago

Get notifed when new similar jobs are uploaded

Similar Skill Jobs

Site Reliability Engineer

Argus Labs

Calgary, Alberta, Canada (Remote)

• 5 Months ago

Data Scientist

Virtusa

Andhra Pradesh, India (On-Site)

• 11 Months ago

Distributed Cloud | Senior AWS Cloud Engineer

DEVOTEAM

Lisbon, Lisbon, Portugal (Remote)

• 10 Months ago

Principal DevOps Engineer - Star Trek Fleet Command

Scopely

United Kingdom (Remote)

• 5 Months ago

Sr. Infrastructure Engineer

Intel Corporation

Hillsboro, Oregon, United States (On-Site)

• 9 Months ago

Senior Data Platform Engineer

Take-Two Interactive

Bengaluru, Karnataka, India (On-Site)

• 8 Months ago

Compliance Engineer

Eleven Labs

(Remote)

• 10 Months ago

QA Engineer

Scopely

Bengaluru, Karnataka, India (Hybrid)

• 5 Months ago

DevOps Engineer

Ajmera Infotech

San Jose, California, United States (On-Site)

• 11 Months ago

Associate Automation Testing

Golden Opportunities

Bengaluru, Karnataka, India (On-Site)

• 1 Year ago

Get notifed when new similar jobs are uploaded

Jobs in Santa Clara, California, United States

Commercial Accounts Administrator

Bally's Interactive

Malta, New York, United States (On-Site)

• 7 Months ago

Senior Lead UX Researcher

Warner Bros Games

Seattle, Washington, United States (Hybrid)

• 6 Months ago

Sales Development Representative - Arizona

AppZen

Phoenix, Arizona, United States (Hybrid)

• 10 Months ago

Infrastructure Engineering Architect - Central Tech

Bungie

United States (Hybrid)

• 8 Months ago

Principal Technical Product Manager - Application Security

Crunchyroll

San Francisco, California, United States (On-Site)

• 6 Months ago

Associate Manager E-Commerce Merchandising

Mattel Inc

El Segundo, California, United States (On-Site)

• 10 Months ago

Features Software Engineer (Senior)

Bonfire Studios

California, United States (On-Site)

• 1 Year ago

Color Assist (Episodic)

Company3 Method Studios

Los Angeles, California, United States (On-Site)

• 6 Months ago

Manager - Robot Platform Safety, Trajectory Generation

Zoox

Foster City, California, United States (Hybrid)

• 10 Months ago

Sales Team Lead

Dmg

Cincinnati, Ohio, United States (On-Site)

• 8 Months ago

Get notifed when new similar jobs are uploaded

Research Development Jobs

Senior C++ Programmer - Machine Learning

Ubisoft

Montreal, Quebec, Canada (On-Site)

• 5 Months ago

AI Researcher

Windranger Labs

Singapore (On-Site)

• 6 Months ago

Research Engineer Intern (Doubao (Seed) - Machine Learning System) - 2025 Summer (MS)

ByteDance

Seattle, Washington, United States (On-Site)

• 10 Months ago

Senior Software Engineer, Machine Learning, Google Ads

Google

Los Angeles, California, United States (On-Site)

• 8 Months ago

Senior Research Engineer, Foundation Model Training Infrastructure

NVIDIA

Santa Clara, California, United States (On-Site)

• 7 Months ago

Research Scientist Intern, Language and Multimodal Research for MetaAI (PhD)

Data Scientist

Enterprise Bot

Bengaluru, Karnataka, India (On-Site)

• 10 Months ago

Instructional Designer and Facilitator

PlayStation Global

London, England, United Kingdom (Hybrid)

• 6 Months ago

Research Engineer (Foundation Model) - Machine Learning Systems

ByteDance

Singapore (On-Site)

• 10 Months ago

Technical Artist - Generative AI

Magnopus

Los Angeles, California, United States (On-Site)

• 10 Months ago

Get notifed when new similar jobs are uploaded

About The Company

NVIDIA

70 Active Jobs

Since its founding in 1993, NVIDIA (NASDAQ: NVDA) has been a pioneer in accelerated computing. The company’s invention of the GPU in 1999 sparked the growth of the PC gaming market, redefined computer graphics, ignited the era of modern AI and is fueling the creation of the metaverse. NVIDIA is now a full-stack computing company with data-center-scale offerings that are reshaping industry.

Get notified when new jobs are added by NVIDIA

Level Up Your Career in Game Development!

Transform Your Passion into Profession with Our Comprehensive Courses for Aspiring Game Developers.

A global community of game builders. Helping people upskill and land jobs in the best gaming studios.

Company

Key Links

hello@outscal.com

Made in INDIA 💛💙

Senior HPC and AI Networking Performance Research and Analysis Engineer

Job Summary

Job Description

14 skills required

14 skills required for this role

Job Details

Similar Jobs

Software Integration Engineer 2 (IDN - 057)

Senior Platform Engineer - Design

Python (Golang) Developer

Solutions Architect

Senior HPC AI Cluster Engineer

AI Artist (Portrait Specialist)

Research Scientist in Foundation Model (Speech & Audio Generation) - 2025 Start (PhD）

Software Engineer, PhD, Early Career, Campus, Machine Learning, Systems and Cloud AI, 2025 start

Artificial Intelligence (AI) Technical Director (TD)

Research Scientist- Foundation Model, Video Generation

Similar Skill Jobs

Site Reliability Engineer

Data Scientist

Distributed Cloud | Senior AWS Cloud Engineer

Principal DevOps Engineer - Star Trek Fleet Command

Sr. Infrastructure Engineer

Senior Data Platform Engineer

Compliance Engineer

QA Engineer

DevOps Engineer

Associate Automation Testing

Jobs in Santa Clara, California, United States

Commercial Accounts Administrator

Senior Lead UX Researcher

Sales Development Representative - Arizona

Infrastructure Engineering Architect - Central Tech

Principal Technical Product Manager - Application Security

Associate Manager E-Commerce Merchandising

Features Software Engineer (Senior)

Color Assist (Episodic)

Manager - Robot Platform Safety, Trajectory Generation

Sales Team Lead

Research Development Jobs

Senior C++ Programmer - Machine Learning

AI Researcher

Research Engineer Intern (Doubao (Seed) - Machine Learning System) - 2025 Summer (MS)

Senior Software Engineer, Machine Learning, Google Ads

Senior Research Engineer, Foundation Model Training Infrastructure

Research Scientist Intern, Language and Multimodal Research for MetaAI (PhD)

Data Scientist

Instructional Designer and Facilitator

Research Engineer (Foundation Model) - Machine Learning Systems

Technical Artist - Generative AI

About The Company

System Design Power Validation Engineer

OEM Account Manager

System Debug Lead Engineer

Network Site Reliability Engineer

ASIC Engineer

Senior ASIC Design Engineer

Physical Design CAD Team Manager

Engineering Farm Engineer

Senior Mixed Signal Design Verification Engineer

Senior Solutions Architect, Cloud Infrastructure and DevOps

Level Up Your Career in Game Development!