Deep Learning Engineer, Datacenters

3 Months ago • 3 Years + • Research Development

Job Summary

Job Description

NVIDIA's Deep Learning Engineer in Datacenters will help develop software infrastructure to analyze deep learning applications, evolve cost-efficient datacenter architectures for LLMs, and work with experts to develop analysis and profiling tools in Python, bash, and C++. Responsibilities involve analyzing system and software characteristics of DL applications, developing analysis tools, and measuring key performance metrics to estimate efficiency improvements. The role requires collaboration with various teams across NVIDIA, from research to silicon architecture. The ideal candidate will have experience with system software, GPU kernels, or DL frameworks and a strong understanding of system architecture and performance.
Must have:
  • Bachelor's degree in EE/CS (Master's/PhD preferred)
  • 3+ years relevant experience
  • System software/Silicon architecture experience
  • C/C++ and Python programming
  • Deep Learning application analysis
Good to have:
  • CUDA, PyTorch, TensorFlow
  • Containerization (Docker), Slurm
  • Performance monitoring tools (perf, gprof)
  • Performance modeling (CPU, GPU, Memory, Network)
  • Multi-site/functional team experience

Job Details

As NVIDIA makes inroads into the Datacenter business, our team plays a central role in getting the most out of our exponentially growing datacenter deployments as well as establishing a data-driven approach to hardware design and system software development. We collaborate with a broad cross section of teams at NVIDIA ranging from DL research teams to CUDA Kernel and DL Framework development teams, to Silicon Architecture Teams. As our team grows, and as we seek to identify and take advantage of long term opportunities, our skillset needs are expanding as well.

Do you want to influence the development of high-performance Datacenters designed for the future of AI? Do you have an interest in system architecture and performance? In this role you will find how CPU, GPU, networking, and IO relate to deep learning (DL) architectures for Natural Language Processing, Computer Vision, Autonomous Driving and other technologies. Come join our team, and bring your interests to help us optimize our next generation systems and Deep Learning Software Stack.

What you'll be doing:

  • Help develop software infrastructure to characterize and analyze a broad range Deep Learning applications
  • Evolve cost-efficient datacenter architectures tailored to meet the needs of Large Language Models (LLMs).
  • Work with experts to help develop analysis and profiling tools in Python, bash and C++ to measure key performance metrics of DL workloads running on Nvidia systems.
  • Analyze system and software characteristics of DL applications.
  • Develop analysis tools and methodologies to measure key performance metrics and to estimate potential for efficiency improvement.

What we need to see:

  • A Bachelor’s degree in Electrical Engineering or Computer Science with 3 years or more of relevant experience (Masters or PhD degree preferred)
  • Experience in at least one of the following:
    • System Software: Operating Systems (Linux), Compilers, GPU kernels (CUDA), DL Frameworks (PyTorch, TensorFlow).
    • Silicon Architecture and Performance Modeling/Analysis: CPU, GPU, Memory or Network Architecture
  • Experience programming in C/C++ and Python. Exposure to Containerization Platforms (docker) and Datacenter Workload Managers (slurm) is a plus
  • Demonstrated ability to work in virtual environments, and a strong drive to own tasks from beginning to end. Prior experience with such environments will make you stand out.

Ways to stand out from the crowd:

  • Background with system software, Operating system intrinsics, GPU kernels (CUDA), or DL Frameworks (PyTorch, TensorFlow).

  • Experience with silicon performance monitoring or profiling tools (e.g. perf, gprof, nvidia-smi, dcgm).

  • In depth performance modeling experience in any one of CPU, GPU, Memory or Network Architecture

  • Exposure to Containerization Platforms (docker) and Datacenter Workload Managers (slurm).

  • Prior experience with multi-site teams or multi-functional teams.

NVIDIA is widely considered to be one of the technology world’s most desirable employers. We have some of the most forward-thinking and hardworking people on the planet working for us. If you're creative and autonomous, we want to hear from you!

#LI-Hybrid

Similar Jobs

Qualcomm - GPU Lead Engineer

Qualcomm

Bengaluru, Karnataka, India (On-Site)
1 Month ago
Brillio - Associate Manager, Finance

Brillio

Bengaluru, Karnataka, India (Hybrid)
4 Months ago
FalconX - Senior Risk Management Associate, Derivatives

FalconX

New York, New York, United States (On-Site)
2 Months ago
WebFX - Inbound Sales Specialist

WebFX

Harrisburg, Pennsylvania, United States (On-Site)
8 Months ago
Wrike - Marketing Operations Program Coordinator

Wrike

Costa Rica (Hybrid)
1 Month ago
Microsoft - Senior Technical Program Manager, Copilot AI

Microsoft

Mountain View, California, United States (Hybrid)
2 Months ago
NVIDIA - Deep Learning Performance Architect

NVIDIA

Bengaluru, Karnataka, India (Hybrid)
5 Months ago
Thales - Quantum-AI Research Scientist

Thales

Montreal, Quebec, Canada (On-Site)
1 Month ago
Meta - Software Engineer, Machine Learning

Meta

Singapore (On-Site)
7 Months ago
QuinStreet - Machine Learning Engineer

QuinStreet

Monterrey, Nuevo Leon, Mexico (Remote)
2 Months ago

Get notifed when new similar jobs are uploaded

Similar Skill Jobs

ISG - Product Implementation Manager

ISG

Bengaluru, Karnataka, India (On-Site)
2 Months ago
PwC - IN-Senior Associate - Pharma Commercial -Health Industries-Advisory

PwC

Mumbai, Maharashtra, India (On-Site)
3 Weeks ago
Nasdaq - Technical Configuration Manager for LATAM and Americas

Nasdaq

Mexico City, Mexico City, Mexico (On-Site)
2 Months ago
Keywords Studios - Senior Business Development Manager

Keywords Studios

England, United Kingdom (Remote)
3 Months ago
Nium - Senior Finance Analyst – Strategic Finance & Corporate Development

Nium

Mumbai, Maharashtra, India (Hybrid)
3 Weeks ago
Paytm - Go-To-Market Lead - Deputy General Manager - Offline Merchants QR

Paytm

Ahmedabad, Gujarat, India (On-Site)
1 Month ago
Sony Music Career - Digital Optimization & CRM Manager

Sony Music Career

Bangkok, Thailand (On-Site)
2 Months ago
PwC - Event Management Consultant

PwC

Singapore (On-Site)
3 Weeks ago
Daily Wire - Social Content Producer

Daily Wire

Nashville, Tennessee, United States (On-Site)
2 Months ago

Get notifed when new similar jobs are uploaded

Jobs in Bengaluru, Karnataka, India

Google - Software Engineer III, Mobile, Android

Google

Bengaluru, Karnataka, India (On-Site)
7 Months ago
Loyalty Juggernaut - Product Engineer (Class of 2026)

Loyalty Juggernaut

Hyderabad, Telangana, India (On-Site)
1 Month ago
Capgemini - ALR Developer

Capgemini

Hyderabad, Telangana, India (On-Site)
3 Weeks ago
frames store - Imaging Support Engineer

frames store

Mumbai, Maharashtra, India (On-Site)
3 Months ago
HCL Tech - Manager - regulatory affairs

HCL Tech

Madurai, Tamil Nadu, India (On-Site)
1 Month ago
Rigi - Scriptwriter / Content Strategist

Rigi

Bengaluru, Karnataka, India (On-Site)
1 Month ago
Capgemini - Storage Administration

Capgemini

Mumbai, Maharashtra, India (On-Site)
1 Month ago
Capgemini - SAP BODS Admin

Capgemini

Hyderabad, Telangana, India (On-Site)
1 Month ago
Enphase Energy - Staff Engineer EV Charger Embedded Software and Energy Management Gateway

Enphase Energy

Bengaluru, Karnataka, India (On-Site)
6 Months ago
PwC - IN-Senior Associate / Senior Associate_ IDAM / Identity Management_Advisory

PwC

Ahmedabad, Gujarat, India (On-Site)
3 Weeks ago

Get notifed when new similar jobs are uploaded

Research Development Jobs

NACON - AI Programmer

NACON

Milan, Lombardy, Italy (On-Site)
1 Month ago
CD PROJEKT RED - Engineer, AI & Navigation

CD PROJEKT RED

Warsaw, Masovian Voivodeship, Poland (On-Site)
2 Months ago
Apple - Machine Learning Engineer, Siri Attention & Invocation

Apple

Seattle, Washington, United States (On-Site)
1 Week ago
Scale AI - Software Engineer, Frontend - Enterprise Gen AI

Scale AI

San Francisco, California, United States (On-Site)
2 Months ago
Match Group - Sr. Software Engineer, Machine Learning

Match Group

Palo Alto, California, United States (Hybrid)
1 Month ago
Scale AI - Applied AI Engineer, Autonomous Agents

Scale AI

San Francisco, California, United States (On-Site)
1 Month ago
Ubisoft - Junior R&D Engineer

Ubisoft

Pune, Maharashtra, India (Hybrid)
3 Weeks ago
Globalization Partners - AI Intern

Globalization Partners

Ireland (Remote)
1 Month ago
Capgemini - ML OPS

Capgemini

Hyderabad, Telangana, India (On-Site)
4 Weeks ago
DNEG - Machine Learning Engineering Lead, Ziva

DNEG

London, England, United Kingdom (Remote)
3 Months ago

Get notifed when new similar jobs are uploaded

About The Company

Since its founding in 1993, NVIDIA (NASDAQ: NVDA) has been a pioneer in accelerated computing. The company’s invention of the GPU in 1999 sparked the growth of the PC gaming market, redefined computer graphics, ignited the era of modern AI and is fueling the creation of the metaverse. NVIDIA is now a full-stack computing company with data-center-scale offerings that are reshaping industry.

Santa Clara, California, United States (On-Site)

Massachusetts, United States (On-Site)

Santa Clara, California, United States (On-Site)

Texas, United States (On-Site)

Santa Clara, California, United States (Hybrid)

Santa Clara, California, United States (Hybrid)

Pune, Maharashtra, India (On-Site)

Taipei City, Taiwan (On-Site)

Beijing, Beijing, China (On-Site)

Santa Clara, California, United States (On-Site)

View All Jobs

Get notified when new jobs are added by NVIDIA

Level Up Your Career in Game Development!

Transform Your Passion into Profession with Our Comprehensive Courses for Aspiring Game Developers.

Job Common Plug