Deep Learning Engineer, Datacenters

2 Months ago • 3 Years + • Research & Development

Job Summary

Job Description

NVIDIA's Deep Learning Engineer in Datacenters will help develop software infrastructure to analyze deep learning applications, evolve cost-efficient datacenter architectures for LLMs, and work with experts to develop analysis and profiling tools in Python, bash, and C++. Responsibilities involve analyzing system and software characteristics of DL applications, developing analysis tools, and measuring key performance metrics to estimate efficiency improvements. The role requires collaboration with various teams across NVIDIA, from research to silicon architecture. The ideal candidate will have experience with system software, GPU kernels, or DL frameworks and a strong understanding of system architecture and performance.
Must have:
  • Bachelor's degree in EE/CS (Master's/PhD preferred)
  • 3+ years relevant experience
  • System software/Silicon architecture experience
  • C/C++ and Python programming
  • Deep Learning application analysis
Good to have:
  • CUDA, PyTorch, TensorFlow
  • Containerization (Docker), Slurm
  • Performance monitoring tools (perf, gprof)
  • Performance modeling (CPU, GPU, Memory, Network)
  • Multi-site/functional team experience

Job Details

As NVIDIA makes inroads into the Datacenter business, our team plays a central role in getting the most out of our exponentially growing datacenter deployments as well as establishing a data-driven approach to hardware design and system software development. We collaborate with a broad cross section of teams at NVIDIA ranging from DL research teams to CUDA Kernel and DL Framework development teams, to Silicon Architecture Teams. As our team grows, and as we seek to identify and take advantage of long term opportunities, our skillset needs are expanding as well.

Do you want to influence the development of high-performance Datacenters designed for the future of AI? Do you have an interest in system architecture and performance? In this role you will find how CPU, GPU, networking, and IO relate to deep learning (DL) architectures for Natural Language Processing, Computer Vision, Autonomous Driving and other technologies. Come join our team, and bring your interests to help us optimize our next generation systems and Deep Learning Software Stack.

What you'll be doing:

  • Help develop software infrastructure to characterize and analyze a broad range Deep Learning applications
  • Evolve cost-efficient datacenter architectures tailored to meet the needs of Large Language Models (LLMs).
  • Work with experts to help develop analysis and profiling tools in Python, bash and C++ to measure key performance metrics of DL workloads running on Nvidia systems.
  • Analyze system and software characteristics of DL applications.
  • Develop analysis tools and methodologies to measure key performance metrics and to estimate potential for efficiency improvement.

What we need to see:

  • A Bachelor’s degree in Electrical Engineering or Computer Science with 3 years or more of relevant experience (Masters or PhD degree preferred)
  • Experience in at least one of the following:
    • System Software: Operating Systems (Linux), Compilers, GPU kernels (CUDA), DL Frameworks (PyTorch, TensorFlow).
    • Silicon Architecture and Performance Modeling/Analysis: CPU, GPU, Memory or Network Architecture
  • Experience programming in C/C++ and Python. Exposure to Containerization Platforms (docker) and Datacenter Workload Managers (slurm) is a plus
  • Demonstrated ability to work in virtual environments, and a strong drive to own tasks from beginning to end. Prior experience with such environments will make you stand out.

Ways to stand out from the crowd:

  • Background with system software, Operating system intrinsics, GPU kernels (CUDA), or DL Frameworks (PyTorch, TensorFlow).

  • Experience with silicon performance monitoring or profiling tools (e.g. perf, gprof, nvidia-smi, dcgm).

  • In depth performance modeling experience in any one of CPU, GPU, Memory or Network Architecture

  • Exposure to Containerization Platforms (docker) and Datacenter Workload Managers (slurm).

  • Prior experience with multi-site teams or multi-functional teams.

NVIDIA is widely considered to be one of the technology world’s most desirable employers. We have some of the most forward-thinking and hardworking people on the planet working for us. If you're creative and autonomous, we want to hear from you!

#LI-Hybrid

Similar Jobs

Wargaming - DevOps Engineer (World of Warships, PC)

Wargaming

Belgrade, Serbia (Hybrid)
1 Month ago
Motorola solutions - Redhat Openshift Virtualization Administrator

Motorola solutions

Bengaluru, Karnataka, India (On-Site)
2 Weeks ago
Rackspace Technology - Senior Systems Engineer HPC

Rackspace Technology

United States (Remote)
1 Month ago
FICO - CCS DevOps - Engineer II

FICO

Guadalajara, Jalisco, Mexico (Remote)
2 Weeks ago
Thales - Storage/Back-up Administrator

Thales

Bucharest, Bucharest, Romania (Hybrid)
2 Weeks ago
bytedance - Software Engineer, Architecture and Infrastructure

bytedance

San Jose, California, United States (On-Site)
7 Months ago
NXP - Senior Principal Software Architect - Platform and RF Software

NXP

Bucharest, Bucharest, Romania (On-Site)
8 Months ago
Google - Student Researcher, BS/MS, Winter/Summer 2025

Google

Mountain View, California, United States (On-Site)
6 Months ago
Sony Interactive Entertainment - PlayStation向けカスタムLSIの開発・評価エンジニア

Sony Interactive Entertainment

Tokyo, Japan (On-Site)
7 Months ago
Krafton - [Corp Dev Div.] Investment Team Member (3년~8년)

Krafton

Seoul, South Korea (On-Site)
7 Months ago

Get notifed when new similar jobs are uploaded

Similar Skill Jobs

Thales - Product Support Engineer- Data Security

Thales

Mexico (Remote)
2 Weeks ago
Ion - Markets Product Security Engineer - UK

Ion

London, England, United Kingdom (On-Site)
7 Months ago
rivos - SOC Static Timing Analysis Engineer - Full Time

rivos

Hsinchu, Hsinchu City, Taiwan (On-Site)
7 Months ago
playrix  - Senior Release Automation Engineer (Gardenscapes)

playrix

Ireland (Remote)
4 Months ago
Aisera - Customer Engineer

Aisera

Palo Alto, California, United States (On-Site)
2 Weeks ago
NVIDIA - Senior Site Reliability Engineer - AI Research Clusters

NVIDIA

Westford, Massachusetts, United States (Hybrid)
3 Months ago
Capgemini - Site Reliability Engineer-Wintel, Linux, Vmware, Redhat Devops CI/CD AWS

Capgemini

Bengaluru, Karnataka, India (On-Site)
3 Weeks ago
luxsoft - Senior Angular Developer

luxsoft

Ukraine (Remote)
2 Weeks ago
Imanage - Site Reliability Engineer

Imanage

Chicago, Illinois, United States (Hybrid)
1 Week ago
GoReel - DevOps Lead

GoReel

Romania (Remote)
2 Months ago

Get notifed when new similar jobs are uploaded

Jobs in Bengaluru, Karnataka, India

PTW - Trainee Test Engineer

PTW

Hyderabad, Telangana, India (On-Site)
4 Months ago
Merqube - Quant Analyst (Financial Engineer)

Merqube

Bengaluru, Karnataka, India (Hybrid)
4 Days ago
Dream Sports - MISE Sales Manager - Meetings, Incentives & Sporting Experiences

Dream Sports

Delhi, India (On-Site)
2 Months ago
fluence - Senior Control Software Engineer - II

fluence

Bengaluru, Karnataka, India (On-Site)
1 Year ago
PwC - Associate - Python Data Engineer - GDC

PwC

Kolkata, West Bengal, India (On-Site)
8 Months ago
Ethernovia - Principal Software Application Engineer

Ethernovia

Pune, Maharashtra, India (On-Site)
2 Weeks ago
Thousand Eyes - Site Reliability Engineering Technical Leader, Network Assurance Data Platform

Thousand Eyes

Bengaluru, Karnataka, India (On-Site)
2 Weeks ago
PhonePe - Account Manager / Manager - Agency Ad Sales

PhonePe

Delhi, India (On-Site)
4 Weeks ago
PhonePe - Marketing Lead

PhonePe

Bengaluru, Karnataka, India (On-Site)
2 Weeks ago
Addepar - Sr. Software Data Engineer

Addepar

Pune, Maharashtra, India (Hybrid)
3 Weeks ago

Get notifed when new similar jobs are uploaded

Research & Development Jobs

Riot Games - Associate Art Director - League of Legends, Game Modes

Riot Games

Sydney, New South Wales, Australia (On-Site)
11 Months ago
The Walt Disney Company - Mechanical Designer, CAD Designer

The Walt Disney Company

Shanghai, Shanghai, China (On-Site)
3 Months ago
Riot Games - Principal Software Engineer, Product Tech-Lead - Unpublished R&D Product

Riot Games

Dublin, County Dublin, Ireland (On-Site)
6 Months ago
Tesla - Electrical Engineer - Motor Design and Multi-Physics Optimization

Tesla

Athens, Greece (On-Site)
3 Months ago
KPIT - Embedded C Expert

KPIT

Bengaluru, Karnataka, India (On-Site)
8 Months ago
bytedance - Research Scientist, Reinforcement Learning

bytedance

Seattle, Washington, United States (On-Site)
7 Months ago
NVIDIA - Layout Design Engineer

NVIDIA

Bengaluru, Karnataka, India (Hybrid)
2 Months ago
Google - Software Engineering Manager, Google Store

Google

Bengaluru, Karnataka, India (On-Site)
1 Month ago
Krafton - Applied Research Scientist/Engineer - LM/Agent

Krafton

Seoul, South Korea (On-Site)
2 Months ago
Daybreak Game Company LLC - Software Development Engineer (Cardset)

Daybreak Game Company LLC

Renton, Washington, United States (Remote)
6 Months ago

Get notifed when new similar jobs are uploaded

About The Company

Since its founding in 1993, NVIDIA (NASDAQ: NVDA) has been a pioneer in accelerated computing. The company’s invention of the GPU in 1999 sparked the growth of the PC gaming market, redefined computer graphics, ignited the era of modern AI and is fueling the creation of the metaverse. NVIDIA is now a full-stack computing company with data-center-scale offerings that are reshaping industry.

Santa Clara, California, United States (On-Site)

Massachusetts, United States (On-Site)

Santa Clara, California, United States (On-Site)

Santa Clara, California, United States (On-Site)

Santa Clara, California, United States (On-Site)

Texas, United States (On-Site)

Santa Clara, California, United States (Hybrid)

Santa Clara, California, United States (Hybrid)

Pune, Maharashtra, India (On-Site)

Taipei City, Taiwan (On-Site)

View All Jobs

Get notified when new jobs are added by NVIDIA

Level Up Your Career in Game Development!

Transform Your Passion into Profession with Our Comprehensive Courses for Aspiring Game Developers.

Job Common Plug