Deep Learning Engineer, Datacenters

4 Weeks ago • 3 Years + • Research & Development

Job Summary

Job Description

NVIDIA's Deep Learning Engineer in Datacenters will help develop software infrastructure to analyze deep learning applications, evolve cost-efficient datacenter architectures for LLMs, and work with experts to develop analysis and profiling tools in Python, bash, and C++. Responsibilities involve analyzing system and software characteristics of DL applications, developing analysis tools, and measuring key performance metrics to estimate efficiency improvements. The role requires collaboration with various teams across NVIDIA, from research to silicon architecture. The ideal candidate will have experience with system software, GPU kernels, or DL frameworks and a strong understanding of system architecture and performance.
Must have:
  • Bachelor's degree in EE/CS (Master's/PhD preferred)
  • 3+ years relevant experience
  • System software/Silicon architecture experience
  • C/C++ and Python programming
  • Deep Learning application analysis
Good to have:
  • CUDA, PyTorch, TensorFlow
  • Containerization (Docker), Slurm
  • Performance monitoring tools (perf, gprof)
  • Performance modeling (CPU, GPU, Memory, Network)
  • Multi-site/functional team experience

Job Details

As NVIDIA makes inroads into the Datacenter business, our team plays a central role in getting the most out of our exponentially growing datacenter deployments as well as establishing a data-driven approach to hardware design and system software development. We collaborate with a broad cross section of teams at NVIDIA ranging from DL research teams to CUDA Kernel and DL Framework development teams, to Silicon Architecture Teams. As our team grows, and as we seek to identify and take advantage of long term opportunities, our skillset needs are expanding as well.

Do you want to influence the development of high-performance Datacenters designed for the future of AI? Do you have an interest in system architecture and performance? In this role you will find how CPU, GPU, networking, and IO relate to deep learning (DL) architectures for Natural Language Processing, Computer Vision, Autonomous Driving and other technologies. Come join our team, and bring your interests to help us optimize our next generation systems and Deep Learning Software Stack.

What you'll be doing:

  • Help develop software infrastructure to characterize and analyze a broad range Deep Learning applications
  • Evolve cost-efficient datacenter architectures tailored to meet the needs of Large Language Models (LLMs).
  • Work with experts to help develop analysis and profiling tools in Python, bash and C++ to measure key performance metrics of DL workloads running on Nvidia systems.
  • Analyze system and software characteristics of DL applications.
  • Develop analysis tools and methodologies to measure key performance metrics and to estimate potential for efficiency improvement.

What we need to see:

  • A Bachelor’s degree in Electrical Engineering or Computer Science with 3 years or more of relevant experience (Masters or PhD degree preferred)
  • Experience in at least one of the following:
    • System Software: Operating Systems (Linux), Compilers, GPU kernels (CUDA), DL Frameworks (PyTorch, TensorFlow).
    • Silicon Architecture and Performance Modeling/Analysis: CPU, GPU, Memory or Network Architecture
  • Experience programming in C/C++ and Python. Exposure to Containerization Platforms (docker) and Datacenter Workload Managers (slurm) is a plus
  • Demonstrated ability to work in virtual environments, and a strong drive to own tasks from beginning to end. Prior experience with such environments will make you stand out.

Ways to stand out from the crowd:

  • Background with system software, Operating system intrinsics, GPU kernels (CUDA), or DL Frameworks (PyTorch, TensorFlow).

  • Experience with silicon performance monitoring or profiling tools (e.g. perf, gprof, nvidia-smi, dcgm).

  • In depth performance modeling experience in any one of CPU, GPU, Memory or Network Architecture

  • Exposure to Containerization Platforms (docker) and Datacenter Workload Managers (slurm).

  • Prior experience with multi-site teams or multi-functional teams.

NVIDIA is widely considered to be one of the technology world’s most desirable employers. We have some of the most forward-thinking and hardworking people on the planet working for us. If you're creative and autonomous, we want to hear from you!

#LI-Hybrid

Similar Jobs

Next Level Business Services - Full Stack Developer

Next Level Business Services

Jersey City, New Jersey, United States (On-Site)
6 Months ago
ION - Senior DevSecOps Engineer, Italy

ION

Collecchio, Emilia-Romagna, Italy (On-Site)
6 Months ago
Google - Systems Development Engineer, Edge Infrastructure Operations

Google

Dublin, County Dublin, Ireland (On-Site)
1 Week ago
SmartBear - Customer Success Engineer - Test Hub

SmartBear

Ahmedabad, Gujarat, India (On-Site)
1 Day ago
Naughty Dog - IT Help Desk Technician

Naughty Dog

Los Angeles, California, United States (On-Site)
1 Week ago
NVIDIA - Principal Autonomous Vehicles Engineer - Mapping and Localization

NVIDIA

Shanghai, Shanghai, China (On-Site)
3 Months ago
NVIDIA - HPC Operations Manager – Hardware Engineering

NVIDIA

Westford, Massachusetts, United States (On-Site)
2 Months ago
NVIDIA - Power Integrity Engineer

NVIDIA

Tel Aviv-Yafo, Tel Aviv District, Israel (On-Site)
2 Months ago
Riot Games - Principal Software Engineer, Foundations Developer Experience & Workflows

Riot Games

Los Angeles, California, United States (On-Site)
6 Months ago
Google - SoC and IP Design Engineer

Google

Haifa, Haifa District, Israel (On-Site)
2 Weeks ago

Get notifed when new similar jobs are uploaded

Similar Skill Jobs

Nintendo - DevOps Engineer

Nintendo

Redmond, Washington, United States (On-Site)
3 Months ago
Next Level Business Services - Linux Scripting and Clear case SME Consultant

Next Level Business Services

Milwaukee, Wisconsin, United States (On-Site)
6 Months ago
ZeniMax Media - DevOps Engineer

ZeniMax Media

Austin, Texas, United States (Remote)
1 Month ago
Tala - Telephony and Network Lead

Tala

Manila, Metro Manila, Philippines (Hybrid)
1 Month ago
JMA - Principal Firmware Engineer - Radio

JMA

Plano, Texas, United States (On-Site)
6 Months ago
Netflix - Senior Software Engineer (L5) - Developer Infrastructure

Netflix

Los Gatos, California, United States (On-Site)
2 Weeks ago
Velotio Technologies - Senior DevOps Engineer (GCP)

Velotio Technologies

Pune, Maharashtra, India (Remote)
1 Month ago
NVIDIA - Senior Network Engineer

NVIDIA

Yokne'am Illit, North District, Israel (On-Site)
3 Months ago
PwC - Senior Associate_Azure Data Engineer_Data & Analytics_Advisory_PAN  India

PwC

Kolkata, West Bengal, India (On-Site)
7 Months ago
N-iX - Senior Performance Test Engineer

N-iX

Ukraine (Remote)
2 Weeks ago

Get notifed when new similar jobs are uploaded

Jobs in Bengaluru, Karnataka, India

Dassault Systèmes - Localization and Translation Specialist

Dassault Systèmes

Mumbai, Maharashtra, India (Hybrid)
6 Months ago
WinZO - Corp Development & Investments

WinZO

New Delhi, Delhi, India (On-Site)
1 Day ago
Knack Studios - 3D Artist

Knack Studios

Chennai, Tamil Nadu, India (On-Site)
10 Months ago
Nagarro - Associate Staff Engineer, Mobile Hybrid

Nagarro

India (Remote)
6 Months ago
PwC - SAP ABAP-Manager

PwC

Kolkata, West Bengal, India (On-Site)
7 Months ago
P99 soft - Data Engineer

P99 soft

Hyderabad, Telangana, India (On-Site)
17 Hours ago
Index Exchange - Staff Software Engineer

Index Exchange

Bengaluru, Karnataka, India (Hybrid)
7 Months ago
JMA - Senior Engineer - ORUC - QA

JMA

Bengaluru, Karnataka, India (Hybrid)
3 Weeks ago
Google - Business Development Consultant, New Business Sales

Google

Bengaluru, Karnataka, India (On-Site)
2 Days ago
Novo - KYC -BSA/AML - Assistant Manager

Novo

Gurugram, Haryana, India (On-Site)
1 Day ago

Get notifed when new similar jobs are uploaded

Research & Development Jobs

Samsung Semiconductor - Staff Engineer, Performance Modeling Architect

Samsung Semiconductor

San Jose, California, United States (Hybrid)
1 Month ago
NXP - <2025 Internship Program> Application Engineer

NXP

Taipei City, Taiwan (On-Site)
6 Months ago
NVIDIA - Senior Mixed Signal Design Engineer

NVIDIA

Hsinchu, Hsinchu City, Taiwan (On-Site)
3 Months ago
Cadence - Sr Principal Product Validation Engineer

Cadence

Noida, Uttar Pradesh, India (On-Site)
7 Months ago
Netflix - Engineering Manager, Compute Controlplane and Capacity

Netflix

United States (Remote)
2 Weeks ago
Google - Software Engineering Manager II, Chrome OS

Google

San Jose, California, United States (On-Site)
2 Weeks ago
NVIDIA - Mixed Signal Analog Circuit Designer (RDSS Intern)

NVIDIA

Hsinchu, Hsinchu City, Taiwan (On-Site)
3 Months ago
ByteDance - Research Scientist, Reinforcement Learning

ByteDance

San Jose, California, United States (On-Site)
6 Months ago
Rockstar Games - Engineering Manager

Rockstar Games

New York, New York, United States (On-Site)
1 Month ago
NVIDIA - Senior Circuit Design Engineer

NVIDIA

Austin, Texas, United States (Hybrid)
2 Weeks ago

Get notifed when new similar jobs are uploaded

About The Company

Since its founding in 1993, NVIDIA (NASDAQ: NVDA) has been a pioneer in accelerated computing. The company’s invention of the GPU in 1999 sparked the growth of the PC gaming market, redefined computer graphics, ignited the era of modern AI and is fueling the creation of the metaverse. NVIDIA is now a full-stack computing company with data-center-scale offerings that are reshaping industry.

Santa Clara, California, United States (On-Site)

Santa Clara, California, United States (On-Site)

Massachusetts, United States (On-Site)

Santa Clara, California, United States (On-Site)

Santa Clara, California, United States (On-Site)

Santa Clara, California, United States (On-Site)

Texas, United States (On-Site)

Santa Clara, California, United States (On-Site)

Santa Clara, California, United States (Hybrid)

Santa Clara, California, United States (Hybrid)

View All Jobs

Get notified when new jobs are added by NVIDIA

Level Up Your Career in Game Development!

Transform Your Passion into Profession with Our Comprehensive Courses for Aspiring Game Developers.

Job Common Plug