Deep Learning Engineer, Datacenters

3 Months ago • 3 Years + • Research & Development

Job Summary

Job Description

NVIDIA's Deep Learning Engineer, Datacenters role focuses on optimizing next-generation systems and the deep learning software stack. Responsibilities include developing software infrastructure for analyzing deep learning applications, evolving cost-efficient datacenter architectures for LLMs, creating analysis and profiling tools (Python, bash, C++), analyzing system and software characteristics of DL applications, and developing methodologies to measure performance metrics. This role requires collaboration with various teams across NVIDIA, impacting the development of high-performance datacenters designed for the future of AI. The engineer will analyze how CPU, GPU, networking, and IO relate to deep learning architectures for various technologies.
Must have:
  • Bachelor's degree in EE or CS
  • 3+ years relevant experience
  • System software/Silicon architecture experience
  • C/C++ and Python programming
  • Strong analytical skills
Good to have:
  • GPU kernels (CUDA)
  • DL Frameworks (PyTorch, TensorFlow)
  • Containerization (Docker)
  • Datacenter Workload Managers (Slurm)
  • Performance modeling experience

Job Details

As NVIDIA makes inroads into the Datacenter business, our team plays a central role in getting the most out of our exponentially growing datacenter deployments as well as establishing a data-driven approach to hardware design and system software development. We collaborate with a broad cross section of teams at NVIDIA ranging from DL research teams to CUDA Kernel and DL Framework development teams, to Silicon Architecture Teams. As our team grows, and as we seek to identify and take advantage of long term opportunities, our skillset needs are expanding as well.

Do you want to influence the development of high-performance Datacenters designed for the future of AI? Do you have an interest in system architecture and performance? In this role you will find how CPU, GPU, networking, and IO relate to deep learning (DL) architectures for Natural Language Processing, Computer Vision, Autonomous Driving and other technologies. Come join our team, and bring your interests to help us optimize our next generation systems and Deep Learning Software Stack.

What you'll be doing:

  • Help develop software infrastructure to characterize and analyze a broad range Deep Learning applications
  • Evolve cost-efficient datacenter architectures tailored to meet the needs of Large Language Models (LLMs).
  • Work with experts to help develop analysis and profiling tools in Python, bash and C++ to measure key performance metrics of DL workloads running on Nvidia systems.
  • Analyze system and software characteristics of DL applications.
  • Develop analysis tools and methodologies to measure key performance metrics and to estimate potential for efficiency improvement.

What we need to see:

  • A Bachelor’s degree in Electrical Engineering or Computer Science with 3 years or more of relevant experience (Masters or PhD degree preferred)
  • Experience in at least one of the following:
    • System Software: Operating Systems (Linux), Compilers, GPU kernels (CUDA), DL Frameworks (PyTorch, TensorFlow).
    • Silicon Architecture and Performance Modeling/Analysis: CPU, GPU, Memory or Network Architecture
  • Experience programming in C/C++ and Python. Exposure to Containerization Platforms (docker) and Datacenter Workload Managers (slurm) is a plus
  • Demonstrated ability to work in virtual environments, and a strong drive to own tasks from beginning to end. Prior experience with such environments will make you stand out.

Ways to stand out from the crowd:

  • Background with system software, Operating system intrinsics, GPU kernels (CUDA), or DL Frameworks (PyTorch, TensorFlow).

  • Experience with silicon performance monitoring or profiling tools (e.g. perf, gprof, nvidia-smi, dcgm).

  • In depth performance modeling experience in any one of CPU, GPU, Memory or Network Architecture

  • Exposure to Containerization Platforms (docker) and Datacenter Workload Managers (slurm).

  • Prior experience with multi-site teams or multi-functional teams.

NVIDIA is widely considered to be one of the technology world’s most desirable employers. We have some of the most forward-thinking and hardworking people on the planet working for us. If you're creative and autonomous, we want to hear from you!

#LI-Hybrid

Similar Jobs

Next Level Business Services - Linux Scripting and Clear case SME Consultant

Next Level Business Services

Milwaukee, Wisconsin, United States (On-Site)
6 Months ago
Kaedim - DevOps Engineer

Kaedim

San Francisco, California, United States (On-Site)
7 Months ago
Rackspace Technology - Cloud Practice Engineer

Rackspace Technology

Bengaluru, Karnataka, India (Hybrid)
6 Months ago
Ubisoft - DevOps Linux Administrator

Ubisoft

Saint-Mandé, Île-de-France, France (On-Site)
2 Months ago
Larian Studios - DEVOPS BUILD ENGINEER

Larian Studios

Quebec, Canada (On-Site)
3 Months ago
NVIDIA - Senior Digital Design Verification Engineer - Hardware

NVIDIA

Hsinchu, Hsinchu City, Taiwan (On-Site)
3 Months ago
NVIDIA - Senior Technical Program Manager – CSP Datacenter Compute Server Software

NVIDIA

Santa Clara, California, United States (On-Site)
2 Months ago
Riot Games - Principal Software Engineer - Riot Client

Riot Games

Dublin, County Dublin, Ireland (On-Site)
5 Months ago
Krafton  - Applied Research Engineer - Reinforcement Learning (Intern)

Krafton

Seoul, South Korea (On-Site)
1 Month ago
NVIDIA - Senior System Software Engineer - SoC Power

NVIDIA

Santa Clara, California, United States (On-Site)
3 Months ago

Get notifed when new similar jobs are uploaded

Similar Skill Jobs

PwC - IN_Associate_Azure Cloud Data Engineer_OneCloud _Advisory _Bangalore

PwC

Gurugram, Haryana, India (On-Site)
4 Months ago
Rackspace Technology - Senior AWS Migration Engineer

Rackspace Technology

Gurugram, Haryana, India (Remote)
2 Months ago
Luxoft - Java/Scala Developer

Luxoft

(Remote)
4 Months ago
Telesign - Site Reliability Engineer (SRE) III

Telesign

Bengaluru, Karnataka, India (On-Site)
6 Months ago
NICE - Senior Cloud SRE

NICE

Pune, Maharashtra, India (Hybrid)
6 Months ago
Polygon Labs - Senior DevOps Engineer

Polygon Labs

United States (Remote)
1 Month ago
The Walt Disney Company - Senior Real Time Pipeline Engineer (PH)

The Walt Disney Company

Glendale, California, United States (On-Site)
5 Months ago
Intel Corporation - Sr. Infrastructure Engineer - Storage

Intel Corporation

Hillsboro, Oregon, United States (On-Site)
4 Months ago
Ness Digital - Sr AWS DevOps Engineer

Ness Digital

Iași, Iași County, Romania (Remote)
2 Months ago
Thatgamecompany - Backend Engineer - Shanghai

Thatgamecompany

Shanghai, Shanghai, China (On-Site)
10 Months ago

Get notifed when new similar jobs are uploaded

Jobs in Bengaluru, Karnataka, India

JBPCO India - Recruiter (Contract)

JBPCO India

Bengaluru, Karnataka, India (On-Site)
6 Months ago
Zazz - Marketing Data Specialist

Zazz

India (On-Site)
3 Months ago
Zones - MDM L2 Support

Zones

Bengaluru, Karnataka, India (On-Site)
5 Months ago
Barracuda Networks  Inc  - Senior Security Engineer

Barracuda Networks Inc

Bengaluru, Karnataka, India (On-Site)
6 Months ago
Phantom FX - Rigging Artist

Phantom FX

Mumbai, Maharashtra, India (On-Site)
2 Months ago
Warner Bros Discovery - Senior Manager, Data Platform & AWS Infrastructure - (Streaming), Hyderabad

Warner Bros Discovery

Hyderabad, Telangana, India (On-Site)
5 Months ago
Glean - Software Engineer, Frontend (India)

Glean

Bengaluru, Karnataka, India (On-Site)
6 Months ago
PwC - IN-Senior Associate_Telecom_ Cities_Advisory _ Mumbai

PwC

Mumbai, Maharashtra, India (On-Site)
4 Months ago
Nagarro - Staff Engineer, Cloud

Nagarro

India (Remote)
6 Months ago
CloudHire - Python Developer

CloudHire

India (Remote)
6 Months ago

Get notifed when new similar jobs are uploaded

Research & Development Jobs

Google - Senior Software Engineer, Machine Learning, YouTube

Google

San Bruno, California, United States (On-Site)
3 Months ago
NVIDIA - Principal System Architect - Tegra

NVIDIA

Bengaluru, Karnataka, India (Hybrid)
3 Months ago
Cadence - Lead Design Engineer ( Layout Design )

Cadence

Bengaluru, Karnataka, India (On-Site)
7 Months ago
Intel Corporation - CPU-SoC Silicon Design Engineering Part Time Intern

Intel Corporation

Kedah, Malaysia (On-Site)
4 Months ago
ByteDance - Research Engineer (Machine Learning Training System) - 2025 Start

ByteDance

Singapore (On-Site)
5 Months ago
Krafton  - 2025 Krafton New Recruitment - AI Game Tech (Client Programming)

Krafton

Seoul, South Korea (On-Site)
1 Month ago
Fabric - Applied Researcher, Cryptography Proof Systems

Fabric

France (Remote)
6 Months ago
NVIDIA - Hardware Application Engineer, Ethernet Switch

NVIDIA

Shanghai, Shanghai, China (Hybrid)
3 Months ago
Riot Games - Principal Researcher

Riot Games

Los Angeles, California, United States (On-Site)
5 Months ago
Krafton  - Game Analyst / Game Researcher

Krafton

Seoul, South Korea (On-Site)
2 Months ago

Get notifed when new similar jobs are uploaded

About The Company

Since its founding in 1993, NVIDIA (NASDAQ: NVDA) has been a pioneer in accelerated computing. The company’s invention of the GPU in 1999 sparked the growth of the PC gaming market, redefined computer graphics, ignited the era of modern AI and is fueling the creation of the metaverse. NVIDIA is now a full-stack computing company with data-center-scale offerings that are reshaping industry.

California, United States (Remote)

Yokne'am Illit, North District, Israel (On-Site)

Tel Aviv-Yafo, Tel Aviv District, Israel (On-Site)

Bengaluru, Karnataka, India (On-Site)

Santa Clara, California, United States (On-Site)

Pune, Maharashtra, India (On-Site)

Taipei City, Taiwan (On-Site)

Taipei City, Taiwan (On-Site)

Beijing, Beijing, China (On-Site)

Tel Aviv-Yafo, Tel Aviv District, Israel (On-Site)

View All Jobs

Get notified when new jobs are added by NVIDIA

Level Up Your Career in Game Development!

Transform Your Passion into Profession with Our Comprehensive Courses for Aspiring Game Developers.

Job Common Plug