Data Center System Software Architect, DGX Cloud

2 Months ago • 10 Years + • DevOps • Research & Development • $184,000 PA - $425,500 PA

Job Summary

Job Description

NVIDIA seeks a Data Center System Software Architect for its DGX Cloud team. Responsibilities include leading the architecture, design, and implementation of next-generation DGX cloud clusters using cutting-edge technologies. This full-stack role encompasses hardware architecture, workload orchestration, and application performance tuning. The ideal candidate possesses 10+ years of experience in system software, strong programming skills (C, C++, Go, Rust), expertise in distributed systems, and excellent communication skills. The role involves collaborating with various engineering teams across NVIDIA to ensure seamless software integration, from hardware to AI training applications. The architect will provide solutions for complex problems and translate requirements into a vision, architecture, and roadmap.
Must have:
  • 10+ years system software experience
  • Strong programming skills (C, C++, Go, Rust)
  • Distributed systems expertise
  • Excellent communication skills
  • Data science/deep learning knowledge
Good to have:
  • TensorFlow/PyTorch experience
  • Docker, Kubernetes, Slurm experience
  • CUDA/NCCL programming
  • HPC programming (MPI, OpenACC)
  • DGX Cloud, NVIDIA AI Enterprise experience
Perks:
  • Equity
  • Benefits

Job Details

NVIDIA is hiring engineers to scale up its AI Infrastructure. We expect you to have a strong programming background, a deep understanding of distributed systems, familiarity with software testing and deployment, and excellent communication and planning abilities. We also welcome out-of-the-box thinkers who can provide new ideas with strong at execution bias. Expect to be constantly challenged, improving, and evolving for the better. You and other engineers in this team will help advance NVIDIA's capacity to build and deploy leading infrastructure solutions for a broad range of AI-based applications that affect core data science. What are you waiting for if you're creative, passionate about what you do, and love having fun apply today!

We’re looking for a highly motivated, creative engineer with strong experience in system software to join the DGX Cloud Software Team. You will lead the architecture, design and implementation of our next generation DGX cloud clusters using latest technologies. On this team, you will do full stack deployment including hardware architecture, workload orchestration and application performance tuning. Are you ready to change the next generation of computing? Join us at the forefront of technological advancement.

What you’ll be doing:

  • Lead technical activities for data centers with focus on hybrid deployments between cloud and on-prem

  • Providing expertise in infrastructure workflows, including hardware, workload orchestration and application tuning

  • Provide fast and creative solutions for complex problems and write effective, clear and reliable architecture specification

  • Translate requirements to vision, architecture and roadmap

  • Work with engineering teams across NVIDIA to ensure your software integrates seamlessly from the hardware all the way up to the AI training applications.

What we need to see:

  • Masters or PhD in Computer Science, Computer Engineering, Physics or equivalent experience

  • 10+ years of experience in this field.

  • Data Sciences, Deep Learning, or Machine Learning coursework

  • Ability to seamlessly shift between Linux system environments to Python programming

  • Programming skills in 1 or more high-level languages (C, C++,Go,Rust etc)

  • System-level experience with both hardware and software

  • Motivated self-starter with an equal balance of strong problem-solving skills and customer-facing communication skills

  • Strong design, coding, analytical, debugging and problem-solving skills

  • Passion for continuous learning and knowledge transfer. Ability to work concurrently with multiple groups locally and abroad in the organization

Ways to stand out from the crowd:

  • Experience with GPU deep learning and data sciences. Experience using TensorFlow, PyTorch or other DL framework. Experience working with Docker containers, Slurm, Terraform and Kubernetes

  • CUDA programming and NCCL experience. HPC programming experience including MPI, OpenACC, or other parallel programming tools

  • Hands-on experience with DGX Cloud, NVIDIA AI Enterprise AI Software, Base Command Manager, NEMO and NVIDIA Inference Microservices.

NVIDIA is widely considered to be one of the technology world’s most desirable employers. We have some of the most forward-thinking and hardworking people on the planet working for us. If you are creative and autonomous, we want to hear from you!

The base salary range is 184,000 USD - 425,500 USD. Your base salary will be determined based on your location, experience, and the pay of employees in similar positions.

You will also be eligible for equity and benefits. NVIDIA accepts applications on an ongoing basis.

NVIDIA is committed to fostering a diverse work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.

Similar Jobs

Velotio Technologies - Data Scientist

Velotio Technologies

Maharashtra, India (Remote)
1 Month ago
NVIDIA - Senior ASIC Design Engineer

NVIDIA

Massachusetts, United States (Hybrid)
2 Months ago
NVIDIA - Research Scientist, Efficient Deep Learning - New College Grad 2025

NVIDIA

Santa Clara, California, United States (On-Site)
2 Months ago
Krafton  - Technical Project Manager, Deep Learning Division

Krafton

Seoul, South Korea (On-Site)
2 Months ago
Microsoft - Technical Support Engineer - Identity & Security (Entra)

Microsoft

Seoul, South Korea (Hybrid)
3 Months ago
Wargaming - Senior Infrastructure Engineer (Internal Development)

Wargaming

Nicosia, Nicosia, Cyprus (Hybrid)
1 Month ago
Microsoft - Principal Software Engineering Manager

Microsoft

Bucharest, Bucharest, Romania (Remote)
2 Months ago
Nagarro - Associate Principal Engineer, DevOps

Nagarro

India (Remote)
5 Months ago

Get notifed when new similar jobs are uploaded

Similar Skill Jobs

Cricketpedia - AI Engineer

Cricketpedia

Gurugram, Haryana, India (Remote)
2 Years ago
NVIDIA - Senior Site Reliability Engineer - AI Research Clusters

NVIDIA

Hyderabad, Telangana, India (Hybrid)
2 Months ago
Meta - Software Engineer, Systems ML - SW/HW Co-design

Meta

Bellevue, Washington, United States (Remote)
4 Months ago
Tencent - Senior Staff Researcher

Tencent

Palo Alto, California, United States (On-Site)
4 Months ago
NVIDIA - Senior Software Product Manager, Nemo LLM Microservices

NVIDIA

California, United States (Hybrid)
2 Months ago
NVIDIA - AI Algorithms SW Engineer (RDSS Intern)

NVIDIA

Hsinchu, Hsinchu City, Taiwan (On-Site)
2 Months ago
ByteDance - Engineering Manager - Applied Machine Learning Algorithm

ByteDance

San Jose, California, United States (On-Site)
4 Months ago
Mistplay - Senior Data Scientist II

Mistplay

Montreal, Quebec, Canada (Hybrid)
1 Month ago
Meta - Software Engineer, Computer Vision (Technical Leadership)

Meta

Bellevue, Washington, United States (Remote)
4 Months ago
Coursera - Machine Learning Scientist

Coursera

India (Remote)
3 Months ago

Get notifed when new similar jobs are uploaded

Jobs in Santa Clara, California, United States

Rockstar Games - Associate Principal Product Manager, Commerce

Rockstar Games

New York, New York, United States (On-Site)
2 Months ago
Salesforce - Key Platform and Agents Account Executive - Walmart

Salesforce

Arkansas, United States (Remote)
2 Months ago
Scientific Games  - Helpdesk Technician I

Scientific Games

Alpharetta, Georgia, United States (On-Site)
2 Months ago
Lionsgate Games - Manager, International Sales Strategy & Planning, LATAM/ASIA

Lionsgate Games

Santa Monica, California, United States (On-Site)
3 Months ago
Sphere Entertainment Co - Employee Service Center Representative (Part-Time)

Sphere Entertainment Co

Las Vegas, Nevada, United States (On-Site)
3 Months ago
Riot Games - Senior Manager, Technical Product Management - League of Legends

Riot Games

Los Angeles, California, United States (On-Site)
3 Months ago
Crunchyroll - DevOps Engineer, Core Infrastructure Engineering

Crunchyroll

San Francisco, California, United States (Hybrid)
4 Weeks ago
The Walt Disney Company - Senior Member Experience Professional - Branch

The Walt Disney Company

Lake Buena Vista, Florida, United States (On-Site)
1 Month ago
ByteDance - Research Scientist Graduate (Foundation Models for Science - ByteDance Research) - 2025 Start (PhD)

ByteDance

San Jose, California, United States (On-Site)
4 Months ago
Next Level Business Services - Android Developer

Next Level Business Services

Redwood City, California, United States (On-Site)
5 Months ago

Get notifed when new similar jobs are uploaded

DevOps Jobs

NVIDIA - Web Application Developer

NVIDIA

Pune, Maharashtra, India (On-Site)
2 Months ago
Fortis Games - Senior Cloud Security Engineer

Fortis Games

Portugal (On-Site)
1 Month ago
Luxoft - Senior Software Support Engineer

Luxoft

Slovakia (Remote)
4 Months ago
Milestone - Lead Data Engineer

Milestone

United States (Remote)
1 Month ago
ION - Cloud Engineer/Architect (DevOps)

ION

Italy (On-Site)
5 Months ago
DEVOTEAM - Distributed Cloud | DevOps Azure Engineer

DEVOTEAM

Lisbon, Lisbon, Portugal (Remote)
5 Months ago
Tencent - Cloud Engineer

Tencent

(On-Site)
4 Months ago
NVIDIA - Senior System Software Engineer, Distributed Systems - DGX Cloud

NVIDIA

Santa Clara, California, United States (Remote)
2 Months ago
SmileGate - SRE Strategy Manager

SmileGate

Seongnam-si, Gyeonggi-do, South Korea (On-Site)
2 Months ago

Get notifed when new similar jobs are uploaded

About The Company

Since its founding in 1993, NVIDIA (NASDAQ: NVDA) has been a pioneer in accelerated computing. The company’s invention of the GPU in 1999 sparked the growth of the PC gaming market, redefined computer graphics, ignited the era of modern AI and is fueling the creation of the metaverse. NVIDIA is now a full-stack computing company with data-center-scale offerings that are reshaping industry.


Yokne'am Illit, North District, Israel (On-Site)

Hyderabad, Telangana, India (On-Site)

Tel Aviv-Yafo, Tel Aviv District, Israel (On-Site)

Santa Clara, California, United States (On-Site)

Texas, United States (On-Site)

Santa Clara, California, United States (On-Site)

Santa Clara, California, United States (On-Site)

Santa Clara, California, United States (On-Site)

Santa Clara, California, United States (On-Site)

Santa Clara, California, United States (On-Site)

View All Jobs

Get notified when new jobs are added by NVIDIA

Level Up Your Career in Game Development!

Transform Your Passion into Profession with Our Comprehensive Courses for Aspiring Game Developers.

Job Common Plug