Senior Solutions Architect, HPC and AI

6 Hours ago • 5 Years + • Research & Development • $148,000 PA - $235,750 PA

Job Summary

Job Description

NVIDIA seeks a Senior Solutions Architect with 5+ years' experience in validating and debugging large-scale GPU clusters for performance. Responsibilities include architecting and scaling high-performance AI infrastructure (on-prem or cloud), resolving performance issues across all levels (bare metal to applications), collaborating with various teams (demos, POCs, documentation), working directly with developers and architects, supporting the account team in debugging customer issues, and building custom product demonstrations. The ideal candidate possesses strong analytical and communication skills, expertise in HPC, deep learning, and cloud technologies (Docker, Kubernetes, cloud APIs), and proficiency in relevant tools and technologies (e.g., CUDA, NCCL, SLURM).
Must have:
  • 5+ years experience in accelerated computing
  • Platform-level server architecture understanding
  • Networking expertise (Ethernet, InfiniBand)
  • DevOps experience (Docker, cloud APIs)
  • SLURM/Kubernetes experience
  • Deep understanding of data center design
  • Strong analytical & problem-solving skills
  • Excellent communication & collaboration skills
Good to have:
  • NCCL experience
  • Excellent customer-facing skills
  • Proficient debugging skills (C/C++, Linux kernel)
  • NVIDIA systems/SDKs experience (CUDA)
  • NVIDIA Networking technologies (RoCE, InfiniBand)
  • Deep Learning/ML framework experience (TensorFlow, PyTorch)
  • LLM, MLOps, DevOps experience
Perks:
  • Highly competitive salary
  • Comprehensive benefits package
  • Excellent engineering culture

Job Details

NVIDIA is looking for a Field Escalation Solution Architect with experience in validation and debugging of large-scale GPU clusters focused on performance. As part of the Solution Architecture organization, we work with the most sophisticated computing hardware and software, driving the latest deep learning and machine learning breakthroughs with NVIDIA’s enterprise customers. This role offers an excellent opportunity to build your career in the rapidly growing field of deep learning while enabling the world's most successful technology companies. Primary responsibilities will be to validate and debug customer cluster performance issues, functional bottlenecks and drive customer technical engagements around NVIDIA products and technologies. Join us in this exciting endeavor!

What you’ll be doing:

  • A considerable part of the day-to-day job is staying up to date on pioneering High Performance Computing, Deep Learning and Machine Learning ecosystems. You'll be called on to help architect and scale high-performance, distributed AI infrastructure on-prem or in the cloud built with the latest NVIDIA GPU supercomputers for new and existing customers.

  • Address and resolve problems starting from the bare metal level, all the way up to the operating system, software stack, and application level.

  • Share knowledge with different teams by delivering demos, assisting with proof-of-concepts, and writing papers and developer blogs. By collaborating with executives and engineering, address sophisticated problems and help bring NVIDIA's premiere technologies to life in the cloud and in the datacenter.

  • Work directly with developers and hardware architects to debug cluster performance issues, identify new requirements, and improve workflows.

  • Will be engaged by the account team when extra analysis is required in debugging customer issues.

  • Provide additional expertise to enable the account team to be more adaptable to the customer and product engineering to get more actionable data at speed of light making them more efficient.

  • Building custom product demonstrations and POCs for solutions that address critical business needs of our customers.

What we need to see:

  • BS, MS, or PhD in Computer Science, Electrical/Computer Engineering, Physics, Mathematics, or other Engineering fields or equivalent experience.

  • 5+ years of work-related experience in NVIDIA and/or accelerated computing technologies.

  • Platform level understanding of server architecture, PCIe topology, GPUs, NICs, Linux OS and kernel drivers.

  • Networking experience, including knowledge of Ethernet, InfiniBand or other networking protocols.

  • Experience working with DevOps on-prem or in cloud environments, including but not limited to Docker/Containers, cloud APIs, IaaS and Data Center deployments.

  • SLURM, Kubernetes, and/or other job scheduler use, deployment, and debugging skills.

  • Deep understanding of dense data center design, including computing, storage, networking, cloud APIs, and IaaS.

  • Effective time management and capable of balancing multiple tasks.

  • Strong analytical and problem-solving skills.

  • Strong communication skills, both written and verbal, with the ability to collaborate and coordinate efficiently across multi-functional teams in engineering, sales, marketing, product, and program management.

Ways to stand out from the crowd:

  • Demonstrated Communication Collectives (NCCL) experience.

  • Excellent customer-facing skills and background.

  • Platform design engineering, coding and proficient debugging skills including experience in C/C++, Linux kernel, virtualization and drivers, profilers/performance analysis tools (NSys).

  • Familiarity with NVIDIA systems/SDKs (e.g. CUDA), NVIDIA Networking technologies (e.g., RoCE, InfiniBand), Switch interconnects and/or ARM CPU solutions through hands-on experience.

  • Understanding of Deep Learning and Machine Learning frameworks (TensorFlow or PyTorch), LLM, MLOps, DevOps, and workflows applying cloud technologies, using Docker/containers, Kubernetes, cloud APIs, and data center deployments, among others.

We make extensive use of conferencing tools, but occasional travel is required for a local on-site visit to customers and data science conferences.

With highly competitive salaries, a comprehensive benefits package, and an excellent engineering culture, NVIDIA is widely considered to be one of the technology industry's most desirable employers. NVIDIA has some of the most innovative people working on significant problems that define the field of ML/DL, data science, and graphics.

#LI-Hybrid

The base salary range is 148,000 USD - 235,750 USD. Your base salary will be determined based on your location, experience, and the pay of employees in similar positions.

You will also be eligible for equity and benefits. NVIDIA accepts applications on an ongoing basis.

NVIDIA is committed to fostering a diverse work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.

Similar Jobs

Beyond Sports  - 3D Artist - Generalist - Blender/Unity

Beyond Sports

Alkmaar, North Holland, Netherlands (On-Site)
1 Week ago
ByteDance - Optical Scientist - Display Optics System

ByteDance

San Jose, California, United States (On-Site)
1 Month ago
Ubisoft - Machine Learning Programmer (Character & Animation)

Ubisoft

Montreal, Quebec, Canada (On-Site)
1 Week ago
Meta - Product Design Engineer, Reality Labs

Meta

Seattle, Washington, United States (On-Site)
4 Months ago
Lakshya Digital - VFX Artist

Lakshya Digital

Haryana, India (On-Site)
2 Weeks ago
ByteDance - XR Embedded Engineer / Architect- Pico Lab - San Jose

ByteDance

San Jose, California, United States (On-Site)
5 Months ago
Samsung Semiconductor - Staff Engineer, DRAM Design

Samsung Semiconductor

San Jose, California, United States (On-Site)
6 Days ago
NVIDIA - Deep Learning Performance Architect

NVIDIA

Shanghai, Shanghai, China (On-Site)
2 Months ago
Meta - Software Engineer (Technical Leadership) - Machine Learning

Meta

Seattle, Washington, United States (On-Site)
4 Months ago
Rivos - Accelerator Verification Intern

Rivos

Santa Clara, California, United States (Hybrid)
5 Months ago

Get notifed when new similar jobs are uploaded

Similar Skill Jobs

Playrix - Lead C++ Software Engineer (Gameplay)

Playrix

Montenegro (Remote)
5 Months ago
NVIDIA - AI Digital Human Development Intern - 2025

NVIDIA

(On-Site)
1 Month ago
DraftKings - Lead Software Engineer, Unity

DraftKings

Tel Aviv-Yafo, Tel Aviv District, Israel (On-Site)
1 Month ago
Playrix - Lead Location Designer

Playrix

Portugal (Remote)
5 Months ago
Playrix - Senior C++ Software Engineer (Tools)

Playrix

Ukraine (Remote)
5 Months ago
Discord - Senior Developer Advocate, Web Games

Discord

San Francisco, California, United States (Remote)
5 Months ago
ION - Junior Product Specialist – IT Sector

ION

Milan, Lombardy, Italy (Hybrid)
5 Months ago
Limit Break - Sr. Mobile Game Designer

Limit Break

Singapore, Singapore (On-Site)
8 Months ago
Fantastic Pixel Castle - Senior Producer

Fantastic Pixel Castle

United States (Remote)
1 Month ago
Ubisoft - Lead UI/UX Designer AAA [The Division Resurgence]

Ubisoft

Saint-Mandé, Île-de-France, France (Hybrid)
1 Month ago

Get notifed when new similar jobs are uploaded

Jobs in Santa Clara, California, United States

The Pokemon Company International - Premier Event Judge Operations and Side Events Manager

The Pokemon Company International

Bellevue, Washington, United States (Hybrid)
1 Month ago
ByteDance - Research Scientist Graduate (Foundation Model, Vision and Language) - 2025 Start (PhD)

ByteDance

San Jose, California, United States (On-Site)
5 Months ago
Inworld AI - Financial Controller

Inworld AI

Mountain View, California, United States (Remote)
5 Days ago
Super - Senior Software Engineer, Payments

Super

United States (Remote)
4 Months ago
The Walt Disney Company - Executive Producer, Digital

The Walt Disney Company

Durham, North Carolina, United States (On-Site)
1 Week ago
Trek - Store Manager

Trek

Plainview, New York, United States (On-Site)
1 Month ago
Patel greene - PD&E Project Manager

Patel greene

Tallahassee, Florida, United States (On-Site)
5 Months ago
Evolution - Online Casino Dealer - Live In-Studio - Philadelphia

Evolution

Philadelphia, Pennsylvania, United States (On-Site)
10 Months ago
ByteDance - Machine Learning Engineer, Tech Lead - Code AI

ByteDance

San Jose, California, United States (On-Site)
1 Week ago

Get notifed when new similar jobs are uploaded

Research & Development Jobs

Krafton  - Gamelab Coach - Studio Supporter Conversion Position (10+ years)

Krafton

Seoul, South Korea (On-Site)
5 Days ago
NVIDIA - Signal and Power Integrity Engineer (RDSS Intern)

NVIDIA

Hsinchu, Hsinchu City, Taiwan (On-Site)
1 Month ago
NXP - Senior Principal Software Architect - Platform and RF Software

NXP

Bucharest, Bucharest, Romania (On-Site)
6 Months ago
NVIDIA - Senior Datacenter GPU Power Architect

NVIDIA

Austin, Texas, United States (On-Site)
1 Month ago
ByteDance - Research Scientist Graduate (High-Performance Computing (Algorithm Acceleration)- Vision AI Platform)

ByteDance

San Jose, California, United States (On-Site)
1 Week ago
Krafton  - IT Strategy Manager

Krafton

Seoul, South Korea (On-Site)
4 Weeks ago
Tesla - Senior Embedded Software/Firmware Engineer - Power Electronics

Tesla

Baden-Württemberg, Germany (On-Site)
1 Month ago
NVIDIA - Senior Developer Technology Engineer, Public Sector

NVIDIA

Santa Clara, California, United States (Remote)
3 Weeks ago
NVIDIA - Clock Design Engineer

NVIDIA

Tel Aviv-Yafo, Tel Aviv District, Israel (Hybrid)
1 Month ago
NVIDIA - Senior Board Design Hardware Engineer

NVIDIA

Yokne'am Illit, North District, Israel (On-Site)
1 Month ago

Get notifed when new similar jobs are uploaded

About The Company

Since its founding in 1993, NVIDIA (NASDAQ: NVDA) has been a pioneer in accelerated computing. The company’s invention of the GPU in 1999 sparked the growth of the PC gaming market, redefined computer graphics, ignited the era of modern AI and is fueling the creation of the metaverse. NVIDIA is now a full-stack computing company with data-center-scale offerings that are reshaping industry.


Santa Clara, California, United States (On-Site)

Santa Clara, California, United States (On-Site)

Santa Clara, California, United States (Hybrid)

California, United States (Hybrid)

Santa Clara, California, United States (On-Site)

Santa Clara, California, United States (On-Site)

Ra'anana, Center District, Israel (On-Site)

Ra'anana, Center District, Israel (On-Site)

Yokne'am Illit, North District, Israel (On-Site)

Tel Aviv-Yafo, Tel Aviv District, Israel (On-Site)

View All Jobs

Get notified when new jobs are added by NVIDIA

Level Up Your Career in Game Development!

Transform Your Passion into Profession with Our Comprehensive Courses for Aspiring Game Developers.

Job Common Plug