Senior Solutions Architect, Infiniband and Networking Ethernet

1 Month ago • 8 Years + • Network Engineering

Job Summary

Job Description

NVIDIA seeks a Senior Networking (ETH/IB) Solutions Architect to design and implement large-scale networking projects for AI/HPC systems. Responsibilities include building infrastructure for new and existing customers, supporting operational reliability, improving service lifecycles, and monitoring system health. The role requires strong collaboration with customers and internal teams, utilizing expertise in Infiniband and Ethernet networks, automation tools, and troubleshooting skills. The ideal candidate possesses deep knowledge of networking protocols, experience with various network platforms, and a proven ability to deliver automated network provisioning solutions.
Must have:
  • 8+ years networking experience
  • InfiniBand & Ethernet expertise
  • EVPN, BGP, OSPF, VXLAN knowledge
  • Automation skills (Ansible, Salt, Python)
  • CI/CD pipeline development
  • Customer communication skills
Good to have:
  • Cloud network familiarity (AWS, GCP, Azure)
  • Linux/Networking certifications
  • HPC architecture understanding
  • Job scheduler (Slurm, PBS) knowledge
  • Lustre management experience
  • GPU hardware/software experience
  • Mandarin communication skills

Job Details

NVIDIA is looking for Senior Networking (ETH/IB) Solutions Architect to join its NVIDIA Infrastructure Specialist Team. Academic and commercial groups around the world are using NVIDIA products to revolutionize deep learning and data analytics, and to power data centers. Join the team building many of the largest and fastest AI/HPC systems in the world! We are looking for someone with the ability to work on a dynamic customer focused team that requires excellent interpersonal skills. This role will be interacting with customers, partners and internal teams, to analyze, define and implement large scale Networking projects. The scope of these efforts includes a combination of Networking, System Design and Automation and being the face to the customer!

What you'll be doing:

  • Primary responsibilities will include building AI/HPC infrastructure for new and existing customers.
  • Support operational and reliability aspects of large-scale AI clusters, focusing on performance at scale, real-time monitoring, logging, and alerting.
  • Engage in and improve the whole lifecycle of services—from inception and design through deployment, operation, and refinement.
  • Maintain services once they are live by measuring and monitoring availability, latency, and overall system health.
  • Provide feedback to internal teams such as opening bugs, documenting workarounds, and suggesting improvements.

What we need to see:

  • BS/MS/PhD or equivalent experience in Computer Science, Electrical/Computer Engineering, Physics, Mathematics, or related fields.
  • At least 8 years of professional experience in networking fundamentals, TCP/IP stack, and data center architecture
  • Proficiency in configuring, testing, validating, and resolving issues in LAN and InfiniBand networks, especially in medium to large-scale HPC/AI environments.
  • Advanced knowledge of EVPN, BGP, OSPF, VXLAN protocols.
  • Hands-on experience with network switch/router platforms like Cumulus Linux, SONiC, IOS, JunosOS, and EOS.
  • Extensive experience delivering automated network provisioning solutions using tools like Ansible, Salt, and Python.
  • Ability to develop CI/CD pipelines for network operations.
  • Strong focus on customer needs and satisfaction.
  • Self-motivated with leadership skills to work collaboratively with customers and internal teams.
  • Ability to communicate technical concepts and collaborate effectively with Mandarin-speaking customers.
  • Strong written, verbal, and listening skills in English are essential.

Ways to stand out from the crowd:

  • Familiarity with cloud networks (AWS, GCP, Azure) is a plus.
  • Linux or Networking Certifications.
  • Experience with High-performance computing architectures. Understanding of how job schedulers(Slurm, PBS) work.
  • luster management technologies knowledge (bonus credit for BCM (Base Command Manager).)
  • Experience with GPU (Graphics Processing Unit) focused hardware/software.

NVIDIA is widely considered to be one of the technology world’s most desirable employers. We have some of the most forward-thinking and hardworking individuals in the world working for us. If you're creative and autonomous, we want to hear from you.

Similar Jobs

Scanline VFX - Houdini FX Artist

Scanline VFX

Vancouver, British Columbia, Canada (Hybrid)
4 Months ago
SuperPlay - QA Manual Engineer

SuperPlay

Warsaw, Masovian Voivodeship, Poland (Remote)
3 Weeks ago
emerald city games - 3D ARTIST

emerald city games

Burnaby, British Columbia, Canada (On-Site)
1 Year ago
playrix  - Lead SDET

playrix

Montenegro (Remote)
7 Months ago
FISHLABS GmbH - (Regular/Senior) UI/UX Designer (m/f/d)

FISHLABS GmbH

Hamburg, Hamburg, Germany (On-Site)
10 Months ago
bytedance - Traffic Access Architectural Engineer - Traffic Infrastructure

bytedance

Singapore (On-Site)
7 Months ago
bytedance - Software Engineer - Data Tech Infrastructure- San Jose

bytedance

San Jose, California, United States (On-Site)
7 Months ago
bytedance - Software Developer (Routing Verification & Emulation)

bytedance

Seattle, Washington, United States (On-Site)
2 Months ago
bytedance - Site Reliability Engineer, Edge Services

bytedance

Boston, Massachusetts, United States (On-Site)
7 Months ago
Google - Network Implementation Engineer, Optical Operations

Google

Bengaluru, Karnataka, India (On-Site)
1 Month ago

Get notifed when new similar jobs are uploaded

Similar Skill Jobs

Oculus VR - Senior Technical Animator

Oculus VR

Canada (Remote)
2 Weeks ago
limit break - Unity Level Integrator

limit break

Tokyo, Japan (On-Site)
10 Months ago
Gameduell studios - Senior 2D Animator (Unity) - Character & Asset Specialist

Gameduell studios

Berlin, Berlin, Germany (Hybrid)
1 Week ago
Virtuos - Software Engineer

Virtuos

Czechia (Hybrid)
1 Month ago
Unity - Grow Programmatic Solutions Developer Support Engineer - (Temporary)

Unity

Tel Aviv-Yafo, Tel Aviv District, Israel (On-Site)
1 Month ago
SciPlay - Unity Developer

SciPlay

Tel Aviv-Yafo, Tel Aviv District, Israel (On-Site)
1 Month ago
luxsoft - Telecommunication Consultant

luxsoft

India (Remote)
3 Weeks ago
Game freak - Graphic Designer [Section Director]

Game freak

Chiyoda City, Tokyo, Japan (On-Site)
2 Weeks ago
Gametaq - Unity Team Lead

Gametaq

(Remote)
3 Weeks ago
Appirits - 2D Illustrator

Appirits

Shibuya, Tokyo, Japan (On-Site)
2 Months ago

Get notifed when new similar jobs are uploaded

Jobs in undefined

Looks like we're out of matches

Set up an alert and we'll send you similar jobs the moment they appear!

Network Engineering Jobs

PwC - Senior Network Engineer (m/f/d)

PwC

Luxembourg (On-Site)
7 Months ago
bytedance - Data Center Commercial Manager - Data Center Development

bytedance

Singapore (On-Site)
7 Months ago
Meta - Software Engineer - Datacenter networking

Meta

New York, New York, United States (On-Site)
6 Months ago
bytedance - Network Automation Engineer

bytedance

Ashburn, Virginia, United States (On-Site)
2 Months ago
Google - Technicus Datacenter

Google

(On-Site)
5 Months ago
Next Level Business Services - Network Architecture and Operations

Next Level Business Services

Mount Laurel Township, New Jersey, United States (On-Site)
7 Months ago
ness digital  - Network Engineer with German

ness digital

Timișoara, Timiș, Romania (On-Site)
4 Months ago
bytedance - Site Reliability Engineer - Game

bytedance

Singapore (On-Site)
7 Months ago
Meta - Technical Program Manager, Net Infra (Backbone)

Meta

Denver, Colorado, United States (On-Site)
6 Months ago

Get notifed when new similar jobs are uploaded

About The Company

Since its founding in 1993, NVIDIA (NASDAQ: NVDA) has been a pioneer in accelerated computing. The company’s invention of the GPU in 1999 sparked the growth of the PC gaming market, redefined computer graphics, ignited the era of modern AI and is fueling the creation of the metaverse. NVIDIA is now a full-stack computing company with data-center-scale offerings that are reshaping industry.

Santa Clara, California, United States (On-Site)

Massachusetts, United States (On-Site)

Santa Clara, California, United States (On-Site)

Santa Clara, California, United States (On-Site)

Santa Clara, California, United States (On-Site)

Texas, United States (On-Site)

Santa Clara, California, United States (Hybrid)

Santa Clara, California, United States (Hybrid)

Pune, Maharashtra, India (On-Site)

Taipei City, Taiwan (On-Site)

View All Jobs

Get notified when new jobs are added by NVIDIA

Level Up Your Career in Game Development!

Transform Your Passion into Profession with Our Comprehensive Courses for Aspiring Game Developers.

Job Common Plug