Senior Solutions Architect, Infiniband and Networking Ethernet

2 Weeks ago • 8 Years + • Network Engineering

Job Summary

Job Description

NVIDIA seeks a Senior Networking (ETH/IB) Solutions Architect to design and implement large-scale networking projects for AI/HPC systems. Responsibilities include building infrastructure for new and existing customers, supporting operational reliability, improving service lifecycles, and monitoring system health. The role requires strong collaboration with customers and internal teams, utilizing expertise in Infiniband and Ethernet networks, automation tools, and troubleshooting skills. The ideal candidate possesses deep knowledge of networking protocols, experience with various network platforms, and a proven ability to deliver automated network provisioning solutions.
Must have:
  • 8+ years networking experience
  • InfiniBand & Ethernet expertise
  • EVPN, BGP, OSPF, VXLAN knowledge
  • Automation skills (Ansible, Salt, Python)
  • CI/CD pipeline development
  • Customer communication skills
Good to have:
  • Cloud network familiarity (AWS, GCP, Azure)
  • Linux/Networking certifications
  • HPC architecture understanding
  • Job scheduler (Slurm, PBS) knowledge
  • Lustre management experience
  • GPU hardware/software experience
  • Mandarin communication skills

Job Details

NVIDIA is looking for Senior Networking (ETH/IB) Solutions Architect to join its NVIDIA Infrastructure Specialist Team. Academic and commercial groups around the world are using NVIDIA products to revolutionize deep learning and data analytics, and to power data centers. Join the team building many of the largest and fastest AI/HPC systems in the world! We are looking for someone with the ability to work on a dynamic customer focused team that requires excellent interpersonal skills. This role will be interacting with customers, partners and internal teams, to analyze, define and implement large scale Networking projects. The scope of these efforts includes a combination of Networking, System Design and Automation and being the face to the customer!

What you'll be doing:

  • Primary responsibilities will include building AI/HPC infrastructure for new and existing customers.
  • Support operational and reliability aspects of large-scale AI clusters, focusing on performance at scale, real-time monitoring, logging, and alerting.
  • Engage in and improve the whole lifecycle of services—from inception and design through deployment, operation, and refinement.
  • Maintain services once they are live by measuring and monitoring availability, latency, and overall system health.
  • Provide feedback to internal teams such as opening bugs, documenting workarounds, and suggesting improvements.

What we need to see:

  • BS/MS/PhD or equivalent experience in Computer Science, Electrical/Computer Engineering, Physics, Mathematics, or related fields.
  • At least 8 years of professional experience in networking fundamentals, TCP/IP stack, and data center architecture
  • Proficiency in configuring, testing, validating, and resolving issues in LAN and InfiniBand networks, especially in medium to large-scale HPC/AI environments.
  • Advanced knowledge of EVPN, BGP, OSPF, VXLAN protocols.
  • Hands-on experience with network switch/router platforms like Cumulus Linux, SONiC, IOS, JunosOS, and EOS.
  • Extensive experience delivering automated network provisioning solutions using tools like Ansible, Salt, and Python.
  • Ability to develop CI/CD pipelines for network operations.
  • Strong focus on customer needs and satisfaction.
  • Self-motivated with leadership skills to work collaboratively with customers and internal teams.
  • Ability to communicate technical concepts and collaborate effectively with Mandarin-speaking customers.
  • Strong written, verbal, and listening skills in English are essential.

Ways to stand out from the crowd:

  • Familiarity with cloud networks (AWS, GCP, Azure) is a plus.
  • Linux or Networking Certifications.
  • Experience with High-performance computing architectures. Understanding of how job schedulers(Slurm, PBS) work.
  • luster management technologies knowledge (bonus credit for BCM (Base Command Manager).)
  • Experience with GPU (Graphics Processing Unit) focused hardware/software.

NVIDIA is widely considered to be one of the technology world’s most desirable employers. We have some of the most forward-thinking and hardworking individuals in the world working for us. If you're creative and autonomous, we want to hear from you.

Similar Jobs

Breach - XR Quality Assurance (QA) Lead

Breach

Trondheim, Trøndelag, Norway (On-Site)
6 Months ago
Tripledot Studios - Product Artist

Tripledot Studios

Warsaw, Masovian Voivodeship, Poland (On-Site)
3 Weeks ago
Nintendo - Lighting Artist [Remote Contract] (Retro Studios)

Nintendo

United States (Remote)
8 Months ago
Dinosour Polo club  - Senior Programmer

Dinosour Polo club

Wellington, Wellington, New Zealand (On-Site)
6 Hours ago
Life church - Senior UX Researcher

Life church

Edmond, Oklahoma, United States (On-Site)
6 Months ago
Rackspace Technology - Network Operations Specialist/Engineer

Rackspace Technology

Bengaluru, Karnataka, India (Remote)
1 Week ago
Google - Staff Software Engineer, Network Management

Google

Sunnyvale, California, United States (On-Site)
2 Days ago
NVIDIA - Senior Networking Architect

NVIDIA

Beijing, Beijing, China (On-Site)
2 Weeks ago

Get notifed when new similar jobs are uploaded

Similar Skill Jobs

Wicresoft - unity开发【玩法】

Wicresoft

Shenzhen, Guangdong Province, China (On-Site)
9 Months ago
NVIDIA - Senior Solutions Architect, Networking

NVIDIA

California, United States (On-Site)
1 Month ago
Plarium - Level Designer

Plarium

Kyiv, Kyiv City, Ukraine (Remote)
2 Weeks ago
Skillz - Manager - FP&A

Skillz

Bengaluru, Karnataka, India (On-Site)
1 Week ago
Playtika - Unity Developer

Playtika

Warsaw, Masovian Voivodeship, Poland (Hybrid)
9 Months ago
Google - Data Scientist, Research, App Safety Engineering

Google

Bengaluru, Karnataka, India (On-Site)
2 Days ago
Playrix - Senior Unity Software Engineer (Gameplay)

Playrix

Georgia (Remote)
6 Months ago
Google - Data Scientist, Chrome

Google

Warsaw, Masovian Voivodeship, Poland (On-Site)
2 Weeks ago
ByteDance - Research Scientist in Machine Learning for Science (AML - AI-for-Science) - 2024 Start (PhD)

ByteDance

Seattle, Washington, United States (On-Site)
6 Months ago
Life church - Associate LifeKids Pastor

Life church

United States (On-Site)
6 Months ago

Get notifed when new similar jobs are uploaded

Jobs in undefined

Looks like we're out of matches

Set up an alert and we'll send you similar jobs the moment they appear!

Network Engineering Jobs

ByteDance - Experienced Technical Lead - Edge Cloud Infrastructure - San Jose / Seattle / Boston

ByteDance

Seattle, Washington, United States (On-Site)
6 Months ago
Activision - Senior Network Programmer

Activision

Warsaw, Masovian Voivodeship, Poland (On-Site)
5 Months ago
ByteDance - Site Reliability Engineer, Edge Services

ByteDance

Seattle, Washington, United States (On-Site)
2 Months ago
ByteDance - Senior Software Engineer - Traffic Infrastructure

ByteDance

Singapore (On-Site)
6 Months ago
ByteDance - Software Development Engineer Graduate (Intent-based networking) - 2025 Start (PhD)

ByteDance

Seattle, Washington, United States (On-Site)
6 Months ago
Google - Network Architect, Software

Google

Ann Arbor, Michigan, United States (On-Site)
2 Weeks ago
Microsoft - ROP - Cloud Network Engineer

Microsoft

Hyderabad, Telangana, India (On-Site)
1 Week ago
NVIDIA - Senior Software Engineer

NVIDIA

Tel Aviv-Yafo, Tel Aviv District, Israel (On-Site)
2 Months ago
Rockstar Games - Senior Network Programmer

Rockstar Games

Edinburgh, Scotland, United Kingdom (On-Site)
2 Months ago
Google - Staff Network Design Engineer, Google Cloud

Google

Tel Aviv-Yafo, Tel Aviv District, Israel (On-Site)
2 Days ago

Get notifed when new similar jobs are uploaded

About The Company

Since its founding in 1993, NVIDIA (NASDAQ: NVDA) has been a pioneer in accelerated computing. The company’s invention of the GPU in 1999 sparked the growth of the PC gaming market, redefined computer graphics, ignited the era of modern AI and is fueling the creation of the metaverse. NVIDIA is now a full-stack computing company with data-center-scale offerings that are reshaping industry.

Santa Clara, California, United States (On-Site)

Santa Clara, California, United States (On-Site)

Massachusetts, United States (On-Site)

Santa Clara, California, United States (On-Site)

Santa Clara, California, United States (On-Site)

Santa Clara, California, United States (On-Site)

Texas, United States (On-Site)

Santa Clara, California, United States (On-Site)

Santa Clara, California, United States (Hybrid)

Santa Clara, California, United States (Hybrid)

View All Jobs

Get notified when new jobs are added by NVIDIA

Level Up Your Career in Game Development!

Transform Your Passion into Profession with Our Comprehensive Courses for Aspiring Game Developers.

Job Common Plug