Senior Solutions Architect, Infiniband and Networking Ethernet

2 Months ago • 8 Years + • Network Engineering

Job Summary

Job Description

NVIDIA seeks a Senior Networking (ETH/IB) Solutions Architect to design and implement large-scale networking projects for AI/HPC systems. Responsibilities include building infrastructure for new and existing customers, supporting operational reliability, improving service lifecycles, and monitoring system health. The role requires strong collaboration with customers and internal teams, utilizing expertise in Infiniband and Ethernet networks, automation tools, and troubleshooting skills. The ideal candidate possesses deep knowledge of networking protocols, experience with various network platforms, and a proven ability to deliver automated network provisioning solutions.
Must have:
  • 8+ years networking experience
  • InfiniBand & Ethernet expertise
  • EVPN, BGP, OSPF, VXLAN knowledge
  • Automation skills (Ansible, Salt, Python)
  • CI/CD pipeline development
  • Customer communication skills
Good to have:
  • Cloud network familiarity (AWS, GCP, Azure)
  • Linux/Networking certifications
  • HPC architecture understanding
  • Job scheduler (Slurm, PBS) knowledge
  • Lustre management experience
  • GPU hardware/software experience
  • Mandarin communication skills

Job Details

NVIDIA is looking for Senior Networking (ETH/IB) Solutions Architect to join its NVIDIA Infrastructure Specialist Team. Academic and commercial groups around the world are using NVIDIA products to revolutionize deep learning and data analytics, and to power data centers. Join the team building many of the largest and fastest AI/HPC systems in the world! We are looking for someone with the ability to work on a dynamic customer focused team that requires excellent interpersonal skills. This role will be interacting with customers, partners and internal teams, to analyze, define and implement large scale Networking projects. The scope of these efforts includes a combination of Networking, System Design and Automation and being the face to the customer!

What you'll be doing:

  • Primary responsibilities will include building AI/HPC infrastructure for new and existing customers.
  • Support operational and reliability aspects of large-scale AI clusters, focusing on performance at scale, real-time monitoring, logging, and alerting.
  • Engage in and improve the whole lifecycle of services—from inception and design through deployment, operation, and refinement.
  • Maintain services once they are live by measuring and monitoring availability, latency, and overall system health.
  • Provide feedback to internal teams such as opening bugs, documenting workarounds, and suggesting improvements.

What we need to see:

  • BS/MS/PhD or equivalent experience in Computer Science, Electrical/Computer Engineering, Physics, Mathematics, or related fields.
  • At least 8 years of professional experience in networking fundamentals, TCP/IP stack, and data center architecture
  • Proficiency in configuring, testing, validating, and resolving issues in LAN and InfiniBand networks, especially in medium to large-scale HPC/AI environments.
  • Advanced knowledge of EVPN, BGP, OSPF, VXLAN protocols.
  • Hands-on experience with network switch/router platforms like Cumulus Linux, SONiC, IOS, JunosOS, and EOS.
  • Extensive experience delivering automated network provisioning solutions using tools like Ansible, Salt, and Python.
  • Ability to develop CI/CD pipelines for network operations.
  • Strong focus on customer needs and satisfaction.
  • Self-motivated with leadership skills to work collaboratively with customers and internal teams.
  • Ability to communicate technical concepts and collaborate effectively with Mandarin-speaking customers.
  • Strong written, verbal, and listening skills in English are essential.

Ways to stand out from the crowd:

  • Familiarity with cloud networks (AWS, GCP, Azure) is a plus.
  • Linux or Networking Certifications.
  • Experience with High-performance computing architectures. Understanding of how job schedulers(Slurm, PBS) work.
  • luster management technologies knowledge (bonus credit for BCM (Base Command Manager).)
  • Experience with GPU (Graphics Processing Unit) focused hardware/software.

NVIDIA is widely considered to be one of the technology world’s most desirable employers. We have some of the most forward-thinking and hardworking individuals in the world working for us. If you're creative and autonomous, we want to hear from you.

Similar Jobs

Tesla - Instructional Designer

Tesla

North Holland, Netherlands (On-Site)
4 Months ago
Accurate - Healthcare Vertical Strategist

Accurate

United States (Remote)
3 Months ago
Daxko - Vice President of Sales

Daxko

Birmingham, Alabama, United States (Remote)
3 Weeks ago
version 1 - Finance Business Partner

version 1

Dublin, County Dublin, Ireland (On-Site)
1 Month ago
CME Group - Manager - JiraAlign Business Office

CME Group

Belfast, Northern Ireland, United Kingdom (On-Site)
1 Month ago
Microsoft - Senior Software Engineer - Networking

Microsoft

Pune, Maharashtra, India (Hybrid)
2 Months ago
bytedance - Network Automation Engineer

bytedance

San Jose, California, United States (On-Site)
3 Months ago
Zscaler - Principal Network Engineer

Zscaler

(Remote)
1 Month ago
Yodlee - Lead - Network Engineering

Yodlee

Thiruvananthapuram, Kerala, India (On-Site)
1 Month ago
Excel Hr solutions - Java developer with Server Experience

Excel Hr solutions

Navi Mumbai, Maharashtra, India (Remote)
1 Year ago

Get notifed when new similar jobs are uploaded

Similar Skill Jobs

Toast - Retail Account Executive

Toast

Midland, Texas, United States (Hybrid)
2 Weeks ago
London stock Exchange - DataScope Senior Software Engineer (Backend C++)

London stock Exchange

Bucharest, Bucharest, Romania (On-Site)
3 Weeks ago
PrizePicks - Director of Engineering

PrizePicks

Atlanta, Georgia, United States (Remote)
1 Month ago
LeoVegas - Product Owner - Trading

LeoVegas

Málaga, Andalusia, Spain (Hybrid)
3 Months ago
Monzo - Senior Data Scientist

Monzo

London, England, United Kingdom (Remote)
1 Month ago
Match Group - Senior Machine Learning Engineer, Trust & Safety

Match Group

New York, New York, United States (Hybrid)
2 Months ago
Realworld one - Director, Client Strategy & Value Creation

Realworld one

Germany (Hybrid)
3 Months ago
HHA Exchange - Customer Success Manager

HHA Exchange

Ohio, United States (On-Site)
1 Month ago
Zelis  - Service Delivery Analyst

Zelis

Hyderabad, Telangana, India (On-Site)
3 Weeks ago
Head Digital Works - Data Scientist - Retention

Head Digital Works

Hyderabad, Telangana, India (On-Site)
11 Months ago

Get notifed when new similar jobs are uploaded

Jobs in undefined

Looks like we're out of matches

Set up an alert and we'll send you similar jobs the moment they appear!

Network Engineering Jobs

Microsoft - Technical Support Engineer - Windows Networking

Microsoft

(Hybrid)
2 Months ago
fluence - Senior Network Monitoring Engineer

fluence

Bengaluru, Karnataka, India (Hybrid)
6 Months ago
Intel  - AI SW Networking/Runtime Engineer

Intel

Tel Aviv-Yafo, Tel Aviv District, Israel (On-Site)
11 Months ago
Meta - Software Engineer - Datacenter networking

Meta

Bellevue, Washington, United States (On-Site)
7 Months ago
Qualcomm - Senior Engineer - Network Stack Development with AI

Qualcomm

Hyderabad, Telangana, India (On-Site)
1 Month ago
Sony Interactive Entertainment - Developer Experience Engineer (PlayStation™Network Server Platform Development)

Sony Interactive Entertainment

Tokyo, Japan (On-Site)
2 Months ago
Zones - Network Engineer L2

Zones

Bengaluru, Karnataka, India (On-Site)
6 Months ago
bytedance - Network Automation Engineer

bytedance

San Jose, California, United States (On-Site)
3 Months ago
bytedance - Cloud Network Engineer

bytedance

Ashburn, Virginia, United States (On-Site)
4 Months ago
CME Group - Staff Network Engineer

CME Group

Belfast, Northern Ireland, United Kingdom (Hybrid)
2 Weeks ago

Get notifed when new similar jobs are uploaded

About The Company

Since its founding in 1993, NVIDIA (NASDAQ: NVDA) has been a pioneer in accelerated computing. The company’s invention of the GPU in 1999 sparked the growth of the PC gaming market, redefined computer graphics, ignited the era of modern AI and is fueling the creation of the metaverse. NVIDIA is now a full-stack computing company with data-center-scale offerings that are reshaping industry.

Santa Clara, California, United States (On-Site)

Massachusetts, United States (On-Site)

Santa Clara, California, United States (On-Site)

Texas, United States (On-Site)

Santa Clara, California, United States (Hybrid)

Santa Clara, California, United States (Hybrid)

Pune, Maharashtra, India (On-Site)

Taipei City, Taiwan (On-Site)

Beijing, Beijing, China (On-Site)

Santa Clara, California, United States (On-Site)

View All Jobs

Get notified when new jobs are added by NVIDIA

Level Up Your Career in Game Development!

Transform Your Passion into Profession with Our Comprehensive Courses for Aspiring Game Developers.

Job Common Plug