Outscal Logooutscal logo

Senior Software Engineer - HPC

1 Month ago • 10 Years + • DevOps • $184,000 PA - $356,500 PA

Job Summary

Job Description

NVIDIA seeks a Senior Software Engineer for its HPC infrastructure team. Responsibilities encompass designing highly available and scalable systems, evaluating new technologies, improving infrastructure provisioning using automation, supporting a multi-cloud environment (AWS, GCP, on-prem), building cross-functional relationships, and ensuring high uptime. The role involves on-call rotation and incident response. Ideal candidates possess strong software development skills, experience with distributed systems, and expertise in cloud computing and CI/CD.
Must have:
  • 10+ years experience in large engineering projects
  • Proficiency in Golang, Java, C/C++, Scala, Python, or Elixir
  • Experience in designing scalable, resilient systems
  • Cloud computing expertise (GCP, AWS, or Azure)
  • CI/CD, GitOps, and IaC proficiency
  • Strong problem-solving skills
Good to have:
  • Experience with Slurm or Kubernetes for HPC clusters
  • Strong understanding of Linux and TCP/IP
Perks:
  • Equity
  • Benefits

Job Details

NVIDIA has continuously reinvented itself over two decades. Our invention of the GPU in 1999 fueled the growth of the PC gaming market, redefined modern computer graphics, and revolutionized parallel computing. More recently, GPU deep learning ignited modern AI and enabled the next era of computing. NVIDIA is a “learning machine” that constantly evolves by adapting to new opportunities that are hard to address, that matters to the world, and that only we can address. This is our life’s work, to amplify human imagination and intelligence, and expand what is possible. We’re seeking strategic, bold, hard-working, and creative individuals who are passionate about helping us tackle challenges no one else can solve. Make the choice to join us today.
 

We are looking for a Senior Software Engineer to join our mission to continue improving our HPC infrastructure. Our team builds and operates sophisticated infrastructure to enable business critical services and AI applications. You will be working with a team of passionate and skilled engineers that are continuously working to provide better tools to build and manage this infrastructure. Ideal candidate is strong in software development, designing and creating reliable distributed systems, and has the ability to implement well thought out long term maintenance strategy.


What you’ll be doing:

  • Design highly available and scalable systems to meet the demands of our HPC clusters

  • Evaluate new and innovative technologies as the landscape evolves

  • Continuously improve infrastructure provisioning and management using automation

  • Support a globally distributed, multi-cloud hybrid environment - AWS, GCP and On-prem

  • Build strong cross functional relationships and align with partners across various business units

  • Ensure the highest level of up-time and Quality of Service (QoS) to our users through operational excellence

  • Participate in team's on-call rotation and be a contact for service incidents


What we need to see:

  • 10+ years of experience in design, implementation, and delivery of large engineering projects

  • Comfortable with at least two of the following programming languages: Golang, Java, C/C++, Scala, Python, Elixir.

  • Understands scalability challenges and performance of server-side code. Able to craft and develop horizontally-scalable, resilient and performing-under-load systems.

  • Versatile technologist with experience in full software development lifecycle – from inception and design to deployment, operation, and iterative development.

  • Proficient in cloud computing and are hands-on in at least one cloud platform: GCP, AWS, or Azure.

  • Proficient in modern CI/CD techniques, GitOps and Infrastructure as Code(IaC)

  • Strong work ethic and a passion for problem solving

  • B.S. degree in Computer Science or related technical field (or equivalent experience)

  • Detail oriented with great communication and collaboration skills


Ways to stand out from the crowd:

  • Prior experience building solutions for HPC clusters based on Slurm or Kubernetes

  • Strong understanding of Linux operation system and TCP/IP fundamentals

The base salary range is 184,000 USD - 356,500 USD. Your base salary will be determined based on your location, experience, and the pay of employees in similar positions.

You will also be eligible for equity and benefits. NVIDIA accepts applications on an ongoing basis.

NVIDIA is committed to fostering a diverse work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.

Similar Jobs

The Walt Disney Company - Lead Software Engineer, Machine Learning - Ad Platforms

The Walt Disney Company

Santa Monica, California, United States (On-Site)
4 Months ago
Nagarro - Senior Staff Engineer, Mobile Android

Nagarro

United Arab Emirates (Remote)
5 Months ago
Axinous - Principal Software Engineer - API Tooling & Frameworks

Axinous

San Jose, California, United States (Hybrid)
4 Months ago
Snowprint Studios - Server Developer

Snowprint Studios

Stockholm, Stockholm County, Sweden (Hybrid)
11 Hours ago
Moon Active - Data Platform Engineer

Moon Active

Tel Aviv-Yafo, Tel Aviv District, Israel (Hybrid)
6 Days ago
Paytm - DevOps Engineer/Senior DevOps-Paytm Money

Paytm

Bengaluru, Karnataka, India (On-Site)
3 Months ago
ByteDance - Backend Software Engineer - Foundational Technology

ByteDance

Singapore (On-Site)
1 Day ago
Onward Search - DevOps/Automation Engineer

Onward Search

New York, New York, United States (Remote)
1 Month ago
ION - Software Architect - Java Multi-Tenant SAAS Cloud Native

ION

Pune, Maharashtra, India (On-Site)
5 Months ago

Get notifed when new similar jobs are uploaded

Similar Skill Jobs

Meta - Software Engineer (Leadership) - Machine Learning

Meta

Bellevue, Washington, United States (Remote)
4 Months ago
CCP Games - Backend Programming Intern

CCP Games

Shanghai, Shanghai, China (On-Site)
2 Weeks ago
Starkflow - Oracle SOA Consultant

Starkflow

New South Wales, Australia (Hybrid)
3 Weeks ago
Glean - Software Engineer (Support Tools Developer)

Glean

Bengaluru, Karnataka, India (On-Site)
4 Weeks ago
ByteDance - Software Engineer Intern (CDN/Edge/Traffic Platform)

ByteDance

San Jose, California, United States (On-Site)
1 Day ago
Meta - Production Engineering

Meta

Boston, Massachusetts, United States (On-Site)
4 Months ago
ByteDance - Software Engineer - Data Tech Infrastructure- San Jose

ByteDance

San Jose, California, United States (On-Site)
4 Months ago
Microsoft - Member of Technical Staff – Machine Learning Engineer

Microsoft

New York, New York, United States (Hybrid)
1 Day ago
Dun & Bradstreet - 2025 Summer Internship Program - Technology

Dun & Bradstreet

Jacksonville, Florida, United States (On-Site)
5 Months ago
The Walt Disney Company - Retail ERP Solution Architect

The Walt Disney Company

Île-de-France, France (On-Site)
21 Hours ago

Get notifed when new similar jobs are uploaded

Jobs in Canada

Epic Games - Senior Outsource Artist

Epic Games

Vancouver, British Columbia, Canada (On-Site)
6 Months ago
Scanline VFX - Compositing Supervisor

Scanline VFX

Vancouver, British Columbia, Canada (Hybrid)
3 Weeks ago
Electronic Arts - Senior Analyst - NHL

Electronic Arts

Vancouver, British Columbia, Canada (Hybrid)
1 Week ago
People Can Fly - Live Operations Technician

People Can Fly

Montreal, Quebec, Canada (Remote)
1 Week ago
Gamemode One  Inc  - Junior Programmer - Summer 2025 Co-op

Gamemode One Inc

Nova Scotia, Canada (Hybrid)
2 Months ago
Epic Games - Senior Gameplay Programmer

Epic Games

Montreal, Quebec, Canada (On-Site)
2 Months ago
Scanline VFX - Lead Software Engineer

Scanline VFX

Vancouver, British Columbia, Canada (Remote)
5 Months ago
Activate Games - Store Leader (Store Manager)

Activate Games

Mississauga, Ontario, Canada (On-Site)
1 Month ago

Get notifed when new similar jobs are uploaded

DevOps Jobs

Teradata - Senior Cloud Engineer

Teradata

Pune, Maharashtra, India (On-Site)
4 Months ago
ByteDance - Site Reliability Engineer Lead, Security Engineering

ByteDance

Singapore (On-Site)
5 Months ago
Enphase Energy - Staff Devops Engineer

Enphase Energy

Bengaluru, Karnataka, India (On-Site)
4 Months ago
Ajmera Infotech - Senior ASP.NET Developer with Azure Expertise

Ajmera Infotech

Austin, Texas, United States (On-Site)
1 Month ago
Nintendo - CONTRACT - Sr Engineer (NTD)

Nintendo

Redmond, Washington, United States (On-Site)
4 Months ago
Interactive Brokers - Senior Platform Engineer - Design

Interactive Brokers

Fort Lauderdale, Florida, United States (Hybrid)
5 Months ago
bosh group india - 2024_MS_EDE3_XC_SRE_DataEngineering

bosh group india

Bengaluru, Karnataka, India (On-Site)
3 Months ago
ION - Microsoft System Engineer, Italy

ION

Italy (Hybrid)
5 Months ago
OneLocal - Senior DevOps Engineer

OneLocal

(Remote)
1 Week ago

Get notifed when new similar jobs are uploaded

About The Company

Since its founding in 1993, NVIDIA (NASDAQ: NVDA) has been a pioneer in accelerated computing. The company’s invention of the GPU in 1999 sparked the growth of the PC gaming market, redefined computer graphics, ignited the era of modern AI and is fueling the creation of the metaverse. NVIDIA is now a full-stack computing company with data-center-scale offerings that are reshaping industry.


Hsinchu, Hsinchu City, Taiwan (On-Site)

Yokne'am Illit, North District, Israel (On-Site)

Seoul, South Korea (Hybrid)

Tel Aviv-Yafo, Tel Aviv District, Israel (On-Site)

Ra'anana, Center District, Israel (On-Site)

Shanghai, Shanghai, China (On-Site)

Tel Aviv-Yafo, Tel Aviv District, Israel (On-Site)

Be'er Sheva, South District, Israel (On-Site)

Santa Clara, California, United States (On-Site)

Santa Clara, California, United States (On-Site)

View All Jobs

Get notified when new jobs are added by NVIDIA

Level Up Your Career in Game Development!

Transform Your Passion into Profession with Our Comprehensive Courses for Aspiring Game Developers.

Job Common Plug