Senior DevOps Infrastructure Engineer, Open-Source CI and CD

3 Days ago • 7 Years + • DevOps • $168,000 PA - $333,500 PA

Job Summary

Job Description

NVIDIA's GitHub Actions Runner team seeks a Senior DevOps Infrastructure Engineer to manage and scale self-hosted GPU-enabled GitHub Actions runners. Responsibilities include managing runners using Kubernetes, expanding runner support for various hardware and OS combinations, using Infrastructure as Code (Terraform and ArgoCD), building and maintaining runner VM images (HashiCorp Packer), securing distributed services (mTLS, PKI, HashiCorp Vault), developing Golang tools for platform observability, configuring monitoring (Prometheus, Grafana), contributing to open-source projects, and addressing CVEs. The role involves collaboration with 100+ developers and requires strong Kubernetes, GitOps, and CI/CD expertise.
Must have:
  • 7+ years experience in infrastructure/DevOps
  • Strong Kubernetes expertise
  • GitOps tools (ArgoCD) experience
  • Terraform/Terragrunt proficiency
  • Golang, Python, TypeScript proficiency
  • Monitoring/logging experience (Prometheus, Grafana)
  • GitHub Actions expertise
Good to have:
  • Telemetry instrumentation for distributed systems
  • GPU workloads on Kubernetes
  • Custom Kubernetes controllers
  • KubeVirt/virtualization experience
  • Open-source Kubernetes project contributions
Perks:
  • Competitive salary
  • Generous benefits package

Job Details

At NVIDIA, we are pushing the boundaries of AI, graphics, and computing. The GitHub Actions Runner team manages self-hosted GPU-enabled GitHub Actions runners, using Actions Runner Controller with KubeVirt to deploy ephemeral VM-based runners for NVIDIA’s open source projects on GitHub. The team operates both on-premise and in the cloud (AWS) to support 100+ developers with whom they collaborate regularly to ensure a seamless CI/CD experience. We are looking for a Senior Infrastructure Engineer to help scale, optimize, and expand our platform.

Our team is fully remote and distributed across multiple time zones. If you're passionate about infrastructure, Kubernetes, automation, and observability, this is an opportunity to work with exciting technology at one of the most innovative companies in the world.

Preferred work location: Eastern/Central time zones

What you'll be doing:

  • Manage and scale self-hosted GitHub Actions runners using Kubernetes

  • Help expand runner support for various hardware and operating system combinations, including Linux, Windows, single-GPU, multi-GPU, NVLink, and more

  • Use Infrastructure as Code (Terraform and ArgoCD) to deploy and maintain infrastructure both on-premise and in AWS

  • Build and maintain runner VM images using HashiCorp Packer

  • Connect distributed services securely using mTLS, PKI, and HashiCorp Vault

  • Develop, package, and deploy custom Golang tools to support platform observability, stability, and efficiency

  • Configure alerting and monitoring to identify and address issues quickly, using tools like Prometheus and Grafana

  • Contribute upstream to open-source tools and libraries that our team depends on

  • Periodically update platform dependencies and address CVEs

What we need to see:

  • B.S. or M.S. in Computer Science, Computer Engineering, or a related field (or equivalent experience)

  • 7+ years of proven experience in infrastructure, DevOps, or platform engineering

  • Strong Kubernetes expertise (running, debugging, and scaling workloads)

  • Experience with GitOps tools (ArgoCD or similar)

  • Proficiency in Linux administration and troubleshooting

  • Experience with Infrastructure as Code using Terraform/Terragrunt

  • Proficiency in Golang, Python, and TypeScript

  • Hands-on experience with monitoring, logging, and tracing (Prometheus, Grafana, OpenTelemetry, etc.)

  • Solid understanding of CI/CD pipelines, particularly GitHub Actions

  • Ability to work and collaborate effectively with a fully remote, distributed team

Ways to stand out from the crowd:

  • Experience instrumenting telemetry for distributed systems

  • Strong background in GPU workloads on Kubernetes with experience writing custom Kubernetes controllers

  • Deep understanding of KubeVirt and/or virtualization

  • Experience with self-hosted GitHub Actions runners

  • Contributions to open-source Kubernetes-related projects

With competitive salaries and a generous benefits package, NVIDIA is considered one of the technology world’s most desirable employers. We have some of the most forward-thinking and hardworking individuals in the industry working for us. Due to unprecedented growth, our exclusive engineering teams are expanding rapidly. If you're a creative and autonomous engineer with a genuine passion for technology, we want to hear from you!

The base salary range is 168,000 USD - 333,500 USD. Your base salary will be determined based on your location, experience, and the pay of employees in similar positions.

You will also be eligible for equity and benefits. NVIDIA accepts applications on an ongoing basis.

NVIDIA is committed to fostering a diverse work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.

Similar Jobs

KBG Blockchain Game Studios - Back-End Developer (NodeJS)

KBG Blockchain Game Studios

Thành Phố Hồ Chí Minh, Vietnam (On-Site)
9 Months ago
Knuddels - Senior Web Developer

Knuddels

Baden-Württemberg, Germany (Remote)
3 Weeks ago
Netflix - Full-Stack Engineer (L5)

Netflix

Warsaw, Masovian Voivodeship, Poland (On-Site)
2 Months ago
Canva - Frontend Software Engineer - Internationalization

Canva

Beijing, Beijing, China (Remote)
1 Month ago
Hacksaw Studios - Game developer

Hacksaw Studios

Stockholm, Stockholm County, Sweden (On-Site)
9 Months ago
Rackspace Technology - Lead Cloud Engineer

Rackspace Technology

United States (Remote)
1 Month ago
Zazz - Cloud Engineer (AWS)

Zazz

(Remote)
2 Months ago
Google - Customer Engineer III, Infrastructure, National Security, Public Sector

Google

Reston, Virginia, United States (On-Site)
4 Days ago
CD PROJEKT RED - ML Ops Engineer

CD PROJEKT RED

Warsaw, Masovian Voivodeship, Poland (On-Site)
3 Weeks ago
Kolibri Games - DevOps Engineer

Kolibri Games

Berlin, Berlin, Germany (Hybrid)
3 Weeks ago

Get notifed when new similar jobs are uploaded

Similar Skill Jobs

Flow - Senior/Staff Web Engineer

Flow

Miami, Florida, United States (Hybrid)
6 Months ago
Notion - Software Engineer, Android

Notion

San Francisco, California, United States (On-Site)
6 Months ago
PlayStation Global - Staff Software Engineer

PlayStation Global

Dublin, County Dublin, Ireland (On-Site)
1 Week ago
Ludeo - Senior Back End Developer

Ludeo

Tel Aviv-Yafo, Tel Aviv District, Israel (On-Site)
2 Weeks ago
Google - Software Engineer III, VM Manager, Google Cloud

Google

Warsaw, Masovian Voivodeship, Poland (On-Site)
4 Days ago
Google - Senior Staff Engineer, Google Distributed Cloud Air-gapped

Google

Sunnyvale, California, United States (On-Site)
4 Days ago
Push Gaming - Senior Game Developer

Push Gaming

Malta (Hybrid)
3 Weeks ago
Every matrix - Game Developer (Slots, Pixi.js)

Every matrix

Stockholm, Stockholm County, Sweden (Hybrid)
3 Months ago
Tesla - Vehicle Software, Service Engineering Internship

Tesla

North Holland, Netherlands (On-Site)
2 Months ago
Arrowhead Game Studios - Full-Stack Engineer

Arrowhead Game Studios

Stockholm, Stockholm County, Sweden (Hybrid)
3 Weeks ago

Get notifed when new similar jobs are uploaded

Jobs in New Jersey, United States

Next Level Business Services - Big Data Engineer

Next Level Business Services

Phoenix, Arizona, United States (On-Site)
6 Months ago
Lionsgate Games - Manager, Tax

Lionsgate Games

Santa Monica, California, United States (On-Site)
4 Weeks ago
Inkittt - Senior Product Manager, Inkitt Product

Inkittt

San Francisco, California, United States (On-Site)
8 Months ago
ByteDance - Senior Software Engineer

ByteDance

San Jose, California, United States (On-Site)
2 Months ago
Thatgamecompany - Senior Software Engineer - Golang

Thatgamecompany

United States (Remote)
3 Weeks ago
Dun & Bradstreet - Account Executive, Outbound - Manage (R-16769)

Dun & Bradstreet

Jacksonville, Florida, United States (On-Site)
6 Months ago
ByteDance - Research Scientist Graduate (High-Performance Computing (Algorithm Acceleration)- Vision AI Platform)

ByteDance

San Jose, California, United States (On-Site)
3 Weeks ago
ByteDance - Software Engineer Intern (Doubao (Seed) - Machine Learning System) - 2025 Summer (MS)

ByteDance

Seattle, Washington, United States (On-Site)
5 Months ago
Blue Yonder - Sr Solution Architect

Blue Yonder

Dallas, Texas, United States (On-Site)
6 Months ago
Universal Music - Manager, Royalty Audits

Universal Music

Los Angeles, California, United States (On-Site)
1 Month ago

Get notifed when new similar jobs are uploaded

DevOps Jobs

NVIDIA - Senior Site Reliability Engineer - AI Research Clusters

NVIDIA

Westford, Massachusetts, United States (Hybrid)
1 Month ago
Google - Software Engineering Manager II, Google Cloud Platform

Google

Warsaw, Masovian Voivodeship, Poland (On-Site)
5 Days ago
Scanline VFX - Senior DevOps Engineer

Scanline VFX

Seoul, South Korea (Hybrid)
2 Months ago
Luxoft - Senior Software Support Engineer

Luxoft

(Remote)
5 Months ago
Lost Boys Interactive - Senior DevOps Engineer

Lost Boys Interactive

(Remote)
3 Months ago
Netflix - Engineering Manager - Edge Gateway & Services

Netflix

United States (Remote)
3 Days ago
Teradata - Senior Cloud Engineer

Teradata

Pune, Maharashtra, India (On-Site)
5 Months ago
Rackspace Technology - OpenStack Cloud Engineer IV

Rackspace Technology

(Remote)
2 Months ago
Ajmera Infotech - Site Reliability Engineer (SRE) - Kubernetes

Ajmera Infotech

Bengaluru, Karnataka, India (On-Site)
1 Week ago

Get notifed when new similar jobs are uploaded

About The Company

Since its founding in 1993, NVIDIA (NASDAQ: NVDA) has been a pioneer in accelerated computing. The company’s invention of the GPU in 1999 sparked the growth of the PC gaming market, redefined computer graphics, ignited the era of modern AI and is fueling the creation of the metaverse. NVIDIA is now a full-stack computing company with data-center-scale offerings that are reshaping industry.

Tel Aviv-Yafo, Tel Aviv District, Israel (On-Site)

Hyderabad, Telangana, India (On-Site)

Pune, Maharashtra, India (On-Site)

Pune, Maharashtra, India (On-Site)

Yokne'am Illit, North District, Israel (On-Site)

Shenzhen, Guangdong Province, China (On-Site)

Taipei City, Taiwan (On-Site)

California, United States (Remote)

Yokne'am Illit, North District, Israel (On-Site)

View All Jobs

Get notified when new jobs are added by NVIDIA

Level Up Your Career in Game Development!

Transform Your Passion into Profession with Our Comprehensive Courses for Aspiring Game Developers.

Job Common Plug