Senior Production Engineer - Storage

1 Month ago • 5 Years + • DevOps • $148,000 PA - $356,500 PA

Job Summary

Job Description

This Senior Production Engineer - Storage role at NVIDIA involves designing, implementing, and supporting large-scale storage clusters, working with AI/ML workloads, and collaborating with peers to improve service lifecycles. Responsibilities include system design consulting, software development, capacity management, monitoring system health, scaling systems through automation and AI/ML, and practicing sustainable incident response. The role requires on-call rotation to support production systems and a strong understanding of SRE principles. The candidate should possess expertise in storage, data management, systems, networking, coding, and cloud technologies like Kubernetes and containers.
Must have:
  • BS in CS or related field
  • 5+ years experience
  • Experience with large-scale Linux systems
  • Proficiency in C/C++, Java, Python, or similar
  • Experience with configuration management tools
  • Experience with observability tools
Good to have:
  • SRE mindset
  • Experience with Git and CI/CD
  • Experience with Kubernetes, OpenStack, and Docker
  • Strong debugging skills
  • Experience with AI/ML frameworks
Perks:
  • Equity
  • Benefits

Job Details

Site Reliability Engineering (SRE) is an engineering discipline that involves designing, building, and maintaining large-scale production systems with high efficiency and availability. It encompasses various areas, including software and systems engineering practices, storage, data management, and services. SRE professionals are highly specialized and possess expertise in different domains such as systems, networking, storage, coding, database management, capacity management, continuous delivery, and deployment, as well as open-source cloud-enabling technologies like Kubernetes, containers, and virtualization. Their responsibilities encompass ensuring reliable storage solutions, managing data efficiently, and providing related services to support the overall stability and performance of the production systems.

SRE at NVIDIA ensures that our internal and external facing GPU cloud services have reliability and uptime as promised to the users and at the same time enables developers to make changes to the existing system through careful preparation and planning while keeping an eye on capacity, latency, and performance. SRE is also a mindset and a set of engineering approaches to running better production systems and optimizations. Much of our software development focuses on eliminating manual work through automation, performance tuning, and growing the efficiency of production systems. As SREs are responsible for the big picture of how our systems relate to each other, we use a breadth of tools and approaches to tackle a broad spectrum of problems. Practices such as limiting time spent on reactive operational work, blameless postmortems, and proactive identification of potential outages factor into iterative improvement that is key to product quality and interesting and dynamic day-to-day work. SRE's culture of diversity, intellectual curiosity, problem-solving, and openness is important to its success. Our organization brings together people with a wide variety of backgrounds, experiences, and perspectives. We encourage them to collaborate, think big, and take risks in a blame-free environment. We promote self-direction to work on meaningful projects while striving to build an environment that provides the support and mentorship needed to learn and grow.

What You Will Be Doing:

  • Assist in the design, implementation, and support of large-scale storage clusters, including monitoring, logging, and alerting.

  • Work with AI/ML workloads to capture and correlate behavior in large clusters and workflows, which are otherwise hard to understand.

  • Work closely with peers on the team to improve the lifecycle of services – from inception and design, through deployment, operation, and refinement.

  • Support services before they go live through activities such as system design consulting, developing software and frameworks, capacity management, and launch reviews.

  • Maintain services once they are live by measuring and monitoring availability, latency, and overall system health, including leveraging machine learning models.

  • Scale systems sustainably through mechanisms like AI/ML and automation, and evolve systems by pushing for changes that improve reliability and velocity.

  • Practice sustainable incident response and blameless postmortems.

  • Be part of an on-call rotation to support production systems.

What We Need To See:

  • BS degree in Computer Science or related technical field involving coding (e.g., physics or mathematics) or equivalent experience.

  • At least 5+ years practical experience.

  • Experience with algorithms, data structures, complexity analysis, software design, and maintaining large-scale Linux-based systems.

  • Experience in one or more of the following: C/C++, Java, Python, Go, Perl or Ruby, AI/ML frameworks and methodologies.

  • Good knowledge of infrastructure configuration management tools like Ansible, Chef, Puppet, and Terraform.

  • Experience in using observability and tracing-related tools like InfluxDB, Prometheus, and Elastic stack.

Ways to stand out from the crowd:

  • Demonstrated experience in having SRE mindset, customer-first approach, and focus on customer satisfaction and passion for ensuring customer success.Experience with Git, code review, pipelines, and CI/CD.

  • Interest in crafting, analyzing, and fixing large-scale distributed systems. Strong debugging skills with a systematic problem-solving approach to identify complex problems.

  • Thrive in collaborative environments and enjoy working with various teams. Experience in using or running large private and public cloud systems based on Kubernetes, OpenStack, and Docker. Flexible in adapting to different working styles.

NVIDIA is widely considered to be one of the technology world’s most desirable employers. We have some of the most forward-thinking and talented people on the planet working for us. If you're creative and autonomous, we want to hear from you!

The base salary range is 148,000 USD - 356,500 USD. Your base salary will be determined based on your location, experience, and the pay of employees in similar positions.

You will also be eligible for equity and benefits. NVIDIA accepts applications on an ongoing basis.

NVIDIA is committed to fostering a diverse work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.

Similar Jobs

Playtech - Java Developer

Playtech

Kyiv, Kyiv City, Ukraine (On-Site)
1 Month ago
NVIDIA - Senior Software Engineer - Conversational AI

NVIDIA

Pune, Maharashtra, India (On-Site)
1 Month ago
Ness Digital - Lead .Net Full-stack Engineer

Ness Digital

Timișoara, Timiș, Romania (Remote)
1 Month ago
The Walt Disney Company - Lead Software Engineer - Ad Platforms

The Walt Disney Company

Santa Monica, California, United States (On-Site)
1 Month ago
Nagarro - Senior Staff Engineer - Python Developer

Nagarro

Colombia (Remote)
2 Months ago
PwC - IN_Senior Associate_Azure Data Engineer _OneCloud _Advisory _Bangalore

PwC

Bengaluru, Karnataka, India (On-Site)
7 Months ago
Tencent - Technical Account Manager

Tencent

Tokyo, Japan (On-Site)
3 Months ago
Scanline VFX - Senior DevOps Engineer

Scanline VFX

Montreal, Quebec, Canada (Hybrid)
2 Months ago
Playrix - Senior Release Automation Engineer (Gardenscapes)

Playrix

Ireland (Remote)
3 Months ago
Warner Bros Games - Staff Software Engineer - Database Engineer with Aurora Postgres

Warner Bros Games

Bengaluru, Karnataka, India (Hybrid)
3 Weeks ago

Get notifed when new similar jobs are uploaded

Similar Skill Jobs

ION - Data Engineer

ION

Budapest, Hungary (On-Site)
6 Months ago
Mixmob - Senior Full-Stack React/Node & NFT Gaming Developer

Mixmob

Vancouver, British Columbia, Canada (Remote)
9 Months ago
Toptracer - Software Engineer

Toptracer

Stockholm, Stockholm County, Sweden (Hybrid)
3 Months ago
Intrepid Studios,  Inc  - DevOps Engineer (Kubernetes & Cloud Services)

Intrepid Studios, Inc

San Diego, California, United States (On-Site)
8 Months ago
ION - Senior Security Architect

ION

Milan, Lombardy, Italy (On-Site)
6 Months ago
N-iX - Senior .NET Full-Stack Engineer

N-iX

Ukraine (Remote)
3 Months ago
Socialpoint - Senior Software Engineer (Full Stack Engineer)

Socialpoint

Barcelona, Catalonia, Spain (Hybrid)
1 Month ago
Every matrix - Senior Backend Developer (NodeJS)

Every matrix

Lviv, Lviv Oblast, Ukraine (Hybrid)
1 Month ago
SmileGate - Platform Engineering Lead

SmileGate

Seongnam-si, Gyeonggi-do, South Korea (On-Site)
2 Months ago
ByteDance - Site Reliability Engineer (Cloud) - Infrastructure Engineering

ByteDance

Singapore (On-Site)
5 Months ago

Get notifed when new similar jobs are uploaded

Jobs in Canada

Maxis Studios - Creative Director

Maxis Studios

Vancouver, British Columbia, Canada (Hybrid)
4 Months ago
IGG - Unity Front-End Programmer

IGG

Vancouver, British Columbia, Canada (On-Site)
1 Month ago
Ubisoft - Lead Technical Programmer - Rainbow Six

Ubisoft

Montreal, Quebec, Canada (On-Site)
1 Month ago
PwC - PwC Private, High Net Worth Tax, Manager

PwC

Toronto, Ontario, Canada (On-Site)
7 Months ago
Scanline VFX - Senior Generalist Technical Animator

Scanline VFX

Quebec, Canada (On-Site)
1 Month ago
Epic Games - Senior DevOps Programmer

Epic Games

Montreal, Quebec, Canada (On-Site)
1 Month ago
Scanline VFX - Project Manager

Scanline VFX

Vancouver, British Columbia, Canada (Hybrid)
2 Months ago
Ubisoft - ML OPS Senior _ Groupe Technologique Création de contenu

Ubisoft

Montreal, Quebec, Canada (On-Site)
3 Months ago
Lionbridge Games - Community Coordinator

Lionbridge Games

Quebec, Canada (Hybrid)
2 Months ago
Ubisoft - Technical Graphic Director (Art)

Ubisoft

Montreal, Quebec, Canada (On-Site)
1 Month ago

Get notifed when new similar jobs are uploaded

DevOps Jobs

Netflix - Full Stack Engineer L5 - Cloud Engineering

Netflix

Los Gatos, California, United States (On-Site)
6 Months ago
Wargaming - DevOps Engineer

Wargaming

Belgrade, Serbia (On-Site)
4 Months ago
CloudHire - AWS Cloud Engineer

CloudHire

India (Remote)
1 Month ago
ION - Site Reliability Engineer

ION

Pisa, Tuscany, Italy (Hybrid)
6 Months ago
SmileGate - SRE Strategy Project Manager

SmileGate

Seongnam-si, Gyeonggi-do, South Korea (On-Site)
1 Month ago
Wargaming - DevOps Engineer

Wargaming

Shanghai, Shanghai, China (On-Site)
1 Month ago
Fandom - Principal DevOps Engineer

Fandom

Poznań, Greater Poland Voivodeship, Poland (Remote)
2 Months ago
Rockstar Games - DevOps Engineer

Rockstar Games

Edinburgh, Scotland, United Kingdom (On-Site)
11 Months ago
Company3 Method Studios - Technical Architect D365 Finance & Operations

Company3 Method Studios

Maharashtra, India (Remote)
3 Months ago
SmileGate - Platform Engineering Lead

SmileGate

Seongnam-si, Gyeonggi-do, South Korea (On-Site)
2 Months ago

Get notifed when new similar jobs are uploaded

About The Company

Since its founding in 1993, NVIDIA (NASDAQ: NVDA) has been a pioneer in accelerated computing. The company’s invention of the GPU in 1999 sparked the growth of the PC gaming market, redefined computer graphics, ignited the era of modern AI and is fueling the creation of the metaverse. NVIDIA is now a full-stack computing company with data-center-scale offerings that are reshaping industry.

Santa Clara, California, United States (On-Site)

Santa Clara, California, United States (On-Site)

Massachusetts, United States (On-Site)

Santa Clara, California, United States (On-Site)

Santa Clara, California, United States (On-Site)

Santa Clara, California, United States (On-Site)

Texas, United States (On-Site)

Santa Clara, California, United States (On-Site)

Santa Clara, California, United States (Hybrid)

Santa Clara, California, United States (Hybrid)

View All Jobs

Get notified when new jobs are added by NVIDIA

Level Up Your Career in Game Development!

Transform Your Passion into Profession with Our Comprehensive Courses for Aspiring Game Developers.

Job Common Plug