Senior Production Engineer - Storage

1 Month ago • 5 Years + • DevOps

Job Summary

Job Description

NVIDIA seeks a Senior Production Engineer - Storage to assist in designing, implementing, and supporting large-scale storage clusters. Responsibilities include monitoring, logging, and alerting; working with AI/ML workloads to understand cluster behavior; improving service lifecycles; supporting pre-live services; maintaining live services by monitoring availability and performance; scaling systems sustainably through automation and AI/ML; practicing sustainable incident response; and participating in on-call rotation. This role requires expertise in large-scale Linux systems, coding (C/C++, Java, Python, Go, Perl, or Ruby), infrastructure management tools, and observability tools. The ideal candidate demonstrates an SRE mindset, a customer-first approach, and strong problem-solving skills.
Must have:
  • 5+ years experience
  • Large-scale Linux systems
  • Coding (C/C++, Java, Python etc.)
  • Infrastructure management tools
  • Observability tools
  • AI/ML experience
Good to have:
  • Kubernetes, OpenStack, Docker
  • Git, CI/CD
  • SRE mindset
  • Customer focus
  • Strong debugging skills

Job Details

Site Reliability Engineering (SRE) is an engineering discipline that involves designing, building, and maintaining large-scale production systems with high efficiency and availability. It encompasses various areas, including software and systems engineering practices, storage, data management, and services. SRE professionals are highly specialized and possess expertise in different domains such as systems, networking, storage, coding, database management, capacity management, continuous delivery, and deployment, as well as open-source cloud-enabling technologies like Kubernetes, containers, and virtualization. Their responsibilities encompass ensuring reliable storage solutions, managing data efficiently, and providing related services to support the overall stability and performance of the production systems.

SRE at NVIDIA ensures that our internal and external facing GPU cloud services have reliability and uptime as promised to the users and at the same time enables developers to make changes to the existing system through careful preparation and planning while keeping an eye on capacity, latency, and performance. SRE is also a mindset and a set of engineering approaches to running better production systems and optimizations. Much of our software development focuses on eliminating manual work through automation, performance tuning, and growing the efficiency of production systems. As SREs are responsible for the big picture of how our systems relate to each other, we use a breadth of tools and approaches to tackle a broad spectrum of problems. Practices such as limiting time spent on reactive operational work, blameless postmortems, and proactive identification of potential outages factor into iterative improvement that is key to product quality and interesting and dynamic day-to-day work. SRE's culture of diversity, intellectual curiosity, problem-solving, and openness is important to its success. Our organization brings together people with a wide variety of backgrounds, experiences, and perspectives. We encourage them to collaborate, think big, and take risks in a blame-free environment. We promote self-direction to work on meaningful projects while striving to build an environment that provides the support and mentorship needed to learn and grow.

What You Will Be Doing:

  • Assist in the design, implementation, and support of large-scale storage clusters, including monitoring, logging, and alerting.

  • Work with AI/ML workloads to capture and correlate behavior in large clusters and workflows, which are otherwise hard to understand.

  • Work closely with peers on the team to improve the lifecycle of services – from inception and design, through deployment, operation, and refinement.

  • Support services before they go live through activities such as system design consulting, developing software and frameworks, capacity management, and launch reviews.

  • Maintain services once they are live by measuring and monitoring availability, latency, and overall system health, including leveraging machine learning models.

  • Scale systems sustainably through mechanisms like AI/ML and automation, and evolve systems by pushing for changes that improve reliability and velocity.

  • Practice sustainable incident response and blameless postmortems.

  • Be part of an on-call rotation to support production systems.

What We Need To See:

  • BS degree in Computer Science or related technical field involving coding (e.g., physics or mathematics) or equivalent experience.

  • At least 5+ years practical experience.

  • Experience with algorithms, data structures, complexity analysis, software design, and maintaining large-scale Linux-based systems.

  • Experience in one or more of the following: C/C++, Java, Python, Go, Perl or Ruby, AI/ML frameworks and methodologies.

  • Good knowledge of infrastructure configuration management tools like Ansible, Chef, Puppet, and Terraform.

  • Experience in using observability and tracing-related tools like InfluxDB, Prometheus, and Elastic stack.

Ways to stand out from the crowd:

  • Demonstrated experience in having SRE mindset, customer-first approach, and focus on customer satisfaction and passion for ensuring customer success.Experience with Git, code review, pipelines, and CI/CD.

  • Interest in crafting, analyzing, and fixing large-scale distributed systems. Strong debugging skills with a systematic problem-solving approach to identify complex problems.

  • Thrive in collaborative environments and enjoy working with various teams. Experience in using or running large private and public cloud systems based on Kubernetes, OpenStack, and Docker. Flexible in adapting to different working styles.

NVIDIA is widely considered to be one of the technology world’s most desirable employers. We have some of the most forward-thinking and talented people on the planet working for us. If you're creative and autonomous, we want to hear from you!

Similar Jobs

Roofstacks - Senior Platform Engineer

Roofstacks

İstanbul, İstanbul, Türkiye (On-Site)
2 Months ago
Aristocrat Gaming - Sr. Tech Lead - Linux

Aristocrat Gaming

Noida, Uttar Pradesh, India (Hybrid)
1 Month ago
Playtech - Dev Ops Engineer

Playtech

London, England, United Kingdom (On-Site)
4 Months ago
Visa - Sr. Site Reliability Engineer, Product Reliability Engineering - Middleware

Visa

Austin, Texas, United States (Hybrid)
4 Months ago
Netflix - Distributed Systems Engineer (L5) - Compute Abstractions

Netflix

United States (Remote)
4 Months ago
GoTo Group - Site Reliability Engineer - EP (SE4)

GoTo Group

Gurugram, Haryana, India (On-Site)
6 Months ago
ByteDance - Backend Software Engineer - Foundational Technology

ByteDance

Singapore (On-Site)
1 Month ago
Tencent - Senior Site Reliability Engineer

Tencent

Shanghai, Shanghai, China (On-Site)
7 Months ago

Get notifed when new similar jobs are uploaded

Similar Skill Jobs

Life church - Ruby Staff Engineer

Life church

Edmond, Oklahoma, United States (On-Site)
6 Months ago
Scopely - Principal DevOps Engineer - Star Trek Fleet Command

Scopely

United Kingdom (Remote)
1 Month ago
Every matrix - Middle Frontend Developer (JavaScript)

Every matrix

Lviv, Lviv Oblast, Ukraine (Hybrid)
1 Month ago
NinjaVan - Senior Data Engineer

NinjaVan

Hyderabad, Telangana, India (On-Site)
6 Months ago
Voodoo - Senior Backend Engineer (Python) - Blitz

Voodoo

Paris, Île-de-France, France (On-Site)
4 Months ago
Lockwood - Cloud Engineer

Lockwood

United Kingdom (Remote)
1 Month ago
Every matrix - Senior Full-stack Developer (Angular/Node.js)

Every matrix

Bucharest, Bucharest, Romania (Hybrid)
2 Months ago
The Walt Disney Company - Senior Data Engineer

The Walt Disney Company

San Francisco, California, United States (On-Site)
3 Months ago
Ludeo - Senior Front End Developer

Ludeo

Tel Aviv-Yafo, Tel Aviv District, Israel (On-Site)
3 Months ago
NVIDIA - Senior Site Reliability Engineer - AI Research Clusters

NVIDIA

Pune, Maharashtra, India (On-Site)
1 Month ago

Get notifed when new similar jobs are uploaded

Jobs in Australia

WongDoody - Senior UX Designer

WongDoody

Melbourne, Victoria, Australia (Hybrid)
7 Months ago
Riot Games - Senior Software Engineer - 2XKO - Social

Riot Games

Sydney, New South Wales, Australia (On-Site)
2 Months ago
Canva - Backend Software Engineer - Product Quality

Canva

Surry Hills, New South Wales, Australia (Remote)
1 Month ago
Canva - Senior Frontend Engineer - Video AI

Canva

Sydney, New South Wales, Australia (Remote)
1 Month ago
Flying Bark Productions - Rigging Artist

Flying Bark Productions

New South Wales, Australia (Hybrid)
1 Month ago
The Walt Disney Company - Facilities Coordinator

The Walt Disney Company

Richmond, Victoria, Australia (On-Site)
2 Months ago
Canva - Frontend Engineer - Editing APIs

Canva

Surry Hills, New South Wales, Australia (Remote)
1 Month ago
Canva - Senior Software Engineer (Cloud FinOps) - remote across ANZ

Canva

Sydney, New South Wales, Australia (Remote)
3 Months ago
Canva - Senior Frontend Software Engineer - Cross Platform

Canva

Sydney, New South Wales, Australia (Remote)
1 Month ago
Canva - Senior Frontend Engineer - Canva for Education

Canva

Brisbane, Queensland, Australia (Remote)
1 Month ago

Get notifed when new similar jobs are uploaded

DevOps Jobs

Razer - Software Engineer (DevOps)

Razer

Shah Alam, Selangor, Malaysia (On-Site)
7 Months ago
Avathon - DevOps Engineer

Avathon

Bengaluru, Karnataka, India (On-Site)
6 Months ago
Tencent - SRE Intern

Tencent

(On-Site)
2 Months ago
Immutable - Senior Site Reliability Engineer

Immutable

Sydney, New South Wales, Australia (Hybrid)
5 Months ago
SmileGate - Platform Engineering Manager

SmileGate

Seongnam-si, Gyeonggi-do, South Korea (On-Site)
3 Months ago
NVIDIA - Senior ASIC Front End Infrastructure Engineer

NVIDIA

Santa Clara, California, United States (Hybrid)
2 Months ago
ByteDance - Software Engineer, SRE - Platform Services

ByteDance

San Jose, California, United States (On-Site)
2 Months ago
Luxoft - DevOps Engineer with Azure

Luxoft

Pune, Maharashtra, India (On-Site)
4 Months ago
Nielsen Holdings - Software Engineer ( Java , Python , SQL , AWS / Oracle)

Nielsen Holdings

Bengaluru, Karnataka, India (Hybrid)
6 Months ago
Hashone Careers - Cloud Engineer

Hashone Careers

Bengaluru, Karnataka, India (Remote)
5 Months ago

Get notifed when new similar jobs are uploaded

About The Company

Since its founding in 1993, NVIDIA (NASDAQ: NVDA) has been a pioneer in accelerated computing. The company’s invention of the GPU in 1999 sparked the growth of the PC gaming market, redefined computer graphics, ignited the era of modern AI and is fueling the creation of the metaverse. NVIDIA is now a full-stack computing company with data-center-scale offerings that are reshaping industry.

Santa Clara, California, United States (On-Site)

Santa Clara, California, United States (On-Site)

Massachusetts, United States (On-Site)

Santa Clara, California, United States (On-Site)

Santa Clara, California, United States (On-Site)

Santa Clara, California, United States (On-Site)

Texas, United States (On-Site)

Santa Clara, California, United States (On-Site)

Santa Clara, California, United States (Hybrid)

Santa Clara, California, United States (Hybrid)

View All Jobs

Get notified when new jobs are added by NVIDIA

Level Up Your Career in Game Development!

Transform Your Passion into Profession with Our Comprehensive Courses for Aspiring Game Developers.

Job Common Plug