Jobs Courses Resources Companies Placements

Home >

Jobs >

Senior Production Engineer - Storage

NVIDIA

California, United States (On-site)

Senior Production Engineer - Storage

5 Months ago • 5 Years + • Devops • $148,000 PA - $356,500 PA

Job Summary

Job Description

This Senior Production Engineer - Storage role at NVIDIA involves designing, implementing, and supporting large-scale storage clusters. Responsibilities include monitoring, logging, alerting, working with AI/ML workloads, improving service lifecycles, and maintaining service health. The position requires expertise in storage solutions, data management, and related services. The ideal candidate possesses strong problem-solving skills, a systematic approach, and experience with large-scale Linux-based systems, various programming languages (C/C++, Java, Python, etc.), and infrastructure management tools. On-call rotation is required. The role emphasizes a customer-first approach and collaboration within diverse teams.

Must have:

5+ years experience
Large-scale Linux systems
C/C++, Java, Python, or similar
Infrastructure management tools
Observability & tracing tools
AI/ML experience
SRE mindset
Problem-solving skills

Good to have:

Kubernetes, OpenStack, Docker
Git, CI/CD
Ansible, Chef, Puppet, Terraform
InfluxDB, Prometheus, Elastic stack

Perks:

Equity
Benefits

15 skills required

15 skills required for this role

Add these skills to join the top 1% applicants for this job

kubernetes

java

ruby

unity

algorithms

ci-cd

github

puppet

containers

chef

data-structures

python

perl

docker

git

Job Details

Site Reliability Engineering (SRE) is an engineering discipline that involves designing, building, and maintaining large-scale production systems with high efficiency and availability. It encompasses various areas, including software and systems engineering practices, storage, data management, and services. SRE professionals are highly specialized and possess expertise in different domains such as systems, networking, storage, coding, database management, capacity management, continuous delivery, and deployment, as well as open-source cloud-enabling technologies like Kubernetes, containers, and virtualization. Their responsibilities encompass ensuring reliable storage solutions, managing data efficiently, and providing related services to support the overall stability and performance of the production systems.

SRE at NVIDIA ensures that our internal and external facing GPU cloud services have reliability and uptime as promised to the users and at the same time enables developers to make changes to the existing system through careful preparation and planning while keeping an eye on capacity, latency, and performance. SRE is also a mindset and a set of engineering approaches to running better production systems and optimizations. Much of our software development focuses on eliminating manual work through automation, performance tuning, and growing the efficiency of production systems. As SREs are responsible for the big picture of how our systems relate to each other, we use a breadth of tools and approaches to tackle a broad spectrum of problems. Practices such as limiting time spent on reactive operational work, blameless postmortems, and proactive identification of potential outages factor into iterative improvement that is key to product quality and interesting and dynamic day-to-day work. SRE's culture of diversity, intellectual curiosity, problem-solving, and openness is important to its success. Our organization brings together people with a wide variety of backgrounds, experiences, and perspectives. We encourage them to collaborate, think big, and take risks in a blame-free environment. We promote self-direction to work on meaningful projects while striving to build an environment that provides the support and mentorship needed to learn and grow.

What You Will Be Doing:

Assist in the design, implementation, and support of large-scale storage clusters, including monitoring, logging, and alerting.
Work with AI/ML workloads to capture and correlate behavior in large clusters and workflows, which are otherwise hard to understand.
Work closely with peers on the team to improve the lifecycle of services – from inception and design, through deployment, operation, and refinement.
Support services before they go live through activities such as system design consulting, developing software and frameworks, capacity management, and launch reviews.
Maintain services once they are live by measuring and monitoring availability, latency, and overall system health, including leveraging machine learning models.
Scale systems sustainably through mechanisms like AI/ML and automation, and evolve systems by pushing for changes that improve reliability and velocity.
Practice sustainable incident response and blameless postmortems.
Be part of an on-call rotation to support production systems.

What We Need To See:

BS degree in Computer Science or related technical field involving coding (e.g., physics or mathematics) or equivalent experience.
At least 5+ years practical experience.
Experience with algorithms, data structures, complexity analysis, software design, and maintaining large-scale Linux-based systems.
Experience in one or more of the following: C/C++, Java, Python, Go, Perl or Ruby, AI/ML frameworks and methodologies.
Good knowledge of infrastructure configuration management tools like Ansible, Chef, Puppet, and Terraform.
Experience in using observability and tracing-related tools like InfluxDB, Prometheus, and Elastic stack.

Ways to stand out from the crowd:

Demonstrated experience in having SRE mindset, customer-first approach, and focus on customer satisfaction and passion for ensuring customer success.Experience with Git, code review, pipelines, and CI/CD.
Interest in crafting, analyzing, and fixing large-scale distributed systems. Strong debugging skills with a systematic problem-solving approach to identify complex problems.
Thrive in collaborative environments and enjoy working with various teams. Experience in using or running large private and public cloud systems based on Kubernetes, OpenStack, and Docker. Flexible in adapting to different working styles.

NVIDIA is widely considered to be one of the technology world’s most desirable employers. We have some of the most forward-thinking and talented people on the planet working for us. If you're creative and autonomous, we want to hear from you!

The base salary range is 148,000 USD - 356,500 USD. Your base salary will be determined based on your location, experience, and the pay of employees in similar positions.

You will also be eligible for equity and benefits. NVIDIA accepts applications on an ongoing basis.

NVIDIA is committed to fostering a diverse work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.

Similar Jobs

Senior Cloud Engineer

Teradata

Pune, Maharashtra, India (On-Site)

• 9 Months ago

Senior DATA/AI SRE Engineer

Playtika

Poland (On-Site)

• 9 Months ago

Principal Technical Product Manager - Application Security

Crunchyroll

San Francisco, California, United States (On-Site)

• 5 Months ago

Senior Cloud Operations Engineer

Revolgy

United Kingdom (Remote)

• 5 Months ago

Experienced Software Engineer - Traffic Platform

ByteDance

San Jose, California, United States (On-Site)

• 9 Months ago

Lead Software Engineer

Virtuos

Singapore (On-Site)

• 5 Months ago

Senior DevOps (Azure) Engineer

Velotio Technologies

Maharashtra, India (Remote)

• 5 Months ago

IT Support Engineer II

PlayStation Global

United Kingdom (Remote)

• 5 Months ago

Cloud Engineering Co-op

Activision

Vancouver, British Columbia, Canada (Hybrid)

• 7 Months ago

Manager, Software Engineering

The Walt Disney Company

Washington, United States (On-Site)

• 6 Months ago

Get notifed when new similar jobs are uploaded

Similar Skill Jobs

Golang Backend Developer

Xsolla

Montreal, Quebec, Canada (On-Site)

• 10 Months ago

Staff Backend Engineer - Infrastructure

Pocket Worlds

Poland (On-Site)

• 5 Months ago

Senior Backend Engineer (Golang)

Voodoo

Paris, Île-de-France, France (Hybrid)

• 5 Months ago

DevOps Engineer

Aristocrat Gaming

Kraków, Lesser Poland Voivodeship, Poland (Hybrid)

• 8 Months ago

Game Developer

Wargaming

Vilnius, Vilnius County, Lithuania (On-Site)

• 5 Months ago

Site Reliability Engineer

ION

Milan, Lombardy, Italy (Hybrid)

• 10 Months ago

Hardware Integration Engineer & Backend Game Developer (.NET)

Blazesoft

Vaughan, Ontario, Canada (On-Site)

• 5 Months ago

Middle/Senior DevOps Engineer

Kefir Games

Cyprus (On-Site)

• 8 Months ago

Principal Full Stack Engineer (Crash Reporting System)

PlayStation Global

London, England, United Kingdom (On-Site)

• 5 Months ago

Scala Engineer

Evolution

Amsterdam, North Holland, Netherlands (On-Site)

• 1 Year ago

Get notifed when new similar jobs are uploaded

Jobs in Santa Clara, California, United States

Software Engineering Manager, Product Infrastructure

Senior Software Engineer, Design System

prizepicks

Atlanta, Georgia, United States (Remote)

• 5 Months ago

Senior Research Scientist, Data Management and Security - Infrastructure System Lab

ByteDance

San Jose, California, United States (On-Site)

• 5 Months ago

Revenue Accounting Manager, Enterprise Sales

Canva

Los Angeles, California, United States (Remote)

• 5 Months ago

Senior Full Stack Engineer (C#/React)

Rockstar Games

New York, New York, United States (On-Site)

• 11 Months ago

Lead Data Engineer, Data Reliability

The Walt Disney Company

Santa Monica, California, United States (On-Site)

• 7 Months ago

Creative Copywriter

Match Group

New York, New York, United States (Hybrid)

• 10 Months ago

Senior Finance Manager (Finance Business Partner)

Tencent

California, United States (On-Site)

• 5 Months ago

Sales Development Representative

Onward Search

Cincinnati, Ohio, United States (On-Site)

• 7 Months ago

Director, Design

Illumination

Santa Monica, California, United States (On-Site)

• 10 Months ago

Get notifed when new similar jobs are uploaded

Devops Jobs

Staff Engineer (C++ Linux)

Omnissa

Bengaluru, Karnataka, India (Hybrid)

• 10 Months ago

Sr. Software Engineer

Egnyte

Mountain View, California, United States (Hybrid)

• 9 Months ago

Software Engineer, SRE - Platform Services

ByteDance

Seattle, Washington, United States (On-Site)

• 5 Months ago

Cloud Solutions Technical Account Manager - BytePlus - Hong Kong

ByteDance

(On-Site)

• 10 Months ago

Cloud Engineer (Azure)

Zazz

(Remote)

• 6 Months ago

Senior Cloud SRE

NICE

Pune, Maharashtra, India (Hybrid)

• 10 Months ago

Software Engineering Manager II, Google Cloud

Google

Hyderabad, Telangana, India (On-Site)

• 9 Months ago

Senior Data Engineer - (Big Data, Spark, Scala, Java, Python, Cassandra, Elasticserach, AWS, Airflow, RDBMS, SQL)

Nielsen Holdings

Bengaluru, Karnataka, India (Hybrid)

• 10 Months ago

Site Reliability Engineer

Playtika

Netherlands (Hybrid)

• 5 Months ago

Site Reliability Engineer

ION

Pisa, Tuscany, Italy (Hybrid)

• 10 Months ago

Get notifed when new similar jobs are uploaded

About The Company

NVIDIA

77 Active Jobs

Since its founding in 1993, NVIDIA (NASDAQ: NVDA) has been a pioneer in accelerated computing. The company’s invention of the GPU in 1999 sparked the growth of the PC gaming market, redefined computer graphics, ignited the era of modern AI and is fueling the creation of the metaverse. NVIDIA is now a full-stack computing company with data-center-scale offerings that are reshaping industry.

Get notified when new jobs are added by NVIDIA

Level Up Your Career in Game Development!

Transform Your Passion into Profession with Our Comprehensive Courses for Aspiring Game Developers.

A global community of game builders. Helping people upskill and land jobs in the best gaming studios.

Company

Key Links

hello@outscal.com

Made in INDIA 💛💙

Senior Production Engineer - Storage

Job Summary

Job Description

15 skills required

15 skills required for this role

Job Details

Similar Jobs

Senior Cloud Engineer

Senior DATA/AI SRE Engineer

Principal Technical Product Manager - Application Security

Senior Cloud Operations Engineer

Experienced Software Engineer - Traffic Platform

Lead Software Engineer

Senior DevOps (Azure) Engineer

IT Support Engineer II

Cloud Engineering Co-op

Manager, Software Engineering

Similar Skill Jobs

Golang Backend Developer

Staff Backend Engineer - Infrastructure

Senior Backend Engineer (Golang)

DevOps Engineer

Game Developer

Site Reliability Engineer

Hardware Integration Engineer & Backend Game Developer (.NET)

Middle/Senior DevOps Engineer

Principal Full Stack Engineer (Crash Reporting System)

Scala Engineer

Jobs in Santa Clara, California, United States

Software Engineering Manager, Product Infrastructure

Senior Software Engineer, Design System

Senior Research Scientist, Data Management and Security - Infrastructure System Lab

Revenue Accounting Manager, Enterprise Sales

Senior Full Stack Engineer (C#/React)

Lead Data Engineer, Data Reliability

Creative Copywriter

Senior Finance Manager (Finance Business Partner)

Sales Development Representative

Director, Design

Devops Jobs

Staff Engineer (C++ Linux)

Sr. Software Engineer

Software Engineer, SRE - Platform Services

Cloud Solutions Technical Account Manager - BytePlus - Hong Kong

Cloud Engineer (Azure)

Senior Cloud SRE

Software Engineering Manager II, Google Cloud

Senior Data Engineer - (Big Data, Spark, Scala, Java, Python, Cassandra, Elasticserach, AWS, Airflow, RDBMS, SQL)

Site Reliability Engineer

Site Reliability Engineer

About The Company

System Design Power Validation Engineer

OEM Account Manager

System Debug Lead Engineer

Network Site Reliability Engineer

ASIC Engineer

Senior ASIC Design Engineer

Physical Design CAD Team Manager

Senior Solutions Architect, Infiniband and Networking Ethernet

Engineering Farm Engineer

Senior Mixed Signal Design Verification Engineer

Level Up Your Career in Game Development!