Jobs Courses Resources Companies Placements

Home >

Jobs >

Senior Site Reliability Engineer, Data Science and ML Platforms

NVIDIA

Shanghai, China (On-site)

Senior Site Reliability Engineer, Data Science and ML Platforms

7 Months ago • 5-8 Years • Devops

Job Summary

Job Description

NVIDIA seeks a Senior Site Reliability Engineer (SRE) for its Data Science & ML Platforms team. Responsibilities include designing, building, and maintaining services for real-time data analytics, streaming, data lakes, observability, and ML/AI training and inferencing. This role requires implementing software and systems engineering practices to ensure high availability, applying SRE principles to improve production systems, and optimizing service SLOs. Collaboration with customers to plan and implement system changes, while monitoring capacity, latency, and performance, is crucial. The successful candidate will develop software solutions for large-scale systems, gain deep understanding of system operations, create automation tools, establish operational frameworks, define reliability metrics, and oversee capacity and performance management. Incident response and blameless postmortems are also key responsibilities.

Must have:

5-8 years SRE/DevOps experience
Strong SRE principles understanding
Proficiency in incident/change management
Experience with streaming data infrastructure
Expertise in large-scale observability platforms
Proficiency in Python/Go/Perl/Ruby
Experience scaling distributed systems

Good to have:

Experience operating large-scale systems with strong SLAs
Excellent Python/Go coding skills and data platform experience
Knowledge of CI/CD systems (Jenkins, GitHub Actions)
Familiarity with IaC methodologies and tools
Excellent interpersonal skills for communicating data-driven insights

15 skills required

15 skills required for this role

Add these skills to join the top 1% applicants for this job

data-science

python

communication

ruby

jenkins

github-actions

ci-cd

github

kubernetes

microservices

spark

elk

perl

incident-response

openstack

Job Details

Are you passionate about building and maintaining large-scale production systems that support advanced data science and machine learning applications? Do you want to join a team at the heart of NVIDIA's data-driven decision-making culture? If so, we have a great opportunity for you! NVIDIA is seeking a Senior Site Reliability Engineer (SRE) for the Data Science & ML Platform(s) team. The role involves designing, building, and maintaining services that enable real-time data analytics, streaming, data lakes, observability and ML/AI training and inferencing. The responsibilities include implementing software and systems engineering practices to ensure high efficiency and availability of the platform, as well as applying SRE principles to improve production systems and optimize service SLOs. Additionally, collaboration with our customers to plan implement changes to the existing system, while monitoring capacity, latency, and performance is part of the role.

To succeed in this position, a strong background in SRE practices, systems, networking, coding, capacity management, cloud operations, continuous delivery and deployment, and open-source cloud enabling technologies like Kubernetes and OpenStack is required. Deep understanding of the challenges and standard methodologies of running large-scale distributed systems in production, solving complex issues, automating repetitive tasks, and proactively identifying potential outages is also necessary. Furthermore, excellent communication and collaboration skills, and a culture of diversity, intellectual curiosity, problem solving, and openness are essential. As a Senior SRE at NVIDIA, you will have the opportunity to work on innovative technologies that power the future of AI and data science, and be part of a dynamic and supportive team that values learning and growth. The role provides the autonomy to work on meaningful projects with the support and mentorship needed to succeed, and contributes to a culture of blameless postmortems, iterative improvement, and risk-taking. If you are seeking an exciting and rewarding career that makes a difference, we invite you to apply now!

What you’ll be doing:

Develop software solutions to ensure reliability and operability of large-scale systems supporting machine-critical use cases.
Gain a deep understanding of our system operations, scalability, interactions, and failures to identify improvement opportunities and risks.
Create tools and automation to reduce operational overhead and eliminate manual tasks.
Establish frameworks, processes, and standard methodologies to enhance operational maturity, team efficiency, and accelerate innovation.
Define meaningful and actionable reliability metrics to track and improve system and service reliability.
Oversee capacity and performance management to facilitate infrastructure scaling across public and private clouds globally.
Build tools to improve our service observability for faster issue resolution.
Practice sustainable incident response and blameless postmortems

What we need to see:

Minimum of 5-8 years of experience in SRE, Cloud platforms, or DevOps with large-scale microservices in production environments.
Master's or Bachelor's degree in Computer Science or Electrical Engineering or CE or equivalent experience.
Strong understanding of SRE principles, including error budgets, SLOs, and SLAs.
Proficiency in incident, change, and problem management processes.
Skilled in problem-solving, root cause analysis, and optimization.
Experience with streaming data infrastructure services, such as Kafka and Spark.
Expertise in building and operating large-scale observability platforms for monitoring and logging (e.g., ELK, Prometheus).
Proficiency in programming languages such as Python, Go, Perl, or Ruby.
Hands-on experience with scaling distributed systems in public, private, or hybrid cloud environments.
Experience in deploying, supporting, and supervising services, platforms, and application stacks.

Ways to stand out from the crowd:

Experience operating large-scale distributed systems with strong SLAs.
Excellent coding skills in Python and Go and extensive experience in operating data platforms.
Knowledge of CI/CD systems, such as Jenkins and GitHub Actions.
Familiarity with Infrastructure as Code (IaC) methodologies and tools.
Excellent interpersonal skills for identifying and communicating data-driven insights.

NVIDIA leads the way in groundbreaking developments in Artificial Intelligence, High-Performance Computing, and Visualization. The GPU, our invention, serves as the visual cortex of modern computers and is at the heart of our products and services. Our work opens up new universes to explore, enables amazing creativity and discovery, and powers what were once science fiction inventions, from artificial intelligence to autonomous cars. NVIDIA is looking for exceptional people like you to help us accelerate the next wave of artificial intelligence.

Similar Jobs

Software Engineer III

Electronic Arts

Hyderabad, Telangana, India (On-Site)

• 5 Months ago

Principal Software Engineer

The Walt Disney Company

Seattle, Washington, United States (On-Site)

• 4 Months ago

Software Engineer - AI/Machine Translation

Snail Games

Beverly Hills, California, United States (Remote)

• 6 Months ago

IT Support Engineer

MIQ Digital

London, England, United Kingdom (On-Site)

• 4 Months ago

Principal Product Manager - Growth Product

Mistplay

Toronto, Ontario, Canada (Hybrid)

• 4 Months ago

Backend Software Engineer - Foundational Technology

ByteDance

Singapore (On-Site)

• 4 Months ago

SOFTWARE ENGINEER (CLOUD)

Britive

Bengaluru, Karnataka, India (Remote)

• 9 Months ago

Senior DevOps Engineer (AWS)

Velotio Technologies

Maharashtra, India (Remote)

• 5 Months ago

Release DevOps Engineer

Scanline VFX

Montreal, Quebec, Canada (Hybrid)

• 5 Months ago

Engineering Manager - Edge Gateway & Services

Netflix

United States (Remote)

• 4 Months ago

Get notifed when new similar jobs are uploaded

Similar Skill Jobs

Cybersecurity Engineer

Barracuda Networks Inc

Chelmsford, Massachusetts, United States (Hybrid)

• 5 Months ago

Elm Team Lead

Well

(Remote)

• 4 Months ago

Senior Product Manager, Fully Automated Messaging

Attentive

San Francisco, California, United States (Hybrid)

• 4 Months ago

Principal Engineer, Data Science

Nagarro

India (Remote)

• 10 Months ago

Senior Data Engineer

Applike Group

Hamburg, Hamburg, Germany (Hybrid)

• 4 Months ago

Associate Distinguished Engineer - Enterprise Data Architect

Nagarro

Allentown, Pennsylvania, United States (Remote)

• 10 Months ago

Senior Analytics Engineer

Mythical Games

United States (Remote)

• 4 Months ago

Senior Data Analyst

GameJobs

Bengaluru, Karnataka, India (On-Site)

• 1 Year ago

Data Scientist, Ads

(Remote)

• 4 Months ago

Data Science II (Marketing Analytics)

Match Group

West Hollywood, California, United States (Hybrid)

• 10 Months ago

Get notifed when new similar jobs are uploaded

Jobs in Shanghai, Shanghai, China

Industrial Engineer

Mattel Inc

Dongguan, Guangdong Province, China (On-Site)

• 9 Months ago

Strategic Investment Manager - AI+Game Tech

Tencent

Shenzhen, Guangdong Province, China (On-Site)

• 4 Months ago

Senior DevOps Engineer (LiveOps)

Thatgamecompany

Shanghai, Shanghai, China (On-Site)

• 5 Months ago

Lead Technical Artist

Virtuos

China (On-Site)

• 5 Months ago

41299-服务器性能测试工程师(北京)

Tencent

Beijing, Beijing, China (On-Site)

• 1 Year ago

GPU Kernel Software Engineering Intern - 2025

NVIDIA

Shanghai, Shanghai, China (On-Site)

• 7 Months ago

Card Game Numerical Planner

Ourpalm

Beijing, Beijing, China (On-Site)

• 7 Months ago

Software Engineer, Fullstack

Grab

Beijing, China (On-Site)

• 4 Months ago

Storyboard Artist - Infinite Warmth

Paper Games

Shanghai, Shanghai, China (On-Site)

• 4 Months ago

Lead 3D Character Artist - AAA Stylized Realistic Shooting Game

Light Speed Studios

Shenzhen, Guangdong Province, China (On-Site)

• 4 Months ago

Get notifed when new similar jobs are uploaded

Devops Jobs

C++ Windows Internals Dev_MTS2/3 (2-7 Yrs)_Horizon Team

Omnissa

Bengaluru, Karnataka, India (Hybrid)

• 11 Months ago

Terraform with DevOps

Hitachi

Pune, Maharashtra, India (On-Site)

• 10 Months ago

Senior DevOps Engineer

UXBERT Labs

Riyadh, Riyadh Province, Saudi Arabia (Hybrid)

• 7 Months ago

Technical Support Engineer - Azure Monitoring

Microsoft

Taipei City, Taiwan (Hybrid)

• 4 Months ago

Senior DevOps Programmer

Epic Games

United States (On-Site)

• 6 Months ago

Technical Support Engineer

Microsoft

Bengaluru, Karnataka, India (Hybrid)

• 4 Months ago

Sr. Site Reliability Engineer, Product Reliability Engineering - Middleware

Visa

Austin, Texas, United States (Hybrid)

• 8 Months ago

Senior DevOps Engineer

Milestone

Copenhagen, Denmark (Hybrid)

• 4 Months ago

Kubernetes Experts

Ajmera Infotech

Bengaluru, Karnataka, India (On-Site)

• 1 Year ago

Integration Engineer

Playtech

Tartu, Tartu County, Estonia (On-Site)

• 5 Months ago

Get notifed when new similar jobs are uploaded

About The Company

NVIDIA

76 Active Jobs

Since its founding in 1993, NVIDIA (NASDAQ: NVDA) has been a pioneer in accelerated computing. The company’s invention of the GPU in 1999 sparked the growth of the PC gaming market, redefined computer graphics, ignited the era of modern AI and is fueling the creation of the metaverse. NVIDIA is now a full-stack computing company with data-center-scale offerings that are reshaping industry.

Get notified when new jobs are added by NVIDIA

Level Up Your Career in Game Development!

Transform Your Passion into Profession with Our Comprehensive Courses for Aspiring Game Developers.

A global community of game builders. Helping people upskill and land jobs in the best gaming studios.

Company

Key Links

hello@outscal.com

Made in INDIA 💛💙

Senior Site Reliability Engineer, Data Science and ML Platforms

Job Summary

Job Description

15 skills required

15 skills required for this role

Job Details

Similar Jobs

Software Engineer III

Principal Software Engineer

Software Engineer - AI/Machine Translation

IT Support Engineer

Principal Product Manager - Growth Product

Backend Software Engineer - Foundational Technology

SOFTWARE ENGINEER (CLOUD)

Senior DevOps Engineer (AWS)

Release DevOps Engineer

Engineering Manager - Edge Gateway & Services

Similar Skill Jobs

Cybersecurity Engineer

Elm Team Lead

Senior Product Manager, Fully Automated Messaging

Principal Engineer, Data Science

Senior Data Engineer

Associate Distinguished Engineer - Enterprise Data Architect

Senior Analytics Engineer

Senior Data Analyst

Data Scientist, Ads

Data Science II (Marketing Analytics)

Jobs in Shanghai, Shanghai, China

Industrial Engineer

Strategic Investment Manager - AI+Game Tech

Senior DevOps Engineer (LiveOps)

Lead Technical Artist

41299-服务器性能测试工程师(北京)

GPU Kernel Software Engineering Intern - 2025

Card Game Numerical Planner

Software Engineer, Fullstack

Storyboard Artist - Infinite Warmth

Lead 3D Character Artist - AAA Stylized Realistic Shooting Game

Devops Jobs

C++ Windows Internals Dev_MTS2/3 (2-7 Yrs)_Horizon Team

Terraform with DevOps

Senior DevOps Engineer

Technical Support Engineer - Azure Monitoring

Senior DevOps Programmer

Technical Support Engineer

Sr. Site Reliability Engineer, Product Reliability Engineering - Middleware

Senior DevOps Engineer

Kubernetes Experts

Integration Engineer

About The Company

System Design Power Validation Engineer

OEM Account Manager

System Debug Lead Engineer

Network Site Reliability Engineer

ASIC Engineer

Senior ASIC Design Engineer

Physical Design CAD Team Manager

Engineering Farm Engineer

Senior Mixed Signal Design Verification Engineer

Senior Solutions Architect, Cloud Infrastructure and DevOps

Level Up Your Career in Game Development!