Senior Site Reliability Engineer, Data Science and ML Platforms

1 Month ago • 5-8 Years • DevOps

Job Summary

Job Description

NVIDIA seeks a Senior Site Reliability Engineer (SRE) for its Data Science & ML Platforms team. Responsibilities include designing, building, and maintaining services for real-time data analytics, streaming, data lakes, observability, and ML/AI training and inferencing. This role requires implementing software and systems engineering practices to ensure high availability, applying SRE principles to improve production systems, and optimizing service SLOs. Collaboration with customers to plan and implement system changes, while monitoring capacity, latency, and performance, is crucial. The successful candidate will develop software solutions for large-scale systems, gain deep understanding of system operations, create automation tools, establish operational frameworks, define reliability metrics, and oversee capacity and performance management. Incident response and blameless postmortems are also key responsibilities.
Must have:
  • 5-8 years SRE/DevOps experience
  • Strong SRE principles understanding
  • Proficiency in incident/change management
  • Experience with streaming data infrastructure
  • Expertise in large-scale observability platforms
  • Proficiency in Python/Go/Perl/Ruby
  • Experience scaling distributed systems
Good to have:
  • Experience operating large-scale systems with strong SLAs
  • Excellent Python/Go coding skills and data platform experience
  • Knowledge of CI/CD systems (Jenkins, GitHub Actions)
  • Familiarity with IaC methodologies and tools
  • Excellent interpersonal skills for communicating data-driven insights

Job Details

Are you passionate about building and maintaining large-scale production systems that support advanced data science and machine learning applications? Do you want to join a team at the heart of NVIDIA's data-driven decision-making culture? If so, we have a great opportunity for you! NVIDIA is seeking a Senior Site Reliability Engineer (SRE) for the Data Science & ML Platform(s) team. The role involves designing, building, and maintaining services that enable real-time data analytics, streaming, data lakes, observability and ML/AI training and inferencing. The responsibilities include implementing software and systems engineering practices to ensure high efficiency and availability of the platform, as well as applying SRE principles to improve production systems and optimize service SLOs. Additionally, collaboration with our customers to plan implement changes to the existing system, while monitoring capacity, latency, and performance is part of the role.

To succeed in this position, a strong background in SRE practices, systems, networking, coding, capacity management, cloud operations, continuous delivery and deployment, and open-source cloud enabling technologies like Kubernetes and OpenStack is required. Deep understanding of the challenges and standard methodologies of running large-scale distributed systems in production, solving complex issues, automating repetitive tasks, and proactively identifying potential outages is also necessary. Furthermore, excellent communication and collaboration skills, and a culture of diversity, intellectual curiosity, problem solving, and openness are essential. As a Senior SRE at NVIDIA, you will have the opportunity to work on innovative technologies that power the future of AI and data science, and be part of a dynamic and supportive team that values learning and growth. The role provides the autonomy to work on meaningful projects with the support and mentorship needed to succeed, and contributes to a culture of blameless postmortems, iterative improvement, and risk-taking. If you are seeking an exciting and rewarding career that makes a difference, we invite you to apply now!

What you’ll be doing:

  • Develop software solutions to ensure reliability and operability of large-scale systems supporting machine-critical use cases.

  • Gain a deep understanding of our system operations, scalability, interactions, and failures to identify improvement opportunities and risks.

  • Create tools and automation to reduce operational overhead and eliminate manual tasks.

  • Establish frameworks, processes, and standard methodologies to enhance operational maturity, team efficiency, and accelerate innovation.

  • Define meaningful and actionable reliability metrics to track and improve system and service reliability.

  • Oversee capacity and performance management to facilitate infrastructure scaling across public and private clouds globally.

  • Build tools to improve our service observability for faster issue resolution.

  • Practice sustainable incident response and blameless postmortems

What we need to see:

  • Minimum of 5-8 years of experience in SRE, Cloud platforms, or DevOps with large-scale microservices in production environments.

  • Master's or Bachelor's degree in Computer Science or Electrical Engineering or CE or equivalent experience.

  • Strong understanding of SRE principles, including error budgets, SLOs, and SLAs.

  • Proficiency in incident, change, and problem management processes.

  • Skilled in problem-solving, root cause analysis, and optimization.

  • Experience with streaming data infrastructure services, such as Kafka and Spark.

  • Expertise in building and operating large-scale observability platforms for monitoring and logging (e.g., ELK, Prometheus).

  • Proficiency in programming languages such as Python, Go, Perl, or Ruby.

  • Hands-on experience with scaling distributed systems in public, private, or hybrid cloud environments.

  • Experience in deploying, supporting, and supervising services, platforms, and application stacks.

Ways to stand out from the crowd:

  • Experience operating large-scale distributed systems with strong SLAs.

  • Excellent coding skills in Python and Go and extensive experience in operating data platforms.

  • Knowledge of CI/CD systems, such as Jenkins and GitHub Actions.

  • Familiarity with Infrastructure as Code (IaC) methodologies and tools.

  • Excellent interpersonal skills for identifying and communicating data-driven insights.

NVIDIA leads the way in groundbreaking developments in Artificial Intelligence, High-Performance Computing, and Visualization. The GPU, our invention, serves as the visual cortex of modern computers and is at the heart of our products and services. Our work opens up new universes to explore, enables amazing creativity and discovery, and powers what were once science fiction inventions, from artificial intelligence to autonomous cars. NVIDIA is looking for exceptional people like you to help us accelerate the next wave of artificial intelligence.

Similar Jobs

Homa games - Senior Machine Learning Engineer

Homa games

Paris, ĂŽle-de-France, France (On-Site)
• 7 Months ago
Canva - Data Lead - Print

Canva

Auckland, Auckland, New Zealand (Hybrid)
• 2 Months ago
ARHS - Data Manager

ARHS

Stockholm, Stockholm County, Sweden (On-Site)
• 4 Months ago
Homa games - Game Data Analyst - All in Hole

Homa games

Paris, ĂŽle-de-France, France (Remote)
• 3 Months ago
Infraveo Technologies - Jr. Data Scientist

Infraveo Technologies

Gurugram, Haryana, India (Remote)
• 3 Months ago
SmileGate - Hybrid Cloud Strategy/Development Lead (CTO Division)

SmileGate

Seongnam-si, Gyeonggi-do, South Korea (On-Site)
• 1 Month ago
Hitachi - Azure Developer

Hitachi

Hyderabad, Telangana, India (Remote)
• 4 Months ago
Sony Interactive Entertainment - Server-Side Engineer (PlayStation™Network Server Application Development)

Sony Interactive Entertainment

Tokyo, Japan (On-Site)
• 1 Month ago
NVIDIA - Solutions Architect, Infrastructure - Research Computing

NVIDIA

New York, New York, United States (Remote)
• 1 Month ago
Luxoft - Infrastructure Engineer with AWS

Luxoft

Poland, Ohio, United States (Remote)
• 3 Months ago

Get notifed when new similar jobs are uploaded

Similar Skill Jobs

Peak - Summer Intern, Data Scientist

Peak

(On-Site)
• 1 Month ago
Wind River Systems - Embedded Software Engineer – RTOS / VxWorks

Wind River Systems

Galați, Județul Galați, Romania (On-Site)
• 3 Months ago
N-iX - Data Engineer

N-iX

Ukraine (Remote)
• 1 Day ago
Dream Sports - Lead ML Scientist

Dream Sports

Mumbai, Maharashtra, India (On-Site)
• 6 Months ago
Netflix - Senior Product Designer - Enterprise Technology & Agent Tools

Netflix

United States (Remote)
• 1 Week ago
Netflix - Engineering Manager, Content Acquisition

Netflix

Los Gatos, California, United States (On-Site)
• 1 Month ago
Visa - Principle Data Scientist, Visa Rules, Compliance and Standards (RCS)

Visa

Austin, Texas, United States (Hybrid)
• 2 Months ago
Microsoft - Principal Product Manager

Microsoft

Norway (On-Site)
• 2 Weeks ago
ARHS - Microsoft Applications Consultant

ARHS

Paris, ĂŽle-de-France, France (On-Site)
• 4 Months ago
King - Data Science Intern

King

Barcelona, Catalonia, Spain (On-Site)
• 2 Weeks ago

Get notifed when new similar jobs are uploaded

Jobs in Shanghai, Shanghai, China

Ourpalm - Lead UI Designer

Ourpalm

Beijing, Beijing, China (On-Site)
• 3 Weeks ago
NVIDIA - Safety Engineer

NVIDIA

Shenzhen, Guangdong Province, China (On-Site)
• 1 Month ago
NVIDIA - Electronics Failure Analysis Hardware Engineer

NVIDIA

Shenzhen, Guangdong Province, China (On-Site)
• 1 Month ago
Tencent - Senior Strategic Investment Manager

Tencent

Shenzhen, Guangdong Province, China (On-Site)
• 2 Months ago
Spin Master - Senior Design Engineer

Spin Master

Guangdong Province, China (On-Site)
• 1 Month ago
Supercell - Senior 3D Environment Artist, Project R.I.S.E

Supercell

Shanghai, Shanghai, China (On-Site)
• 4 Months ago
Hasbro - Global Security Auditor

Hasbro

Shenzhen, Guangdong Province, China (On-Site)
• 2 Months ago
Canva - Enterprise Account Executive

Canva

Beijing, Beijing, China (Remote)
• 1 Month ago
Zengame Technology - Advertising Optimization Specialist

Zengame Technology

Shenzhen, Guangdong Province, China (On-Site)
• 1 Month ago
Tencent - Security Operations - PUBG Mobile

Tencent

Shenzhen, Guangdong Province, China (On-Site)
• 2 Days ago

Get notifed when new similar jobs are uploaded

DevOps Jobs

CLO Virtual Fashion  Inc  - DevOps Engineer

CLO Virtual Fashion Inc

Bengaluru, Karnataka, India (On-Site)
• 4 Months ago
PwC - IN_Associate_Azure Cloud Data Engineer_OneCloud _Advisory _Bangalore

PwC

Gurugram, Haryana, India (On-Site)
• 2 Months ago
Microsoft - Sr. FastTrack Solution Architect

Microsoft

Hyderabad, Telangana, India (On-Site)
• 1 Month ago
NVIDIA - GenAI and MLOps Intern - Spring 2025

NVIDIA

Taipei City, Taiwan (On-Site)
• 2 Weeks ago
Scorewarrior - CI/CD Engineer

Scorewarrior

Limassol, Limassol, Cyprus (On-Site)
• 1 Month ago
Google - Senior Software Engineering Manager, Google Cloud

Google

Hyderabad, Telangana, India (On-Site)
• 1 Month ago
Blazesoft - DevOps engineer

Blazesoft

Vaughan, Ontario, Canada (On-Site)
• 2 Months ago
Nielsen Holdings - SENIOR DEVOPS ENGINEER

Nielsen Holdings

Gurugram, Haryana, India (Hybrid)
• 4 Months ago
Microsoft - ROP-Senior Software Engineer

Microsoft

Hyderabad, Telangana, India (On-Site)
• 1 Month ago
Revolgy - Junior Cloud Ops Engineer (Intern)

Revolgy

(Remote)
• 1 Month ago

Get notifed when new similar jobs are uploaded

About The Company

Since its founding in 1993, NVIDIA (NASDAQ: NVDA) has been a pioneer in accelerated computing. The company’s invention of the GPU in 1999 sparked the growth of the PC gaming market, redefined computer graphics, ignited the era of modern AI and is fueling the creation of the metaverse. NVIDIA is now a full-stack computing company with data-center-scale offerings that are reshaping industry.


Yokne'am Illit, North District, Israel (On-Site)

Santa Clara, California, United States (Hybrid)

Santa Clara, California, United States (Hybrid)

Santa Clara, California, United States (On-Site)

United States (Remote)

Santa Clara, California, United States (On-Site)

Santa Clara, California, United States (On-Site)

Bengaluru, Karnataka, India (Hybrid)

Bengaluru, Karnataka, India (Hybrid)

View All Jobs

Get notified when new jobs are added by NVIDIA

Level Up Your Career in Game Development!

Transform Your Passion into Profession with Our Comprehensive Courses for Aspiring Game Developers.

Job Common Plug