Senior Observability Architect, AI and HPC

1 Week ago • 12 Years + • Research & Development • $224,000 PA - $425,500 PA

Job Summary

Job Description

NVIDIA's Hardware Infrastructure team seeks a Senior Observability Architect to define and implement observability systems for large-scale AI and HPC clusters. Responsibilities include architecting data collection, aggregation, and visualization systems; collaborating with engineering and research teams; defining data policies; and improving system efficiency and performance. The role requires experience with distributed observability systems, data analysis, and tools like Apache Spark, Elastic/OpenSearch, Grafana, and Prometheus. The ideal candidate will have strong technical leadership skills and a passion for improving productivity.
Must have:
  • Design & build large-scale observability systems
  • Collaborate with data scientists & engineers
  • Turn raw data into actionable reports
  • Experience with observability platforms
  • Python programming & API calls
  • Improve team productivity
Good to have:
  • Background in machine learning, deep learning
  • Experience in infrastructure software, DevOps
  • Datacenter & distributed computing experience
  • Working with AI researchers or EDA developers
  • Process improvement and efficiency measurement
Perks:
  • Equity
  • Benefits

Job Details

 NVIDIA’s Hardware Infrastructure organization is seeking a Senior or Principal Data and Observability Architect. We serve and collaborate directly with NVIDIA’s rapidly growing AI, HW, and SW engineering and research teams across the company. We are looking for a technical leader to define a vision and roadmap for distributed observability systems for large-scale AI and HPC clusters and workloads and guide implementation towards this vision. You will architect systems for data collection, aggregation, enrichment, storage, retrieval, and visualization to spectacularly improve efficiency, performance, and productivity of AI and HPC workloads. You will lead technical teams to develop, deploy, and operate observability solutions for multiple compute clusters around the world. 

What You’ll Be Doing: 

  • Collaborate with AI, HW, and SW engineering and research teams to define a vision and roadmap for AI/HPC cluster observability. 

  • Architect and lead teams to develop, test, and deploy data collectors, pipelines, visualization and retrieval services. 

  • Define data collection and retention polices to balance network bandwidth, system load, and storage capacity costs with data analysis requirements. 

  • Work in a diverse team to provide operational and strategic data to empower our engineers and researchers to improve performance, productivity, and efficiency. 

  • Continuously improve quality, workloads, and processes through better observability. 

What We Need to See: 

  • Experience designing and building large scale, distributed observability systems. 

  • Ability to collaborate with data scientists, researchers, and engineering teams to identify high value data for collection and analysis. 

  • Experience with turning raw data into actionable reports 

  • Experience with observability platforms such as Apache Spark, Elastic/Open Search, Grafana, Prometheus, and other similar open-source tools 

  • Technical lead level Python programming experience and use of API calls 

  • Passion for improving the productivity of others 

  • Excellent planning and interpersonal skills 

  • Flexibility/adaptability working in a dynamic environment with changing requirements  

  • MS (preferred) or BS in Computer Science, Electrical Engineering, or related field or equivalent experience

  • 12+ years of relevant experience. 

Ways To Stand Out from The Crowd: 

  • Background in computer science, machine learning, deep learning, open-source software, infrastructure technologies, and GPU technology. 

  • Prior experience in infrastructure software, production application software development, software development, release and support methodology and devops 

  • Experience in the management of datacenters and large-scale distributed computing 

  • Experience in working with AI researchers and/or EDA developers 

  • Consistent track record of driving process improvements and measuring efficiency and a passion for sharing knowledge and experience driving complex projects end-to-end. 

The base salary range is 224,000 USD - 425,500 USD. Your base salary will be determined based on your location, experience, and the pay of employees in similar positions.

You will also be eligible for equity and benefits. NVIDIA accepts applications on an ongoing basis.

NVIDIA is committed to fostering a diverse work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.

Similar Jobs

King - Senior Software Engineer (Data)

King

Barcelona, Catalonia, Spain (On-Site)
1 Week ago
Epic Games - Technical Director, Machine Learning Programmer

Epic Games

Vancouver, British Columbia, Canada (On-Site)
2 Weeks ago
Epic Games - Senior Data Analyst, Unreal Engine & Creator Products

Epic Games

Cary, North Carolina, United States (On-Site)
1 Month ago
Inworld AI - Staff Platform Engineer - USA

Inworld AI

Mountain View, California, United States (On-Site)
2 Months ago
Epic Games - Senior Data Scientist - Product Analytics

Epic Games

Cary, North Carolina, United States (On-Site)
1 Month ago
Krafton  - Publishing Member of Pricing & Monetization Strategy

Krafton

Seoul, South Korea (On-Site)
1 Week ago
Assystems - Référent expérimenté en calculs dynamiques mécaniques H/F

Assystems

Marseille, Provence-Alpes-Côte D'Azur, France (On-Site)
3 Months ago
HP - Senior Technical Lead - MS Dynamics

HP

Bengaluru, Karnataka, India (On-Site)
5 Months ago
GlobalHunt - Design Engineer

GlobalHunt

Bengaluru, Karnataka, India (On-Site)
5 Months ago
NVIDIA - DFT Engineer - Hardware

NVIDIA

Bengaluru, Karnataka, India (Hybrid)
1 Month ago

Get notifed when new similar jobs are uploaded

Similar Skill Jobs

Ziff Davis - Editor, Content Updates

Ziff Davis

United States (Remote)
1 Month ago
Aera Technology - Senior Software Engineer (Backend Java)

Aera Technology

Pune, Maharashtra, India (On-Site)
4 Months ago
The Walt Disney Company - Lead Data Engineer

The Walt Disney Company

New York, New York, United States (On-Site)
2 Months ago
The Walt Disney Company - Lead Data Scientist

The Walt Disney Company

Burbank, California, United States (On-Site)
3 Months ago
Next Level Business Services - Big data Architect with Azure Experience

Next Level Business Services

Columbus, Indiana, United States (On-Site)
4 Months ago
Accenture in India - GN - Song - MT - Brand and Creative Strategy- Jr. Art Director- Analyst

Accenture in India

Maharashtra, India (Hybrid)
7 Months ago
Dream Sports - SDE 2 - Frontend

Dream Sports

Mumbai, Maharashtra, India (On-Site)
4 Months ago
Playtika - Senior Data/AI SRE Engineer

Playtika

Ukraine (On-Site)
3 Months ago
Patterned Learning Career - Senior Software Engineer, Data

Patterned Learning Career

(Remote)
1 Week ago

Get notifed when new similar jobs are uploaded

Jobs in Santa Clara, California, United States

Corsair - Global Supply Manager

Corsair

Milpitas, California, United States (On-Site)
1 Month ago
AppLovin - PIPELINE Software Engineer 2, Backend

AppLovin

Palo Alto, California, United States (Hybrid)
8 Months ago
Nintendo - Intern - Business Intelligence

Nintendo

Redmond, Washington, United States (On-Site)
3 Months ago
The Walt Disney Company - Senior Product Designer

The Walt Disney Company

New York, New York, United States (On-Site)
3 Weeks ago
ByteDance - Research Scientist, Multimodality

ByteDance

San Jose, California, United States (On-Site)
3 Months ago
Scope AR - ABX Marketing Manager

Scope AR

San Francisco, California, United States (Remote)
2 Months ago
Hasbro - District Manager RMO

Hasbro

United States (On-Site)
2 Weeks ago
Firaxis Games - Join Our Talent Community

Firaxis Games

Sparks Glencoe, Maryland, United States (On-Site)
3 Months ago
Onward Search - Producer III

Onward Search

Los Angeles, California, United States (Hybrid)
1 Week ago
Nintendo - Product Tester (Retro Studios)

Nintendo

Austin, Texas, United States (On-Site)
6 Months ago

Get notifed when new similar jobs are uploaded

Research & Development Jobs

NVIDIA - Senior Manager, High-Speed Optical Transceiver Design

NVIDIA

Yokne'am Illit, North District, Israel (On-Site)
6 Days ago
Riot Games - Senior Manager, QA - VALORANT Experience Team

Riot Games

Dublin, County Dublin, Ireland (On-Site)
3 Months ago
Luxoft - Design Verification Engineer - Data Fabric Systems

Luxoft

Bucharest, Bucharest, Romania (On-Site)
2 Months ago
NVIDIA - Senior Verification Engineer - GPU Fullchip

NVIDIA

Bengaluru, Karnataka, India (Hybrid)
1 Month ago
Power Integrations - Staff Automotive Reliability Engineer

Power Integrations

Penang, Malaysia (On-Site)
4 Months ago
BestEx Research - Senior Software Engineer

BestEx Research

Bengaluru, Karnataka, India (On-Site)
4 Months ago
NVIDIA - Senior CPU Verification Engineer

NVIDIA

Hyderabad, Telangana, India (On-Site)
1 Month ago
NVIDIA - Software Engineer

NVIDIA

Ra'anana, Center District, Israel (On-Site)
3 Weeks ago
Fluence - Sr. Software Architect (m/f/d)

Fluence

Berlin, Berlin, Germany (On-Site)
3 Months ago
Social Discovery Group - Team Lead/Principal NLP Engineer

Social Discovery Group

(Remote)
1 Week ago

Get notifed when new similar jobs are uploaded

About The Company

Since its founding in 1993, NVIDIA (NASDAQ: NVDA) has been a pioneer in accelerated computing. The company’s invention of the GPU in 1999 sparked the growth of the PC gaming market, redefined computer graphics, ignited the era of modern AI and is fueling the creation of the metaverse. NVIDIA is now a full-stack computing company with data-center-scale offerings that are reshaping industry.


Santa Clara, California, United States (On-Site)

Santa Clara, California, United States (On-Site)

Shenzhen, Guangdong Province, China (On-Site)

Bengaluru, Karnataka, India (On-Site)

Taipei City, Taiwan (On-Site)

Taipei City, Taiwan (On-Site)

Shanghai, Shanghai, China (On-Site)

Shanghai, Shanghai, China (On-Site)

Yokne'am Illit, North District, Israel (On-Site)

View All Jobs

Get notified when new jobs are added by NVIDIA

Level Up Your Career in Game Development!

Transform Your Passion into Profession with Our Comprehensive Courses for Aspiring Game Developers.

Job Common Plug