Outscal Logooutscal logo

Senior Observability Architect, AI and HPC

1 Month ago • 12 Years + • Research & Development • $224,000 PA - $425,500 PA

Job Summary

Job Description

NVIDIA's Hardware Infrastructure team seeks a Senior Observability Architect to define and implement observability systems for large-scale AI and HPC clusters. Responsibilities include architecting data collection, aggregation, and visualization systems; collaborating with AI, HW, and SW teams; defining data policies; leading technical teams; and continuously improving observability. The ideal candidate possesses extensive experience building large-scale distributed systems and strong collaboration skills.
Must have:
  • Experience with large-scale distributed observability systems
  • Collaboration with data scientists and engineering teams
  • Turning raw data into actionable reports
  • Experience with observability platforms (e.g., Spark, Elasticsearch, Grafana, Prometheus)
  • Python programming and API calls
  • Improving team productivity
Good to have:
  • Background in computer science, machine learning, deep learning, open-source software, infrastructure technologies, and GPU technology
  • Experience in infrastructure software, production application software development, and DevOps
  • Experience managing data centers and large-scale distributed computing
  • Experience with AI researchers or EDA developers
  • Track record of driving process improvements and measuring efficiency
Perks:
  • Equity
  • Benefits

Job Details

 NVIDIA’s Hardware Infrastructure organization is seeking a Senior or Principal Data and Observability Architect. We serve and collaborate directly with NVIDIA’s rapidly growing AI, HW, and SW engineering and research teams across the company. We are looking for a technical leader to define a vision and roadmap for distributed observability systems for large-scale AI and HPC clusters and workloads and guide implementation towards this vision. You will architect systems for data collection, aggregation, enrichment, storage, retrieval, and visualization to spectacularly improve efficiency, performance, and productivity of AI and HPC workloads. You will lead technical teams to develop, deploy, and operate observability solutions for multiple compute clusters around the world. 

What You’ll Be Doing: 

  • Collaborate with AI, HW, and SW engineering and research teams to define a vision and roadmap for AI/HPC cluster observability. 

  • Architect and lead teams to develop, test, and deploy data collectors, pipelines, visualization and retrieval services. 

  • Define data collection and retention polices to balance network bandwidth, system load, and storage capacity costs with data analysis requirements. 

  • Work in a diverse team to provide operational and strategic data to empower our engineers and researchers to improve performance, productivity, and efficiency. 

  • Continuously improve quality, workloads, and processes through better observability. 

What We Need to See: 

  • Experience designing and building large scale, distributed observability systems. 

  • Ability to collaborate with data scientists, researchers, and engineering teams to identify high value data for collection and analysis. 

  • Experience with turning raw data into actionable reports 

  • Experience with observability platforms such as Apache Spark, Elastic/Open Search, Grafana, Prometheus, and other similar open-source tools 

  • Technical lead level Python programming experience and use of API calls 

  • Passion for improving the productivity of others 

  • Excellent planning and interpersonal skills 

  • Flexibility/adaptability working in a dynamic environment with changing requirements  

  • MS (preferred) or BS in Computer Science, Electrical Engineering, or related field or equivalent experience

  • 12+ years of relevant experience. 

Ways To Stand Out from The Crowd: 

  • Background in computer science, machine learning, deep learning, open-source software, infrastructure technologies, and GPU technology. 

  • Prior experience in infrastructure software, production application software development, software development, release and support methodology and devops 

  • Experience in the management of datacenters and large-scale distributed computing 

  • Experience in working with AI researchers and/or EDA developers 

  • Consistent track record of driving process improvements and measuring efficiency and a passion for sharing knowledge and experience driving complex projects end-to-end. 

The base salary range is 224,000 USD - 425,500 USD. Your base salary will be determined based on your location, experience, and the pay of employees in similar positions.

You will also be eligible for equity and benefits. NVIDIA accepts applications on an ongoing basis.

NVIDIA is committed to fostering a diverse work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.

Similar Jobs

The Walt Disney Company - Senior Data Engineer

The Walt Disney Company

Santa Monica, California, United States (On-Site)
1 Day ago
Netflix - Research Engineer L4/L5 -LLMs for Search, Recommendations, and Personalization

Netflix

Los Gatos, California, United States (On-Site)
5 Months ago
Sourcegraph  Inc  - Customer Engineer (Pre-sales) [IC2]

Sourcegraph Inc

San Francisco, California, United States (On-Site)
4 Months ago
Dream Sports - Director - DevOps

Dream Sports

Mumbai, Maharashtra, India (On-Site)
4 Weeks ago
Omnissa - Staff Engineer (Data Science)

Omnissa

Bengaluru, Karnataka, India (Hybrid)
4 Months ago
Passive Logic - Weather Simulation Engineer

Passive Logic

Salt Lake City, Utah, United States (On-Site)
3 Months ago
ByteDance - Research Scientist Graduate (Foundation Model - Vision and Language)

ByteDance

Seattle, Washington, United States (On-Site)
2 Months ago
NVIDIA - Senior Software Architect, Advanced Development

NVIDIA

Yokne'am Illit, North District, Israel (On-Site)
2 Months ago
NVIDIA - Senior System Software Engineer, Robotics Simulation

NVIDIA

Canada (Hybrid)
3 Weeks ago
Rivos - Silicon Microarchitecture & Logic Design - Intern

Rivos

Santa Clara, California, United States (On-Site)
5 Months ago

Get notifed when new similar jobs are uploaded

Similar Skill Jobs

Dream Sports - Engineering Manager

Dream Sports

Mumbai, Maharashtra, India (On-Site)
12 Hours ago
Epic Games - Data Analyst - Product Analytics

Epic Games

Vancouver, British Columbia, Canada (On-Site)
2 Months ago
Dream Sports - ML Engineer

Dream Sports

Mumbai, Maharashtra, India (On-Site)
3 Months ago
ByteDance - Data Engineer, Cloud and System

ByteDance

Seattle, Washington, United States (On-Site)
1 Day ago
Rackspace Technology - Principal MLOps Engineer

Rackspace Technology

(Remote)
6 Days ago
Sandsoft Games - Director of Data Science and Engineering

Sandsoft Games

Riyadh, Riyadh Province, Saudi Arabia (On-Site)
12 Hours ago
Highspot - Sr. Full Stack Engineer, Training & Coaching

Highspot

Hyderabad, Telangana, India (Hybrid)
5 Months ago
NinjaVan - Staff Data Engineer

NinjaVan

Hyderabad, Telangana, India (On-Site)
5 Months ago
PwC - Data Architect – Technology Consulting

PwC

Prague, Prague, Czechia (On-Site)
5 Months ago
PhonePe - SRE - Big Data (OnPrem)

PhonePe

Bengaluru, Karnataka, India (On-Site)
4 Months ago

Get notifed when new similar jobs are uploaded

Jobs in Canada

Granicus - SLED Enterprise Account Executive - Western Canada (Local)

Granicus

Canada (Remote)
5 Months ago
Epic Games - Art Director

Epic Games

Montreal, Quebec, Canada (On-Site)
1 Day ago
Ubisoft - Technical Animation Director

Ubisoft

Montreal, Quebec, Canada (Hybrid)
1 Day ago
Turbulent - Senior DevOps Engineer

Turbulent

Montreal, Quebec, Canada (On-Site)
1 Week ago
Keywords Studios (Player Support) - Expert Subtitle Translator/QCer: English to Canadian French

Keywords Studios (Player Support)

Québec City, Quebec, Canada (Remote)
1 Week ago
Scanline VFX - Lead Software Engineer (Production Tools)

Scanline VFX

Vancouver, British Columbia, Canada (Remote)
5 Months ago
Epic Games - Concept Artist

Epic Games

Montreal, Quebec, Canada (On-Site)
1 Day ago
PwC - Corporate Tax Real Estate, Manager

PwC

Toronto, Ontario, Canada (On-Site)
5 Months ago
2K - Expert Gameplay Animation Engineer

2K

Vancouver, British Columbia, Canada (Hybrid)
5 Months ago
Electronic Arts - Technical Artist - User Interface

Electronic Arts

Vancouver, British Columbia, Canada (On-Site)
2 Months ago

Get notifed when new similar jobs are uploaded

Research & Development Jobs

NVIDIA - Mixed Signal Circuit Designer (RDSS Intern)

NVIDIA

Taipei City, Taiwan (On-Site)
1 Month ago
NVIDIA - Senior Boot Reset Silicon Hardware Engineer

NVIDIA

Santa Clara, California, United States (Hybrid)
2 Months ago
Passive Logic - Senior Electrical Engineer

Passive Logic

Salt Lake City, Utah, United States (On-Site)
5 Months ago
Netflix - Research Scientist 4 - Content and Studio

Netflix

Los Gatos, California, United States (On-Site)
5 Months ago
NVIDIA - SOC Clock Distribution Engineer

NVIDIA

Tel Aviv-Yafo, Tel Aviv District, Israel (On-Site)
2 Months ago
Krafton  - [Publishing Platform Div.] Global Publishing Platform QA (5년 ~ 10년)

Krafton

Seoul, South Korea (On-Site)
4 Months ago
NVIDIA - System Software Engineer - CUDA Driver

NVIDIA

Santa Clara, California, United States (On-Site)
2 Months ago
NVIDIA - Senior Manager, Hardware Engineering

NVIDIA

Canada (Hybrid)
1 Month ago
ByteDance - Research Scientist Graduate (Foundation Model - Vision and Language)

ByteDance

Seattle, Washington, United States (On-Site)
2 Months ago
NVIDIA - Physical Design Backend Engineer

NVIDIA

Tel Aviv-Yafo, Tel Aviv District, Israel (On-Site)
2 Months ago

Get notifed when new similar jobs are uploaded

About The Company

Since its founding in 1993, NVIDIA (NASDAQ: NVDA) has been a pioneer in accelerated computing. The company’s invention of the GPU in 1999 sparked the growth of the PC gaming market, redefined computer graphics, ignited the era of modern AI and is fueling the creation of the metaverse. NVIDIA is now a full-stack computing company with data-center-scale offerings that are reshaping industry.


Hsinchu, Hsinchu City, Taiwan (On-Site)

Yokne'am Illit, North District, Israel (On-Site)

Seoul, South Korea (Hybrid)

Tel Aviv-Yafo, Tel Aviv District, Israel (On-Site)

Ra'anana, Center District, Israel (On-Site)

Shanghai, Shanghai, China (On-Site)

Tel Aviv-Yafo, Tel Aviv District, Israel (On-Site)

Be'er Sheva, South District, Israel (On-Site)

Santa Clara, California, United States (On-Site)

Santa Clara, California, United States (On-Site)

View All Jobs

Get notified when new jobs are added by NVIDIA

Level Up Your Career in Game Development!

Transform Your Passion into Profession with Our Comprehensive Courses for Aspiring Game Developers.

Job Common Plug