Senior Staff Engineer, Memory Fault Management Architect

3 Months ago • 10-15 Years • Research & Development • $180,950 PA - $289,050 PA

Job Summary

Job Description

As a Senior Staff Engineer, Memory Fault Management Architect, you will be part of an incubation team focused on transforming the customer quality experience for Samsung memory products. This role involves analyzing massive datasets from memory fleet telemetry to identify failure modes, project failure rates, and develop proactive solutions to minimize system downtime. You will collaborate with customers, contribute to industry standardization efforts (OCP), and design/develop RAS algorithms (page offlining, hPPR). Responsibilities include recommending solutions to mitigate DRAM failure rates, communicating ECC schemes, and establishing the value of in-field fault management architecture. The position requires deep knowledge of SOC controllers, memory operations, RAS features, ECC design, and Linux kernel experience.
Must have:
  • 10+ years experience in hardware fault management
  • Knowledge of platform memory subsystem and RAS
  • ECC design, verification, and reverse engineering
  • Understanding of DRAM and HBM failure modes
  • Excellent communication and collaboration skills
Good to have:
  • Linux kernel commit experience
  • Memory controller register modification
Perks:
  • 4+ weeks paid time off
  • Medical/Dental/Vision/401k
  • Fertility care or adoption stipend
  • Medical travel support
  • On-site gym and cafe
  • Virtual classes
  • Flexible work environment

Job Details

Please Note:

To provide the best candidate experience amidst our high application volumes, each candidate is limited to 10 applications across all open jobs within a 6-month period. 

Advancing the World’s Technology Together
Our technology solutions power the tools you use every day--including smartphones, electric vehicles, hyperscale data centers, IoT devices, and so much more. Here, you’ll have an opportunity to be part of a global leader whose innovative designs are pushing the boundaries of what’s possible and powering the future. 

We believe innovation and growth are driven by an inclusive culture and a diverse workforce. We’re dedicated to empowering people to be their true selves. Together, we’re building a better tomorrow for our employees, customers, partners, and communities.

Conventional DRAM failure analysis was physical electrical FA and physical FA. But, in the era of Data center, it is easier to track the field failure information. With this data set, Fault management team’s role is finding DRAM failure mode, abnormality and failure rate projection.

You will be part of an incubation team working on in-field telemetry intended to transform the Customer Quality Experience for Samsung memory products. Fault Management is the future of quality to minimize system downtime within AI/ML hardware deployments and workloads of the future. We analyze trends and patterns from enormous memory fleet telemetry to bucketize failures and perform virtual root-cause analysis. Telemetry analysis helps us design solutions to proactively avoid system downtime. We conduct research and develop both in-house and collaboratively in the industry with the opportunity to publish our findings through whitepapers and conferences. We are looking for innovative and passionate thinkers who can work in a start-up environment and are excited to shape the future of data centers around the world. Join us in our mission!

What You'll Do

  • Based on the knowledge of  SOC controller and memory operation including RAS feature, find and recommends better solution to mitigate the field DRAM failure rate.
  • Needs to communicate better ECC scheme to customers based on Samsung DRAM failure mode(DQ and burst)
  • Interface with customers to establish the value add of enabling in-field fault management architecture
  • Contribute to the standardization of DRAM/HBM failure logging in the OCP.
  • Propose and develop platform RAS (Reliability Availability Serviceability) algorithms for memory fault management such as page offlining, hPPR and conduct POC with known failure DIMMs in the real server and application.

Location: Hybrid with at least 3 days in office in San Jose, CA office location remainder of time to work remotely

Job ID: 42448

 What You Bring

  • Bachelors with 15+ years of relevant industry experience, or Masters with 13+ years or PhD with 10+ years hardware fault management, reliability, data center fleet management experience or related technical field preferred
  • Knowledge of platform memory subsystem, platform RAS (Reliability Availability Serviceability) such as ECC, page offlining, hPPR and hardware sparing.
  • ECC design and verification and reverse engineering experience.
  • Understanding on the address mapping between CPU and memory.
  • Memory controller register modification.
  • Linux kernel commit experience.
  • DRAM and HBM failure mode understanding.
  • Excellent communication and interpersonal skills.
  • Ability to work independently and as part of a team.
  • You’re inclusive, adapting your style to the situation and diverse global norms of our people.
  • An avid learner, you approach challenges with curiosity and resilience, seeking data to help build understanding.
  • You’re collaborative, building relationships, humbly offering support and openly welcoming approaches.
  • Innovative and creative, you proactively explore new ideas and adapt quickly to change.

#LI-SF1

 

 

 

What We Offer
The pay range below is for all roles at this level across all US locations and functions. Individual pay rates depend on a number of factors—including the role’s function and location, as well as the individual’s knowledge, skills, experience, education, and training. We also offer incentive opportunities that reward employees based on individual and company performance. 

This is in addition to our diverse package of benefits centered around the wellbeing of our employees and their loved ones. In addition to the usual Medical/Dental/Vision/401k, our inclusive rewards plan empowers our people to care for their whole selves. An investment in your future is an investment in ours.

Give Back With a charitable giving match and frequent opportunities to get involved, we take an active role in supporting the community.
Enjoy Time Away You’ll start with 4+ weeks of paid time off a year, plus holidays and sick leave, to rest and recharge.
Care for Family Whatever family means to you, we want to support you along the way—including a stipend for fertility care or adoption, medical travel support, and an errand service.
Prioritize Emotional Wellness With on-demand apps and paid therapy sessions, you’ll have support no matter where you are.
Stay Fit Eating well and being active are important parts of a healthy life. Our onsite Café and gym, plus virtual classes, make it easier.
Embrace Flexibility Benefits are best when you have the space to use them. That’s why we facilitate a flexible environment so you can find the right balance for you.

Base Pay Range

$180,950 - $289,050 USD

Equal Opportunity Employment Policy 

Samsung Semiconductor takes pride in being an equal opportunity workplace dedicated to fostering an environment where all individuals feel valued and empowered to excel, regardless of race, religion, color, age, disability, sex, gender identity, sexual orientation, ancestry, genetic information, marital status, national origin, political affiliation, or veteran status.

When selecting team members, we prioritize talent and qualities such as humility, kindness, and dedication. We extend comprehensive accommodations throughout our recruiting processes for candidates with disabilities, long-term conditions, neurodivergent individuals, or those requiring pregnancy-related support. All candidates scheduled for an interview will receive guidance on requesting accommodations.

Recruiting Agency Policy

We do not accept unsolicited resumes. Only authorized recruitment agencies that have a current and valid agreement with Samsung Semiconductor, Inc. are permitted to submit resumes for any job openings.

Covid-19 Policy
To help keep our employees, customers, and communities safe, we’ve developed guidelines for our teams. Currently, we encourage vaccination for all employees and may require it depending on job functions (e.g., traveling for business, meeting with customers). While visiting our offices or attending team events, we ask employees to complete a daily health questionnaire and complete a weekly COVID test. Our COVID policies are subject to change depending on public health, regulatory and business circumstances. 

Applicant Privacy Policy
https://semiconductor.samsung.com/us/careers/privacy

 

Similar Jobs

Google - Software Engineer II, Chrome Enterprise Core

Google

Warsaw, Masovian Voivodeship, Poland (On-Site)
1 Week ago
Zazz - Artificial Intelligence Engineer

Zazz

(Remote)
2 Months ago
Google - Software Engineer, Performance Modeling

Google

Raleigh, North Carolina, United States (On-Site)
1 Week ago
Microsoft - Research Intern - Quantum Computing

Microsoft

California, United States (On-Site)
1 Day ago
Google - Software Engineer III, Site Reliability Engineering, Network Management

Google

Dublin, County Dublin, Ireland (On-Site)
1 Week ago
Google - Staff Software Engineer, YouTube

Google

Mountain View, California, United States (On-Site)
1 Week ago
Google - Data Center Cooling Engineer

Google

Sunnyvale, California, United States (On-Site)
4 Days ago
Google - Senior SoC Power Engineer

Google

Tel Aviv-Yafo, Tel Aviv District, Israel (On-Site)
1 Week ago
NVIDIA - Senior Physical Design Engineer

NVIDIA

Hsinchu, Hsinchu City, Taiwan (On-Site)
3 Months ago
Google - Electronics Technician

Google

Chicago, Illinois, United States (On-Site)
1 Week ago

Get notifed when new similar jobs are uploaded

Similar Skill Jobs

Google - Software Engineer III, Artificial Intelligence/Machine Learning

Google

Hyderabad, Telangana, India (On-Site)
4 Days ago
Arkose Labs - Senior Machine Learning Researcher

Arkose Labs

Pune, Maharashtra, India (Hybrid)
6 Months ago
Wargaming - Gameplay Developer (World of Tanks)

Wargaming

Warsaw, Masovian Voivodeship, Poland (Hybrid)
2 Days ago
CD PROJEKT RED - Senior UI Programmer

CD PROJEKT RED

Boston, Massachusetts, United States (Hybrid)
16 Hours ago
ByteDance - Software Development Engineer Graduate (Distributed NoSQL Database Systems)

ByteDance

Seattle, Washington, United States (On-Site)
2 Months ago
Genies - Lead Applied ML Engineer, Real-time 3D Asset Optimization

Genies

San Mateo, California, United States (On-Site)
4 Weeks ago
The Walt Disney Company - Principal Machine Learning Engineer

The Walt Disney Company

San Francisco, California, United States (On-Site)
1 Week ago
ByteDance - Research Scientist - Multimedia Lab

ByteDance

San Diego, California, United States (On-Site)
1 Month ago
Google - Staff Software Engineer, Google Cloud Business Platforms

Google

Kirkland, Washington, United States (On-Site)
6 Days ago

Get notifed when new similar jobs are uploaded

Jobs in San Jose, California, United States

Universal Music - Creative Director- Creative Services

Universal Music

Philadelphia, Pennsylvania, United States (On-Site)
2 Months ago
Google - Staff Software Engineer, Machine Learning

Google

Los Angeles, California, United States (On-Site)
6 Days ago
ByteDance - Architect - AML Engine

ByteDance

San Jose, California, United States (On-Site)
6 Months ago
Google - Security Sales Specialist, Google Cloud

Google

Chicago, Illinois, United States (On-Site)
1 Week ago
IGN - Senior Full Stack Software Engineer

IGN

New York, New York, United States (Hybrid)
5 Months ago
Evolution - In-Studio Online Casino Dealer- Overnight ONLY 11pm-7am

Evolution

Philadelphia, Pennsylvania, United States (On-Site)
11 Months ago
Nagarro - Associate Distinguished Engineer - Enterprise Data Architect

Nagarro

Allentown, Pennsylvania, United States (Remote)
6 Months ago
The Walt Disney Company - Disney Store Lead Cast Member

The Walt Disney Company

Oklahoma City, Oklahoma, United States (On-Site)
1 Week ago
Nintendo - Assistant Manager - Nintendo San Francisco Store

Nintendo

San Francisco, California, United States (On-Site)
8 Months ago
Microsoft - Member of Technical Staff - Prompt Engineer, Product

Microsoft

Mountain View, California, United States (On-Site)
1 Week ago

Get notifed when new similar jobs are uploaded

Research & Development Jobs

NVIDIA - Senior Physical Design Full Chip STA Engineer

NVIDIA

Yokne'am Illit, North District, Israel (On-Site)
2 Months ago
Rivos - Logic Equivalence Check (LEC) Engineer

Rivos

Hsinchu, Hsinchu City, Taiwan (Hybrid)
6 Months ago
ByteDance - Senior Research Scientist, Foundation Model, Speech Understanding

ByteDance

San Jose, California, United States (On-Site)
5 Months ago
Riot Games - Staff Research Scientist - Tech Research

Riot Games

Los Angeles, California, United States (On-Site)
1 Month ago
NVIDIA - CSP Hardware Application Engineer

NVIDIA

Shenzhen, Guangdong Province, China (On-Site)
1 Month ago
ByteDance - Software Engineer, Model Inference

ByteDance

Seattle, Washington, United States (On-Site)
1 Month ago
ByteDance - Senior Technical Lead - Edge Cloud Infrastructure - San Jose / Seattle / Boston

ByteDance

San Jose, California, United States (On-Site)
5 Months ago
Google - SoC and IP Design Engineer

Google

Haifa, Haifa District, Israel (On-Site)
6 Days ago
Riot Games - Staff Software Engineer, Game Build - Teamfight Tactics

Riot Games

Los Angeles, California, United States (On-Site)
1 Week ago
NVIDIA - Senior Chip Design Methodologies Engineer

NVIDIA

Tel Aviv-Yafo, Tel Aviv District, Israel (On-Site)
3 Weeks ago

Get notifed when new similar jobs are uploaded

About The Company

San Jose, California, United States (On-Site)

San Jose, California, United States (On-Site)

San Jose, California, United States (On-Site)

San Jose, California, United States (On-Site)

San Jose, California, United States (On-Site)

San Jose, California, United States (Hybrid)

San Jose, California, United States (On-Site)

San Jose, California, United States (On-Site)

San Jose, California, United States (On-Site)

San Jose, California, United States (On-Site)

View All Jobs

Get notified when new jobs are added by Samsung Semiconductor

Level Up Your Career in Game Development!

Transform Your Passion into Profession with Our Comprehensive Courses for Aspiring Game Developers.

Job Common Plug