Senior Site Reliability Engineer

2 Months ago • 5 Years + • Devops

Job Summary

Job Description

Reddit SRE is seeking a Senior Site Reliability Engineer to innovate and improve the reliability and performance of Reddit's infrastructure and services. This role involves working at the intersection of infrastructure and software development, focusing on compute, traffic, and observability teams. You will own tools for engineers to understand their creations, utilizing open-source solutions like Prometheus, Thanos, Grafana, and Vector. Key responsibilities include risk management, collaborating with teams to mitigate risks, enhancing system resilience, and driving proactive measures for uptime and service delivery. The ideal candidate will contribute to building the future of Reddit.

Must have:

5+ years of experience in SRE/DevOps
Proficiency in Go or Python
Experience with Kubernetes and Cloud systems
Experience with distributed systems development
Debug, fix, and optimize code
Troubleshoot applications, networking, and systems
Strong Linux and container knowledge
Excellent communication and collaboration skills

Good to have:

Familiarity with Prometheus, Thanos, Grafana, Vector, Clickhouse, Otel, Loki
Experience with high-traffic backend systems

Perks:

Pension Scheme
Private Medical and Dental Scheme
Life Assurance
Income Protection
Workspace benefit for home office
Personal & Professional development funds
Family Planning Support
Commuter Benefits
Flexible Vacation & Reddit Global Days Off

13 skills required

13 skills required for this role

Add these skills to join the top 1% applicants for this job

cross-functional

communication

problem-solving

risk-management

game-texts

quality-control

networking

linux

incident-response

prometheus

grafana

kubernetes

python

Job Details

Reddit is a community of communities. It’s built on shared interests, passion, and trust and is home to the most open and authentic conversations on the internet. Every day, Reddit users submit, vote, and comment on the topics they care most about. With 100,000+ active communities and approximately 101M+ daily active unique visitors, Reddit is one of the internet’s largest sources of information. For more information, visit redditinc.com.

Reddit SRE is rapidly innovating and our teams are working to meet the needs of infrastructure and development teams as they evolve our product faster than ever before. This is a unique opportunity to leave your mark on one of the most influential and trafficked corners of the internet.

As a Senior Site Reliability Engineer on Reddit’s Infrastructure SRE team, you’ll use your knowledge of distributed systems and architecture to improve the reliability and performance of Reddit’s engineering platforms and services. We are looking for someone who thrives at the intersection of infrastructure and software development. This team will work very closely with the Compute, Traffic, and Observability infrastructure teams. They will own a suite of tools for allowing engineers to understand their creations, based primarily on open-source solutions at scale. We’re active users of and contributors to Prometheus, Thanos, Grafana, Vector and more.

In this role, you will also take ownership of risk management, ensuring the reliability and performance of our systems. You will collaborate with cross-functional teams to identify, assess, and mitigate risks, implementing best practices to enhance system resilience. Your expertise will drive proactive measures to maintain uptime and optimize service delivery, making a significant impact on our operational excellence.

Join us and help build the future of Reddit!

Responsibilities:

Advise:

Work closely with engineering teams in designing and developing systems that are resilient and highly performant at a tremendous scale, and maintaining the foundational platform for running Reddit’s infrastructure.

Amplify:

Identify and build capabilities into our foundational Infrastructure and Platform services, which are used by Reddit engineering teams to build, deploy, and operate Reddit.
Deliver software to improve the availability, scalability, latency, and efficiency of observability components.
Identify and engineer away risk across Reddit’s systems.

Automate:

Take repetitive, manual, or risky tasks and automate them out of existence. Build tools and integrate systems to support Reddit’s evolution.
Automate critical aspects of the event driven development process

Diagnose:

Draw on your knowledge of distributed systems to identify and fix network, system, and service-level issues. Practice sustainable incident response, and drive structural improvement with blameless postmortem.
Share on-call responsibilities.

Optimize:

Observe and improve performance, reduce cost, and improve the experience for millions of users
Contribute upstream changes to the open source projects we use

Qualifications

5+ years of experience in Software Engineering, Site Reliability Engineering, or a development-focused DevOps role.
Proficiency in one or more programming languages. We’re predominantly writing code in Go and Python.
Experience with Kubernetes and Cloud systems.
Familiarity with distributed systems development, bonus if familiar with any of the specific tools (Prometheus, Thanos, Grafana, Vector, Clickhouse, Otel, Loki)
Experience with the development and operation of high-traffic backend systems.
A demonstrated ability to debug, fix, and optimize code.
Troubleshooting skills that span applications, networking (TCP/IP), and systems.
Strong working knowledge of Linux and containers.
Excellent communication and collaborative skills.

Benefits:

Pension Scheme
Private Medical and Dental Scheme
Life Assurance, Income Protection
Workspace benefit for your home office
Personal & Professional development funds
Family Planning Support
Commuter Benefits
Flexible Vacation & Reddit Global Days Off

Reddit is proud to be an equal opportunity employer, and is committed to building a workforce representative of the diverse communities we serve. Reddit is committed to providing reasonable accommodations for qualified individuals with disabilities and disabled veterans in our job application procedures. If, due to a disability, you need an accommodation during the interview process, please let your recruiter know.

Similar Jobs

Packaging Operations Designer

hogarth

Sunnyvale, California, United States (Hybrid)

• 3 Months ago

Senior Product Manager - Token Services

Adyen

Amsterdam, North Holland, Netherlands (On-Site)

• 2 Months ago

Lead Business Analyst

endava

Brisbane, Queensland, Australia (On-Site)

• 3 Months ago

Finance Systems & Processes Analyst

Betson Group

Malta (On-Site)

• 4 Months ago

SoC Power Analysis and Optimization Engineer

Apple

San Diego, California, United States (On-Site)

• 3 Months ago

AWS Cloud Architect

Brillio

Jersey City, New Jersey, United States (Remote)

• 2 Months ago

Staff Engineer – DevSecOps

extreme network

Ontario, Canada (Hybrid)

• 2 Months ago

Solutions Architect, Conversational AI & Prompt Engineering

Quentus

United States (Remote)

• 5 Months ago

Software Engineer, Multi Cloud CDN - San Jose / Seattle / Boston

bytedance

Seattle, Washington, United States (On-Site)

• 8 Months ago

Senior Cloud Engineer

Blinkhealth

Pittsburgh, Pennsylvania, United States (On-Site)

• 2 Months ago

Get notifed when new similar jobs are uploaded

Similar Skill Jobs

Music Product Counsel - Global Legal

bytedance

Los Angeles, California, United States (On-Site)

• 10 Months ago

Strategic Account Manager

Nordson Corporation

(Remote)

• 4 Months ago

Product Compliance Lead (Growth/Partnerships)

OKX

New York, United States (On-Site)

• 1 Month ago

Senior Director – PC Product Marketing

Qualcomm

Beijing, China (On-Site)

• 3 Months ago

Reconciliation Analyst

Betson Group

Buenos Aires, Buenos Aires, Argentina (On-Site)

• 1 Month ago

Junior Slot Mathematician

MURKA

(Remote)

• 7 Months ago

Video Game UI/UX Designer

VOID Interactive

Dublin, County Dublin, Ireland (Remote)

• 3 Months ago

Senior Marketing Manager

VVater

Austin, Texas, United States (On-Site)

• 4 Months ago

Product Manager, Ad Insertion (Ads)

Netflix

Seattle, Washington, United States (On-Site)

• 4 Months ago

Senior Growth Manager

Rippling

Toronto, Ontario, Canada (On-Site)

• 4 Months ago

Get notifed when new similar jobs are uploaded

Jobs in London, England, United Kingdom

Lead Data Scientist

Monzo

London, England, United Kingdom (Remote)

• 3 Months ago

Senior Legal Counsel, Privacy & Compliance

fortis games

United Kingdom (Remote)

• 3 Months ago

Internship

Ninja theory

Cambridge, England, United Kingdom (On-Site)

• 2 Months ago

UX Manager

Playstation

London, England, United Kingdom (Hybrid)

• 3 Months ago

Go Software Engineer

Vidsy

London, England, United Kingdom (Hybrid)

• 2 Months ago

Field Service Engineer

Illumina

Scotland, United Kingdom (On-Site)

• 1 Year ago

Channel Business Manager

Barracuda

Reading, England, United Kingdom (On-Site)

• 2 Months ago

Avaloq Software Engineer (Data)

luxsoft

London, England, United Kingdom (On-Site)

• 4 Months ago

AI Safety Operations

ElevenLabs

United Kingdom (Remote)

• 5 Months ago

Brand Marketing Specialist

Dream Games

London, England, United Kingdom (On-Site)

• 2 Years ago

Get notifed when new similar jobs are uploaded

Devops Jobs

Senior Cloud Solutions Engineer

Sonar Source

Austin, Texas, United States (On-Site)

• 3 Months ago

Enterprise Solutions Architect

Progress

United States (Remote)

• 4 Months ago

Lead Solutions Engineer

Cadence

Bengaluru, Karnataka, India (On-Site)

• 11 Months ago

Senior Staff Software Engineer, Google Cloud

Google

Pune, Maharashtra, India (On-Site)

• 4 Months ago

Infrastructure Engineer

Colo pl

Minato City, Tokyo, Japan (On-Site)

• 4 Months ago

Senior Backend Engineer - Support Automation and AI Enablement

Canva

Melbourne, Victoria, Australia (Remote)

• 5 Months ago

SRE – Python Developer

Synechron

Montreal, Quebec, Canada (On-Site)

• 2 Months ago

DevOps Engineer

Apple

Austin, Texas, United States (On-Site)

• 3 Months ago

Director/Lead Platform Support Engineer

SSC Technologies

Union, New Jersey, United States (Hybrid)

• 3 Months ago

Cloud Monitoring SRE

Apple

Seattle, Washington, United States (On-Site)

• 3 Months ago

Get notifed when new similar jobs are uploaded

About The Company

75 Active Jobs

Get notified when new jobs are added by Reddit

Level Up Your Career in Game Development!

Transform Your Passion into Profession with Our Comprehensive Courses for Aspiring Game Developers.

A global community of game builders. Helping people upskill and land jobs in the best gaming studios.

Company

Key Links

hello@outscal.com

Made in INDIA 💛💙

Senior Site Reliability Engineer

Job Summary

Job Description

13 skills required

13 skills required for this role

Job Details

Responsibilities:

Qualifications

Similar Jobs

Packaging Operations Designer

Senior Product Manager - Token Services

Lead Business Analyst

Finance Systems & Processes Analyst

SoC Power Analysis and Optimization Engineer

AWS Cloud Architect

Staff Engineer – DevSecOps

Solutions Architect, Conversational AI & Prompt Engineering

Software Engineer, Multi Cloud CDN - San Jose / Seattle / Boston

Senior Cloud Engineer

Similar Skill Jobs

Music Product Counsel - Global Legal

Strategic Account Manager

Product Compliance Lead (Growth/Partnerships)

Senior Director – PC Product Marketing

Reconciliation Analyst

Junior Slot Mathematician

Video Game UI/UX Designer

Senior Marketing Manager

Product Manager, Ad Insertion (Ads)

Senior Growth Manager

Jobs in London, England, United Kingdom

Lead Data Scientist

Senior Legal Counsel, Privacy & Compliance

Internship

UX Manager

Go Software Engineer

Field Service Engineer

Channel Business Manager

Avaloq Software Engineer (Data)

AI Safety Operations

Brand Marketing Specialist

Devops Jobs

Senior Cloud Solutions Engineer

Enterprise Solutions Architect

Lead Solutions Engineer

Senior Staff Software Engineer, Google Cloud

Infrastructure Engineer

Senior Backend Engineer - Support Automation and AI Enablement

SRE – Python Developer

DevOps Engineer

Director/Lead Platform Support Engineer

Cloud Monitoring SRE

About The Company

Ads Engineering Manager, SMB Activation

Senior Community Manager (contract)

Community Manager - France (contract)

Senior iOS Engineer - Advertiser Growth

Senior iOS Engineer - Advertiser Growth

Senior iOS Engineer - Advertiser Growth

Senior Software Engineer, Ads ML Features Platform

Senior Software Engineer, Ads ML Features Platform

Software Engineer, Ads ML Features Platform

Senior Machine Learning Engineer, Ads Training Platform

Level Up Your Career in Game Development!