Senior Site Reliability Engineer

1 Month ago • 5 Years + • Devops

Job Summary

Job Description

As a Senior Site Reliability Engineer at Reddit, you will be responsible for improving the reliability and performance of Reddit’s engineering platforms and services. You will collaborate with teams to identify and mitigate risks, ensuring system resilience and optimizing service delivery. Responsibilities include advising engineering teams on system design, building capabilities into infrastructure services, automating tasks, diagnosing and fixing issues, and optimizing performance for millions of users. This role involves working with distributed systems, open-source tools, and contributing to the overall operational excellence of Reddit. The team will work very closely with the Compute, Traffic, and Observability infrastructure teams.
Must have:
  • 5+ years of experience in SRE or DevOps.
  • Proficiency in programming languages like Go and Python.
  • Experience with Kubernetes and Cloud systems.
  • Familiarity with distributed systems development.
  • Experience with high-traffic backend systems.
  • Ability to debug, fix, and optimize code.
  • Troubleshooting skills across applications, networking, and systems.
  • Strong working knowledge of Linux and containers.
Good to have:
  • Familiarity with Prometheus, Thanos, Grafana, Vector, Clickhouse, Otel, Loki
Perks:
  • Private Medical, Dental and Vision Benefits
  • Retirement Savings plan with matching contributions
  • Workspace benefits for your home office
  • Personal & Professional development funds
  • Family Planning Support
  • Commuter Benefits
  • Flexible Vacation & Reddit Global Days Off

Job Details

Reddit is a community of communities. It’s built on shared interests, passion, and trust and is home to the most open and authentic conversations on the internet. Every day, Reddit users submit, vote, and comment on the topics they care most about. With 100,000+ active communities and approximately 101M+ daily active unique visitors, Reddit is one of the internet’s largest sources of information. For more information, visit redditinc.com.

Reddit SRE is rapidly innovating and our teams are working to meet the needs of infrastructure and development teams as they evolve our product faster than ever before. This is a unique opportunity to leave your mark on one of the most influential and trafficked corners of the internet.

As a Senior Site Reliability Engineer on Reddit’s Infrastructure SRE team, you’ll use your knowledge of distributed systems and architecture to improve the reliability and performance of Reddit’s engineering platforms and services. We are looking for someone who thrives at the intersection of infrastructure and software development. This team will work very closely with the Compute, Traffic, and Observability infrastructure teams. They will own a suite of tools for allowing engineers to understand their creations, based primarily on open-source solutions at scale. We’re active users of and contributors to Prometheus, Thanos, Grafana, Vector and more.

In this role, you will also take ownership of risk management, ensuring the reliability and performance of our systems. You will collaborate with cross-functional teams to identify, assess, and mitigate risks, implementing best practices to enhance system resilience. Your expertise will drive proactive measures to maintain uptime and optimize service delivery, making a significant impact on our operational excellence.

Join us and help build the future of Reddit!

Responsibilities:

  • Advise
    • Work closely with engineering teams in designing and developing systems that are resilient and highly performant at a tremendous scale, and maintaining the foundational platform for running Reddit’s infrastructure.
  • Amplify
    • Identify and build capabilities into our foundational Infrastructure and Platform services, which are used by Reddit engineering teams to build, deploy, and operate Reddit. 
    • Deliver software to improve the availability, scalability, latency, and efficiency of observability components.
    • Identify and engineer away risk across Reddit’s systems.
  • Automate
    • Take repetitive, manual, or risky tasks and automate them out of existence. Build tools and integrate systems to support Reddit’s evolution.
    • Automate critical aspects of the event driven development process
  • Diagnose
    • Draw on your knowledge of distributed systems to identify and fix network, system, and service-level issues. Practice sustainable incident response, and drive structural improvement with blameless postmortem.
    • Share on-call responsibilities. 
  • Optimize:
    • Observe and improve performance, reduce cost, and improve the experience for millions of users
    • Contribute upstream changes to the open source projects we use

Qualifications

  • 5+ years of experience in Software Engineering, Site Reliability Engineering, or a development-focused DevOps role.
  • Proficiency in one or more programming languages. We’re predominantly writing code in Go and Python.
  • Experience with Kubernetes and Cloud systems.
  • Familiarity with distributed systems development, bonus if familiar with any of the specific tools (Prometheus, Thanos, Grafana, Vector, Clickhouse, Otel, Loki)
  • Experience with the development and operation of high-traffic backend systems.
  • A demonstrated ability to debug, fix, and optimize code.
  • Troubleshooting skills that span applications, networking (TCP/IP), and systems.
  • Strong working knowledge of Linux and containers.
  • Excellent communication and collaborative skills.

Benefits:

  • Private Medical, Dental and Vision Benefits 
  • Retirement Savings plan with matching contributions
  • Workspace benefits for your home office
  • Personal & Professional development funds
  • Family Planning Support
  • Commuter Benefits  
  • Flexible Vacation & Reddit Global Days Off

Reddit is proud to be an equal opportunity employer, and is committed to building a workforce representative of the diverse communities we serve.  Reddit is committed to providing reasonable accommodations for qualified individuals with disabilities and disabled veterans in our job application procedures. If, due to a disability, you need an accommodation during the interview process, please let your recruiter know.

Similar Jobs

The Walt Disney Company - Director, PMO

The Walt Disney Company

Burbank, California, United States (On-Site)
2 Months ago
Autodesk - Senior Machine Learning Engineer

Autodesk

Bengaluru, Karnataka, India (On-Site)
5 Days ago
Betson Group - PR Manager

Betson Group

Buenos Aires, Buenos Aires, Argentina (On-Site)
1 Month ago
Veeam Software - Product Director, Data Resilience Strategy

Veeam Software

Washington, United States (Remote)
2 Weeks ago
bytedance - Software Development Engineer - Distributed NoSQL Database Systems

bytedance

Seattle, Washington, United States (On-Site)
5 Months ago

Get notifed when new similar jobs are uploaded

Similar Skill Jobs

Zuora - Sr Software Engineer

Zuora

Bengaluru, Karnataka, India (On-Site)
1 Week ago
Tide - SEO Marketing Copywriter

Tide

Berlin, Berlin, Germany (On-Site)
1 Month ago
BioFire - Industrialization Scientist, Molecular Biology

BioFire

Philadelphia, Pennsylvania, United States (On-Site)
3 Weeks ago
Alten Technology - Senior Display Systems Engineer

Alten Technology

Newark, California, United States (On-Site)
1 Month ago
Adyen - Global Head of Strategic Growth

Adyen

New York, United States (Hybrid)
1 Month ago
Doola - Director of Influencer Marketing

Doola

New York, New York, United States (On-Site)
1 Month ago
Elyzio - 3D Artist

Elyzio

(Remote)
1 Month ago
PlayStation Global - Technical Product Manager II

PlayStation Global

Carlsbad, California, United States (Hybrid)
2 Months ago
bohemia interactive - Python Programmer

bohemia interactive

Brno, South Moravian Region, Czechia (On-Site)
1 Month ago
Maliyo Games - Data Scientist

Maliyo Games

San Francisco, California, United States (On-Site)
1 Week ago

Get notifed when new similar jobs are uploaded

Jobs in Dublin, County Dublin, Ireland

playrix  - Engineering Manager (Golang)

playrix

Ireland (Remote)
2 Months ago
Cadence - Sr. Software Security Engineer

Cadence

Cork, County Cork, Ireland (On-Site)
1 Week ago
VOID Interactive - Video Game UI/UX Designer

VOID Interactive

Dublin, County Dublin, Ireland (Remote)
1 Week ago
Whatnot - Trust & Risk Agent (French Speaking)

Whatnot

Dublin, County Dublin, Ireland (Remote)
3 Weeks ago
Riot Games - Vendor Enablement Specialist

Riot Games

Dublin, County Dublin, Ireland (On-Site)
2 Weeks ago
Romero games - Multiplayer Gameplay Programmer

Romero games

Galway, County Galway, Ireland (Hybrid)
8 Months ago
Scopely - Principal Game Server Engineer - Unannounced Project

Scopely

Dublin, County Dublin, Ireland (Hybrid)
3 Months ago
whoop - Talent & Ambassador Marketing Specialist

whoop

Ireland (Remote)
1 Month ago
playrix  - Location Game Designer

playrix

Ireland (Remote)
7 Months ago
Take-Two Interactive - Product Security Architect

Take-Two Interactive

Dublin, County Dublin, Ireland (On-Site)
1 Month ago

Get notifed when new similar jobs are uploaded

Devops Jobs

Looks like we're out of matches

Set up an alert and we'll send you similar jobs the moment they appear!

About The Company

France (Remote)

United States (Remote)

United States (Remote)

Vancouver, British Columbia, Canada (Remote)

India (Remote)

United Kingdom (Remote)

United States (Remote)

New York, New York, United States (On-Site)

View All Jobs

Get notified when new jobs are added by Reddit

Level Up Your Career in Game Development!

Transform Your Passion into Profession with Our Comprehensive Courses for Aspiring Game Developers.

Job Common Plug