Senior Site Reliability Engineer

2 Hours ago • 5 Years +

Job Summary

Job Description

As a Senior Site Reliability Engineer at Reddit, you will improve the reliability and performance of the company's engineering platforms and services using your knowledge of distributed systems and architecture. You will work closely with infrastructure teams to build and maintain tools, focusing on open-source solutions like Prometheus and Grafana. Responsibilities include advising on system design, identifying risks, automating tasks, diagnosing and fixing issues, and optimizing performance for millions of users. You will also be expected to share on-call responsibilities and contribute to open-source projects.
Must have:
  • 5+ years experience in SRE or related roles
  • Proficiency in programming languages like Go or Python
  • Experience with Kubernetes and Cloud systems
  • Familiarity with distributed systems and related tools
  • Experience with high-traffic backend systems
  • Demonstrated ability to debug and optimize code
  • Troubleshooting skills spanning applications and systems
  • Strong working knowledge of Linux and containers
  • Excellent communication and collaborative skills
Good to have:
  • Familiarity with Prometheus, Thanos, Grafana, Vector, Clickhouse, Otel, Loki
Perks:
  • Pension Scheme
  • Private Medical and Dental Scheme
  • Life Assurance, Income Protection
  • Workspace benefit for your home office
  • Personal & Professional development funds
  • Family Planning Support
  • Commuter Benefits
  • Flexible Vacation & Reddit Global Days Off

Job Details

Reddit is a community of communities. It’s built on shared interests, passion, and trust and is home to the most open and authentic conversations on the internet. Every day, Reddit users submit, vote, and comment on the topics they care most about. With 100,000+ active communities and approximately 101M+ daily active unique visitors, Reddit is one of the internet’s largest sources of information. For more information, visit redditinc.com.

Reddit SRE is rapidly innovating and our teams are working to meet the needs of infrastructure and development teams as they evolve our product faster than ever before. This is a unique opportunity to leave your mark on one of the most influential and trafficked corners of the internet.

As a Senior Site Reliability Engineer on Reddit’s Infrastructure SRE team, you’ll use your knowledge of distributed systems and architecture to improve the reliability and performance of Reddit’s engineering platforms and services. We are looking for someone who thrives at the intersection of infrastructure and software development. This team will work very closely with the Compute, Traffic, and Observability infrastructure teams. They will own a suite of tools for allowing engineers to understand their creations, based primarily on open-source solutions at scale. We’re active users of and contributors to Prometheus, Thanos, Grafana, Vector and more.

In this role, you will also take ownership of risk management, ensuring the reliability and performance of our systems. You will collaborate with cross-functional teams to identify, assess, and mitigate risks, implementing best practices to enhance system resilience. Your expertise will drive proactive measures to maintain uptime and optimize service delivery, making a significant impact on our operational excellence.

Join us and help build the future of Reddit!

Responsibilities:

  • Advise
    • Work closely with engineering teams in designing and developing systems that are resilient and highly performant at a tremendous scale, and maintaining the foundational platform for running Reddit’s infrastructure.
  • Amplify
    • Identify and build capabilities into our foundational Infrastructure and Platform services, which are used by Reddit engineering teams to build, deploy, and operate Reddit. 
    • Deliver software to improve the availability, scalability, latency, and efficiency of observability components.
    • Identify and engineer away risk across Reddit’s systems.
  • Automate
    • Take repetitive, manual, or risky tasks and automate them out of existence. Build tools and integrate systems to support Reddit’s evolution.
    • Automate critical aspects of the event driven development process
  • Diagnose
    • Draw on your knowledge of distributed systems to identify and fix network, system, and service-level issues. Practice sustainable incident response, and drive structural improvement with blameless postmortem.
    • Share on-call responsibilities. 
  • Optimize:
    • Observe and improve performance, reduce cost, and improve the experience for millions of users
    • Contribute upstream changes to the open source projects we use

Qualifications

  • 5+ years of experience in Software Engineering, Site Reliability Engineering, or a development-focused DevOps role.
  • Proficiency in one or more programming languages. We’re predominantly writing code in Go and Python.
  • Experience with Kubernetes and Cloud systems.
  • Familiarity with distributed systems development, bonus if familiar with any of the specific tools (Prometheus, Thanos, Grafana, Vector, Clickhouse, Otel, Loki)
  • Experience with the development and operation of high-traffic backend systems.
  • A demonstrated ability to debug, fix, and optimize code.
  • Troubleshooting skills that span applications, networking (TCP/IP), and systems.
  • Strong working knowledge of Linux and containers.
  • Excellent communication and collaborative skills.

Benefits:

  • Pension Scheme
  • Private Medical and Dental Scheme
  • Life Assurance, Income Protection
  • Workspace benefit for your home office 
  • Personal & Professional development funds
  • Family Planning Support 
  • Commuter Benefits
  • Flexible Vacation & Reddit Global Days Off

Reddit is proud to be an equal opportunity employer, and is committed to building a workforce representative of the diverse communities we serve.  Reddit is committed to providing reasonable accommodations for qualified individuals with disabilities and disabled veterans in our job application procedures. If, due to a disability, you need an accommodation during the interview process, please let your recruiter know.

Similar Jobs

Zynga - Principal Software Engineer - Compliance

Zynga

Toronto, Ontario, Canada (On-Site)
2 Weeks ago
Canonical - Product Manager

Canonical

(Remote)
2 Weeks ago
Gitlab - Intermediate Support Engineer (US Federal)

Gitlab

(Remote)
7 Months ago
Canonical - Cloud Professional Services Manager

Canonical

(Remote)
2 Weeks ago
ByteDance - Senior Machine Learning Ops Engineer, ML System - Foundation Model

ByteDance

San Jose, California, United States (On-Site)
3 Months ago

Get notifed when new similar jobs are uploaded

Similar Skill Jobs

Egnyte - Senior Cloud Security Engineer

Egnyte

(Remote)
1 Week ago
Conga - Sr. Software Engineer

Conga

Ahmedabad, Gujarat, India (On-Site)
2 Weeks ago
Paytm - Devops - Senior DevOps Engineer

Paytm

Noida, Uttar Pradesh, India (On-Site)
6 Months ago
NCR Atleos - Software Engineer II

NCR Atleos

Hyderabad, Telangana, India (On-Site)
6 Days ago
DOTSOFT SA - Technical Project Manager & Systems Architect

DOTSOFT SA

Greece (Remote)
1 Month ago
LSEG (London Stock Exchange Group) - DevOps Engineer

LSEG (London Stock Exchange Group)

Bengaluru, Karnataka, India (Hybrid)
7 Months ago
CLO Virtual Fashion  Inc  - DevOps Engineer

CLO Virtual Fashion Inc

Bengaluru, Karnataka, India (On-Site)
7 Months ago
Blackshark - Senior Software Engineer

Blackshark

Graz, Styria, Austria (Hybrid)
1 Week ago
Microsoft - Member of Technical Staff, AI Data

Microsoft

London, England, United Kingdom (On-Site)
1 Month ago
Maersk Careers - Elixir Software Engineer - Energy Transition

Maersk Careers

Porto, Porto District, Portugal (Remote)
5 Months ago

Get notifed when new similar jobs are uploaded

Jobs in London, England, United Kingdom

Glowmade - Tools Programmer

Glowmade

England, United Kingdom (On-Site)
1 Month ago
Tencent - Technical Director (Games)

Tencent

London, England, United Kingdom (On-Site)
2 Months ago
NBC universal - Director of Product, Relationship Management & Localization

NBC universal

Brentford, England, United Kingdom (On-Site)
3 Days ago
AppZen - Solutions Consultant

AppZen

London, England, United Kingdom (Hybrid)
4 Months ago
Epic Games - Product Director, LiveOps

Epic Games

London, England, United Kingdom (On-Site)
3 Weeks ago
Corsair - Channel Marketing Assistant

Corsair

Wokingham, England, United Kingdom (On-Site)
1 Month ago
Universally Speaking - Brazilian Portuguese Games Tester

Universally Speaking

England, United Kingdom (On-Site)
1 Month ago
Cloud Imperium Games - Game Designer (Vehicle Specialist)

Cloud Imperium Games

Manchester, England, United Kingdom (On-Site)
1 Month ago
TiMi Studio Group - TiMi Europe- Senior business development manager

TiMi Studio Group

London, England, United Kingdom (On-Site)
6 Months ago
Maverick Games - Technical Animator

Maverick Games

Warwick, England, United Kingdom (Hybrid)
3 Months ago

Get notifed when new similar jobs are uploaded

Similar Category Jobs

Looks like we're out of matches

Set up an alert and we'll send you similar jobs the moment they appear!