Senior Site Reliability Engineer

17 Hours ago • 5 Years + • Devops

Job Summary

Job Description

Reddit SRE is seeking a Senior Site Reliability Engineer to innovate and improve the reliability and performance of Reddit's infrastructure and services. This role involves working at the intersection of infrastructure and software development, focusing on compute, traffic, and observability teams. You will own tools for engineers to understand their creations, utilizing open-source solutions like Prometheus, Thanos, Grafana, and Vector. Key responsibilities include risk management, collaborating with teams to mitigate risks, enhancing system resilience, and driving proactive measures for uptime and service delivery. The ideal candidate will contribute to building the future of Reddit.
Must have:
  • 5+ years of experience in SRE/DevOps
  • Proficiency in Go or Python
  • Experience with Kubernetes and Cloud systems
  • Experience with distributed systems development
  • Debug, fix, and optimize code
  • Troubleshoot applications, networking, and systems
  • Strong Linux and container knowledge
  • Excellent communication and collaboration skills
Good to have:
  • Familiarity with Prometheus, Thanos, Grafana, Vector, Clickhouse, Otel, Loki
  • Experience with high-traffic backend systems
Perks:
  • Pension Scheme
  • Private Medical and Dental Scheme
  • Life Assurance
  • Income Protection
  • Workspace benefit for home office
  • Personal & Professional development funds
  • Family Planning Support
  • Commuter Benefits
  • Flexible Vacation & Reddit Global Days Off

Job Details

Reddit is a community of communities. It’s built on shared interests, passion, and trust and is home to the most open and authentic conversations on the internet. Every day, Reddit users submit, vote, and comment on the topics they care most about. With 100,000+ active communities and approximately 101M+ daily active unique visitors, Reddit is one of the internet’s largest sources of information. For more information, visit redditinc.com.

Reddit SRE is rapidly innovating and our teams are working to meet the needs of infrastructure and development teams as they evolve our product faster than ever before. This is a unique opportunity to leave your mark on one of the most influential and trafficked corners of the internet.

As a Senior Site Reliability Engineer on Reddit’s Infrastructure SRE team, you’ll use your knowledge of distributed systems and architecture to improve the reliability and performance of Reddit’s engineering platforms and services. We are looking for someone who thrives at the intersection of infrastructure and software development. This team will work very closely with the Compute, Traffic, and Observability infrastructure teams. They will own a suite of tools for allowing engineers to understand their creations, based primarily on open-source solutions at scale. We’re active users of and contributors to Prometheus, Thanos, Grafana, Vector and more.

In this role, you will also take ownership of risk management, ensuring the reliability and performance of our systems. You will collaborate with cross-functional teams to identify, assess, and mitigate risks, implementing best practices to enhance system resilience. Your expertise will drive proactive measures to maintain uptime and optimize service delivery, making a significant impact on our operational excellence.

Join us and help build the future of Reddit!

Responsibilities:

  • Advise
    • Work closely with engineering teams in designing and developing systems that are resilient and highly performant at a tremendous scale, and maintaining the foundational platform for running Reddit’s infrastructure.
  • Amplify
    • Identify and build capabilities into our foundational Infrastructure and Platform services, which are used by Reddit engineering teams to build, deploy, and operate Reddit. 
    • Deliver software to improve the availability, scalability, latency, and efficiency of observability components.
    • Identify and engineer away risk across Reddit’s systems.
  • Automate
    • Take repetitive, manual, or risky tasks and automate them out of existence. Build tools and integrate systems to support Reddit’s evolution.
    • Automate critical aspects of the event driven development process
  • Diagnose
    • Draw on your knowledge of distributed systems to identify and fix network, system, and service-level issues. Practice sustainable incident response, and drive structural improvement with blameless postmortem.
    • Share on-call responsibilities. 
  • Optimize:
    • Observe and improve performance, reduce cost, and improve the experience for millions of users
    • Contribute upstream changes to the open source projects we use

Qualifications

  • 5+ years of experience in Software Engineering, Site Reliability Engineering, or a development-focused DevOps role.
  • Proficiency in one or more programming languages. We’re predominantly writing code in Go and Python.
  • Experience with Kubernetes and Cloud systems.
  • Familiarity with distributed systems development, bonus if familiar with any of the specific tools (Prometheus, Thanos, Grafana, Vector, Clickhouse, Otel, Loki)
  • Experience with the development and operation of high-traffic backend systems.
  • A demonstrated ability to debug, fix, and optimize code.
  • Troubleshooting skills that span applications, networking (TCP/IP), and systems.
  • Strong working knowledge of Linux and containers.
  • Excellent communication and collaborative skills.

Benefits:

  • Pension Scheme
  • Private Medical and Dental Scheme
  • Life Assurance, Income Protection
  • Workspace benefit for your home office 
  • Personal & Professional development funds
  • Family Planning Support 
  • Commuter Benefits
  • Flexible Vacation & Reddit Global Days Off

Reddit is proud to be an equal opportunity employer, and is committed to building a workforce representative of the diverse communities we serve.  Reddit is committed to providing reasonable accommodations for qualified individuals with disabilities and disabled veterans in our job application procedures. If, due to a disability, you need an accommodation during the interview process, please let your recruiter know.

Similar Jobs

Qualcomm - Senior IP Physical Design Engineer

Qualcomm

Noida, Uttar Pradesh, India (On-Site)
1 Month ago
Regent craft - Senior Perception Software Engineer - Sensor Fusion

Regent craft

North Kingstown, Rhode Island, United States (On-Site)
2 Weeks ago
Snyk - Senior Software Engineer

Snyk

Tel Aviv-Yafo, Tel Aviv District, Israel (On-Site)
1 Month ago
Trend Micro - UX Designer

Trend Micro

Ottawa, Ontario, Canada (On-Site)
1 Month ago
Springer Group - Enterprise Architect

Springer Group

Berlin, Berlin, Germany (Remote)
1 Month ago
Zenoti - Lead Site Reliability Engineer - DBA

Zenoti

Hyderabad, Telangana, India (On-Site)
2 Months ago
Colo pl - Kubernetes Engineer

Colo pl

Tokyo, Japan (On-Site)
1 Month ago
Capgemini - NodeJS Developer with DevOps

Capgemini

Gurugram, Haryana, India (On-Site)
1 Month ago
Tencent - Principal Cloud Solution Architect

Tencent

California, United States (On-Site)
3 Months ago
Cadence - Lead Full Stack Cloud Engineer

Cadence

Noida, Uttar Pradesh, India (On-Site)
9 Months ago

Get notifed when new similar jobs are uploaded

Similar Skill Jobs

Qualcomm - Staff Camera SW Customer Engineer

Qualcomm

Shanghai, China (On-Site)
4 Weeks ago
Ethernovia - BSP Embedded S/W

Ethernovia

Pune, Maharashtra, India (On-Site)
1 Month ago
Yahoo - Account Director

Yahoo

Toronto, Ontario, Canada (Hybrid)
4 Weeks ago
Integrant - Principal iOS Developer

Integrant

Cairo, Cairo Governorate, Egypt (Hybrid)
2 Months ago
hogarth - Content Project Manager

hogarth

Singapore (On-Site)
1 Month ago
Illumina - Senior Accounting Analyst II - LATAM Accounting

Illumina

State Of São Paulo, Brazil (Hybrid)
1 Week ago
WebMD - Audio Visual Technician (m/w/d)

WebMD

United Kingdom (On-Site)
8 Months ago
Paytm - Technical Product Management - Senior Product Manager - Telco

Paytm

Noida, Uttar Pradesh, India (On-Site)
2 Weeks ago
Capgemini - Business Analyst (BFSI)

Capgemini

Bengaluru, Karnataka, India (On-Site)
4 Weeks ago

Get notifed when new similar jobs are uploaded

Jobs in London, England, United Kingdom

Rebellion - Graduate Build Engineer

Rebellion

Oxford, England, United Kingdom (On-Site)
1 Month ago
Ion - Senior Security Architect

Ion

London, England, United Kingdom (On-Site)
8 Months ago
Haleon - Dental Territory Manager

Haleon

United Kingdom (On-Site)
3 Weeks ago
Rocket Science - Software Engineer - Unreal

Rocket Science

Brighton And Hove, England, United Kingdom (Hybrid)
3 Months ago
SSC Technologies - Senior Product Manager, Mobile

SSC Technologies

London, England, United Kingdom (Hybrid)
1 Month ago
Diligent Corporation - Director, Product Management, AI

Diligent Corporation

London, England, United Kingdom (On-Site)
1 Month ago
Universally Speaking - Czech Games Tester

Universally Speaking

England, United Kingdom (On-Site)
3 Months ago
Tesla - Security Operations Center (SOC) Operator

Tesla

Milton Keynes, England, United Kingdom (On-Site)
4 Months ago
oni - Stock Controller

oni

Oxford, England, United Kingdom (On-Site)
2 Months ago
Perplexity - UK Internship Program

Perplexity

London, England, United Kingdom (Hybrid)
1 Month ago

Get notifed when new similar jobs are uploaded

Devops Jobs

InMobiInMobi - SDE III - Devops

InMobiInMobi

Bengaluru, Karnataka, India (On-Site)
2 Months ago
10 Chambers - Senior Build Engineer

10 Chambers

Stockholm, Stockholm County, Sweden (On-Site)
2 Weeks ago
Sonar Source - Senior Cloud Solutions Engineer

Sonar Source

Austin, Texas, United States (On-Site)
1 Month ago
Apple - Cloud Traffic Engineer, Apple Pay

Apple

New York, New York, United States (On-Site)
2 Weeks ago
Ion - Software Architect - Java Multi-Tenant SAAS Cloud Native

Ion

Pune, Maharashtra, India (On-Site)
8 Months ago
Devoteam - Software Architect

Devoteam

Frankfurt Am Main, Hessen, Germany (On-Site)
2 Weeks ago
Reddit - Principal Software Engineer, ML Feature Platform

Reddit

United States (Remote)
1 Month ago
Zelis  - Senior Azure Automation Engineer

Zelis

(Remote)
1 Month ago
bytedance - Senior Software Engineer - Development Infrastructure Team

bytedance

Mountain View, California, United States (On-Site)
8 Months ago
Unity - DevOps Tech Lead

Unity

Tel Aviv-Yafo, Tel Aviv District, Israel (On-Site)
2 Months ago

Get notifed when new similar jobs are uploaded