Senior Site Reliability Engineer

17 Hours ago • 5 Years + • Devops

Job Summary

Job Description

Reddit SRE is rapidly innovating to meet the evolving needs of infrastructure and development teams. As a Senior Site Reliability Engineer on the Infrastructure SRE team, you will leverage your knowledge of distributed systems and architecture to enhance the reliability and performance of Reddit's engineering platforms and services. This role involves working at the intersection of infrastructure and software development, collaborating closely with Compute, Traffic, and Observability infrastructure teams. You will own a suite of tools for engineers to understand their creations, primarily using open-source solutions at scale, and actively contribute to projects like Prometheus, Thanos, Grafana, and Vector. Additionally, you will take ownership of risk management, ensuring system reliability and performance, and collaborating with cross-functional teams to mitigate risks and implement best practices for system resilience, driving proactive measures for uptime and service delivery optimization.
Must have:
  • 5+ years of experience in Software Engineering, Site Reliability Engineering, or DevOps.
  • Proficiency in Go and/or Python.
  • Experience with Kubernetes and Cloud systems.
  • Experience with development and operation of high-traffic backend systems.
  • Ability to debug, fix, and optimize code.
  • Troubleshooting skills across applications, networking (TCP/IP), and systems.
  • Strong working knowledge of Linux and containers.
  • Excellent communication and collaborative skills.
Good to have:
  • Familiarity with distributed systems development.
  • Familiarity with Prometheus, Thanos, Grafana, Vector, Clickhouse, Otel, Loki.
Perks:
  • Retirement Savings plan
  • Workspace benefits for your home office
  • Personal & Professional development funds
  • Family Planning Support
  • Flexible Vacation & Reddit Global Days Off

Job Details

Reddit is a community of communities. It’s built on shared interests, passion, and trust and is home to the most open and authentic conversations on the internet. Every day, Reddit users submit, vote, and comment on the topics they care most about. With 100,000+ active communities and approximately 101M+ daily active unique visitors, Reddit is one of the internet’s largest sources of information. For more information, visit redditinc.com.

Reddit SRE is rapidly innovating and our teams are working to meet the needs of infrastructure and development teams as they evolve our product faster than ever before. This is a unique opportunity to leave your mark on one of the most influential and trafficked corners of the internet.

As a Senior Site Reliability Engineer on Reddit’s Infrastructure SRE team, you’ll use your knowledge of distributed systems and architecture to improve the reliability and performance of Reddit’s engineering platforms and services. We are looking for someone who thrives at the intersection of infrastructure and software development. This team will work very closely with the Compute, Traffic, and Observability infrastructure teams. They will own a suite of tools for allowing engineers to understand their creations, based primarily on open-source solutions at scale. We’re active users of and contributors to Prometheus, Thanos, Grafana, Vector and more.

In this role, you will also take ownership of risk management, ensuring the reliability and performance of our systems. You will collaborate with cross-functional teams to identify, assess, and mitigate risks, implementing best practices to enhance system resilience. Your expertise will drive proactive measures to maintain uptime and optimize service delivery, making a significant impact on our operational excellence.

Join us and help build the future of Reddit!

Responsibilities:

  • Advise
    • Work closely with engineering teams in designing and developing systems that are resilient and highly performant at a tremendous scale, and maintaining the foundational platform for running Reddit’s infrastructure.
  • Amplify
    • Identify and build capabilities into our foundational Infrastructure and Platform services, which are used by Reddit engineering teams to build, deploy, and operate Reddit. 
    • Deliver software to improve the availability, scalability, latency, and efficiency of observability components.
    • Identify and engineer away risk across Reddit’s systems.
  • Automate
    • Take repetitive, manual, or risky tasks and automate them out of existence. Build tools and integrate systems to support Reddit’s evolution.
    • Automate critical aspects of the event driven development process
  • Diagnose
    • Draw on your knowledge of distributed systems to identify and fix network, system, and service-level issues. Practice sustainable incident response, and drive structural improvement with blameless postmortem.
    • Share on-call responsibilities. 
  • Optimize:
    • Observe and improve performance, reduce cost, and improve the experience for millions of users
    • Contribute upstream changes to the open source projects we use

Qualifications

  • 5+ years of experience in Software Engineering, Site Reliability Engineering, or a development-focused DevOps role.
  • Proficiency in one or more programming languages. We’re predominantly writing code in Go and Python.
  • Experience with Kubernetes and Cloud systems.
  • Familiarity with distributed systems development, bonus if familiar with any of the specific tools (Prometheus, Thanos, Grafana, Vector, Clickhouse, Otel, Loki)
  • Experience with the development and operation of high-traffic backend systems.
  • A demonstrated ability to debug, fix, and optimize code.
  • Troubleshooting skills that span applications, networking (TCP/IP), and systems.
  • Strong working knowledge of Linux and containers.
  • Excellent communication and collaborative skills.

Benefits:

  • Retirement Savings plan 
  • Workspace benefits for your home office 
  • Personal & Professional development funds
  • Family Planning Support 
  • Flexible Vacation & Reddit Global Days Off

Reddit is proud to be an equal opportunity employer, and is committed to building a workforce representative of the diverse communities we serve.  Reddit is committed to providing reasonable accommodations for qualified individuals with disabilities and disabled veterans in our job application procedures. If, due to a disability, you need an accommodation during the interview process, please let your recruiter know.

Similar Jobs

PayPal - Senior Quality Assurance Engineer

PayPal

Chennai, Tamil Nadu, India (Hybrid)
4 Weeks ago
Palo Alto Networks - Sr. Technical Support Engineer, Focused Services, XSIAM

Palo Alto Networks

Plano, Texas, United States (On-Site)
3 Weeks ago
Bosch Group - Electronic Technician Supervisor - D Shift

Bosch Group

Roseville, California, United States (On-Site)
2 Weeks ago
Addepar - Lead Product Designer

Addepar

Pune, Maharashtra, India (On-Site)
1 Month ago
Scout - Engineer, Thermal Systems

Scout

Novi, Michigan, United States (On-Site)
1 Month ago
Canonical - Software Architect - Containers / Virtualisation

Canonical

(Remote)
1 Month ago
playrix  - Senior C++ Software Engineer (Build System)

playrix

Montenegro (Remote)
7 Months ago
miniclip - Senior Cloud Engineer

miniclip

Lisbon, Lisbon, Portugal (On-Site)
1 Month ago
Insight Software - Solution Architect

Insight Software

Berlin, Berlin, Germany (Remote)
3 Months ago
London stock Exchange - Platform Principal Engineer

London stock Exchange

New York, United States (On-Site)
1 Month ago

Get notifed when new similar jobs are uploaded

Similar Skill Jobs

Apple - Sports Business Optimization & League Relations

Apple

New York, New York, United States (On-Site)
1 Month ago
Glean - Deal Desk Manager

Glean

San Francisco, California, United States (Hybrid)
1 Month ago
Hasbro - Lead Narrative Designer

Hasbro

Durham, North Carolina, United States (On-Site)
3 Weeks ago
Reddit - Senior iOS Software Engineer

Reddit

San Francisco, California, United States (On-Site)
1 Month ago
Nagarro - Associate Principal Engineer, Delivery

Nagarro

Colombia (Remote)
8 Months ago
Lingo Kids LLC - Senior CRM Specialist

Lingo Kids LLC

Madrid, Community Of Madrid, Spain (Remote)
5 Months ago
WebMD - Medical Advisor

WebMD

Konstanz, Baden-Württemberg, Germany (Hybrid)
1 Month ago
Scopely - Senior SQL Developer (Oracle)

Scopely

Mexico City, Mexico (On-Site)
1 Week ago
Yodo1 - Game Marketing Lead

Yodo1

(Remote)
5 Months ago
Zenoti - Manager - Product Support

Zenoti

Hyderabad, Telangana, India (On-Site)
1 Month ago

Get notifed when new similar jobs are uploaded

Jobs in Berlin, Berlin, Germany

caliogo - FAST Scheduling & Content Planner

caliogo

Berlin, Berlin, Germany (Hybrid)
1 Month ago
Zscaler - Account Executive, Enterprise

Zscaler

Frankfurt Am Main, Hessen, Germany (Hybrid)
1 Month ago
Tesla - Senior Cell Mechanical Design Engineer

Tesla

Brandenburg, Germany (On-Site)
4 Months ago
Plaion - Working Student Acquisitions (all genders)

Plaion

Munich, Bavaria, Germany (Hybrid)
1 Month ago
Corsair gaming - Software Engineer, Stream Deck

Corsair gaming

Munich, Bavaria, Germany (On-Site)
2 Weeks ago
Tesla - Mechanical Design Engineer - Exterior Engineering

Tesla

Berlin, Berlin, Germany (On-Site)
4 Months ago
Altagram Group - Senior French Language Specialist

Altagram Group

Berlin, Berlin, Germany (On-Site)
1 Month ago
Plaion - Intern/YouTube Channel Management (m/f/d) - Japanese

Plaion

Berlin, Berlin, Germany (On-Site)
1 Month ago
Sunday games - Marketing Artist

Sunday games

Hamburg, Hamburg, Germany (Hybrid)
7 Months ago
Globaloc studios - Localization Project Manager – Audio Productions

Globaloc studios

Berlin, Berlin, Germany (On-Site)
2 Weeks ago

Get notifed when new similar jobs are uploaded

Devops Jobs

Google - Software Engineer III, Full Stack, Google Cloud Business Platforms

Google

Kirkland, Washington, United States (On-Site)
2 Months ago
Rackspace Technology - Cloud Engineer IV (Java Dev Google Cloud Practice Engineer)

Rackspace Technology

Gurugram, Haryana, India (Remote)
3 Months ago
Homa Games - DevOps / SRE

Homa Games

Paris, Île-de-France, France (Hybrid)
1 Month ago
Spaulding Ridge - OneStream Solution Architect

Spaulding Ridge

Chicago, Illinois, United States (On-Site)
2 Months ago
Google - Software Engineer III, Full Stack, Google Cloud Business Platforms

Google

Sunnyvale, California, United States (On-Site)
2 Months ago
Trend Micro - (Sr.) Cloud Backend Engineer

Trend Micro

Taipei City, Taiwan (On-Site)
9 Months ago
Notion - Solutions Engineer

Notion

Seoul, South Korea (On-Site)
1 Month ago
Aristocrat - DevOps Engineer

Aristocrat

Kraków, Lesser Poland Voivodeship, Poland (Hybrid)
1 Month ago
Capgemini - Site Reliability Engineer-Wintel, Linux, Vmware, Redhat Devops CI/CD AWS

Capgemini

Bengaluru, Karnataka, India (On-Site)
1 Month ago
Reltio - Intern - DevOps

Reltio

Lisbon, Lisbon, Portugal (On-Site)
1 Month ago

Get notifed when new similar jobs are uploaded