Senior Site Reliability Engineer

9 Hours ago • 5 Years + • Devops

Job Summary

Job Description

Reddit is seeking a Senior Site Reliability Engineer to join their Infrastructure SRE team. This role involves improving the reliability and performance of Reddit's engineering platforms and services by leveraging knowledge of distributed systems and architecture. The engineer will focus on tools for enabling engineers to understand their creations, utilizing open-source solutions like Prometheus, Thanos, Grafana, and Vector. Key responsibilities include advising on resilient system design, amplifying infrastructure capabilities, automating repetitive tasks, diagnosing and resolving system issues, and optimizing performance and cost. The role also involves risk management and collaboration with cross-functional teams to maintain system resilience and uptime.
Must have:
  • 5+ years of experience in SRE/DevOps
  • Proficiency in Go or Python
  • Experience with Kubernetes and Cloud systems
  • Troubleshooting skills (applications, networking, systems)
  • Strong Linux and container knowledge
  • Excellent communication and collaboration skills
Good to have:
  • Familiarity with distributed systems
  • Experience with Prometheus, Thanos, Grafana, Vector, Clickhouse, Otel, Loki
  • Experience with high-traffic backend systems
  • Ability to debug, fix, and optimize code
  • Contribute upstream changes to open-source projects
Perks:
  • Private Medical, Dental and Vision Benefits
  • Retirement Savings plan with matching contributions
  • Workspace benefits for your home office
  • Personal & Professional development funds
  • Family Planning Support
  • Commuter Benefits
  • Flexible Vacation & Reddit Global Days Off

Job Details

Back to jobs

Senior Site Reliability Engineer

Dublin, Ireland
Reddit is a community of communities. It’s built on shared interests, passion, and trust and is home to the most open and authentic conversations on the internet. Every day, Reddit users submit, vote, and comment on the topics they care most about. With 100,000+ active communities and approximately 101M+ daily active unique visitors, Reddit is one of the internet’s largest sources of information. For more information, visit redditinc.com.

Reddit SRE is rapidly innovating and our teams are working to meet the needs of infrastructure and development teams as they evolve our product faster than ever before. This is a unique opportunity to leave your mark on one of the most influential and trafficked corners of the internet.

As a Senior Site Reliability Engineer on Reddit’s Infrastructure SRE team, you’ll use your knowledge of distributed systems and architecture to improve the reliability and performance of Reddit’s engineering platforms and services. We are looking for someone who thrives at the intersection of infrastructure and software development. This team will work very closely with the Compute, Traffic, and Observability infrastructure teams. They will own a suite of tools for allowing engineers to understand their creations, based primarily on open-source solutions at scale. We’re active users of and contributors to Prometheus, Thanos, Grafana, Vector and more.

In this role, you will also take ownership of risk management, ensuring the reliability and performance of our systems. You will collaborate with cross-functional teams to identify, assess, and mitigate risks, implementing best practices to enhance system resilience. Your expertise will drive proactive measures to maintain uptime and optimize service delivery, making a significant impact on our operational excellence.

Join us and help build the future of Reddit!

Responsibilities:

  • Advise
    • Work closely with engineering teams in designing and developing systems that are resilient and highly performant at a tremendous scale, and maintaining the foundational platform for running Reddit’s infrastructure.
  • Amplify
    • Identify and build capabilities into our foundational Infrastructure and Platform services, which are used by Reddit engineering teams to build, deploy, and operate Reddit. 
    • Deliver software to improve the availability, scalability, latency, and efficiency of observability components.
    • Identify and engineer away risk across Reddit’s systems.
  • Automate
    • Take repetitive, manual, or risky tasks and automate them out of existence. Build tools and integrate systems to support Reddit’s evolution.
    • Automate critical aspects of the event driven development process
  • Diagnose
    • Draw on your knowledge of distributed systems to identify and fix network, system, and service-level issues. Practice sustainable incident response, and drive structural improvement with blameless postmortem.
    • Share on-call responsibilities. 
  • Optimize:
    • Observe and improve performance, reduce cost, and improve the experience for millions of users
    • Contribute upstream changes to the open source projects we use

Qualifications

  • 5+ years of experience in Software Engineering, Site Reliability Engineering, or a development-focused DevOps role.
  • Proficiency in one or more programming languages. We’re predominantly writing code in Go and Python.
  • Experience with Kubernetes and Cloud systems.
  • Familiarity with distributed systems development, bonus if familiar with any of the specific tools (Prometheus, Thanos, Grafana, Vector, Clickhouse, Otel, Loki)
  • Experience with the development and operation of high-traffic backend systems.
  • A demonstrated ability to debug, fix, and optimize code.
  • Troubleshooting skills that span applications, networking (TCP/IP), and systems.
  • Strong working knowledge of Linux and containers.
  • Excellent communication and collaborative skills.

Benefits:

  • Private Medical, Dental and Vision Benefits 
  • Retirement Savings plan with matching contributions
  • Workspace benefits for your home office
  • Personal & Professional development funds
  • Family Planning Support
  • Commuter Benefits  
  • Flexible Vacation & Reddit Global Days Off

Reddit is proud to be an equal opportunity employer, and is committed to building a workforce representative of the diverse communities we serve.  Reddit is committed to providing reasonable accommodations for qualified individuals with disabilities and disabled veterans in our job application procedures. If, due to a disability, you need an accommodation during the interview process, please let your recruiter know.

Apply for this job

*

indicates a required field

Resume/CV*

Accepted file types: pdf, doc, docx, txt, rtf

Cover Letter

Accepted file types: pdf, doc, docx, txt, rtf


Select...
Select...
Select...

Similar Jobs

Illumina - Customer Care Manager, Global Commercial Operations

Illumina

State Of São Paulo, Brazil (On-Site)
2 Months ago
Riot Games - Manager, Accounting

Riot Games

Seoul, South Korea (On-Site)
1 Month ago
Qualcomm - Premium Modem - Senior Physical Design Engineer

Qualcomm

Bengaluru, Karnataka, India (On-Site)
2 Weeks ago
Wolters Kluwer - Consulting Manager

Wolters Kluwer

Beijing, China (Hybrid)
1 Month ago
Beta Craft - Python Developer

Beta Craft

Pune, Maharashtra, India (Remote)
7 Months ago
Apple - Cloud Infrastructure Software Developer

Apple

Seattle, Washington, United States (On-Site)
1 Month ago
Sword Health - Senior Site Reliability Engineer (SRE)

Sword Health

Porto, Porto District, Portugal (Hybrid)
11 Months ago
Enphase Energy - Sr. Staff Engineer Cloud

Enphase Energy

Bengaluru, Karnataka, India (On-Site)
6 Months ago
Workato - Senior Software Engineer (Platform, Ruby)

Workato

Lisbon, Lisbon, Portugal (On-Site)
1 Month ago
luxsoft - Solution Architect

luxsoft

Egypt (Remote)
1 Week ago

Get notifed when new similar jobs are uploaded

Similar Skill Jobs

Google - Account Strategist, Engage, Google Customer Solutions

Google

Hong Kong (On-Site)
2 Months ago
Saxo Bank - Head of Client Funding, Reporting & Referrals Experience

Saxo Bank

Tokyo, Japan (On-Site)
2 Months ago
Penumbrainc - Therapy Development Specialist

Penumbrainc

Oklahoma City, Oklahoma, United States (Remote)
2 Months ago
Ansys - Senior Field Marketing Specialist

Ansys

Milan, Lombardy, Italy (Hybrid)
1 Month ago
Coupa - Sr. Manager, Internal Audit

Coupa

Ann Arbor, Michigan, United States (Remote)
3 Weeks ago
Paytm - Technical Program Manager

Paytm

Bengaluru, Karnataka, India (On-Site)
2 Weeks ago
Niantic - Software Engineer, Server

Niantic

Tokyo, Japan (Hybrid)
2 Months ago
Like Card - Customer Service Senior Supervisor – Chat, Social Media & Call Center

Like Card

Istanbul, İstanbul, Türkiye (On-Site)
1 Month ago
Ubisoft - Game Designer

Ubisoft

Pune, Maharashtra, India (On-Site)
4 Months ago
NXP - Principal System Application Engineer

NXP

San Jose, California, United States (On-Site)
1 Month ago

Get notifed when new similar jobs are uploaded

Jobs in Dublin, County Dublin, Ireland

playrix  - Senior QA Engineer (VSO Engine)

playrix

Ireland (Remote)
3 Months ago
Google - Customer Growth Associate

Google

Dublin, County Dublin, Ireland (On-Site)
2 Months ago
Globalization Partners - Senior Product Designer – Mobile (AI Native App)

Globalization Partners

Ireland (Remote)
1 Month ago
Coupa - Operations Analyst, Partner Management

Coupa

Dublin, County Dublin, Ireland (Hybrid)
2 Months ago
version 1 - Business Systems Analyst

version 1

Dublin, County Dublin, Ireland (Hybrid)
1 Month ago
Trellix - Site Reliability Engineer

Trellix

Cork, County Cork, Ireland (On-Site)
1 Month ago
Riot Games - Senior Principal Technical Artist

Riot Games

Dublin, County Dublin, Ireland (On-Site)
7 Months ago
playrix  - Senior Playable Ads Developer

playrix

Ireland (Remote)
3 Months ago
Forcepoint - Sr. Mobile Developer

Forcepoint

Cork, County Cork, Ireland (On-Site)
3 Weeks ago
Cadence - DRC, LVS, 3DIC PV Solutions Engineer

Cadence

Cork, County Cork, Ireland (On-Site)
2 Months ago

Get notifed when new similar jobs are uploaded

Devops Jobs

Glean - Solutions Engineer

Glean

Palo Alto, California, United States (On-Site)
1 Month ago
Adyen - Platform Monitoring Engineer (Incident Management)

Adyen

Bengaluru, Karnataka, India (On-Site)
1 Month ago
Qualcomm - Engineer, Staff -Devops

Qualcomm

Hyderabad, Telangana, India (On-Site)
2 Months ago
Intel  - Senior Infrastructure Engineer - Virtualization and Cloud Platforms

Intel

Phoenix, Arizona, United States (On-Site)
2 Weeks ago
PwC - SAP Lead Solution/ Enterprise Architect - NCR region

PwC

Bengaluru, Karnataka, India (On-Site)
9 Months ago
Flexera - Member Technical Staff - Site Reliability Engineer

Flexera

Bengaluru, Karnataka, India (Hybrid)
9 Months ago
C3 IoT - AI Solution Architect/Senior AI Solution Architect (Post-Sales)

C3 IoT

Redwood City, California, United States (On-Site)
4 Weeks ago
Google - Senior Staff Software Engineer, Google Cloud Compute

Google

Seattle, Washington, United States (On-Site)
2 Months ago
BigID - Site Reliability Engineer

BigID

Buenos Aires, Buenos Aires, Argentina (Remote)
2 Weeks ago
luxsoft - Solution Architect

luxsoft

Germany (Remote)
1 Week ago

Get notifed when new similar jobs are uploaded