Senior Site Reliability Engineer

17 Hours ago • 5 Years + • Devops

Job Summary

Job Description

Reddit is seeking a Senior Site Reliability Engineer to join their Infrastructure SRE team. The role involves improving the reliability and performance of Reddit's engineering platforms and services by leveraging knowledge of distributed systems and architecture. Responsibilities include advising engineering teams on system design, amplifying capabilities of infrastructure and platform services, automating repetitive tasks, diagnosing and fixing system issues, and optimizing performance and cost. The engineer will also own risk management, ensuring system resilience and implementing best practices. This position offers an opportunity to impact one of the internet's largest sources of information.
Must have:
  • 5+ years of experience in SRE or DevOps
  • Proficiency in Go or Python
  • Experience with Kubernetes and Cloud systems
  • Knowledge of distributed systems
  • Experience debugging and optimizing code
  • Troubleshooting skills (applications, networking, systems)
  • Strong Linux and container knowledge
  • Excellent communication and collaboration skills
Good to have:
  • Familiarity with Prometheus, Thanos, Grafana, Vector, Clickhouse, Otel, Loki
  • Experience with high-traffic backend systems
Perks:
  • Pension Savings plan
  • Medical Plan
  • Short term sickness benefits
  • WIA excess and WGA gap insurance
  • Workspace benefits for your home office
  • Personal & Professional development funds
  • Family Planning Support
  • Flexible Vacation & Reddit Global Days Off

Job Details

Reddit is a community of communities. It’s built on shared interests, passion, and trust and is home to the most open and authentic conversations on the internet. Every day, Reddit users submit, vote, and comment on the topics they care most about. With 100,000+ active communities and approximately 101M+ daily active unique visitors, Reddit is one of the internet’s largest sources of information. For more information, visit redditinc.com.

Reddit SRE is rapidly innovating and our teams are working to meet the needs of infrastructure and development teams as they evolve our product faster than ever before. This is a unique opportunity to leave your mark on one of the most influential and trafficked corners of the internet.

As a Senior Site Reliability Engineer on Reddit’s Infrastructure SRE team, you’ll use your knowledge of distributed systems and architecture to improve the reliability and performance of Reddit’s engineering platforms and services. We are looking for someone who thrives at the intersection of infrastructure and software development. This team will work very closely with the Compute, Traffic, and Observability infrastructure teams. They will own a suite of tools for allowing engineers to understand their creations, based primarily on open-source solutions at scale. We’re active users of and contributors to Prometheus, Thanos, Grafana, Vector and more.

In this role, you will also take ownership of risk management, ensuring the reliability and performance of our systems. You will collaborate with cross-functional teams to identify, assess, and mitigate risks, implementing best practices to enhance system resilience. Your expertise will drive proactive measures to maintain uptime and optimize service delivery, making a significant impact on our operational excellence.

Join us and help build the future of Reddit!

Responsibilities:

  • Advise
    • Work closely with engineering teams in designing and developing systems that are resilient and highly performant at a tremendous scale, and maintaining the foundational platform for running Reddit’s infrastructure.
  • Amplify
    • Identify and build capabilities into our foundational Infrastructure and Platform services, which are used by Reddit engineering teams to build, deploy, and operate Reddit. 
    • Deliver software to improve the availability, scalability, latency, and efficiency of observability components.
    • Identify and engineer away risk across Reddit’s systems.
  • Automate
    • Take repetitive, manual, or risky tasks and automate them out of existence. Build tools and integrate systems to support Reddit’s evolution.
    • Automate critical aspects of the event driven development process
  • Diagnose
    • Draw on your knowledge of distributed systems to identify and fix network, system, and service-level issues. Practice sustainable incident response, and drive structural improvement with blameless postmortem.
    • Share on-call responsibilities. 
  • Optimize:
    • Observe and improve performance, reduce cost, and improve the experience for millions of users
    • Contribute upstream changes to the open source projects we use

Qualifications

  • 5+ years of experience in Software Engineering, Site Reliability Engineering, or a development-focused DevOps role.
  • Proficiency in one or more programming languages. We’re predominantly writing code in Go and Python.
  • Experience with Kubernetes and Cloud systems.
  • Familiarity with distributed systems development, bonus if familiar with any of the specific tools (Prometheus, Thanos, Grafana, Vector, Clickhouse, Otel, Loki)
  • Experience with the development and operation of high-traffic backend systems.
  • A demonstrated ability to debug, fix, and optimize code.
  • Troubleshooting skills that span applications, networking (TCP/IP), and systems.
  • Strong working knowledge of Linux and containers.
  • Excellent communication and collaborative skills.

Benefits:

  • Pension Savings plan 
  • Medical Plan
  • Short term sickness benefits 
  • WIA excess and WGA gap insurance 
  • Workspace benefits for your home office 
  • Personal & Professional development funds
  • Family Planning Support 
  • Flexible Vacation & Reddit Global Days Off

Reddit is proud to be an equal opportunity employer, and is committed to building a workforce representative of the diverse communities we serve.  Reddit is committed to providing reasonable accommodations for qualified individuals with disabilities and disabled veterans in our job application procedures. If, due to a disability, you need an accommodation during the interview process, please let your recruiter know.

Similar Jobs

Digital agents - XR Developer

Digital agents

Gurugram, India (On-Site)
1 Year ago
Loft Orbital - Designer

Loft Orbital

Toulouse, Occitanie, France (On-Site)
2 Months ago
warner bros games - Principal Engineer - Backend (MSC Team)

warner bros games

(Hybrid)
4 Months ago
Flow - Engineering Manager

Flow

Palo Alto, California, United States (Hybrid)
8 Months ago
Playtika - Product Manager

Playtika

Israel (On-Site)
8 Months ago
Axon - Sr. Solutions Architect, Fusus

Axon

Atlanta, Georgia, United States (Hybrid)
1 Month ago
Zuora - Sr Enterprise Solution Architect-Zuora Billing & CPQ

Zuora

United States (Remote)
1 Month ago
Rackspace Technology - Site Reliability Engineer III

Rackspace Technology

India (Remote)
4 Months ago
NVIDIA - Senior Platform Software Engineer, PCIe

NVIDIA

Santa Clara, California, United States (On-Site)
2 Months ago
Spellbrush - AI Infrastructure Engineer

Spellbrush

San Francisco, California, United States (On-Site)
3 Months ago

Get notifed when new similar jobs are uploaded

Similar Skill Jobs

PhonePe - Server Administrator (Patching)

PhonePe

Bengaluru, Karnataka, India (On-Site)
1 Month ago
Steer Studios - Senior Growth Director

Steer Studios

Riyadh, Riyadh Province, Saudi Arabia (On-Site)
3 Months ago
Aristocrat - Senior Data Science Director

Aristocrat

London, England, United Kingdom (Hybrid)
3 Months ago
Amber - Senior Unreal Game Engineer (Project Based)

Amber

Guadalajara, Jalisco, Mexico (On-Site)
1 Year ago
Capgemini - Commercial and Pricing - C1

Capgemini

Bengaluru, Karnataka, India (On-Site)
1 Month ago
Riot Games - Principal Analytics Engineer - Product Insights

Riot Games

Los Angeles, California, United States (On-Site)
3 Months ago
Zones - Transportation Support Specialist

Zones

Islamabad, Islamabad Capital Territory, Pakistan (On-Site)
2 Months ago
Rackspace Technology - Marketing Operations Senior Manager/Lead

Rackspace Technology

Gurugram, Haryana, India (Remote)
3 Months ago
sitetracker - Senior Software Project Manager

sitetracker

Bengaluru, Karnataka, India (Remote)
3 Weeks ago
Ion - System Engineer Jr, Italy

Ion

Italy (Hybrid)
8 Months ago

Get notifed when new similar jobs are uploaded

Jobs in Amsterdam, North Holland, Netherlands

submarine career - ICT/AV Support Employee

submarine career

Amsterdam, North Holland, Netherlands (On-Site)
1 Month ago
Abstraction Games - Technical Designer

Abstraction Games

Eindhoven, North Brabant, Netherlands (Hybrid)
3 Months ago
storytq - Senior Account Executive

storytq

Amsterdam, North Holland, Netherlands (Hybrid)
2 Months ago
Palo Alto Networks - Consulting Director - Security Operations - Proactive Services (Unit 42)

Palo Alto Networks

Netherlands (Remote)
1 Month ago
Coupa - Senior Account Executive

Coupa

Netherlands (Remote)
1 Month ago
undefined - Senior Product Manager, Optimization Insights & Performance

Amsterdam, North Holland, Netherlands (On-Site)
8 Months ago
Tesla - Service Advisor

Tesla

Purmerend, North Holland, Netherlands (On-Site)
4 Months ago
Adyen - Product Marketing Manager

Adyen

Amsterdam, North Holland, Netherlands (On-Site)
1 Month ago
Tesla - Software Distributed Systems Engineer

Tesla

North Holland, Netherlands (On-Site)
4 Months ago
Abstraction Games - Office Manager (Part-Time & Hybrid)

Abstraction Games

Eindhoven, North Brabant, Netherlands (Hybrid)
3 Months ago

Get notifed when new similar jobs are uploaded

Devops Jobs

Take-Two Interactive - Senior Director, Site Reliability Engineering, Technical Operations Center & Observability

Take-Two Interactive

Austin, Texas, United States (On-Site)
2 Weeks ago
extreme network - Cloud Operations Engineer

extreme network

Toronto, Ontario, Canada (Hybrid)
2 Weeks ago
Square enix Japan - Corporate Infrastructure Engineer

Square enix Japan

Shibuya, Tokyo, Japan (On-Site)
1 Month ago
bytedance - Senior Software Engineer - Compute Infrastructure (Orchestration & Scheduling)

bytedance

San Jose, California, United States (On-Site)
2 Months ago
GoDaddy - Full Stack Software Engineer - AWS

GoDaddy

Serbia (Remote)
1 Month ago
appier - Senior Software Engineer, Machine Learning (Enterprise Solution)

appier

Taipei City, Taiwan (On-Site)
1 Month ago
CyberArk - Solutions Engineer, Enterprise Accounts - Central

CyberArk

United States (On-Site)
1 Month ago
Figma - Software Engineer, AI Infrastructure

Figma

San Francisco, California, United States (Remote)
2 Weeks ago
Capgemini - Ansible Automation Engineer

Capgemini

Pune, Maharashtra, India (On-Site)
1 Month ago

Get notifed when new similar jobs are uploaded