Infrastructure Capacity Engineer

1 Month ago • 4 Years + • Devops • $225,000 PA - $300,000 PA

Job Summary

Job Description

Perplexity is an AI-powered answer engine founded in December 2022, rapidly growing as a leading AI platform. They are seeking an experienced Infrastructure Capacity Engineer to own infrastructure scaling, capacity planning, and resource optimization across their AI/ML infrastructure. The role involves designing and implementing capacity planning models, building automated management systems, leading cross-functional initiatives, and optimizing resource utilization to support a rapidly growing AI product and user base.
Must have:
  • Design and implement comprehensive capacity planning models and forecasting systems.
  • Build and maintain automated capacity management systems.
  • Lead cross-functional capacity planning initiatives.
  • Develop sophisticated monitoring and alerting systems.
  • Create and maintain detailed infrastructure capacity models.
  • Optimize resource utilization and cost efficiency.
  • Design and implement disaster recovery and business continuity plans.
  • Collaborate with Site Reliability Engineering and Platform teams.
  • Play a leading role in defining the capacity engineering discipline.
  • Minimum of 4+ years of experience in infrastructure capacity planning or systems engineering.
  • Proven experience managing infrastructure capacity for high-growth technology companies.
  • Strong background in distributed systems architecture.
  • Experience with cloud infrastructure (AWS/GCP/Azure) and container orchestration (Kubernetes).
  • Experience with capacity modeling tools, forecasting methodologies, and statistical analysis.
  • Proficiency in programming languages such as Python, Go, or similar.
  • Deep understanding of infrastructure monitoring, observability, and performance optimization techniques.
  • Experience with infrastructure-as-code tools (Terraform, Ansible) and CI/CD pipelines.
  • Strong analytical and problem-solving skills.
  • Excellent cross-functional collaboration skills.
Good to have:
  • Experience with large-scale database systems
  • Experience with caching layers
  • Experience with content delivery networks
  • Background in AI/ML infrastructure
  • Experience with LLM inference
  • Experience with GPU cluster management
  • Experience with high-performance computing
Perks:
  • Comprehensive health insurance
  • Comprehensive dental insurance
  • Comprehensive vision insurance
  • 401(k) plan

Job Details

Perplexity is an AI-powered answer engine founded in December 2022 and growing rapidly as one of the world’s leading AI platforms. Perplexity has raised over $1B in venture investment from some of the world’s most visionary and successful leaders, including Elad Gil, Daniel Gross, Jeff Bezos, Accel, IVP, NEA, NVIDIA, Samsung, and many more. Our objective is to build accurate, trustworthy AI that powers decision-making for people and assistive AI wherever decisions are being made. Throughout human history, change and innovation have always been driven by curious people. Today, curious people use Perplexity to answer more than 780 million queries every month–a number that’s growing rapidly for one simple reason: everyone can be curious.

Perplexity is seeking an experienced Infrastructure Capacity Engineer to own our infrastructure scaling, capacity planning, and resource optimization across our AI/ML infrastructure. The ideal candidate will have deep experience in large-scale distributed systems, capacity modeling, and infrastructure efficiency optimization to support our rapidly growing AI products and user base.

Responsibilities

  • Design and implement comprehensive capacity planning models and forecasting systems that predict infrastructure needs across compute, storage, and network resources for our AI/ML workloads
  • Build and maintain automated capacity management systems that dynamically scale our infrastructure based on real-time demand patterns and usage forecasts
  • Lead cross-functional capacity planning initiatives including hardware procurement, data center expansion, and cloud resource optimization
  • Develop sophisticated monitoring and alerting systems that provide early warning indicators for capacity constraints and performance degradation
  • Create and maintain detailed infrastructure capacity models that account for seasonal patterns, product launches, and scaling efficiency across different workload types
  • Optimize resource utilization and cost efficiency through advanced placement algorithms, load balancing strategies, and infrastructure rightsizing
  • Design and implement disaster recovery and business continuity plans that ensure service availability during infrastructure failures or capacity emergencies
  • Collaborate with Site Reliability Engineering and Platform teams to establish capacity-aware deployment strategies and infrastructure automation
  • Play a leading role in defining the capacity engineering discipline within Perplexity’s engineering organization

Qualifications

  • Minimum of 4+ years of experience in infrastructure capacity planning, systems engineering, or related technical roles at scale
  • Proven experience managing infrastructure capacity for high-growth technology companies, preferably with AI/ML workloads or real-time systems
  • Strong background in distributed systems architecture, cloud infrastructure (AWS/GCP/Azure), and container orchestration (Kubernetes)
  • Experience with capacity modeling tools, forecasting methodologies, and statistical analysis for infrastructure planning
  • Proficiency in programming languages such as Python, Go, or similar for automation and tooling development
  • Deep understanding of infrastructure monitoring, observability, and performance optimization techniques
  • Experience with infrastructure-as-code tools (Terraform, Ansible) and CI/CD pipelines for infrastructure management
  • Strong analytical and problem-solving skills with the ability to make data-driven decisions under uncertainty
  • Excellent cross-functional collaboration skills and experience working with engineering, product, and business stakeholders
  • Experience with large-scale database systems, caching layers, and content delivery networks preferred
  • Background in AI/ML infrastructure, LLM inference, GPU cluster management, or high-performance computing is a plus

Our cash compensation range for this role is $225,000 - $300,000.

Final offer amounts are determined by multiple factors, including, experience and expertise, and may vary from the amounts listed above.

Equity: In addition to the base salary, equity may be part of the total compensation package.

Benefits: Comprehensive health, dental, and vision insurance for you and your dependents. Includes a 401(k) plan.

Similar Jobs

Neolytix - Marketing and Branding Intern

Neolytix

(Remote)
1 Month ago
GoMotive - Data Scientist, People Analytics

GoMotive

India (Remote)
1 Month ago
Hawkeye Innovations - Director of Technical Sales and Business Development

Hawkeye Innovations

Atlanta, Georgia, United States (Remote)
5 Months ago
Sabre India - Head of Strategic Account Management – Corporate Travel

Sabre India

Sydney, New South Wales, Australia (Hybrid)
1 Month ago
Windranger - Business Development Lead

Windranger

United Kingdom (Remote)
2 Months ago
GHX - Automation Engineer III

GHX

Hyderabad, Telangana, India (On-Site)
3 Months ago
PwC - Cloud Security Specialist - Associate

PwC

Turin, Piedmont, Italy (On-Site)
10 Months ago
Plaid  - Experienced Infrastructure Engineer

Plaid

United States (On-Site)
1 Month ago
Varonis  - Cloud Security Research Team Leader

Varonis

Herzliya, Tel Aviv District, Israel (On-Site)
10 Months ago
ISS Stoxx - Senior Software Engineer in C#/.NET and AWS

ISS Stoxx

Mumbai, Maharashtra, India (On-Site)
2 Months ago

Get notifed when new similar jobs are uploaded

Similar Skill Jobs

Technicolor Creative Studios - Senior GL Accountant (French Speaking Expert - B2 Level)

Technicolor Creative Studios

Bengaluru, Karnataka, India (On-Site)
9 Months ago
Aristocrat - DevOps Lead

Aristocrat

Austin, Texas, United States (Hybrid)
2 Months ago
endava - Data Engineer (Azure)

endava

Cali, Valle Del Cauca, Colombia (On-Site)
2 Months ago
Wind River - Manager, Engineering - Sys

Wind River

Bengaluru, Karnataka, India (On-Site)
1 Month ago
bytedance - Site Reliability Engineer, ML System

bytedance

Seattle, Washington, United States (On-Site)
9 Months ago
Playtika - Software Architect

Playtika

Israel (On-Site)
7 Months ago
bytedance - Human Resources Apprenticeship Program

bytedance

Gurugram, Haryana, India (On-Site)
5 Months ago
hogarth - Technical Business Analyst

hogarth

London, England, United Kingdom (Hybrid)
3 Months ago
Liquid nitro games - Recruiter

Liquid nitro games

Hyderabad, Telangana, India (On-Site)
4 Months ago
Moon Active - Delivery Manager

Moon Active

Tel Aviv-Yafo, Tel Aviv District, Israel (On-Site)
1 Month ago

Get notifed when new similar jobs are uploaded

Jobs in Palo Alto, California, United States

rivos - Senior Member of Technical Staff

rivos

Austin, Texas, United States (Hybrid)
3 Months ago
Nagarro - Principal Engineer - Project Manager (Salesforce)

Nagarro

Boston, Massachusetts, United States (On-Site)
7 Months ago
PlayStation Global - Director, People Technology Services

PlayStation Global

Aliso Viejo, California, United States (On-Site)
4 Months ago
Inkittt - Creative Producer

Inkittt

San Francisco, California, United States (Hybrid)
2 Months ago
Apple - CAD Engineer - RTL Construction

Apple

San Diego, California, United States (On-Site)
3 Months ago
lifechruh - Product Marketing Strategist

lifechruh

Edmond, Oklahoma, United States (On-Site)
10 Months ago
WPI - Grant Writer

WPI

Worcester, Massachusetts, United States (Hybrid)
3 Months ago
Bungie - Deployment Operations Manager

Bungie

Bellevue, Washington, United States (Hybrid)
1 Month ago
Apple - Senior Engineering Program Manager, FM Evaluation

Apple

Cupertino, California, United States (On-Site)
2 Months ago
Zelis  - Senior Marketing Technology Specialist

Zelis

New Jersey, United States (Remote)
1 Month ago

Get notifed when new similar jobs are uploaded

Devops Jobs

bytedance - Senior Software Engineer - Traffic Infrastructure

bytedance

Singapore (On-Site)
9 Months ago
Interactive Brokers - Platform Engineer - Support

Interactive Brokers

Mumbai, Maharashtra, India (On-Site)
3 Months ago
dbt Labs - Customer Solutions Engineer

dbt Labs

Philippines (Remote)
1 Month ago
Loyalty Juggernaut - Solutions Engineer

Loyalty Juggernaut

Hyderabad, Telangana, India (On-Site)
1 Year ago
London stock Exchange - Tech Lead -Database SRE

London stock Exchange

Bengaluru, Karnataka, India (On-Site)
2 Months ago
Wargaming - Senior Build Engineer (Unannounced Project)

Wargaming

Warsaw, Masovian Voivodeship, Poland (Hybrid)
1 Month ago
Safe security - Software Development Engineer II - Platform

Safe security

Bengaluru, Karnataka, India (On-Site)
6 Months ago
ISS Stoxx - Principal Platform Engineer

ISS Stoxx

London, England, United Kingdom (On-Site)
2 Months ago
ShyftLabs - Cloud Engineer

ShyftLabs

Atlanta, Georgia, United States (Hybrid)
1 Month ago
bytedance - Site Reliability Engineer Graduate (Technical Infrastructure) - 2025 Start (BS/MS)

bytedance

Seattle, Washington, United States (On-Site)
9 Months ago

Get notifed when new similar jobs are uploaded

About The Company

San Francisco, California, United States (Hybrid)

San Francisco, California, United States (Hybrid)

San Francisco, California, United States (On-Site)

San Francisco, California, United States (Hybrid)

San Francisco, California, United States (On-Site)

San Francisco, California, United States (On-Site)

New York, New York, United States (On-Site)

Belgrade, Serbia (Hybrid)

Palo Alto, California, United States (On-Site)

View All Jobs

Get notified when new jobs are added by Perplexity

Level Up Your Career in Game Development!

Transform Your Passion into Profession with Our Comprehensive Courses for Aspiring Game Developers.

Job Common Plug