Software Engineer, Cloud Infrastructure

1 Month ago • 5 Years + • Devops • $205,000 PA - $240,000 PA

Job Summary

Job Description

Fireworks is building the future of generative AI infrastructure, offering a platform with high-quality models and fast, scalable inference. As a Software Engineer on the Cloud Infrastructure team, you will architect and build foundational systems for a virtual cloud, serving AI workloads across global cloud providers. Your mission is to ensure unparalleled reliability, efficiency, and scalability. This technical role demands expertise in distributed systems, cloud-native infrastructure, and machine learning platforms. You will collaborate with engineering partners, product teams, and infrastructure stakeholders to design solutions for compute, storage, and networking layers, balancing performance, cost, and operational simplicity. Responsibilities include architecting scalable backend infrastructure for distributed training and inference, leading technical design discussions, mentoring engineers, designing core backend services, driving infrastructure optimization, and integrating cloud-native technologies.
Must have:
  • 5+ years of experience designing backend infrastructure in cloud environments
  • Experience in ML infrastructure and tooling (e.g., PyTorch, TensorFlow, Kubernetes)
  • Strong software development skills in Python or C++
  • Deep understanding of distributed systems fundamentals
Good to have:
  • Master's or PhD in Computer Science
  • Experience leading infrastructure projects for ML/AI workloads
  • Familiarity with infrastructure-as-code and CI/CD tooling
  • Contributions to open-source cloud or ML infrastructure projects
Perks:
  • Meaningful equity
  • Competitive salary
  • Comprehensive benefits package

Job Details

About Us:

Here at Fireworks, we’re building the future of generative AI infrastructure. Fireworks offers the generative AI platform with the highest-quality models and the fastest, most scalable inference. We’ve been independently benchmarked to have the fastest LLM inference and have been getting great traction with innovative research projects, like our own function calling and multi-modal models. Fireworks is funded by top investors, like Benchmark and Sequoia, and we’re an ambitious, fun team composed primarily of veterans from Pytorch and Google Vertex AI.

The Role:

As a Software Engineer on our Cloud Infrastructure team, you'll be at the forefront, architecting and building the foundational systems that power Fireworks AI's revolutionary generative AI platform. You'll spearhead the creation of one of the world's first virtual clouds, seamlessly serving AI workloads across the globe and every cloud provider. Your mission: to deliver unparalleled reliability, efficiency, and scalability, fueling the world's most innovative AI products.This is a highly technical role requiring deep expertise in distributed systems, cloud-native infrastructure, and machine learning platforms. You’ll partner closely with engineering partners, product teams, and infrastructure stakeholders to design solutions that balance performance, cost-efficiency, and operational simplicity across compute, storage, and networking layers.

Key Responsibilities:

  • Architect and build scalable, resilient, and high-performance backend infrastructure to support distributed training, inference, and data processing pipelines.
  • Lead technical design discussions, mentor other engineers, and establish best practices for building and operating large-scale ML infrastructure.
  • Design and implement core backend services (e.g., job schedulers, resource managers, autoscalers, model serving layers) with a focus on efficiency and low latency.
  • Drive infrastructure optimization initiatives, including compute cost reduction, storage lifecycle management, and network performance tuning.
  • Collaborate cross-functionally with ML, DevOps, and product teams to translate research and product needs into robust infrastructure solutions.
  • Continuously evaluate and integrate cloud-native and open-source technologies (e.g., Kubernetes, Ray, Kubeflow, MLFlow) to enhance our platform’s capabilities and reliability.
  • Own end-to-end systems from design to deployment and observability, with a strong emphasis on reliability, fault tolerance, and operational excellence.

Minimum qualifications:

  • Bachelor’s degree in Computer Science, Engineering, or a related technical field (or equivalent practical experience).
  • 5+ years of experience designing and building backend infrastructure in cloud environments (e.g., AWS, GCP, Azure).
  • Proven experience in ML infrastructure and tooling (e.g., PyTorch, TensorFlow, Vertex AI, SageMaker, Kubernetes, etc.).
  • Strong software development skills in languages like Python, or C++.
  • Deep understanding of distributed systems fundamentals: scheduling, orchestration, storage, networking, and compute optimization.

Preferred qualifications:

  • Master’s or PhD in Computer Science or related field.
  • Experience leading infrastructure projects supporting large-scale ML/AI workloads or high-throughput systems.
  • Familiarity with infrastructure-as-code and CI/CD tooling (e.g., Terraform, ArgoCD, GitOps).
  • Track record of driving system performance, reliability, and cost-efficiency improvements.
  • Contributions to open-source cloud or ML infrastructure projects a plus.

Total compensation for this role also includes meaningful equity in a fast-growing startup, along with a competitive salary and comprehensive benefits package. Base salary is determined by a range of factors including individual qualifications, experience, skills, interview performance, market data, and work location. The listed salary range is intended as a guideline and may be adjusted.

Base Pay Range (Plus Equity)

$205,000 - $240,000 USD

Why Fireworks AI?

  • Solve Hard Problems: Tackle challenges at the forefront of AI infrastructure, from low-latency inference to scalable model serving.
  • Build What’s Next: Work with bleeding-edge technology that impacts how businesses and developers harness AI globally.
  • Ownership & Impact: Join a fast-growing, passionate team where your work directly shapes the future of AI—no bureaucracy, just results.
  • Learn from the Best: Collaborate with world-class engineers and AI researchers who thrive on curiosity and innovation.

Fireworks AI is an equal-opportunity employer. We celebrate diversity and are committed to creating an inclusive environment for all innovators.

Similar Jobs

 Pearl Abyss - Summer Internship: Art_Shader/Procedural Modeling

Pearl Abyss

(On-Site)
4 Months ago
Quantic Dream - Engine Programmer

Quantic Dream

Paris, Île-de-France, France (Hybrid)
4 Months ago
rivos - Data Parallel Accelerator Post-Silicon Performance Lead

rivos

Santa Clara, California, United States (Hybrid)
1 Month ago
bytedance - Application Security Engineer - Global Monetization

bytedance

Singapore (On-Site)
4 Months ago
Saronic Technologies - Senior Forward Deployed Engineer

Saronic Technologies

Austin, Texas, United States (On-Site)
2 Weeks ago
Toast - DevOps Engineer II

Toast

Dublin, County Dublin, Ireland (Hybrid)
2 Months ago
appier - Senior Software Engineer, Backend Development (Ad Cloud Serving Services)

appier

Taipei City, Taiwan (On-Site)
1 Month ago
Veeam Software - Devops Engineer

Veeam Software

Prague, Czechia (Hybrid)
2 Months ago
Motorola solutions - Sr Solution Architect

Motorola solutions

Bengaluru, Karnataka, India (On-Site)
1 Year ago
HCL Tech - Solution Architect

HCL Tech

California, United States (On-Site)
2 Months ago

Get notifed when new similar jobs are uploaded

Similar Skill Jobs

Epic Games - Lead UE Tools Engineer

Epic Games

Stockholm, Stockholm County, Sweden (On-Site)
7 Months ago
Canonical - Performance Engineer - Open Source

Canonical

(Remote)
3 Months ago
Larian Studios - Junior Gameplay Programmer

Larian Studios

Kuala Lumpur, Federal Territory Of Kuala Lumpur, Malaysia (On-Site)
4 Months ago
Qualcomm - CPU Server Physical Design Engineer

Qualcomm

Santa Clara, California, United States (On-Site)
2 Months ago
Ion - Quality Assurance Engineer

Ion

Milan, Lombardy, Italy (On-Site)
4 Months ago
Miratech - Conversational Designer (Google Dialogflow)

Miratech

Ahmedabad, Gujarat, India (On-Site)
1 Month ago
Apple - AR/VR Software Development Engineer

Apple

Cupertino, California, United States (On-Site)
3 Months ago
yellow brick games - Gameplay Programmer

yellow brick games

Montreal, Quebec, Canada (Remote)
3 Months ago
Mapbox - Technical Support Engineer

Mapbox

United Kingdom (Remote)
1 Month ago
Power Integrations - Senior Test Engineer

Power Integrations

Penang, Malaysia (On-Site)
3 Months ago

Get notifed when new similar jobs are uploaded

Jobs in New York, United States

Scale AI - Senior Software Engineer, Enterprise GenAI

Scale AI

San Francisco, California, United States (On-Site)
2 Months ago
Saviynt - Lead Site Reliability Engineer - Federal Team

Saviynt

Los Angeles, California, United States (Hybrid)
3 Weeks ago
ChainGuard - Strategic Account Executive - TOLA

ChainGuard

Texas, United States (Remote)
3 Weeks ago
Zones - Field Services Technician

Zones

Memphis, Tennessee, United States (On-Site)
3 Weeks ago
FORTUNE - UI/UX Designer

FORTUNE

New York, New York, United States (On-Site)
3 Months ago
Reddit - Staff Software Engineer, Ads Creative

Reddit

United States (Remote)
3 Months ago
Apple - Sports Business Optimization & League Relations

Apple

Culver City, California, United States (On-Site)
2 Months ago
Philips - Sr. Cardiac Wireless Monitoring Technician

Philips

Norfolk, Virginia, United States (On-Site)
2 Months ago
Next Level Business Services - iOS Mobile Architect

Next Level Business Services

Owings Mills, Maryland, United States (On-Site)
9 Months ago

Get notifed when new similar jobs are uploaded

Devops Jobs

deel. - Senior Backend Engineer, Node.js + AWS

deel.

Italy (Remote)
2 Weeks ago
Reltio - Intern - DevOps

Reltio

Lisbon, Lisbon, Portugal (On-Site)
3 Months ago
Rackspace Technology - Senior Cloud Infrastructure Engineer (Azure)

Rackspace Technology

Germany (Remote)
2 Weeks ago
Nice - Cloud Site Reliability Engineer

Nice

Pune, Maharashtra, India (On-Site)
1 Month ago
Nagarro - Associate Principal Engineer, DevOps

Nagarro

India (Remote)
9 Months ago
Spaulding Ridge - Oracle EPM Solution Architect

Spaulding Ridge

Toronto, Ontario, Canada (On-Site)
3 Months ago
Netomi - Devops Engineer - II

Netomi

Toronto, Ontario, Canada (Remote)
2 Months ago
Rackspace Technology - Senior Platform Engineer (Azure)

Rackspace Technology

Germany (Remote)
9 Months ago
Monolith - Cloud Playout Systems Engineer

Monolith

Sterling, Virginia, United States (On-Site)
5 Months ago

Get notifed when new similar jobs are uploaded

About The Company

Redwood City, California, United States (Hybrid)

Redwood City, California, United States (Hybrid)

Redwood City, California, United States (Hybrid)

Redwood City, California, United States (Hybrid)

New York, United States (Hybrid)

Redwood City, California, United States (Hybrid)

Redwood City, California, United States (Remote)

Redwood City, California, United States (Hybrid)

Redwood City, California, United States (Hybrid)

Redwood City, California, United States (On-Site)

View All Jobs

Get notified when new jobs are added by Fireworks AI

Level Up Your Career in Game Development!

Transform Your Passion into Profession with Our Comprehensive Courses for Aspiring Game Developers.

Job Common Plug