Let’s get started
By clicking ‘Next’, I agree to the Terms of Service
and Privacy Policy, and consent to receive emails from Rise
Jobs / Job page
Senior Site Reliability Engineer - Control Plane image - Rise Careers
Job details

Senior Site Reliability Engineer - Control Plane

In 2012, Lambda started with a crew of AI engineers publishing research at top machine-learning conferences. We began as an AI company built by AI engineers. That hasn't changed. Today, we're on a mission to be the world's top AI computing platform. We equip engineers with the tools to deploy AI that is fast, secure, affordable, and built to scale. Whether they need powerhouse GPU hardware on-site or the flexibility of cloud-based solutions, we've got the horsepower to make it happen. Lambda’s AI Cloud has been adopted by the world’s leading companies and research institutions including Anyscale, Rakuten, The AI Institute, and multiple enterprises with over a trillion dollars of market capitalization. Our goal is to make computation as effortless and ubiquitous as electricity.


If you'd like to build the world's best deep learning cloud, join us. 


*Note: This position requires presence in our San Francisco office location 4 days per week; Lambda’s designated work from home day is currently Tuesday.

What You’ll Do

  • Design and implement cloud-native architectures that deliver the "four nines" (99.99%) of reliability while balancing performance and cost efficiency

  • Develop comprehensive monitoring and alerting systems with actionable dashboards that provide real-time visibility into system health

  • Implement SLIs, SLOs, and SLAs across services and maintain error budgets to guide development priorities

  • Automate deployments using tools like Argo and Terraform

  • Create robust incident management processes, escalation paths, and documentation

  • Architect fault-tolerant systems with graceful degradation capabilities to handle component failures

  • Design and implement disaster recovery solutions with regular testing procedures

  • Lead post-incident reviews that focus on systemic improvements rather than individual blame

  • Champion reliability best practices and system design principles

  • Build automated, auditable, and compliant processes to improve efficiency and productivity

You

  • 5+ years of experience in Site Reliability Engineering or DevOps roles

  • Strong understanding of cloud platforms (AWS, GCP, Azure) and their core services

  • Experience designing and implementing monitoring and observability solutions at scale

  • Proven track record managing production incidents and driving root cause analysis

  • Proficiency with Infrastructure as Code tools and CI/CD pipeline implementation

  • Strong understanding of network architecture, load balancing, and content delivery

  • Expertise in performance tuning and system optimization techniques

  • Experience with container orchestration platforms like Kubernetes

  • Knowledge of database administration and optimization strategies

  • Solid coding skills in at least one language (Python, Go, Bash) for automation

Nice to Have

  • Experience with high-throughput, low-latency systems

  • Knowledge of security best practices and implementing defense-in-depth strategies

  • Experience with multi-region and distributed systems and solving consistency/availability challenges

  • Background in chaos engineering or similar reliability testing methodologies

  • Understanding of compliance frameworks (SOC 2, ISO 27001, etc.)

  • Familiarity with message queuing systems and event-driven architectures

  • Background working with specialized computing hardware

Salary Range Information 

Based on market data and other factors, the annual salary range for this position is $245,000 - $385,000. However, a salary higher or lower than this range may be appropriate for a candidate whose qualifications differ meaningfully from those listed in the job description.

About Lambda

  • Founded in 2012, ~350 employees (2024) and growing fast

  • We offer generous cash & equity compensation

  • Our investors include Andra Capital, SGW, Andrej Karpathy, ARK Invest, Fincadia Advisors, G Squared, In-Q-Tel (IQT), KHK & Partners, NVIDIA, Pegatron, Supermicro, Wistron, Wiwynn, US Innovative Technology, Gradient Ventures, Mercato Partners, SVB, 1517, Crescent Cove.

  • We are experiencing extremely high demand for our systems, with quarter over quarter, year over year profitability

  • Our research papers have been accepted into top machine learning and graphics conferences, including NeurIPS, ICCV, SIGGRAPH, and TOG

  • Health, dental, and vision coverage for you and your dependents

  • Commuter/Work from home stipends for select roles

  • 401k Plan with 2% company match (USA employees)

  • Flexible Paid Time Off Plan that we all actually use

A Final Note:

You do not need to match all of the listed expectations to apply for this position. We are committed to building a team with a variety of backgrounds, experiences, and skills.

Equal Opportunity Employer

Lambda is an Equal Opportunity employer. Applicants are considered without regard to race, color, religion, creed, national origin, age, sex, gender, marital status, sexual orientation and identity, genetic information, veteran status, citizenship, or any other factors prohibited by local, state, or federal law.

Lambda Glassdoor Company Review
3.4 Glassdoor star iconGlassdoor star iconGlassdoor star icon Glassdoor star icon Glassdoor star icon
Lambda DE&I Review
No rating Glassdoor star iconGlassdoor star iconGlassdoor star iconGlassdoor star iconGlassdoor star icon
CEO of Lambda
Lambda CEO photo
Stephen Balaban
Approve of CEO

Average salary estimate

$315000 / YEARLY (est.)
min
max
$245000K
$385000K

If an employer mentions a salary or salary range on their job, we display it as an "Employer Estimate". If a job has no salary data, Rise displays an estimate if available.

Similar Jobs
Photo of the Rise User

Join Lambda as a Senior Software Engineer to drive the development of cutting-edge AI applications in a vibrant team setting.

Photo of the Rise User

Join Lambda as a Data Center Capacity & Forecasting Strategist to shape the future of AI infrastructure planning and capacity modeling.

Image Associates Inc. Hybrid abc, Parkersburg, West Virginia, United States
Posted 4 days ago

Join a dynamic team as an Electrical Project Manager for a heavy manufacturing facility in West Virginia, overseeing critical electrical projects.

Photo of the Rise User
Figure Hybrid San Jose, California, United States
Posted 4 days ago
Dental Insurance
Disability Insurance
Vision Insurance
Flexible Spending Account (FSA)
Performance Bonus
Family Medical Leave
Paid Holidays

Join Figure, an innovative AI Robotics company, as a Manufacturing Test Engineer to enhance the testing processes of our cutting-edge humanoid robots.

Posted 10 days ago

Join Activate Interactive as a DevOps Engineer and help drive innovation in mobile and web technologies within a collaborative environment.

Join Solidigm as an Undergrad Intern in Validation, Quality & Reliability Engineering, contributing to innovative memory solutions.

Photo of the Rise User

Become a pivotal part of Penumbra's team as a Manufacturing Engineer I, contributing to innovative medical solutions in a collaborative work environment.

Photo of the Rise User
Netskope Hybrid Santa Clara, California, United States
Posted 12 days ago
Inclusive & Diverse
Work/Life Harmony
Collaboration over Competition
Growth & Learning
Transparent & Candid
Feedback Forward

Join Netskope as a Sr. Staff Engineer and play a key role in building scalable, high-performance cloud data security solutions.

Photo of the Rise User
Posted 13 days ago

Join LLNL as a Mid-Senior Mechanical Engineer and contribute to innovative projects that enhance national security and research capabilities.

Photo of the Rise User
ITAC Hybrid No location specified
Posted 3 hours ago

Join ITAC as a Senior Civil Engineer, where your expertise will help shape innovative engineering solutions for a variety of industry challenges.

Photo of the Rise User
Posted 7 days ago

Lead a talented team in the mechanical engineering department at Fluxergy, a pioneering in vitro diagnostics company, to create cutting-edge medical devices.

Posted 6 days ago

Join Fluidstack as a Principal Networking Engineer and help optimize networks for cutting-edge AI deployments.

Photo of the Rise User

Join Woven by Toyota to shape the future of mobility as a Site Reliability Engineer focused on innovative data solutions.

Photo of the Rise User
Mattel Hybrid 333 Continental Blvd, El Segundo, CALIFORNIA
Posted yesterday
Inclusive & Diverse
Empathetic
Collaboration over Competition
Growth & Learning

Join Mattel as a Product Development Engineer I to drive innovative toy concepts from design to production, ensuring quality and safety standards are met.

Posted 10 days ago

We're looking for an experienced Integrations Engineer to drive our compliance solutions integration initiatives and shape strategic partnerships with industry leaders.

Lambda provides Artificial Intelligence and Machine Learning infrastructure to companies like Apple, Intel, Microsoft, MIT, Harvard, the Federal Government, and the DOD. Were headquartered in the Dogpatch and are a short walk from the 22nd Street ...

74 jobs
MATCH
Calculating your matching score...
FUNDING
DEPARTMENTS
SENIORITY LEVEL REQUIREMENT
TEAM SIZE
EMPLOYMENT TYPE
Full-time, hybrid
DATE POSTED
April 20, 2025

Subscribe to Rise newsletter

Risa star 🔮 Hi, I'm Risa! Your AI
Career Copilot
Want to see a list of jobs tailored to
you, just ask me below!