Why This Job is Featured on The SaaS Jobs
At scale, SaaS reliability becomes a product feature, and this Senior Site Reliability Engineer role sits directly on that boundary. Supporting API first Search and Discovery services that handle extremely high query volume puts day to day work in the core concerns of modern SaaS: availability, latency, and predictable performance for a large customer base. The remit spans both cloud and bare metal, which is less common in SaaS listings and signals a platform with real infrastructure breadth.
For a long term SaaS career, the role builds durable operating instincts: defining and measuring SLOs, managing error budgets, and using automation to reduce toil while keeping systems observable. The mix of incident response, customer facing reliability work, and building internal components mirrors how mature SaaS organisations scale platform capability without slowing product delivery. Experience across Kubernetes, IaC, and CI/CD also transfers cleanly to other subscription software environments.
This position will suit engineers who prefer ownership over production outcomes and who like turning operational pain into repeatable systems. It fits someone comfortable collaborating across product and engineering groups, and who values pragmatic trade offs between reliability, cost, and delivery. A background in scalable architectures and a craft oriented approach to code will map well to the expectations here.
The section above is editorial commentary from The SaaS Jobs, provided to help SaaS professionals understand the role in a broader industry context.
Job Description
Algolia is set to enable every company to create world-class Search and Discovery experiences with an API-first approach. Performance and Scalability is at the heart of our mission: we power 1.5 trillion searches a year, for 10K+ customers all over the world.
If you're a problem solver, able to think outside the box and eager to nurture others and learn from them, then this is your challenge!
The Team
The Fleet team is a Site Reliability Engineering team focusing on one goal: the Search products should always be available. To make this possible, the Fleet team creates pragmatic solutions to optimize the Search products availability and costs at scale, taking into account the needs of customers, the product teams, and the many engineering teams involved in delivering a unique Search Experience to our customers.
The Opportunity
The team is looking for an individual who has a first experience of building and operating scalable architectures. You will contribute to the delivery of solutions that support other engineering teams and will have a direct impact on the success of Algolia's Search products.
In this role, you'll help design and implement systems focused on reliability, scalability, and cost efficiency, while also having opportunities to grow your skills and collaborate with team members.
Your role will include
- Operating the Search products, building self-healing and automated incident response mechanisms
- Building components that improve reliability and performance
- Monitoring and computing the SLO and the error budget of the product you operate
- Reducing the toil and the technical debt by automating tasks and increasing the quality of existing components
- Managing Incidents and Customer Requests
You might be a good fit if you have
- 5 years experience in a scalable environment
- Knowledge of at least one programming language (Python, Golang, Ruby) and you are familiar with software craftsmanship
- Experience working with APIs
- A focus on designing reliable, operable, and highly available applications
- Familiarity with at least Public Cloud Providers like GCP, AWS, or Microsoft Azure, and their Kubernetes service
- A good understanding of Linux system administration, networking, and troubleshooting
- Strong communication and organizational skills
Team’s current stack:
- Programming languages: Golang, Python, Ruby
- CI/CD:Github Actions, CircleCI
- IaC & configuration management:Terraform, Chef
- Platform: Linux, Kubernetes
- Hosting: Bare Metal Servers & Cloud on AWS & Azure
- Monitoring: Datadog & custom monitoring stack for our Search infrastructure
#LI-Remote