Senior Site Reliability Engineer (SRE)
Mission – Why We Exist, What We Do, and Why We Need You
SpotMe is a leading B2B event platform that helps enterprises increase the impact of their events by delivering CRM-connected, high-quality experiences across in-person, virtual, hybrid events, and webinars. With a strong focus on life sciences, SpotMe powers Onomi, an HCP engagement product that enables medical and commercial teams to run impactful congresses, symposia, advisory boards, and webinars. Together, SpotMe and Onomi turn events into a company’s most effective engagement channel. This role is for a hands-on technical expert who keeps high-performance, scalable systems running and who makes sure a 24/7 SaaS platform operates smoothly under all conditions. If you have a strong background in cloud infrastructure, automation, and incident response, this is your opportunity to take full ownership of our platform’s reliability and scalability, in an environment where both are critical to the business. We build with AI-assisted development tools as a core part of how we work, and you'll have real latitude to use them.. You will not just monitor the platform; you will drive lasting improvements to its uptime, performance, and resilience.
You will report to the Infrastructure Lead and work closely with the engineering and product teams. You will maintain and optimize the platform’s infrastructure and build solutions that improve its reliability and scalability, making sure the platform scales smoothly to handle peak traffic and stays resilient during high-stakes live events. Your time will be spent on:
40% Infrastructure development
- Develop and deploy scalable infrastructure using Terraform and cloud-native AWS services.
- Contribute to critical full-stack work that spans the back end and infrastructure or cloud development.
- Build automation for infrastructure provisioning and CI/CD pipelines (Jenkins), including image builds with Packer and Docker.
40% Platform optimization and resilience
- Optimize the platform's cloud infrastructure for high availability and cost efficiency
- Monitor and update infrastructure to follow security best practices, applying necessary patches and upgrades.
- Make sure the infrastructure can handle peak loads, scaling smoothly during high-traffic events.
20% Support and observability
- Take your turn in the on-call infrastructure rotation, responding to incidents and resolving them quickly.
- Strengthen the platform’s monitoring and observability (Datadog, Pingdom, Elastic) to catch issues before they reach end users.
- Handle infrastructure support requests and drive continuous improvement in incident resolution.
Objectives – The Problems You Will Solve
In Your First Month:
- Understand the current platform architecture and complete 3 infrastructure-as-code change request reviews.
- Learn our monthly patching procedures and deploy critical security patches.
- Get hands-on with our load testing framework (Locust, Gatling) and run one release-validation load test.
- Take part in the weekly risk-analysis meeting and run a scheduled database scaling exercise.
- Get set up with our AI-assisted development toolchain, including Anthropic’s Claude, and use it in your daily work.
- Handle and resolve at least 3 infrastructure support requests.
- Build a report on what surprised you: what looked fragile, and what was hard to find documented.
After 3 Months:
- Own a security-hardening improvement, such as tightening firewall and network rules across an environment.
- Design and deliver one infrastructure-as-code project in Terraform, from proposal to production.
- Build and ship one Python-based AWS Lambda that automates an operational task.
- Reach the level of system knowledge needed to operate autonomously in the on-call rotation, resolving incidents without escalation.
After 6 Months:
- Lead the resolution of a critical infrastructure incident and drive lasting improvements in incident response and recovery times.
- Lead a significant reliability or tech-debt project, such as moving a complex on-premises build system to the cloud.
- Identify repetitive engineering workflows and automate them end-to-end with AI-assisted tooling, so the team's time goes to the hard problems instead of the recurring ones.
- Implement and own observability for one critical service end-to-end: instrumentation, alerting, and dashboards that let the team detect and diagnose issues before customers report them.
- Own and measurably improve one reliability metric for a critical service, such as time to detect or mean time to recovery, against a baseline you establish in your first month, with the target agreed with your manager.
After 12 Months:
- You have set the standard for how infrastructure is built here: your Terraform patterns, review practices, and documentation are what other engineers work from by default.
- You own peak-traffic readiness end to end, with load testing, capacity review, and scaling runbooks running to a schedule you set and refine.
- AI-assisted tooling is embedded in how the team works, with recurring operational work measurably reduced against a baseline you establish and the team's judgment applied where it matters.
- You are the person others come to on infrastructure decisions, and reliability work gets prioritised across teams because you have made the case for it credibly.
What We Are Looking For
We are looking for a senior engineer who has spent years keeping large, business-critical SaaS platforms running, and who wants to take full ownership of reliability rather than simply keeping the lights on. In practice, that means:
- Reliability engineering at scale. You bring around five or more years in a site reliability role, built on earlier experience as a system administrator or software developer, and a track record of keeping large-scale, 24/7 SaaS platforms running. You diagnose and resolve complex system issues under pressure, and you design for resilience before incidents happen.
- Cloud-native, distributed systems. You are hands-on with cloud-native architectures, distributed systems, and high-availability platforms, and you are comfortable across both document-oriented and relational databases. You have strong, production-grade experience with AWS, and knowledge of Azure is a bonus.
- Infrastructure automation and delivery. You automate and manage infrastructure with Terraform, and you build and maintain CI/CD pipelines with Jenkins as well as GitHub actions, including image builds with Packer and Docker. You write real code to do it: Python is essential, and experience with JavaScript, Node.js, or Go is an asset.
- Observability and operations. You instrument systems for visibility and act on what they tell you, using tools such as Datadog, Pingdom, and Elastic to find and resolve issues before they affect end users.
- AI-assisted development. You use AI coding tools such as Anthropic's Claude as a natural part of building and operating infrastructure, you apply good judgment about where they help and where they don't, with concrete examples of both from your own infra work, and you push the team to get more out of them.
- Ownership and collaboration. You take full ownership of reliability, you communicate clearly with engineering and product, you can persuade engineers outside your own team to prioritise reliability work, and you raise the standard of the systems and the teams you work with.
SpotMe recruits, compensates, and promotes regardless of race, color, religion, gender, gender identity or expression, sexual orientation, national origin, genetics, disability, age, parental status, or veteran status.
Apply for this job
*
indicates a required field