Back to jobs
New

Site Reliability Engineer

Bengaluru

Razorpay is one of India’s leading full-stack financial technology companies, powering the way businesses move, manage, and grow money. Founded in 2014 by Harshil Mathur and Shashank Kumar with a simple vision - to simplify payments for Indian businesses - we’ve since grown into a fintech powerhouse driving India’s digital payment revolution.

Razorpay powers millions of businesses with a smarter, scalable stack that goes beyond transactions to help them truly build and grow.

From building AI-native agentic payments, to AI-assisted fraud detection and real-time risk intelligence to automated reconciliation, smart payouts, and predictive financial insights, we are embedding intelligence across our stack to make money movement faster, safer, and more efficient. In close collaboration with ecosystem partners - including banks, networks, regulators - we are pioneering industry-first solutions that are shaping the next era of fintech

Across India, Singapore and Malaysia, our products span everything from seamless checkouts to payroll automation - powering a fintech ecosystem that’s redefining how money moves across Asia.

Today, that ecosystem supports everyone from early-stage startups to some of India’s largest enterprises, enabling them to accept, process, and disburse payments at scale while expanding into new ways of managing money more efficiently.

Our scale speaks volumes: Razorpay processes $180+ billion in annualized transactions, powering leading businesses like Airbnb, Facebook, WhatsApp, Airtel, CRED, BookmyShow, Zomato, Swiggy, Lenskart, Mirae Asset Capital markets, Indian Oil, National Pension Scheme - and over 100 of India’s unicorns. With strong roots in India and growing operations in Southeast Asia, we are shaping the next chapter of financial technology across the region.

We are backed by global investors including GIC, Peak XV Partners (formerly Sequoia Capital India & SEA), Tiger Global, Ribbit Capital, Matrix Partners, MasterCard, and Salesforce Ventures, having raised over $740 million to date. Strategic acquisitions - including Ezetap (POS and offline payments), Curlec (Malaysia expansion), BillMe (digital invoicing), and POP (rewards-first UPI) - along with earlier moves in fraud prevention, payroll, and lending, have further strengthened our platform and widened our footprint across Asia.

But what truly sets Razorpay apart is our culture. At Razorpay, ownership is our oxygen - you own what you build, with no micromanagement or red tape, just the runway to make your ideas fly. Learning is a lifestyle - if you’re curious, you’ll feel at home here. People > Pedigree - we hire for attitude, hustle, and hunger more than degrees. Transparency thrives over titles - this is where interns question CXOs and CXOs say “thank you.” Guided by our values of Customer First, Autonomy & Ownership, Agility with Integrity, Transparency, Challenging the status quo and a strong belief that Razorpay grows with Razors,  you’ll be part of a 3000+ strong team building not just products, but the financial infrastructure of the future.

About the role

You will be one of the founding SREs at Razorpay, embedded with the payment platform teams that move money for millions of businesses. Your mandate is to take our payment flows from three nines to four and five nines of availability. In payments, a failed request is not a retry, it is a customer's money in limbo. You will define what reliability means here, build the systems that enforce it, and set the standard every future SRE is measured against.

What you will do

  • Define SLIs, SLOs, and error budgets for critical payment flows (authorization, capture, refunds, settlements, webhooks) and make them the shared language between product and platform teams.
  • Own the release lifecycle for payment services: design progressive rollout pipelines (canary, staged, feature-flagged), automated rollback triggers, and make "can we roll back in under 5 minutes" a launch-blocking question.
  • Carry the pager for payment-critical services, lead incident command during outages, and drive blameless postmortems where action items actually ship.
  • Eliminate toil through software: build automation for failover, capacity management, load shedding, and degradation so that known failure classes cannot recur.
  • Harden payment flows against distributed systems failure modes: retry storms, thundering herds, cascading failures, partial outages of banks and network partners, idempotency violations, and reconciliation gaps.
  • Run production readiness reviews for new payment services and hold the line on launch gates using error budget data, not opinion.
  • Instrument what matters: design alerting that pages on customer-facing symptoms, not noise, and cut mean time to detection and recovery quarter over quarter.
  • Practice failure on purpose: game days, chaos experiments, and failure injection against payment-critical paths.

What we are looking for

  • 10+ years of engineering experience, with at least 5 years operating large-scale distributed systems in production (high QPS, multi-region, or systems where sub-1 percent error rates were business-critical).
  • Strong software engineering skills in at least one of Go, Java, or Python. You have built tools and services, not just configured them.
  • Deep understanding of distributed systems failure modes and the patterns that contain them: circuit breakers, backpressure, bulkheading, graceful degradation, idempotency.
  • Solid fundamentals in Linux internals, networking, and databases under load (replication, failover, connection pool exhaustion, lock contention).
  • Hands-on experience designing or significantly improving deployment pipelines: canary analysis, automated rollback, feature flags.
  • Genuine on-call ownership: you have carried a pager for systems that mattered, led incidents, and can walk us through a specific outage you handled and what you changed afterward.
  • Fluency with modern observability (metrics, tracing, structured logging; e.g. Prometheus, Grafana, OpenTelemetry, Datadog, Coralogix, Clickhouse or similar) and experience reducing alert noise.
  • Experience defining SLOs and error budget policies from scratch. Contributions to reliability tooling, open source or internal, that other teams adopted.
  • The judgment and communication skills to tell a product team "not yet" with data, and the pragmatism to help them get to "yes" quickly.

Nice to have

  • Experience in payments, fintech, banking, trading, or another domain where correctness and money are coupled (transactional consistency, exactly-once semantics, reconciliation).
  • Experience with Kubernetes at scale, service mesh, and traffic management.
Razorpay believes in and follows an equal employment opportunity policy that doesn't discriminate on gender, religion, sexual orientation, colour, nationality, age, etc. We welcome interests and applications from all groups and communities across the globe.
 
Follow us on LinkedIn & Twitter

Create a Job Alert

Interested in building your career at Razorpay Software Private Limited? Get future opportunities sent straight to your email.

Apply for this job

*

indicates a required field

Phone
Resume/CV*

Accepted file types: pdf, doc, docx, txt, rtf


Employment

Select...
Select...

Select...
Select...
Select...
Select...