← Back to results

Site Reliability Engineer

Join Razorpay as a founding SRE to enhance payment reliability for millions of businesses.

Location
Bengaluru, India
Compensation
Not disclosed
Level
senior
Type
full time

Posted by employer 1 day ago

First seen on Joblaze 1 day ago

Last verified on the company career page 1 day ago

Apply at Razorpay → Save job Scanned from razorpay.com

What you'll build

  • Define SLIs, SLOs, and error budgets for payment flows
  • Own the release lifecycle for payment services
  • Carry the pager for payment-critical services
  • Eliminate toil through software
  • Run production readiness reviews for new payment services

Must have

  • 10+ years of engineering experience
  • 5+ years operating large-scale distributed systems
  • Strong software engineering skills in Go, Java, or Python
  • Deep understanding of distributed systems failure modes
  • Hands-on experience designing deployment pipelines
  • Genuine on-call ownership

Nice to have

  • Experience in payments, fintech, or banking
  • Experience with Kubernetes at scale
  • Experience with service mesh and traffic management

Requirements

Experience
10+ years

Not disclosed in this posting: compensation, work arrangement, visa sponsorship.

Joblaze summary

The Site Reliability Engineer at Razorpay focuses on enhancing the reliability of payment systems, aiming to elevate service availability from three nines to four or five nines. This role requires strong software engineering skills in languages like Go, Java, or Python, along with a deep understanding of distributed systems and deployment pipelines. Ideal candidates have over a decade of engineering experience, particularly in high-stakes environments where system reliability is critical. Razorpay's culture emphasizes ownership and transparency, making it a fitting environment for those who thrive on autonomy and collaboration.

Joblaze insights

  • Listed yesterday — first seen on Joblaze September 25, 2026. Last confirmed on Razorpay's careers page September 25, 2026.

Quick facts

How much experience is required?
At least 10 years of relevant experience for this Site Reliability Engineer role.
What's the tech stack?
Joblaze extracted these technologies from the posting: Datadog, Go, Grafana, Java, Kubernetes, Linux.
What seniority level is this role?
Razorpay targets senior candidates for this position.
Is this full-time or contract?
Full-time for this Site Reliability Engineer role at Razorpay.

From the original posting

Razorpay is one of India’s leading full-stack financial technology companies, powering the way businesses move, manage, and grow money. Founded in 2014 by Harshil Mathur and Shashank Kumar with a simple vision - to simplify payments for Indian businesses - we’ve since grown into a fintech powerhouse driving India’s digital payment revolution.

About the role

You will be one of the founding SREs at Razorpay, embedded with the payment platform teams that move money for millions of businesses. Your mandate is to take our payment flows from three nines to four and five nines of availability. In payments, a failed request is not a retry, it is a customer's money in limbo. You will define what reliability means here, build the systems that enforce it, and set the standard every future SRE is measured against.

What you will do

  • Define SLIs, SLOs, and error budgets for critical payment flows (authorization, capture, refunds, settlements, webhooks) and make them the shared language between product and platform teams.
  • Own the release lifecycle for payment services: design progressive rollout pipelines (canary, staged, feature-flagged), automated rollback triggers, and make "can we roll back in under 5 minutes" a launch-blocking question.
  • Carry the pager for payment-critical services, lead incident command during outages, and drive blameless postmortems where action items actually ship.
  • Eliminate toil through software: build automation for failover, capacity management, load shedding, and degradation so that known failure classes cannot recur.
  • Harden payment flows against distributed systems failure modes: retry storms, thundering herds, cascading failures, partial outages of banks and network partners, idempotency violations, and reconciliation gaps.
  • Run production readiness reviews for new payment services and hold the line on launch gates using error budget data, not opinion.
  • Instrument what matters: design alerting that pages on customer-facing symptoms, not noise, and cut mean time to detection and recovery quarter over quarter.
  • Practice failure on purpose: game days, chaos experiments, and failure injection against payment-critical paths.

What we are looking for

  • 10+ years of engineering experience, with at least 5 years operating large-scale distributed systems in production (high QPS, multi-region, or systems where sub-1 percent error rates were business-critical).
  • Strong software engineering skills in at least one of Go, Java, or Python. You have built tools and services, not just configured them.
  • Deep understanding of distributed systems failure modes and the patterns that contain them: circuit breakers, backpressure, bulkheading, graceful degradation, idempotency.
  • Solid fundamentals in Linux internals, networking, and databases under load (replication, failover, connection pool exhaustion, lock contention).
  • Hands-on experience designing or significantly improving deployment pipelines: canary analysis, automated rollback, feature flags.
  • Genuine on-call ownership: you have carried a pager for systems that mattered, led incidents, and can walk us through a specific outage you handled and what you changed afterward.
  • Fluency with modern observability (metrics, tracing, structured logging; e.g. Prometheus, Grafana, OpenTelemetry, Datadog, Coralogix, Clickhouse or similar) and experience reducing alert noise.
  • Experience defining SLOs and error budget policies from scratch. Contributions to reliability tooling, open source or internal, that other teams adopted.
  • The judgment and communication skills to tell a product team "not yet" with data, and the pragmatism to help them get to "yes" quickly.

Nice to have

  • Experience in payments, fintech, banking, trading, or another domain where correctness and money are coupled (transactional consistency, exactly-once semantics, reconciliation).
  • Experience with Kubernetes at scale, service mesh, and traffic management.
Razorpay believes in and follows an equal employment opportunity policy that doesn't discriminate on gender, religion, sexual orientation, colour, nationality, age, etc. We welcome interests and applications from all groups and communities across the globe.
Follow us on LinkedIn & Twitter

Standard company text repeated across Razorpay's postings is omitted here.

Similar positions