← Back to results

Staff Software Engineer (L4)

Join Twilio as a Staff Site Reliability Engineer to enhance the reliability of production services in a remote-first environment.

Location
Ireland
Compensation
Not disclosed
Level
staff
Type
full time · Remote

Posted by employer 2 days ago

First seen on Joblaze 1 day ago

Last verified on the company career page 1 day ago

What you'll build

  • Own the reliability posture of production services
  • Define, instrument, and operate against SLIs and SLOs
  • Drive down repair items and prevent classes of incidents
  • Write post-mortems that identify true root causes
  • Orchestrate complex changes across systems and services

Must have

  • 8+ years of related engineering experience
  • Demonstrated accountability for production systems
  • Strong software engineering fundamentals
  • Experience defining and operating against SLIs and SLOs
  • Depth in production operations
  • Experience driving changes that span multiple teams

Nice to have

  • Familiarity with infrastructure-as-code
  • Experience with multi-region architecture
  • Background in chaos engineering

Practical constraints

  • May be required to travel occasionally

Role intensity

40% coding

AI in the day-to-day

We use Artificial Intelligence (AI) to help make our hiring process efficient.

Requirements

Experience
8+ years

Not disclosed in this posting: compensation, visa sponsorship.

Benefits

Remote Work Health Insurance Parental Leave

Joblaze summary

In the role of Staff Site Reliability Engineer at Twilio, the individual is responsible for ensuring the reliability and performance of production services, focusing on system architecture and automation. Key skills include strong software engineering fundamentals, experience with SLIs and SLOs, and a background in incident management and observability. This position is suited for seasoned engineers with over eight years of experience in reliability or platform engineering, who can lead cross-team initiatives and mentor peers. Twilio's remote-first culture fosters a collaborative environment, allowing for diverse contributions to the team's goals.

Joblaze insights

  • Listed yesterday — first seen on Joblaze October 1, 2026. Last confirmed on Twilio's careers page October 1, 2026.
  • This exact title is also open at 1 other location at Twilio: Remote - Canada.

Quick facts

Is the Staff Software Engineer (L4) role remote?
Yes — Twilio lists this as a fully remote position.
How much experience is required?
At least 8 years of relevant experience for this Staff Software Engineer (L4) role.
What's the tech stack?
Joblaze extracted these technologies from the posting: Chaos Engineering, Cloud, Container Orchestration, GitOps, infrastructure as code.
What seniority level is this role?
Twilio targets staff-level candidates for this position.
Is this full-time or contract?
Full-time for this Staff Software Engineer (L4) role at Twilio.

From the original posting

Who we are

See yourself at Twilio

Join the team as Twilio’s next Staff SRE, Platform Engineering

About the job

Twilio is looking for a Staff Site Reliability Engineer to join our Platform Engineering organization. SRE owns production health and resiliency at Twilio — we are accountable for whether our services stay available, performant, and recoverable for the customers who build their businesses on us. SREs here are software engineers who focus on reliability: you will architect systems with reliability designed in from the outset, define the SLIs and SLOs that determine whether we are meeting customer expectations, and build the automation that keeps our infrastructure efficiently ahead of capacity and performance demand.

At the Staff SRE you work without day-to-day guidance, applying deep subject-matter knowledge and industry-leading practice to improve the products, processes, and services that Twilio runs on. You will own the reliability posture of significant parts of our production estate. Your impact will be felt across multiple teams rather than within one.

This is a hands-on engineering role. You will write and deploy code that improves service reliability, orchestrate complex changes across systems, lead the response when production is degraded, and raise the quality bar for the engineers around you.

Responsibilities

In this role, you’ll:

  • Own the reliability posture of production services in your area — availability, latency, capacity, efficiency, performance, and the monitoring and alerting that makes them visible
  • Define, instrument, and operate against SLIs and SLOs, and use error budgets to drive engineering priorities
  • Identify trends and problem areas that threaten stability, and provide a path forward to mitigate risk before it reaches customers
  • Drive down repair items and prevent classes of incidents rather than resolving them one at a time
  • Improve detection, response, and recovery — reducing time to acknowledge, engage, mitigate, and restore, with fewer people pulled in
  • Design for failure: strengthen failure domains, validate recovery paths, and make production changes safer to ship and safer to roll back
  • Participate in on-call for the services you support, and lead the response when production is degraded
  • Write post-mortems that identify true root causes, and drive the follow-up work to completion
  • Oversee efforts to identify, diagnose, report, and document production problems across all reliability dimensions
  • Write, configure, and deploy code that measurably improves service reliability — maintainable, reviewed, documented, and well tested
  • Orchestrate complex changes across systems and services, documenting design changes, technical decisions, migration plans, and upgrades
  • Lead debugging, troubleshooting, and analysis of service architecture and design
  • Use code review to drive up the quality of your coworkers' code
  • Reduce the operational overhead required to run infrastructure and services
  • Drive projects from conception to completion for efforts spanning the concerns of your team
  • Coordinate across programs, collaborating with others to estimate and communicate delivery timelines
  • Break projects into milestones and tasks, track progress, and communicate updates to stakeholders
  • Identify and communicate changes that may impact stability

*Required:

  • 8+ years of related engineering experience, with a substantial portion focused on reliability, infrastructure, or platform engineering
  • Demonstrated accountability for production systems — you have carried a pager for services that mattered and owned the outcome when they failed
  • Strong software engineering fundamentals: you build and ship production code, not only configure tooling
  • Experience defining and operating against SLIs and SLOs, and using error budgets to inform engineering priorities
  • Depth in production operations: incident command, post-mortem analysis, capacity planning, and observability
  • A record of preventing recurrence — reducing incident classes and operational toil, not just closing tickets
  • Experience driving changes that span multiple teams, and the communication skills to build alignment without formal authority
  • A track record of improving the engineers around you through code review, design feedback, and mentorship
  • Experience with large-scale distributed systems in a cloud environment

Desired:

  • Familiarity with infrastructure-as-code, container orchestration, and GitOps-style delivery
  • Experience with multi-region architecture, failure-domain design, or regional expansion work
  • Background in chaos engineering, game days, or other proactive resilience validation

Location

This role will be remote, and based in Ireland.

Twilio thinks big. Do you?

.

.

Standard company text repeated across Twilio's postings is omitted here.

Similar positions

Twilio
Staff Software Engineer
Twilio · Remote - Ireland
Twilio
Technical Support Engineer
Twilio · Remote - Ireland
Twilio
Staff Software Engineer
Twilio · Remote - US
Twilio
Twilio
Senior Software Engineer
Twilio · Remote - US