← Back to results

Senior Software Engineer - Incident Insights & Readiness

Own and improve the on-call experience while leading incident response initiatives in a hybrid workplace at Datadog.

Location
Boston, Massachusetts, USA; New York, New York, USA
Compensation
$192k–$240k/yr
Level
senior
Type
full time · Hybrid

Posted by employer 22 hours ago

First seen on Joblaze 17 hours ago

Last verified on the company career page 5 hours ago

Skills & Technologies

What you'll build

  • Own and improve the on-call experience
  • Define incident response processes
  • Contribute to post-mortem processes
  • Support teams in incident reviews
  • Provide technical leadership and coaching

Must have

  • 5+ years of experience building software
  • Experience with Go and Python
  • Experience building or operating distributed systems
  • Experience analyzing incidents
  • Experience participating in on-call rotations

Nice to have

  • Experience serving as an incident commander
  • Experience mentoring engineers

Requirements

Experience
5+ years

Not disclosed in this posting: visa sponsorship.

Benefits

401k Match Equity/Stock Options Remote Work Health Insurance Parental Leave

Joblaze summary

In this role, the Senior Software Engineer focuses on enhancing the incident response process by developing software and operational frameworks that support Datadog's engineering teams during incidents. Proficiency in Go and Python, along with experience in distributed systems and incident management, are essential for success. This position is well-suited for candidates with a strong background in software or site reliability engineering, particularly those who have experience mentoring others and driving cross-functional initiatives. The team emphasizes a culture of learning and collaboration, aiming to improve resilience across the organization.

Joblaze insights

  • Listed today — first seen on Joblaze October 2, 2026. Last confirmed on Datadog's careers page October 2, 2026.
  • This exact title is also open at 1 other location at Datadog: Paris, France.
  • Salary band is above the typical range for DevOps/SRE roles (median ~$165,000).
  • Starts above 81% of 80 comparable senior devops/sre roles in United States that list Kubernetes we track (median $165,000 across 39 companies). See Kubernetes salary trends
  • Kubernetes appears in 67.9% of 187 comparable senior devops/sre roles in United States; TypeScript appears in 10.7% of 187 comparable senior devops/sre roles in United States.

Quick facts

Is the Senior Software Engineer - Incident Insights & Readiness role remote?
It's hybrid — Datadog expects some on-site time in Boston, Massachusetts, USA; New York, New York, USA.
What's the salary range?
Datadog lists $192,000–$240,000 for this role.
How much experience is required?
At least 5 years of relevant experience for this Senior Software Engineer - Incident Insights & Readiness role.
Where is the role based?
Datadog is hiring for this position in Boston, Massachusetts, USA; New York, New York, USA.
What's the tech stack?
Joblaze extracted these technologies from the posting: Go, Kubernetes, Python, TypeScript.
What seniority level is this role?
Datadog targets senior candidates for this position.
Is this full-time or contract?
Full-time for this Senior Software Engineer - Incident Insights & Readiness role at Datadog.

From the original posting

We're on a mission to build the best platform in the world for engineers to understand and scale their systems, applications, and teams. We operate at high scale—trillions of data points per day—providing always-on alerting, metrics visualization, logs, and application tracing for tens of thousands of companies. Our engineering culture values pragmatism, honesty, and simplicity to solve hard problems the right way

The Incident Insights & Readiness SRE team at Datadog fosters a resilient culture by using incidents as learning opportunities and catalysts for growth. Our users are Datadog engineers, and we build the software, tooling, and operational frameworks that help them prepare for, respond to, and learn from incidents. We work closely with engineering teams across Datadog to analyze incidents and turn those insights into better tools, stronger incident response, and organizational learning. Our efforts empower Datadog to navigate unexpected failures confidently, efficiently, and with a commitment to continuous learning and systems improvement.

At Datadog, we place value in our office culture - the relationships and collaboration it builds and the creativity it brings to the table. We operate as a hybrid workplace to ensure our Datadogs can create a work-life harmony that best fits them.

What You’ll Do:

  • Own and improve the on-call experience for the company by establishing best practices and building platforms to support on-call rotations and compensation.

  • Define how we respond to incidents, lead the design and implementation of software to streamline the process, and collaborate with product teams to improve incident response across Datadog. Our aim is to fully support our incident responders in dealing with complexity.

  • Contribute to the post-mortem process for the company, collaborating with teams on writing them, and identifying opportunities to reduce friction and enhance learning value for the organization. Our team also runs a weekly postmortem reading group.

  • Support various teams in facilitating incident reviews that emphasize learning and blamelessness. Help them share their learnings across the organization to improve the resilience of our people.

  • Provide technical leadership and day-to-day coaching to team members, accelerating their growth through design reviews, collaborative problem-solving and operational excellence best practices.

  • Train our on-callers in incident and post-mortem processes, sharing expertise in incident management best practices. This involves both introducing newcomers to on-call responsibilities and refreshing the knowledge of existing engineers.

  • Lead cross-functional initiatives in engineering organizations across Datadog, embedding with teams to understand their challenges and drive lasting improvements to reliability and operational excellence.

Who You Are:

  • At least 5 years of experience building software that solves real user problems. Experience designing new features and collaborating on code and technical design reviews. We primarily develop in Go and Python, with a bit of TypeScript.

  • Experience building or operating distributed systems, with familiarity with Kubernetes and an understanding of complex failure modes.

  • Demonstrated ability to independently own ambiguous technical problems from design through delivery while balancing long-term engineering quality with pragmatic execution.

  • Experience analyzing incidents, identifying systemic risks, and driving engineering improvements informed by operational learnings.

  • Experience participating in on-call rotations and improving incident response processes. Experience serving as an incident commander or incident coordinator is a plus.

  • Empathy, collaboration, and communication skills in English to cultivate strong relationships across various teams in the organization

  • Experience mentoring engineers, driving cross-functional initiatives, and influencing technical direction without relying on organizational authority.

  • We welcome candidates from a variety of backgrounds, including software engineering, site reliability engineering, production engineering, infrastructure, and other roles focused on building reliable systems or improving incident response.

Benefits and Growth:

  • New hire stock equity (RSUs) and employee stock purchase plan (ESPP)

  • Continuous professional development, product training, and career pathing

  • Intradepartmental mentor and buddy program for in-house networking

  • An inclusive company culture, ability to join our Community Guilds (Datadog employee resource groups)

  • Access to Inclusion Talks, our internal panel discussions

  • Free, global mental health benefits for employees and dependents age 6+

  • Competitive global benefits

Datadog offers a competitive salary and equity package, and may include variable compensation. Actual compensation is based on factors such as the candidate's skills, qualifications, and experience. In addition, Datadog offers a wide range of best in class, comprehensive and inclusive employee benefits for this role including healthcare, dental, parental planning, and mental health benefits, a 401(k) plan and match, paid time off, fitness reimbursements, and a discounted employee stock purchase plan.

$192,000—$240,000 USD

About Datadog:

Standard company text repeated across Datadog's postings is omitted here.

Similar positions

Datadog
Developer Advocate - Service Management
Datadog · California, USA, Remote; New York, USA, Remote
Datadog
Developer Advocate - Service Management EMEA
Datadog · France, Remote; Spain, Remote; The Netherlands, Remote
Datadog
Senior Product Marketing Manager - Incident Response
Datadog · New York, New York, USA; San Francisco, California, USA
Datadog
Senior Staff Software Engineer
Datadog · New York, New York, USA