Join Datadog as a Senior Software Engineer to enhance reliability through chaos engineering and automation in a hybrid work environment.
Posted by employer 17 hours ago
First seen on Joblaze 6 hours ago
Last verified on the company career page 6 hours ago
Not disclosed in this posting: compensation, years of experience, visa sponsorship.
Joblaze summary
In the role of Senior Software Engineer on Datadog's Chaos Engineering team, the individual will focus on developing automation for zonal resilience, ensuring services can effectively evacuate and recover from failures. Key skills include a strong understanding of distributed systems, Kubernetes, and experience with fault injection and reliability engineering. This position is ideal for someone with a solid background in production systems and a collaborative mindset, as it involves working across various engineering teams to enhance system resilience. Datadog fosters a culture of collaboration and creativity, operating in a hybrid work environment.
Quick facts
- Is the Senior Software Engineer, Chaos Engineering role remote?
- It's hybrid — Datadog expects some on-site time in Paris, France.
- Where is the role based?
- Datadog is hiring for this position in Paris, France.
- What's the tech stack?
- Joblaze extracted these technologies from the posting: Kubernetes, gRPC.
- What seniority level is this role?
- Datadog targets senior candidates for this position.
- Is this full-time or contract?
- Full-time for this Senior Software Engineer, Chaos Engineering role at Datadog.
From the original posting
Datadog’s Chaos Engineering team builds systems that surface reliability weaknesses before they become outages. As a Senior Software Engineer, you will initially focus on zonal resilience, building automation that helps services safely evacuate and recover from zonal failures, while also contributing to fault injection, incident replay, gameday orchestration, and reliability tooling. You will work across engineering teams to design systems that safely exercise production failure modes and turn findings into verified remediation. You will also help advance the use of AI and automation to identify, test, and close resilience gaps as Datadog’s software and infrastructure evolve.
What You’ll Do:
- Build zonal-resilience automation that coordinates safe workload evacuations, switchovers, and recovery in partnership with the teams that own affected services.
- Design and build fault-injection systems for production environments, including infrastructure- and application-level testing, incident replay, and controlled resilience experiments.
- Develop safeguards such as blast-radius controls, kill switches, validation mechanisms, and rollback paths that keep production experiments contained and reversible.
- Build agents and automation that help propose failure scenarios, triage experiment results, and connect reliability findings to tracked remediation and verification.
- Lead gamedays from hypothesis and scenario design through execution, documented findings, remediation tracking, and validation of completed fixes.
- Design and implement reliable distributed systems, including gRPC services, Kubernetes controllers, and shared platform components, while contributing to technical design and mentoring other engineers.
Who You Are:
- You have strong distributed systems fundamentals and can reason about consistency, failure modes, backpressure, idempotency, quorum, retries, and failure recovery.
- You understand Kubernetes workload lifecycles, including how pods, controllers, scheduling, draining, and eviction interact with resilient system design.
- You have experience designing, building, or operating production systems where safety, availability, and controlled failure handling are important.
- You communicate complex technical decisions clearly through design documents, runbooks, postmortems, and cross-functional technical discussions.
- You are comfortable collaborating across engineering teams to understand unfamiliar systems, identify failure modes, and drive resilience improvements.
- Experience with reliability engineering, chaos engineering, zonal failover, AI-assisted operational workflows, traffic interception, or large-scale observability systems is beneficial but not required.
Datadog values people from all walks of life. We know not everyone will meet all the above qualifications on day one. That’s okay. If you’re passionate about technology and want to grow your experience, we encourage you to apply.
Benefits and Growth:
- Develop deep expertise in distributed systems, production resilience, Kubernetes, and large-scale infrastructure.
- Work on reliability systems that operate across Datadog’s production environment and influence how engineering teams design for failure.
- Grow your experience designing safe, automated approaches to fault injection, zonal resilience, and incident reproduction.
- Explore practical applications of AI and automation to reliability engineering and operational workflows.
- Collaborate with engineers across infrastructure, databases, observability, and service teams on complex systems challenges.
- Mentor other engineers and contribute to technical designs, engineering practices, and platform strategy.
- Benefits and Growth listed above may vary based on the country of your employment and the nature of your employment with Datadog.
#LI-Hybrid
Standard company text repeated across Datadog's postings is omitted here.