← Back to results

Incident Manager

Lead critical production incidents at Databricks, ensuring timely communication and operational resilience during high-impact events.

Location
United States
Compensation
$103.9k–$145.5k/yr
Level
senior
Type
full time · Remote

Posted by employer 3 months ago

First seen on Joblaze 3 months ago

Last verified on the company career page 14 hours ago

Requirements

Experience
5+ years
Education
Bachelor's degree

Not disclosed in this posting: visa sponsorship.

Joblaze summary

In the role of Incident Manager at Databricks, the individual will oversee critical production incidents, coordinating responses across multiple teams to swiftly mitigate issues and restore services. This position requires a strong background in cloud infrastructure and incident management, along with proficiency in log analysis and programming for automation. Ideal candidates will have over five years of experience in site reliability engineering or production operations, making this role suitable for seasoned professionals looking to enhance operational resilience in a fast-paced environment.

Joblaze insights

  • Listed about 3 months ago — first seen on Joblaze July 1, 2026. Last confirmed on Databricks's careers page October 9, 2026.
  • Salary band is below the typical range for DevOps/SRE roles (median ~$165,000).
  • Starts at or below all 81 comparable senior devops/sre roles in United States that list Python we track (median $165,000 across 38 companies). See Python salary trends
  • Python appears in 60.8% of 189 comparable senior devops/sre roles in United States; Elasticsearch appears in 0.5% of 189 comparable senior devops/sre roles in United States.

Quick facts

Is the Incident Manager role remote?
Yes — Databricks lists this as a fully remote position.
What's the salary range?
Databricks lists $103,900–$145,525 for this role.
How much experience is required?
At least 5 years of relevant experience for this Incident Manager role.
What's the tech stack?
Joblaze extracted these technologies from the posting: AWS, Azure, Cloud Logging, Datadog, Elasticsearch, GCP.
What seniority level is this role?
Databricks targets senior candidates for this position.
Is this full-time or contract?
Full-time for this Incident Manager role at Databricks.

From the original posting

Incident Manager

US Remote

CSQ127R151

At Databricks, we are passionate about empowering data teams to tackle the world’s most challenging problems — from bringing the next mode of transportation to reality to accelerating the development of medical breakthroughs. We achieve this by building and operating the world’s best data and AI infrastructure platform, enabling our customers to leverage deep data insights and enhance their business. Founded by engineers — and customer-obsessed — we leap at every opportunity to tackle technical challenges, from designing next-gen UI/UX for interfacing with data to scaling our services and infrastructure across millions of virtual machines. And we're only getting started.

As an Incident Manager, you will lead Databricks’ most critical production incidents while providing clear, accurate, and timely communication to customers, executives, and engineers. You’ll serve as both incident commander and reliability engineer; orchestrating multi-team responses, driving real-time status updates, and partnering with engineering to analyze and prevent failures. Your work will ensure Databricks maintains not only technical resilience but also customer and stakeholder confidence during high-impact events.

This role combines operational leadership, technical systems knowledge, and exceptional communication skills. You will be at the intersection of engineering depth and operational clarity, ensuring that every major incident is managed with precision, transparency, and continuous improvement.

The impact you will have here:

  • Lead critical incidents — coordinate multi-disciplinary response efforts across Databricks’ cloud-based services to rapidly mitigate impact and restore operations.
  • Drive technical root cause analysis and reliability improvements:
    • collaborate with engineering teams to trace and document underlying causes across distributed systems, services, and data stores.
    • Summarize key learnings, clearly communicate action items, and ensure that technical and procedural improvements are followed through.
  • Own communications during incidents — deliver frequent, high-quality updates to internal stakeholders (executives, engineering leadership, support) and compose and publish customer-facing notifications that are accurate, timely, and empathetic.
  • Mentor and train peers in both incident communication and technical response disciplines to raise the overall quality of Databricks’ incident response.

What are we looking for:

  • 5+ years of experience in incident management, site reliability engineering, or production operations supporting large-scale, cloud-native systems.
  • Proven ability to lead and coordinate high-severity incidents, including identifying impact, isolating fault domains, and managing multi-team response efforts.
  • Strong understanding of cloud infrastructure (AWS, Azure, or GCP) — including compute, networking, storage, and observability components.
  • Deep expertise in log analysis and debugging:
    • Familiarity with log aggregation and search tools (e.g., Datadog, Elasticsearch, Splunk, Cloud Logging, or OpenTelemetry).
    • Hands-on experience with observability systems — metrics, logging, and tracing frameworks (Prometheus, Grafana, OpenTelemetry, etc.).
  • Proficiency in at least one major programming or scripting language (Python, Go, or Bash) for automating diagnostics, data collection, or analysis.
  • Experience developing and maintaining incident playbooks and communication templates to ensure consistent, timely updates.
  • Excellent contextual interpretation and writing skills, as well as the ability to effectively summarize and communicate to both technical and business audiences, are required.
  • BS, Master's or other advanced degree in Computer Science or Computer Engineering, or related Engineering field.

Pay Range Transparency

Databricks is committed to fair and equitable compensation practices. The pay range(s) for this role is listed below and represents the expected base salary range for non-commissionable roles or on-target earnings for commissionable roles. Actual compensation packages are based on several factors that are unique to each candidate, including but not limited to job-related skills, depth of experience, relevant certifications and training, and specific work location. Based on the factors above, Databricks anticipated utilizing the full width of the range. The total compensation package for this position may also include eligibility for annual performance bonus, equity, and the benefits listed above. For more information regarding which range your location is in visit our page here.

Zone 1 Pay Range
$144,600—$198,900 USD
Zone 2 Pay Range
$130,200—$178,950 USD
Zone 3 Pay Range
$123,000—$169,050 USD
Zone 4 Pay Range
$115,700—$159,050 USD

About Databricks

Benefits

Standard company text repeated across Databricks's postings is omitted here.

Similar positions

Databricks
Staff Security Engineer, Incident Response
Databricks · Remote - California
Databricks
Sr Platform Monitoring Engineer
Databricks · United States
Databricks
Staff Security Engineer, Incident Response
Databricks · Belgium; Finland; Remote - Denmark; Remote - France; Remote - Germany; Remote - Italy; Remote - Netherlands; Remote - Spain; Remote - Sweden; Remote - United Kingdom; Switzerland
Databricks
Senior Manager, Infrastructure Data Science
Databricks · Mountain View, California
Databricks
Sr. Manager, Engineering
Databricks · Amsterdam, Netherlands