Lead critical production incidents at Databricks, ensuring timely communication and operational resilience during high-impact events.
Posted by employer 3 months ago
First seen on Joblaze 3 months ago
Last verified on the company career page 14 hours ago
Skills & Technologies
Requirements
Not disclosed in this posting: visa sponsorship.
Joblaze summary
In the role of Incident Manager at Databricks, the individual will oversee critical production incidents, coordinating responses across multiple teams to swiftly mitigate issues and restore services. This position requires a strong background in cloud infrastructure and incident management, along with proficiency in log analysis and programming for automation. Ideal candidates will have over five years of experience in site reliability engineering or production operations, making this role suitable for seasoned professionals looking to enhance operational resilience in a fast-paced environment.
Joblaze insights
Quick facts
From the original posting
At Databricks, we are passionate about empowering data teams to tackle the world’s most challenging problems — from bringing the next mode of transportation to reality to accelerating the development of medical breakthroughs. We achieve this by building and operating the world’s best data and AI infrastructure platform, enabling our customers to leverage deep data insights and enhance their business. Founded by engineers — and customer-obsessed — we leap at every opportunity to tackle technical challenges, from designing next-gen UI/UX for interfacing with data to scaling our services and infrastructure across millions of virtual machines. And we're only getting started.
As an Incident Manager, you will lead Databricks’ most critical production incidents while providing clear, accurate, and timely communication to customers, executives, and engineers. You’ll serve as both incident commander and reliability engineer; orchestrating multi-team responses, driving real-time status updates, and partnering with engineering to analyze and prevent failures. Your work will ensure Databricks maintains not only technical resilience but also customer and stakeholder confidence during high-impact events.
This role combines operational leadership, technical systems knowledge, and exceptional communication skills. You will be at the intersection of engineering depth and operational clarity, ensuring that every major incident is managed with precision, transparency, and continuous improvement.
Pay Range Transparency
Databricks is committed to fair and equitable compensation practices. The pay range(s) for this role is listed below and represents the expected base salary range for non-commissionable roles or on-target earnings for commissionable roles. Actual compensation packages are based on several factors that are unique to each candidate, including but not limited to job-related skills, depth of experience, relevant certifications and training, and specific work location. Based on the factors above, Databricks anticipated utilizing the full width of the range. The total compensation package for this position may also include eligibility for annual performance bonus, equity, and the benefits listed above. For more information regarding which range your location is in visit our page here.
About Databricks
BenefitsStandard company text repeated across Databricks's postings is omitted here.