Contribute to the reliability and scalability of Databricks' realtime products as a Senior Production Engineer.
Posted by employer 1 day ago
First seen on Joblaze 1 day ago
Last verified on the company career page 5 hours ago
Not disclosed in this posting: compensation, work arrangement, visa sponsorship.
Joblaze summary
In the role of Senior Production Engineer for realtime products at Databricks, the individual will focus on enhancing the reliability and performance of systems like Lakebase and Neon through advanced monitoring and incident response strategies. Key skills include expertise in programming languages such as Python or Java, along with experience in infrastructure automation and incident management. This position is suited for professionals with over five years in production or site reliability engineering, who can navigate complex systems and collaborate effectively across teams.
Quick facts
- How much experience is required?
- At least 5 years of relevant experience for this Senior Production Engineer - Realtime Products role.
- What's the tech stack?
- Joblaze extracted these technologies from the posting: AWS, Azure, GCP, Go, Java, Kubernetes.
- What seniority level is this role?
- Databricks targets senior candidates for this position.
- Is this full-time or contract?
- Full-time for this Senior Production Engineer - Realtime Products role at Databricks.
From the original posting
CSQ127R154
As a Senior Production Engineer working on Databricks’ realtime products you will be directly contributing to our customers’ success. You will build advanced monitoring and incident mitigation tooling, and drive changes across the stack to proactively make the realtime products like Lakebase, Neon and Model Serving reliable, secure, and scalable in production.
Our production engineers understand the Databricks platform from end to end, and partner across engineering teams, building durable solutions for a platform operating across AWS, Azure, and GCP.
The Impact You’ll Have
Observability
- Build advanced monitoring and detection capabilities that give early warning of issues and deep insights into workload performance and customer experience.
- Build reliable, observable automation for debugging and incident mitigation.
Reliability
- Improve service reliability, scalability, security, and operational efficiency.
- Develop dependable, safe, mitigations for production issues.
On-Call & Incident Response
- Participate in a follow-the-sun on-call rotation and lead incident response and mitigation.
- Perform root-cause analysis, identify and implement lasting corrective actions.
- Partner with Product Engineering, Security, Support, and other infrastructure teams on follow up actions.
What We Look For
Experience
- 5+ years of experience in Production Engineering, Customer Reliability Engineering (CRE), Site Reliability Engineering (SRE), infrastructure engineering, backend software engineering, or a related field.
- Experience in holistic monitoring and alerting of complex stateful systems, applying multiple strategies like workload alerting, anomaly detection and probing.
- Experience with PostgreSQL, or related managed databases or distributed systems.
- Experience in incident management and participating in on-call rotations for critical infrastructure.
Skillset
- A mindset focused on automation, root-cause resolution, and continuous improvement.
- Strong programming skills in one or more languages such as Python, Go, Java, Scala, or similar.
- Proficiency with infrastructure automation and Infrastructure as Code.
- Ability to work across system boundaries and collaborate effectively during complex incidents, and engage with customers’ infrastructure teams.
- Experience in using AI to address production and operational challenges.
Bonus
- Experience with AWS, Azure, or GCP.
- Experience with Lakebase or Neon.
- Experience with Kubernetes, Terraform.
- Experience building internal platforms, operational tooling, or developer productivity systems.
Education
- BS degree (or higher) in Computer Science, or a related field.
Standard company text repeated across Databricks's postings is omitted here.