← Back to results

Staff Reliability Engineer

Join Okta as a Staff Reliability Engineer to enhance the resilience and availability of our global corporate network.

Location
Bengaluru, India
Compensation
Not disclosed
Level
staff
Type
full time · Hybrid

Posted by employer 3 days ago

First seen on Joblaze 3 days ago

Last verified on the company career page 4 hours ago

What you'll build

  • Design and own the resilience, health and availability of the global corporate network
  • Drive strategic reduction of systemic toil and technical debt
  • Collaborate with cross-functional stakeholders
  • Lead resolution of complex network operations issues
  • Foster learning and talent within the team

Must have

  • 8+ years of related experience
  • Deep expertise in AWS Networking and Palo Alto Networks solutions
  • Comprehensive operational experience in Distributed Systems & Networking fundamentals
  • Strong proficiency in Cloud Platforms, IaC, Observability tools, Programming, and Service Reliability Management
  • Proven track record of managing operational availability

Nice to have

  • Experience with Juniper/JUNOS switching/routing
  • Experience with Palo Alto Networks NGFWs
  • Experience with enterprise office build and construction processes

Requirements

Experience
8+ years
Education
Bachelor's degree

Not disclosed in this posting: compensation, visa sponsorship.

Joblaze summary

In this role, the Staff Reliability Engineer at Okta focuses on ensuring the resilience and availability of the global corporate network, managing operational tasks like responding to alerts and executing reliability projects. Key skills include expertise in AWS Networking, Distributed Systems, and Infrastructure as Code, along with proficiency in programming and observability tools. This position is suited for experienced professionals with a strong operational background who can navigate complex challenges and collaborate effectively with cross-functional teams. The role emphasizes a culture of learning and innovation, aiming to reduce manual toil and enhance network security.

Joblaze insights

  • Listed 3 days ago — first seen on Joblaze September 17, 2026. Last confirmed on Okta's careers page September 20, 2026.

Quick facts

Is the Staff Reliability Engineer role remote?
It's hybrid — Okta expects some on-site time in Bengaluru, India.
How much experience is required?
At least 8 years of relevant experience for this Staff Reliability Engineer role.
Where is the role based?
Okta is hiring for this position in Bengaluru, India.
What's the tech stack?
Joblaze extracted these technologies from the posting: AWS, Ansible, Go, Grafana, Networking, Observability.
What seniority level is this role?
Okta targets staff-level candidates for this position.
Is this full-time or contract?
Full-time for this Staff Reliability Engineer role at Okta.

From the original posting

Secure Every Identity, from AI to Human

Okta’s TDI Network Engineering team is responsible for the global corporate network, building and supporting a high-performing, reliable network at scale. As a member of this team, you will have a direct impact on network design, deployment, and reliability, enabling our employees to work effectively from any location globally. Your role ensures the overall security and integrity of our corporate network by leveraging network security best practices, innovative products, and rigorous security validation.

Reporting to the Network Engineering Manager, this operations-focused role is distinct from core Network Engineering and Network Security, centering primarily on operational execution—including responding to alerts, maintaining service availability, and ensuring system health across our global enterprise network. You will drive the strategic reduction of systemic toil and technical debt across multiple teams, applying a systems-level perspective and leveraging deep expertise in Distributed Systems, Networking fundamentals, Infrastructure as Code, and observability to architect scalable platforms and lead technical efforts to ensure an "Always Secure. Always On." environment. You will own multi-quarter objectives and establish long-term strategies for network reliability.

What you'll be doing :

  • Design and Own the resilience, health and availability of our entire global corporate network domain, managing operational responsibilities such as responding to alerts, monitoring health indicators, and executing reliability projects to ensure an "Always Secure. Always On." environment.
  • Drive Strategic Reduction of systemic toil and technical debt across multiple teams by introducing process efficiencies, automating network operations, and building scalable self-service operational tooling.
  • Collaborate and Influence closely with cross-functional stakeholders—including Business Technology, Workplace, Security, and executive leaders—challenging assumptions with grace and cascading relevant information to project teams.
  • Make Critical Decisions and lead the resolution of complex network operations issues and alerts from a systems perspective, anticipating potential business challenges, monitoring leading indicators, and preventing future outages.
  • Foster Learning and Talent by defining success for the whole team, cultivating an open and transparent environment, and actively mentoring team members through the P4 level to develop their skills and operational engineering best practices.

What you'll bring to the role:

  • Typically requires 8+ years of related experience in a professional role with a Bachelor’s degree; or 6+ years with a Master’s degree; or 3+ years with a PhD; or equivalent experience.
  • Deep expertise in AWS Networking and Palo Alto Networks solutions as core required technical competencies.
  • Comprehensive operational experience in Distributed Systems & Networking fundamentals, including quick incident response to alerts, monitoring system health, and managing protocols such as WiFi, DNS, DHCP, VLANs, VPN, ACLs, Routing, and Firewall Policies.
  • Strong proficiency in core technical skills: Cloud Platforms, IaC (e.g., Terraform/Ansible), Observability tools (e.g., Prometheus/Grafana), Programming (Python/Go), and Service Reliability Management (SLOs/SLIs) for large-scale enterprise environments.
  • Proven track record of managing operational availability, delivering multi-quarter objectives, and executing technical projects within defined budgets and strategic VMTs (Vision, Mission, Targets).
  • Demonstrated ability to navigate high levels of ambiguity, establish credibility with executive stakeholders, and model resilience during major system transitions or production incidents.

Why join us:

  • Desirable Technical Skills: Experience with Juniper/JUNOS switching/routing, Palo Alto Networks NGFWs, and enterprise office build and construction processes is highly desired and will help you hit the ground running.
  • Direct Customer Impact: You will have the opportunity to showcase your strong focus on customer and technology experiences, ensuring a secure, consistent end-user experience across a global enterprise network.
  • Growth and Support: Join a collaborative environment that values systemic learning over personal errors, offering you the chance to eliminate manual toil through innovation, mentor peers, and occasionally travel to build impactful connections.

#P10445_3513586
#LI_Hybrid

The Okta Experience

Standard company text repeated across Okta's postings is omitted here.

Similar positions