← Back to results

Site Reliability Engineering Manager

Lead the Site Reliability Engineering team to ensure reliable and secure deployments in demanding environments.

Location
Colorado Springs, CO
Compensation
Not disclosed
Level
lead
Type
full time · Hybrid

Posted by employer 4 days ago

First seen on Joblaze 4 days ago

Last verified on the company career page 6 hours ago

Apply at Onebrief → Save job Scanned from onebrief.com

What you'll build

  • Lead and develop the SRE team
  • Own capacity and work planning
  • Coordinate delivery across teams
  • Set the reliability roadmap
  • Support team-led incident response

Must have

  • Active Secret clearance
  • 5+ years in Site Reliability Engineering, Platform Engineering, DevOps, or a related role
  • Experience directly managing engineers
  • Experience planning team capacity
  • Experience with incident response

Nice to have

  • Experience leading teams supporting mission-critical customer deployments
  • Experience managing engineering programs
  • Experience operating in DoD, classified, or air-gapped environments
  • Familiarity with RMF, STIGs, and ICD 503
  • Experience implementing SLIs, SLOs, and error budgets

Practical constraints

  • Active Secret clearance required
  • Must be willing to relocate

Role intensity

10% coding — mostly leadership/strategy

Requirements

Experience
5+ years

Not disclosed in this posting: compensation, visa sponsorship.

Benefits

Relocation Assistance

Joblaze summary

The Site Reliability Engineering Manager at Onebrief oversees a team responsible for ensuring the reliability and security of mission-critical software deployments in both on-prem and AWS environments. This role requires a strong technical background in infrastructure, automation, and incident response, along with effective people management skills to guide engineers and coordinate cross-team efforts. Ideal candidates have significant experience in SRE or related fields, particularly in high-stakes environments, and possess an active Secret clearance. Onebrief's focus on military planning adds a unique context to the team's operational challenges.

Joblaze insights

  • Listed 4 days ago — first seen on Joblaze September 19, 2026. Last confirmed on Onebrief's careers page September 23, 2026.

Quick facts

Is the Site Reliability Engineering Manager role remote?
It's hybrid — Onebrief expects some on-site time in Colorado Springs, CO.
How much experience is required?
At least 5 years of relevant experience for this Site Reliability Engineering Manager role.
Where is the role based?
Onebrief is hiring for this position in Colorado Springs, CO.
What's the tech stack?
Joblaze extracted these technologies from the posting: AWS, AWS GovCloud, Ansible, CI/CD, Datadog, ELK.
What seniority level is this role?
Onebrief targets lead candidates for this position.
Is this full-time or contract?
Full-time for this Site Reliability Engineering Manager role at Onebrief.

From the original posting

Consequential Work. Dedicated People.

Security Clearance, Location, and Onsite Notice

This role requires regularly working on-site at customer locations.

If you are not currently within commuting distance, you must be willing to relocate. Onebrief provides relocation assistance.

Active Secret clearance required.

About The Role

We're hiring a Site Reliability Engineering Manager to lead our SRE team within Infrastructure & Security. You'll work closely with platform engineering, application engineering, security, and customer success to ensure Onebrief's mission-critical deployments are reliable, secure, and well supported across on-prem DoD and AWS environments.

You'll lead a team whose work spans customer-facing operations, infrastructure, observability, automation, and application reliability. Most of the team focuses on deploying and operating Onebrief in demanding customer environments.

You'll own the team's priorities, planning, execution, and development. A significant part of this role is coordinating work: understanding demand, balancing capacity, sequencing tasks, managing dependencies, and keeping commitments realistic as customer needs change. You'll help the team deliver immediate operational support while making steady progress on improvements that reduce future support demands.

You'll bring the technical grounding to evaluate risks, ask useful questions, and guide decisions. Your engineers will own technical implementation and lead incident response. You'll provide direction, remove blockers, and create the conditions for them to succeed.

About You

You care deeply about reliability and understand the challenges of operating software in environments where connectivity, access, and deployment options can be constrained. You treat infrastructure and operability as products that deserve clear ownership, thoughtful design, and continuous improvement.

You're an effective people manager who sets clear expectations, gives useful feedback, and helps engineers grow. You build accountability through clear priorities and meaningful ownership, and you recognize when your team needs direction, support, or room to solve a problem.

You're comfortable managing a changing workload. You can turn competing requests into an actionable plan, account for operational interruptions, and explain what the team can commit to with its available capacity. You surface tradeoffs early and work with stakeholders to make deliberate decisions about scope and timing.

You bring calm and structure when priorities shift or incidents occur. You support engineers leading the response, help resolve escalations, and coordinate with customer-facing partners. You build a culture where engineers can surface risks early and examine failures honestly.

You have the technical judgment to help the team determine whether a recurring problem needs an infrastructure change, better automation, an application fix, or a clearer process. You bring the right people together to address it and ensure they have time to follow through.

What You'll Do

  • Lead and develop the SRE team: Hire, coach, and support engineers across infrastructure, operations, and application reliability. Set expectations, manage performance, support career development, and build the skills and coverage the team needs.

  • Own capacity and work planning: Maintain a clear view of incoming requests, ongoing support needs, and planned engineering work. Break initiatives into manageable tasks with the team, establish ownership, sequence work, and adjust commitments as priorities or capacity change.

  • Coordinate delivery across teams: Manage dependencies with platform engineering, application engineering, security, and customer success. Identify blockers early, resolve competing priorities, and communicate progress, risks, and decisions to stakeholders.

  • Set the reliability roadmap: Translate customer needs, production data, incident patterns, and operational risks into a prioritized improvement plan. Protect capacity for work that reduces recurring failures and makes deployments easier to operate.

  • Establish operational ownership: Ensure production deployments have clear support responsibilities, escalation paths, and readiness criteria. Plan with partner teams for new deployments, releases, and ongoing customer support.

  • Support team-led incident response: Establish sustainable on-call coverage and clear incident response expectations. Coach engineers who serve as incident commanders and lead blameless postmortems / After Action Reviews (AARs). Help the team assess corrective actions, assign ownership, and schedule follow-through.

  • Guide technical priorities: Work with engineers and technical leads to evaluate approaches to infrastructure, automation, observability, and application reliability. Ensure plans account for operability, security requirements, and the constraints of on-prem and air-gapped environments.

  • Make reliability and workload visible: Guide the team's use of SLIs, SLOs, and operational metrics. Use service health, support demand, and delivery progress to explain where investment is needed and whether improvements are working.

  • Reduce operational toil: Give engineers time and support to automate repetitive deployment, maintenance, troubleshooting, and recovery work. Help turn lessons from individual customer environments into reusable improvements.

What We Look For

  • An active Secret clearance

  • 5+ years in Site Reliability Engineering, Platform Engineering, DevOps, or a related role, with substantial infrastructure and operations experience

  • Experience directly managing engineers, including coaching, performance management, career development, and hiring

  • Experience planning team capacity, prioritizing competing requests, and coordinating engineering work across teams

  • A track record of delivering reliability improvements while managing ongoing operational responsibilities

  • Experience with incident response and post-incident review practices, including helping engineers develop leadership and ownership

  • Technical judgment sufficient to evaluate engineering proposals, understand operational risks, and guide prioritization

  • Clear communication, including the ability to explain constraints and negotiate scope, timing, and commitments with stakeholders

Technical background:

You should have practical experience operating production systems and enough breadth to guide engineers working across these areas:

  • Infrastructure and automation: Infrastructure as Code, configuration management, and scripting, using tools such as Terraform, Ansible, Python, Go, or Bash

  • Containers and orchestration: Kubernetes deployment, troubleshooting, and operations

  • Delivery practices: CI/CD pipelines, release safety, and repeatable deployments

  • Cloud and on-prem environments: Operating software across AWS or AWS GovCloud and customer-managed infrastructure

  • Observability: Monitoring, logging, and actionable alerting using tools such as the Grafana stack, ELK, or Datadog

  • Networking and security: Core protocols, secure configuration, and connectivity troubleshooting

  • Application reliability: Working with software engineers to diagnose application failures and evaluate infrastructure or code changes

We value depth in relevant areas and the ability to guide specialists across the rest.

Bonus points (nice to have):

  • Experience leading teams supporting mission-critical customer deployments

  • Experience managing engineering programs or coordinating delivery across multiple customer environments

  • Experience operating in DoD, classified, or air-gapped environments

  • Familiarity with RMF, STIGs, and ICD 503

  • Experience implementing SLIs, SLOs, and error budgets for distributed systems

  • GitOps practices and toolchains

  • On-prem virtualization experience with VMware, Proxmox, Nutanix, Hyper-V, or similar platforms

  • Application development experience, especially TypeScript or Node.js

  • Relevant certifications, such as AWS DevOps Engineer or CKA/CKAD

  • Active Security+ or another DoD 8570.01-approved security credential, or the ability to obtain one within 3 months of employment

Standard company text repeated across Onebrief's postings is omitted here.