Own the AWS infrastructure and reliability at Tabs, shaping engineering decisions in a high-growth environment.
Posted by employer 1 day ago
First seen on Joblaze 1 hour ago
Last verified on the company career page 1 hour ago
Skills & Technologies
What you'll build
Must have
Nice to have
Role intensity
40% coding
Requirements
Not disclosed in this posting: compensation, visa sponsorship.
Benefits
Joblaze summary
The Staff Infrastructure Engineer at Tabs is responsible for shaping the company's AWS infrastructure and ensuring the reliability of its software delivery processes. This role requires expertise in infrastructure as code, particularly with Terraform and Docker, as well as a strong background in CI/CD systems and observability practices. Ideal candidates will have over eight years of experience in software engineering or SRE roles, with a proven track record in managing production systems and mentoring other engineers. Tabs is focused on building a robust infrastructure that supports its growth, making this position critical for the engineering team's success.
Joblaze insights
Quick facts
From the original posting
We're looking for a Staff Infrastructure & Reliability Engineer to own the foundation Tabs runs on: our AWS environment, how we ship software, and how we know when something is wrong. You'll set infrastructure direction for the company as a hands-on individual contributor, partnering with our platform team, our product engineers, and the product teams who build on what you build.
Tabs is building, not maintaining. We're at the point where infrastructure is becoming a real investment area, and the decisions made in this seat will shape how the whole engineering org ships for years. Payments, billing, and revenue for high-growth companies come with real compliance requirements. The goal is to build it correctly, keep it easy to maintain, and evolve it as we grow.
This is not a corner seat. You'll be expected to shape engineering and product decisions, and you'll have engineering leadership that understands infrastructure work and will pressure-test your calls. You won't be working alone: you'll own the outcomes, with high-caliber engineers around you.
AWS infrastructure direction and platform evolution, including the migration from ECS/Fargate toward a more modern, scalable runtime
Infrastructure as code and container foundations, with Terraform and Docker at the core
CI/CD systems with a strong emphasis on developer experience, safety, and automation (GitHub Actions today; maturing CD tomorrow)
Ephemeral environments and preview deploys to speed iteration and increase confidence in changes
Observability standards across metrics, logs, and tracing, including alert hygiene, dashboards, and SLO development
Incident response, on-call, postmortems, and the reliability culture that surrounds them
Define and evolve reliability standards, SLIs, SLOs, and error budgets
Improve observability, alerting, and incident processes across services
Lead high-severity incidents hands-on and drive clear, actionable follow-ups
Partner with engineering teams to design resilient, scalable systems
Write production-quality code and automation to reduce toil and lower operational risk, so repeated problems get solved once
Mentor engineers and influence best practices across teams
You're a software engineer first, and your infrastructure expertise is built on that foundation
You've run production systems on AWS and can lead platform-level change
You think in systems: risk, rollback strategy, blast radius, and feedback loops
You treat CI/CD and environments as products that should be fast, reliable, and self-serve
You dig into logs and data yourself when something breaks, especially under pressure
You influence through trust and clarity rather than control
You balance pragmatism with long-term system health
You value learning from failure and improving processes over assigning blame
You communicate clearly and work well across teams
8+ years in software engineering, infrastructure, or SRE roles
Experience in one or more modern languages such as TypeScript with a track record of writing production-quality scripts, tools, and services, and still hands-on today
Deep hands-on experience running production workloads on AWS, including container platforms such as ECS/Fargate
Expertise with infrastructure as code using Terraform, and ownership of Docker and Git workflows in production
Solid working knowledge of Kubernetes and Helm
Experience designing and running CI/CD systems such as GitHub Actions, including build parallelization and developer experience improvements
Deep experience with observability tooling across metrics, logs, tracing, and alerting, including defining SLIs, SLOs, and error budgets
Expertise operating distributed systems in production at scale, with an implementation-level understanding of messaging systems, partitioning, deploy strategies, and failure modes
A track record of leading high-severity incidents, debugging live production issues from logs and data, and running blameless postmortems
Experience proposing and evaluating multiple architectures, making trade-offs that fit the company's stage, and driving infrastructure decisions across teams
Experience across more than one architecture or company environment, ideally including both larger companies and high-growth startups
Comfortable navigating ambiguity and setting direction in a fast-moving environment
Experience mentoring engineers and shaping infrastructure practices across an engineering org
Experience owning broad infrastructure surface area at a high-growth startup, including as the primary infrastructure or SRE owner
Experience operating Kafka or a similar distributed messaging system at scale
Experience building developer tooling that engineers adopt and rely on
Prisma expertise
This role is based onsite in our Soho office in New York City
Competitive compensation and equity
Unlimited PTO
Parental leave up to 12 weeks
Tax free commuter and parking benefits
Voluntary insurances (Life, Hospital, Critical Illness, Accident)
Employee Assistance Program (Rightway)
Free One Medical Membership
401k
Standard company text repeated across Tabs's postings is omitted here.