Join Faire as a Staff Data Infrastructure Engineer to design and build the next version of our data platform.
Posted by employer 1 day ago
First seen on Joblaze 1 hour ago
Last verified on the company career page 1 hour ago
Skills & Technologies
What you'll build
Must have
Nice to have
Practical constraints
Role intensity
40% coding
Not disclosed in this posting: years of experience, visa sponsorship.
Benefits
Joblaze summary
The Staff Data Infrastructure Engineer at Faire is responsible for designing and implementing the data pipeline that moves information from production databases to analytical stores, ensuring data quality and reliability. This role requires expertise in change data capture, streaming ingestion, and lakehouse architectures, particularly with tools like Kafka, Airflow, and Snowflake. Ideal candidates have significant experience leading data infrastructure initiatives and mentoring engineers, making this position suitable for seasoned professionals in data engineering. The role is integral to supporting various teams across the company, emphasizing collaboration and technical leadership.
Joblaze insights
Quick facts
From the original posting
About Faire
About this role:
Our Engineering organization owns the software that makes our marketplace work. The Data Platform group supports everyone at Faire who depends on data: Product Engineering, Data Science, Machine Learning, Analytics, Strategy, Finance, and Product. Our job is to make sure the data is there, it's right, and people can find it and query it without having to think about the plumbing underneath.
We are hiring a Staff Engineer to own that plumbing. Concretely, this means the path data takes out of our production databases (CockroachDB and MySQL) and into a place where analysts and data scientists can query it. Today that involves Fivetran, Kafka, Spark, and Airflow landing data in Snowflake and Databricks. It works, but it grew up over time and it shows. We want someone who can design the next version of it, build the hard parts personally, and bring the rest of the company along.
This is a hands-on role. You'll also be the person other teams come to when they need to know how data should move at Faire.
What you'll do:
Set the technical direction for how data moves from production systems into our analytical stores, and own the roadmap to get there over the next couple of years.
Build the CDC and streaming ingestion layer: CockroachDB changefeeds and MySQL binlogs into Kafka, then into Iceberg tables on S3. You'll be responsible for the hard details like ordering, deduplication, late data, schema changes, and backfills.
Implement data contracts and quality checks throughout our platform
Put real ownership and SLAs on the datasets the business runs on, and wire quality checks into the platform with tools like Anomalo and Monte Carlo so we hear about broken data before a dashboard or a model does.
Run Airflow and Fivetran well, and have an opinion about what we should keep buying versus what we should build.
Own reliability for the platform: SLOs, on-call, incident reviews, and the follow-through so the same thing doesn't break twice.
Work with the senior engineers, data scientists, and analysts who depend on this platform, and lead the migration of existing pipelines onto the new one without breaking what they rely on.
Mentor the engineers around you. We want the team's data engineering practice to be better because you were here.
Qualifications:
You've built and run data infrastructure that other teams depended on, at meaningful scale, and you've been the person setting direction for it, not just working on it.
Deep experience with change data capture and streaming ingestion from operational databases through Kafka. You know what goes wrong with ordering, duplicates, snapshots, and schema evolution because you've dealt with it.
Hands-on experience with lakehouse architectures on an open table format. Iceberg on S3 is what we use, so that's especially valuable. You should be comfortable talking about partitioning, compaction, catalogs, and copy-on-write versus merge-on-read.
Strong Spark skills, and experience running Databricks and Snowflake against shared storage.
Experience with data quality and observability in practice, including data contracts, SLAs, and tools like Anomalo or Monte Carlo.
Experience operating Airflow at scale and working with managed ingestion like Fivetran.
Strong SQL, and good instincts for how to model data so analysts and data scientists can actually use it.
Solid Python plus at least one of Kotlin, Java, Scala, or Go. Experience shipping infrastructure on AWS with Terraform.
A working understanding of data governance: access control, PII, retention and deletion, lineage, and audit.
A track record of leading cross-team data initiatives and migrations, and of mentoring senior engineers.
You can explain a technical tradeoff to a leadership team and to a new grad, and you can get people who disagree with each other to a decision.
You take ownership of things that are broken or unowned, and you're willing to be on call for the systems you build.
Experience in a marketplace, e-commerce, or other transaction-heavy business is a plus.
Technologies we use and teach:
Python, Kotlin, SQL
Kafka, Fivetran, Airflow
S3, Apache Iceberg, Snowflake, Databricks, Apache Spark
AWS, Terraform, Kubernetes
CockroachDB, MySQL, Scylla and DynamoDB
Salary range:
Canada: the pay range for this role is $216,000 to $297,000 per year.
Hybrid Faire employees currently go into the office 3 days per week on Tuesdays, Thursdays, and a third flex day of their choosing (Monday, Wednesday, or Friday). Additionally, hybrid in-office roles will have the flexibility to work remotely up to 4 weeks per year. Specific Workplace and Information Technology positions may require onsite attendance 5 days per week as will be indicated in the job posting.
Standard company text repeated across Faire's postings is omitted here.