Build and run internal tools for LILT's AI delivery teams, ensuring reliable, production-ready systems in a high-impact role.
Posted by employer 2 days ago
First seen on Joblaze 1 day ago
Last verified on the company career page 1 day ago
Skills & Technologies
What you'll build
Must have
Nice to have
Requirements
Not disclosed in this posting: compensation, work arrangement, visa sponsorship.
Joblaze summary
In this role, the Staff Fullstack Engineer will develop and maintain internal tools that support Lilt's Applied AI benchmarking business, ensuring reliable operations for delivery and contributor teams. Key skills include proficiency in Python and React, along with experience in building scalable production systems and integrating various APIs. This position is ideal for a seasoned engineer with a strong background in full-stack development and a focus on internal platforms rather than customer-facing products. The team operates in a high-impact environment, directly influencing the quality and timeliness of deliverables for major AI labs.
Joblaze insights
Quick facts
From the original posting
As part of LILT’s Internal Tools group, you will build and run the tools that LILT's own delivery, contributor-ops, and engineering teams rely on to operate LILT's Applied AI (AAI) benchmarking business — which builds and delivers multilingual benchmarks and evaluation data to frontier AI labs using LILT's global network of subject-matter experts. As the AAI business grows, expect the scope of these platforms to grow in parallel. This platform sits directly on the critical path of paid customer deliverables with real SLAs, maintained by a small, high-leverage engineering team that produces reliable, production-ready tooling. You will own significant surface area across both end-to-end: architecture, data-pipeline design, and the hands-on engineering that keeps this infrastructure reliable. This is a high-impact, high-visibility role: the reliability and craft you bring directly enables on-time, high-quality delivery for some of the most prominent AI labs in the world.
In this role, you will drive the long-term technical strategy for internal platforms while working directly with the delivery, ops, and engineering teams across the business who depend on them daily. This team ships cross-service automation in careful stages — advisory first, then assistive, only later decision-relevant, always with a human fallback — and you'll be expected to hold that same bar as you extend these systems. You will partner with a small existing team to raise the engineering bar across two very different runtimes, set technical direction, and make the calls that determine how this infrastructure scales as the business grows.
Partner with the benchmarking business's researchers and TPMs to translate new benchmark and data-quality requirements into scoped technical designs
Own the internal workforce-management tools on our internal platform for AI delivery for hundreds of external contributors: vetting flows, candidate assessment, QC, payment/delivery tracking, and roster/reporting exports
Own architecture and long-term technical direction across multiple services and the platform
Extend IAA and audio-QA pipelines: annotator outlier detection, ASR sidecar enhancements, LLM-based QC, and DNSMOS/librosa audio-quality scoring
Design and ship new modules on the platform that plug new benchmark and vetting workflows into the existing multi-stage review lifecycle
Own and extend the platform's API key provisioning and budget-governance system
Build self-serve ChatOps-style automation for the internal engineering org and contributor base to enable accelerated annotator workflows and query resolution
Harden background worker and job-processing infrastructure
Set and enforce testing, CI/CD, and deployment practices across both codebases — from unit/integration testing through infrastructure-as-code and release automation
Raise the technical bar for a small, high-leverage team through code review, design docs, and mentoring as the surface area grows
5+ years of professional full-stack software engineering experience, with a track record of owning production systems end-to-end across more than one runtime/language
Deep experience with a Python backend framework (FastAPI or comparable) plus async SQLAlchemy/Postgres, alongside production experience in at least one statically-typed backend language (Go, Java, or similar)
Strong React/TypeScript frontend experience — component architecture, state management (Zustand, Redux, or similar), and a modern data-fetching layer (TanStack Query or comparable)
Experience building and operating background job/worker systems (queue-driven or polling-based) with failure tolerance and idempotency in mind
Experience integrating with third-party and platform APIs — including the GitHub API, OAuth/OIDC SSO, and at least one LLM API (Gemini, OpenAI, or similar) — handling auth, rate limits, and webhook-driven sync
Experience building Slack (or comparable chat-platform) bot integrations that automate internal workflows — resource provisioning, approvals, budget/TTL enforcement — with real operational guardrails, not just CRUD features
Comfort reading and extending applied-statistics or ML-adjacent code (agreement metrics, audio-quality scoring, or comparable data-quality tooling)
Solid grasp of CI/CD, containerized deployment (Docker, Helm, ArgoCD/GitOps or comparable), and infrastructure-as-code (Terraform or comparable)
Experience debugging and optimizing native-library (numpy/scipy/onnxruntime-class) memory growth in long-running Python worker processes via safe, boundary-aware process recycling — not just raising memory limits
Experience building and owning internal platforms/tools that increase leverage for a non-engineering team (research, operations, support, data/workforce management, or similar) — not solely external-customer-facing product work
Experience with audio/speech pipelines: ASR (Whisper or similar) or audio-quality metrics (DNSMOS, librosa)
Experience building internal tools for managing a data-labeling, annotation, or crowdsourced-contributor workforce (vetting, QC, payments)
Experience with inter-annotator agreement or statistical agreement metrics
Notification and delivery systems experience — Slack bot integrations, transactional email, and idempotent delivery guarantees
Experience designing abstractions over heterogeneous data sources with different consistency guarantees — e.g. a fully-replayable event history vs. an observe-only current-state API requiring synthesized diffing — behind one common interface
Experience building ChatOps-style automation — Slack or GitHub PR-comment bot commands that trigger backend workflows or CI/CD runs
Experience implementing short-lived, rotatable service-to-service JWT auth (key-ID-based rotation, replay-protected tokens, fail-fast config validation) alongside a separate human-facing SSO flow in a paired service
Comfort owning both sides of a system with genuinely different runtimes without a large team to lean on
Prior experience as the primary or sole engineer on a small, high-leverage internal platform
Experience with LLM-as-judge or LLM-based QA/review pipelines
Familiarity with OpenTelemetry or comparable observability instrumentation in Go services
Familiarity with LLM provider gateway/routing services (OpenRouter or comparable) — model aliasing, rate-limit and timeout handling, and budget enforcement
Experience with data export/reporting tools (Excel generation, CSV pipelines, or BI-style dashboards)
What sets our platform apart:
Brand-aware AI that learns your voice, tone, and terminology to ensure every translation is accurate and consistent
Featured in The Software Report’s Top 100 Software Companies!
Standard company text repeated across Lilt's postings is omitted here.