Join Arena as a Software Engineer to build core infrastructure for AI model evaluation in a fast-paced startup environment.
Posted by employer 3 days ago
First seen on Joblaze 2 days ago
Last verified on the company career page 5 hours ago
Skills & Technologies
What you'll build
Must have
Nice to have
Practical constraints
Role intensity
70% hands-on coding
AI in the day-to-day
Arena evaluates AI models in real-world scenarios, focusing on performance and user experience.
Requirements
Not disclosed in this posting: compensation, visa sponsorship.
Benefits
Joblaze summary
The Software Engineer - Platform at Arena focuses on developing the foundational infrastructure that supports the company's AI evaluation systems, ensuring reliability and scalability. Key skills include proficiency in Go and experience with distributed systems, particularly in building API-based products and handling streaming challenges. This role is suited for senior engineers with a strong background in backend development and a product-oriented mindset, ready to navigate the dynamic environment of a startup. Arena's team comprises diverse experts from top institutions, fostering a culture of collaboration and innovation.
Joblaze insights
Quick facts
From the original posting
Arena Intelligence is looking for a Software Engineer - Platform to build the core infrastructure that sits beneath our online evaluation systems — the AI gateways, automated arena runtimes, and serving layers that make real-world model evaluation possible at scale.
This is a critical part of the Arena Service. Arenas are live, online systems: they route traffic across frontier models from many providers, handle bursty and unpredictable load, need to fail gracefully when upstream models do, and have to remain fair and consistent under all of it. We exist to build foundational infrastructure for our users that scales, is reliable, and makes the complexities of operating this infrastructure at scale disappear. We need a practitioner who's shipped this kind of infrastructure before and knows where the sharp edges are.
Our AI gateway is currently in private, gated launch with individual developers as our primary users today; enterprise use cases will follow later. Next month we're shipping direct model access and new infrastructure features, and Arena may begin collecting usage traces and building leaderboards shared with lab partners — so there's a lot of near-term, zero-to-one work ahead.
You'll be an early member of our infrastructure team, working closely with researchers, engineers, and product leadership. The work is zero-to-one in places and scale-it-up in others. We move fast and stay rigorous.
This is a hands-on individual contributor role — we're not hiring for a tech lead or SRE function right now; everyone on the team is heads-down building.
Build API-based products from the ground up. Design and implement low-latency, high-reliability APIs for leaderboards, models, and arenas.
Solve hard streaming problems. Handle SSE/streaming responses across heterogeneous providers, including partial failure recovery, mid-stream fallback, and consistent response normalization.
Ship enterprise-grade infrastructure. Build the systems enterprise customers will eventually expect — rate limiting, authentication, usage metering, cost attribution, audit logging, and SOC 2 compliance — as we grow beyond our current individual-developer user base.
Build deep observability. Instrument infrastructure with distributed tracing, latency breakdowns, token-level usage tracking, and real-time dashboards so customers (and we) can see exactly what's happening.
Build AI-centered products. Integrate with our core evaluation platform, Arena data, and customer-specific benchmarks. Collaborate with the research team to turn novel ideas into full-featured products. (This role does not involve data labeling or third-party data-verification work.)
Flex across the stack. Contribute to the backend of our Leaderboards and Evals platforms when needed, helping unify our public and private data architectures.
~5+ years of backend engineering experience, with meaningful time spent on distributed systems, infrastructure, or developer-facing platforms.
Senior: proven delivery on meaningful backend work.
Staff: extensive, deep experience owning backend systems end to end.
Strong proficiency in Go — this is our primary backend language and a must-have for the role.
Experience with LLM provider APIs (OpenAI, Anthropic, Google, etc.) and a working understanding of the challenges: streaming, token management, rate limits, model-specific quirks.
A product-oriented mindset. You think about the developer experience of your APIs, not just the implementation. You ask "why" before "how."
Comfort with ambiguity. We're a startup. Scope is fluid, context shifts, and you'll wear many hats. That should sound exciting, not stressful.
Cloud infrastructure experience (AWS, GCP, or Azure), Kubernetes, Terraform, and database systems like Postgres and Redis — helpful, but not a hard filter.
Experience building API gateways, proxies, or developer tools (Bifrost, Kong, Envoy, Tyk, or custom).
Background in AI/ML infrastructure, model serving, inference, or evaluation frameworks.
Experience building enterprise-ready features: SSO, RBAC, audit logs, multi-tenancy.
Experience building billing infrastructure around systems like Stripe, Metronome, and Orb.
Familiarity with the modern AI infra stack (vLLM, LiteLLM, LangChain, etc.).
This role is based in San Francisco, with a minimum of 3 days/week onsite. Fully remote candidates will only be considered with a very strong endorsement.
We offer competitive compensation and equity aligned to the markets where our team members are based. The base salary range will depend on the candidate’s permanent work location.
Standard company text repeated across Arena's postings is omitted here.