Careers · San Francisco

Two products.
One small team in San Francisco.

We place pre-vetted AI-native engineers, and we run the Agent Usability Index — the benchmark for how well AI agents use real-world APIs. We're hiring senior engineers, on-site in SF, to build both.

See open roles San Francisco · On-site
§ Open roles

4 roles, all on-site in San Francisco.

  • Research · Full-time

    AI Engineer

    San Francisco · On-site

    Own the AI systems behind both products. You build the eval harness that drives frontier coding agents against real third-party APIs for the Agent Usability Index, and the AI-assisted assessments that vet the engineers we place. Agent orchestration and rigorous evaluation are the core of how this company measures anything.

    You will
    • Build and tune the harness that runs frontier models (Claude, GPT, Gemini, Cursor agents) against real API-integration tasks.
    • Build the AI-assisted coding assessments that screen the engineers we recruit — with un-gameable scoring.
    • Design eval pipelines that capture every tool call, token, and failure mode, and make them reproducible.
    • Push the methodology forward and publish research notes when the data surprises us.
    Fit
    • Strong Python and/or TypeScript. You have shipped real software with LLM APIs and agent frameworks.
    • You think in evals — ground truth, reproducibility, un-gameable graders.
    • Bias toward measurable results over framework astronautics.
    • In SF and able to work on-site.
  • Platform · Full-time

    Backend Engineer

    San Francisco · On-site

    Own the platform behind both products — the recruiting system that sources, vets, and matches engineers to clients, and the Index that turns thousands of agent runs into a fast, reproducible public leaderboard. Traditional, excellent backend engineering: the foundation every other track stands on.

    You will
    • Build the services and data model behind the recruiting platform — candidate profiles, matching, the hiring dashboard.
    • Build the services behind the Index — run storage, scoring, vendor pages, and the publish pipeline.
    • Keep everything reproducible and reliable: Docker, CI, queues, sub-process orchestration.
    • Scale the systems as we add clients, candidates, vendors, and model providers.
    Fit
    • Senior backend chops. You have shipped and operated real production systems.
    • Strong with Docker, CI/CD, databases, and clean API design.
    • Bias toward boring, reproducible infrastructure over magic.
    • In SF and able to work on-site.
  • Product · Full-time

    Frontend Engineer, Agent Experience

    San Francisco · On-site

    Build the product UIs for both businesses — the recruiting platform clients and candidates use, and the public Index — and run the experiments that figure out what UI agents, not just humans, navigate fastest. We know what humans like; nobody knows what agents like. You would be one of the first people doing that seriously, while shipping real product.

    You will
    • Build and polish the frontends of both products: the recruiting platform and the Index.
    • Design and run experiments measuring how well coding and browsing agents complete tasks across different UI patterns.
    • Develop a taxonomy of what makes an interface agent-readable: semantics, affordances, structure, latency.
    • Publish findings — you will help name a field that does not have a name yet.
    Fit
    • Excellent React and TypeScript, with a real eye for interface craft.
    • Genuine curiosity about agents as users — you want to A/B test against models, not just people.
    • Comfortable being publicly named on novel research.
    • In SF and able to work on-site.
  • Engineering · Internship

    Software Engineering Intern

    San Francisco · On-site

    One intern seat. Work directly with the founding engineers across both products — the recruiting platform and the Index harness. Real code in production from week one, not busywork.

    You will
    • Ship features across the recruiting platform, the harness, and the leaderboard.
    • Author a benchmark or an assessment end to end under mentorship.
    • Help run model evals and triage results.
    Fit
    • Strong coder in Python or TypeScript.
    • Built things with LLM APIs, or hungry to.
    • In SF and able to work on-site.
§ How we work
01

Senior, small, in person

A small team of senior engineers, all on-site in San Francisco. Fast iteration loops, whiteboards, and shipping together beat async for the work we do.

02

Two products, one engine

We place AI-native engineers and we run the Agent Usability Index. Each sharpens the other — the Index shows us what great looks like, and recruiting puts those people to work.

03

Measurable over magic

We bias toward shipped, reproducible results over framework astronautics. If we cannot measure it, we do not claim it — in our benchmarks or our placements.

04

Straight with everyone

Honest with candidates, clients, and vendors. We publish what the data shows, and we never sell a score or oversell a hire.

Don't see your role?

If you are in San Francisco, want to build AI-native recruiting and the Agent Usability Index, and think you would be obviously great here, email us anyway. Include a short note on what you would want to build first.

← Back to the Index

NextdevNextdev

AI-native engineer recruiting
and the Agent Usability Index.

© 2026 NextDev, Inc. All rights reserved.
Online
v3.0.1
Nextdev — Careers