Six dimensions
Each dimension is scored 0 to 100 from the round's recording. The overall score is their weighted sum, with a letter grade (A+ from 90, A from 80, B from 70, C from 60, D from 45).
Did the fix work? Measured by the task's hidden tests, which the candidate never sees.
- ·Share of the hidden failing tests (FAIL_TO_PASS) that now pass: up to 75 points.
- ·Share of the existing tests (PASS_TO_PASS) that still pass: up to 25 points.
- ·A submission with no change scores near zero; code that does not run scores zero.
- ·Take-home builds: automated checks, plus a rubric review when one was run (the report says which).
How the candidate directed the agent.
- ·Each prompt is rated from its text: specific files or identifiers, error context, a clear action, constraints or acceptance criteria.
- ·Pasted material (raw errors, code dumps, the issue text) is rated on the words the candidate wrote around it.
- ·Points off for vague prompts, re-sending the same request, and work split into too many or too few prompts.
- ·Points on for correcting the agent's course.
Whether the candidate checked the work instead of trusting the agent.
- ·Running the tests, and how often.
- ·Reading the code before and while changing it.
- ·Testing the final version: edits after the last test run cost points.
- ·Whether the last local run before submitting was green.
What the result cost in agent usage.
- ·Tokens used against the round's budget.
- ·Context bloat (the conversation growing fast per request).
- ·Tokens burned inside failing agent loops.
Held to what the outcome earned: thrift on work that did not land is only partly credited.
Time used against the time allowed.
- ·Finishing early scores higher; running over scores lower.
- ·Credited only for work that landed: fast and wrong never scores.
Fully gated on the outcome. A round with nothing attempted scores zero here.
What happened after something went wrong.
- ·Error loops: three or more failing agent steps in a row without a change of approach.
- ·Repeating an identical prompt, or 'fix it'-style nudges while stuck.
- ·Getting back to a green test run after a loop.
Held to what the outcome earned, down to a floor that still shows the behavior.
When the hidden tests cannot run (the task's environment failed to build), the round is marked not graded with no score, rather than a misleading number.
The written verdict
The headline on a report (Strong hire signal, Promising, Mixed signal, Concerns) is chosen by rules from the same numbers: for example "Strong" needs the hidden tests to pass with nothing broken, an overall score of 75 or more, solid process scores, no high-severity flag, and prompts mostly in the candidate's own words. The strengths, risks and interview questions under it are templates filled from the evidence, each pointing at the moment it refers to. Candidates see their own copy of an assessment without the hiring team's verdict wording or interview questions.
Hireability
At the top of every report, the hiring side's question: from an AI-coding point of view, how strong is this candidate? Seven behaviors are scored 0 to 100 by itemized rules over the same recording, each point tied to a sentence and, where there is one, a moment in the round. The hireability score is their weighted average (75%) blended with the hidden-test outcome (25%). It is a signal for the person deciding, never a decision.
- Problem understandingweight 15
- Read the code before changing it, reproduced the failure first, framed the task instead of pasting the issue, changed the right file.
- Steering the AIweight 20
- Prompt quality, plus points for correcting the agent and asking it to prove its work; off for “fix it” nudges, raw error pastes, code dumps and long unsupervised runs.
- Verificationweight 20
- How often the tests ran, whether the final version was tested and green, and catching the agent's wrong turns.
- Debuggingweight 10
- How failures were handled: failing loops left running, retries without a diagnosis, getting back to green. Not measured when nothing failed.
- Code qualityweight 15
- Size and scope of the diff, files unrelated to the fix, regressions, regression tests added, test files edited while failing.
- Efficiencyweight 10
- Token economy and pace from the score breakdown (60/40), both held to what the outcome earned.
- Ownership and communicationweight 10
- The debrief and the in-round answers: how much was explained, whether it names the actual change, and whether answers were pasted.
The call: Strong hire from 85, Hire from 70, Lean no from 50, No hire below. Strong hire needs the issue solved and no behavior under 50; any red flag (pasted debrief answers, tests edited while failing, an unsupervised agent, tests never run) holds the call at Hire, and a high-severity one (only test files changed, an integrity flag) at Lean no. A round that could not be graded, or recorded too little, gets no call. A behavior that was not observed is left out of the average rather than scored zero.
Attention flags
Raised from editor and proxy events, shown with timestamps, and never applied to the score automatically except where a dimension above says so.
- Long focus loss
- The candidate left the workspace tab for more than a minute.
- Large paste
- More than 500 characters pasted at once.
- Paste after focus loss
- A paste soon after returning to the tab. Marked low severity when the text came from the round's own terminal.
- Paste-heavy prompting
- Half or more of the prompts were mostly pasted material.
- Repeated prompts
- The same prompt sent more than once.
- Budget exhausted
- The agent stopped answering because the token budget ran out.
Known limitations
- The weights are a design choice
- 35/15/15/15/10/10 reflects what we think matters. It has not yet been validated against later job performance, and we have not yet run an independent bias audit. Treat the overall number as a summary of the evidence, not a measurement of a person.
- Prompt rating reads text
- Terse prompts score lower. People writing in a second language, or who prefer short instructions, may be rated lower for reasons unrelated to skill. Read the prompts themselves before relying on this dimension.
- Public tasks can be memorized
- Bug-fix tasks come from SWE-bench, whose issues and merged fixes are public on GitHub and likely in AI models' training data. An agent may reproduce a fix from memory, which inflates the outcome. Hiring teams can use tasks held back from the public practice site, and every task carries a contamination label; neither makes an upstream issue unknown to models.
- One harness, one agent
- Candidates use the agent built into the workspace, which may not be the tool they use every day. Unfamiliarity can cost time early in a round.
- Attention signals are signals
- Focus and paste flags have innocent explanations (a second monitor, reading documentation, pasting from the round's own terminal). They are for a person to look at, never an automatic penalty.
- Percentiles need volume
- A percentile compares against graded rounds so far. With few rounds it is noisy; the report states the cohort size.