The problem
A two-person engineering team is hiring its third backend engineer. The posting draws a couple of hundred applicants, most of whom list AI coding tools on their resume.
A take-home for everyone is unreviewable at that volume, and a live first round for everyone costs more engineer time than the team has. Resumes alone cannot separate people who use an agent well from people who paste whatever it says.
The approach
- 01
One link in the job post
The team creates a general link from the portal, sets max uses to the size of the pool and a two-week expiration, and pastes it into the application confirmation email. Candidates sign in with Google, see what is recorded, and start when they are ready.
- 02
Rotating tasks from a difficulty band
Instead of one fixed issue that could leak, the link draws from the medium band of the SWE-bench bank, so candidates get comparable but different tasks from real repositories such as Django, Flask or SymPy.
- 03
Sort the board, read the top of it
Every round lands on the candidate board with status, score, solved or not, tokens used, time taken and any integrity flags. The team sorts by score, opens the full reports for the top group, and compares finalists side by side.
What the report shows
The signals a team would read in this scenario, all recorded in every round. See them in the sample report.
- Outcome
- Did the hidden failing tests pass after the change, without breaking the tests that already passed?
- Verification
- Did the candidate run tests themselves before submitting, or trust the agent's word?
- Prompt quality
- Were prompts specific, with context and constraints, or vague one-liners repeated until something worked?
- Token economy
- How much agent usage the fix took, so two equal outcomes can still be told apart.
- Integrity flags
- Long focus losses, large pastes and pastes after leaving the tab are flagged for a human to look at, not auto-rejected.
The takeaway
The first round becomes something the team reads rather than runs. Engineer time goes into the final interviews, and those interviews can start from what the report shows about each finalist.