The problem
A team's senior take-home says "about four hours", but some candidates clearly spend a weekend on it and others submit something an agent produced in twenty minutes. Neither is visible from the final repository.
Calibrating the exercise is guesswork without knowing how the time was spent and how the work was done.
The approach
- 01
A brief with the rubric shown up front
The project bank's briefs come with their rubric and automated checks visible to the candidate, so the exercise is fair: everyone knows how it will be judged.
- 02
Recorded, not policed
The candidate builds in the same recorded environment with an AI agent. Time on the clock, prompts, agent steps, edits and test runs are all captured, so the team sees how the work was done as well as what was submitted.
- 03
Calibrate the brief from real rounds
After a few finalists, the team reads time taken and token usage across reports and adjusts the brief or the time limit so it fits the time they actually want to ask for.
What the report shows
The signals a team would read in this scenario, all recorded in every round. See them in the sample report.
- Time on task
- Real time spent, from the session timestamps, instead of a self-reported estimate.
- Agent steps and tokens
- How much of the build was delegated, and whether that delegation was directed or open-ended.
- Prompt log
- Design decisions in the candidate's own words, as they were made.
- Edits and revisits
- Files edited and how often the candidate came back to them, a proxy for iteration and refactoring.
The takeaway
The take-home becomes a bounded, fair exercise with evidence attached, and the team can tune it to the effort it actually wants to ask of senior candidates.