What exactly does a candidate do?+
Either fix a real bug in a real open-source repository (taken from the SWE-bench benchmark, checked out at the failing commit) within 30 to 120 minutes, or build a small product from a written brief within up to 24 hours. Both happen in a browser VS Code with an AI coding agent they can use freely.
Isn't using AI cheating?+
No, it is the point. Engineers now build with agents every day, so the assessment measures how well they direct one: how specific their prompts are, whether they verify the agent's work, how they recover when it goes wrong, and what it costs. We record outside help signals such as pastes and long focus losses and flag them for you to review.
How is a round scored?+
Bug-fix outcomes are checked by separate grading tests. Take-home outcomes use automated checks and, when available, a model review against the shown rubric. Process dimensions use fixed rules for prompt quality, verification, token economy, speed and recovery, with recorded evidence. If the grading environment cannot run, no overall score is shown. Read the scoring method for weights and limitations.
Could a candidate have seen the task before?+
Bug-fix tasks are real SWE-bench issues, so the upstream issue and its fix are public on GitHub and may be known to AI models. We say so on every task. You can choose tasks held back from our public practice site so candidates cannot rehearse them here, and every report shows the process evidence alongside the outcome, which is much harder to fake than a passing test.
Does the score decide who gets hired?+
No. Scores support your reviewers' judgment and every report says so. Nobody is rejected automatically. Candidates are told before they start what is recorded, how the result is used, and how to ask for an accommodation or a different way to be assessed.
How does pricing work?+
Your team's agreement sets the price per assessment. AI usage within the round's token limit and the report are included. Buy more from your team portal at the agreed price. Invitations alone use no assessments; one is used when a candidate's workspace is provisioned, before they press Start inside it. A round marked failed by the platform returns that assessment; a submission that fails grading tests still uses one.
How quickly can we start?+
Contact us to discuss your roles, assessment format and setup. Once your team's workspace and agreement are ready, an owner can add colleagues and create assessment links.