Skip to content
PraxisAI for hiring teams

Hire engineers for how they build with AI.

Candidates fix a real bug in a real repository, or build a small product in 24 hours, with a live coding agent. Grading tests check the outcome and recorded evidence supports process review. Set token budgets for the round and agree your team's price per assessment.

Per-assessment pricing agreed with your team

A PraxisAI candidate report: overall score with six scored dimensions, token usage, the prompt log and a timeline of the round
Why PraxisAI

Evidence from the work, with budgets you set.

Process scoring you can explain

Prompt forensics, verification, recovery and token economy are scored by fixed, published rules from the round's recording, with evidence attached. These observations support your reviewers' judgment.

Read the scoring method →

Hidden tests you don't write

Bug-fix tasks come with recorded required tests and regression checks from the upstream project. Your team can review the graded outcome without authoring a question bank or a rubric.

Assessment pricing with AI included

One provisioned workspace uses one assessment, with AI usage included within its token limit. Your team's agreement sets the price; you can buy more from the team portal.

Real repositories and 24-hour builds

Fix a genuine issue in projects like Django or SymPy at the failing commit, or build a small product from a brief over up to 24 hours. Work that looks like the job, not a puzzle.

How it works

From link to shortlist in three steps.

  1. 01

    Create an assessment link

    Pick a real-repo bug fix or a take-home build, set the duration and a token limit, and choose a general link for a whole pool or a personal link for one candidate.

  2. 02

    The candidate builds with AI

    They sign in, read what is recorded, and work in a browser VS Code with an AI coding agent on a real repository. Nothing to install; the environment is isolated per candidate.

  3. 03

    You read the report

    Hidden tests grade the outcome; the process is scored with evidence. Rounds land on your candidate board, ready to sort, export and compare side by side.

The candidate's workspace: a browser VS Code with the repository, an AI agent chat panel, a countdown timer and a token budget
The report

A score you can argue with, because the evidence is attached.

Bug-fix reports combine the test outcome with fixed process rules and recorded evidence: prompts, test results and retry loops. Take-home outcomes can include a model review of the rubric. Plus token usage and cost, the full prompt log and a minute-by-minute timeline.

Outcome
Did the hidden failing tests pass, without breaking the ones that already did.
Prompt quality
Specific, contextual instructions to the agent, or vague retries.
Verification
Whether they ran the tests themselves before calling it done.
Token economy
How much agent usage the result took, against comparable rounds.
Speed
Pace on work that landed. Gated on the outcome, so fast and wrong never scores.
Recovery
What they did after a failing run or a wrong turn by the agent.
Assessment types

Two ways to see real work.

30 to 120 minutes

Real-repo bug fix

A genuine GitHub issue from the SWE-bench benchmark, in projects like Django, Flask and SymPy, checked out at the exact failing commit. Pin one issue so every candidate is compared like for like, or let the link rotate through a difficulty band. Choose tasks held back from our public practice site so candidates cannot rehearse them here.

  • · Graded by hidden tests you don't have to write
  • · Easy, medium and hard bands
  • · Best for first rounds and high volume
Up to 24 hours

Take-home build

A small product built from an empty repository against a written brief, with the rubric and automated checks shown up front. Recorded the same way, so you see how the time was spent and how much was the agent.

  • · A fair, bounded take-home with evidence attached
  • · Rubric visible to the candidate
  • · Best for senior and final rounds
Integrity and security

Recorded openly. Isolated by default.

How recording and data handling work →

Recorded work evidence

Captured prompts, agent steps, tool calls, edits, terminal commands and test runs, on one timeline.

Focus and paste signals

Long tab switches, large pastes and pastes right after leaving the tab are flagged for a human to review.

Isolated environments

Each round runs in its own container with its own workspace; nothing is shared between candidates.

Bounded token use

A token limit per candidate, set by you and enforced by the platform.

Candidates are told what is recorded before they start, on the page they join from, along with how the score is used and how to ask for an accommodation. Flags are signals for a person to review, never an automatic rejection.

Pricing

Choose the assessment volume your team needs.

Agree your team's price per assessment, then buy more from the portal as needed. One provisioned workspace uses one assessment, including AI usage within its token limit and the report. A round marked failed by the platform returns its assessment.

  • Assessment-based

    One count for your whole team, visible in the portal with every one used.

  • Per assessment

    One assessment per provisioned workspace, with AI usage included within its token limit. No seat licenses.

  • One agreed price

    Every assessment costs the price agreed with your team. No metered bill.

  • Self-serve top-ups

    Buy more from your portal at your agreed price whenever you need them, with an invoice.

FAQ

Questions hiring teams ask.

Something else? Contact us. Engineers looking to practice can head to the practice site.

What exactly does a candidate do?

Either fix a real bug in a real open-source repository (taken from the SWE-bench benchmark, checked out at the failing commit) within 30 to 120 minutes, or build a small product from a written brief within up to 24 hours. Both happen in a browser VS Code with an AI coding agent they can use freely.

Isn't using AI cheating?

No, it is the point. Engineers now build with agents every day, so the assessment measures how well they direct one: how specific their prompts are, whether they verify the agent's work, how they recover when it goes wrong, and what it costs. We record outside help signals such as pastes and long focus losses and flag them for you to review.

How is a round scored?

Bug-fix outcomes are checked by separate grading tests. Take-home outcomes use automated checks and, when available, a model review against the shown rubric. Process dimensions use fixed rules for prompt quality, verification, token economy, speed and recovery, with recorded evidence. If the grading environment cannot run, no overall score is shown. Read the scoring method for weights and limitations.

Could a candidate have seen the task before?

Bug-fix tasks are real SWE-bench issues, so the upstream issue and its fix are public on GitHub and may be known to AI models. We say so on every task. You can choose tasks held back from our public practice site so candidates cannot rehearse them here, and every report shows the process evidence alongside the outcome, which is much harder to fake than a passing test.

Does the score decide who gets hired?

No. Scores support your reviewers' judgment and every report says so. Nobody is rejected automatically. Candidates are told before they start what is recorded, how the result is used, and how to ask for an accommodation or a different way to be assessed.

How does pricing work?

Your team's agreement sets the price per assessment. AI usage within the round's token limit and the report are included. Buy more from your team portal at the agreed price. Invitations alone use no assessments; one is used when a candidate's workspace is provisioned, before they press Start inside it. A round marked failed by the platform returns that assessment; a submission that fails grading tests still uses one.

How quickly can we start?

Contact us to discuss your roles, assessment format and setup. Once your team's workspace and agreement are ready, an owner can add colleagues and create assessment links.

See it on your own hiring loop.

Tell us about the roles you hire for and the assessment format you want to improve.