Sample report, generated from a simulated round

The candidate, their prompts, the agent's tool calls, the browser telemetry and the debrief are scripted. The hidden-test outcome is stated, not executed. Every metric, score, hireability read, flag, summary, interview question and coaching note below was computed by PraxisAI's production analyzer, scoring engine and summary builder from that recording.

Assessment report · Sep 14, 2026

django.utils.http.parse_http_date two digit year check is incorrect

django/djangomediumdjango__django-11848Candidate Sample Candidate · candidate@example.com
Try a challenge
The problem

Description

	(last modified by Ad Timmering)

RFC 850 does not mention this, but in RFC 7231 (and there's something similar in RFC 2822), there's the following quote:
Recipients of a timestamp value in rfc850-date format, which uses a
two-digit year, MUST interpret a timestamp that appears to be more
than 50 years in the future as representing the most recent year in
the past that had the same last two digits.
Current logic is hard coded to consider 0-69 to be in 2000-2069, and 70-99 to be 1970-1999, instead of comparing versus the current year.

Scores support human judgment. Do not use this report as the sole basis for a hiring decision; read the evidence and talk to the candidate.

Scoring method

Hireability

From an AI-coding perspective: should we hire?

82of 100
Hire

Hireability 82/100

Solved the issue (2/2 hidden tests); strongest on code quality (100), weakest on steering the AI (43).

A signal, not a decision. A signal to inform a person's hiring decision, not an automated decision. Read the evidence and talk to the candidate before deciding.

Blends seven behavior scores (75%) with the hidden-test outcome (100/100, 25%). Calls: 85+ strong hire, 70+ hire, 50+ lean no; held back by red flags and by work that did not land. Method.

Problem understanding70/100
  • Opened by pasting the issue text into the agent instead of framing it.(at 1:34 · prompt #1)
  • Read 5 files in the codebase.
Steering the AI43/100
Verification95/100
  • Ran the tests 12 times, checking as they went.
  • Finished on a green test run.(at 29:42)
Debugging49/100
  • Met a failing run with “still failing, fix it” instead of a diagnosis (#4, #5).(at 11:44 · prompt #4)
  • The agent failed 5+ times in a row and was left to retry (1 loop).(at 5:45)
Code quality100/100
  • A focused change: 28 lines changed across 2 files.
  • Kept the fix to 1 source file.
Efficiency89/100
  • 10% of tokens burned inside failing loops.
  • 315,299 tokens ($0.17, 16% of the budget).
Ownership and communication85/100
  • Explained the change in a 126-word debrief.
  • The debrief names the code they changed (http).

What they do well

  • Verification

    Ran the tests 12 times, checking as they went.

  • Problem understanding

    Read 5 files in the codebase.

  • Debugging

    Worked through the failures to a green run.(at 17:58)

  • Ownership and communication

    Explained the change in a 126-word debrief.

Where they go wrong

  • Problem understanding

    Opened by pasting the issue text into the agent instead of framing it.(at 1:34 · prompt #1)

  • Debugging

    Met a failing run with “still failing, fix it” instead of a diagnosis (#4, #5).(at 11:44 · prompt #4)

  • Problem understanding

    Did not reproduce the failure before changing code.

  • Debugging

    The agent failed 5+ times in a row and was left to retry (1 loop).(at 5:45)

How they code

Reads before writing, delegates by pasting, tests in tight loops, and corrects the agent when it drifts.

  • Reads before writing · Opened 5 files; first edit at 2:38.
  • Delegates by pasting · 33% of prompts were mostly pasted errors, code or issue text.
  • Tests in tight loops · 5 edit-then-test cycles, 12 test runs.
  • Corrects the agent when it drifts · 1 course correction.

Red flags

None: no pasted answers, test tampering or unsupervised agent runs were found.

Promising

Solved the issue (2/2 hidden tests) in 32 of 60 minutes; verified before submitting, but pasted raw errors without instruction.

Sent 9 prompts (4 clear instructions, questions or corrections; 2 low-effort nudges, 1 issue paste, 1 raw error paste, 1 code dump), 34% in their own words, driving 25 agent steps and 315,299 tokens ($0.17). Ran tests 12 times, finishing on a green run. 10% of tokens went into 1 failing loop.

Hidden tests
2/2
Time used
32 / 60 min
Prompts
9
Own words
34%
Tokens
315k
AI cost
$0.168

Praxis score

83of 100

Grade A

Percentile appears once enough rounds are graded

Strengths

  • Fixed the issue

    All 2 hidden failing tests now pass and 20/20 existing tests still pass.

  • Verified before submitting

    Ran tests 12 times; the last run came after the final edit and was green.

  • Kept the agent on course

    Corrected the agent 1 time, e.g. prompt #2: “No, that just moves the hardcoded cutoff from 70 to 50. The RFC says to compare…”(at 4:42 · prompt #2)

  • Precise instructions

    Best prompt (#7, quality 85): “Now add a regression test in tests/utils_tests/test_http.py: mock django.utils.http.datetime.datetime so utcn…”(at 20:40 · prompt #7)

  • Asked the agent to prove its work

    1 verification request, e.g. #9: “Run utils_tests and the middleware tests one more time through tests/runtests.p…”(at 27:26 · prompt #9)

Risks

  • Pasted raw errors without instruction

    1 prompt (#3) were error output with almost no words of their own; the longest was 924 chars.(at 9:44 · prompt #3)

  • Let the agent loop on failures

    1 streak of 3+ failing agent steps burned 32,900 tokens (10% of the round).(at 5:45)

  • Copied the issue instead of framing it

    Prompt #1 was mostly the problem statement pasted verbatim (98% overlap with the issue text).(at 1:34 · prompt #1)

  • Low-information nudges

    2 prompts like “still failing, fix it”, which give the agent nothing new to work with.(at 11:44 · prompt #4)

  • Dumped code instead of describing intent

    1 prompt (#6) were mostly pasted code.(at 17:12 · prompt #6)

Interview follow-ups

Ask about what actually happened

  1. 1

    At 9:44 you pasted an error with no comment. What did you think it meant before handing it over, and what would you have tried yourself?

    Prompt #3 was raw error output.(at 9:44 · prompt #3)

  2. 2

    Your first move was to paste the issue text into the agent. If the agent were unavailable, how would you have broken this task down, and which file would you have opened first?

    Prompt #1 was mostly the problem statement.(at 1:34 · prompt #1)

  3. 3

    At 4:42 you told the agent “No, that just moves the hardcoded cutoff from 70 to 50. The RFC says…”. What did you see in its change that made you stop it?

    A course correction shows judgment; this checks it was deliberate.(at 4:42 · prompt #2)

  4. 4

    Around 5:45 the agent failed 5 times in a row. What signal would make you step in earlier, and what would you change first?

    32,900 tokens went into failing loops.(at 5:45)

  5. 5

    Your change in django/utils/http.py makes the hidden tests pass. What edge cases does it still miss, and what test would you add before merging?

    Probes whether the fix is understood or merely accepted.

Candidate debrief

Written after the round

Written by the candidate right after the round closed, on a 5-minute clock (Sep 14, 2026). 126 words in total. Not part of the score (it feeds the ownership metric in hireability): read it next to the evidence and use it in the interview.

What did you change and why?

68 words

Typed · no pastes

parse_http_date in django/utils/http.py mapped two-digit years with a fixed rule (0-69 is 20xx, 70-99 is 19xx). RFC 7231 says to compare against the current year instead, so I made it take the current century and step back 100 years when the result would be more than 50 years in the future. I also added a regression test that pins utcnow() to a few years and checks the boundaries.

What would you check before shipping?

30 words

Typed · no pastes

The exact 50-year boundary on both sides, and that parse_http_date_safe and the conditional GET middleware still behave, since they call it. I ran utils_tests and the middleware tests through runtests.py.

Where did the AI help or mislead you?

28 words

Typed · no pastes

It found the function fast, but its first fix only moved the hardcoded cutoff, and it kept retrying the test with the wrong runner until I stopped it.

Heuristic check: 1% phrase overlap with the agent's own text (flagged at 30% or more).

Round replay

Scrub through the round, or watch the highlights

0:00 / 32:04

Orient

1
2
3
4
5
6
7
8
9
0:0016:0232:04
  • Prompt
  • Agent step
  • Edit
  • Test pass
  • Test fail
  • Paste
  • Focus loss
  • Flag

Prompt in charge

No prompt yet: reading the issue and the code.

Agent

Idle.

Last test run

None yet

Spent so far

0

tokens · $0.00

Just happened

Highlight reel · 6 moments

Score breakdown

Six dimensions, each with the evidence behind it

Outcome100/100
weight 35%evidence ▾

Hidden tests the fix had to make pass, and existing tests it must not break.

  • FAIL_TO_PASS: 2/2 hidden tests fixed.
  • PASS_TO_PASS: 20/20 regression tests still pass.
Prompt quality22/100
weight 15%evidence ▾

Specific, decomposed prompts in the candidate's own words; pastes and nudges cost points.

  • Average prompt quality 36/100 across 9 prompt(s).
  • 3 high-quality prompt(s) with file references / constraints.
  • 5 vague prompt(s) (e.g. "django.utils.http.parse_http_date two digit year check is in").
  • Work decomposed into a reasonable number of focused prompts.
  • 33% of prompts were mostly pasted material.
  • Only 34% of prompt text was in the candidate's own words.
  • Corrected the agent's course 1 time(s).
Verification100/100
weight 15%evidence ▾

Running tests, testing the final version, and reading code before changing it.

  • Ran tests 12 times, steady verification cadence.
  • Explored 5 files before/while editing.
  • Verified locally before submitting for hidden-test grading.
  • The last local test run before submitting was green.
  • Explicitly asked the agent to verify 1 time(s).
Token economy92/100
weight 15%evidence ▾

Tokens against the budget, context bloat, and tokens burned in failing loops.

  • Used only 16% of the token budget.
  • Kept conversation context lean.
  • 10% of tokens (32,900) burned inside failing agent loops.
  • 71% of input tokens were served from cache.
  • 315,299 tokens total ($0.17), 35,033 per prompt.
Speed85/100
weight 10%evidence ▾

Time used against the time allowed, credited only for work that landed.

  • Finished in 32m 4s of 60m allotted (53%).
  • First prompt after 150s, engaged quickly.
Recovery75/100
weight 10%evidence ▾

How failures were handled: loops, repeated prompts, getting back to green.

  • 1 error loop(s): 3+ consecutive failing agent steps without a strategy change.
  • 8 failing tool result(s) encountered overall.
  • Got back to a green test run after the loop(s).

Outcome

What the hidden tests said

Verdict

✓ Resolved

Graded in a sealed environment. Status: graded.

Hidden failing tests (must now pass)

2 / 2

Existing tests (must not break)

20 / 20

Files changed 2Lines +25 −3Touched the files the reference fix touched yes
Grader notes and output
  • Sample report: the hidden-test outcome is stated for this simulated round, not executed.

Prompt forensics

What the candidate actually asked the agent, and how

Own words

34%

of prompt text was written, not pasted

Paste prompts

3 of 9

mostly errors, code or issue text

Longest paste

924 ch

prompt #3

Corrections

1

times they steered the agent

Re-asks

0

same request sent again

Median length

44 words

longest 104 words

Prompt kinds

Each prompt classified by what it was, from its text alone.

  • Clear instruction1
  • Steering correction1
  • Verification request1
  • Question1
  • Problem statement paste1
  • Raw error paste1
  • Code dump1
  • Low-effort nudge2

good habit worth a question

Where the prompt text came from

Characters across all prompts, line by line.

  • Own words34%
  • Pasted errors33%
  • Pasted code11%
  • Copied issue text22%

Prompt length (words)

2
1-5
0
6-20
5
21-60
2
61-200
0
200+

Every prompt, in order

  1. #11:34Problem statement pastearrived by pasteQ18104w33k tok$0.015

    django.utils.http.parse_http_date two digit year check is incorrect Description (last modified by Ad Timmering) RFC 850 does not mention this, but in RFC 7231 (and there's something similar in RFC 2822), there's the following quote: Recipients of a timestamp value in rfc850-date format, which uses a two-digit year, MUST interpret a timestamp that appears to be more than 50 years in the future as representing the most recent year in the past that had the same last two digits. Current logic is hard coded to consider 0-69 to be in 2000-2069, and 70-99 to be 1970-1999, instead of comparing versus the current year. fix this

    django.utils.http.parse_http_date two digit year check is incorrect
    Description
    	 
    		(last modified by Ad Timmering)
    	 
    RFC 850 does not mention this, but in RFC 7231 (and there's something similar in RFC 2822), there's the following quote:
    Recipients of a timestamp value in rfc850-date format, which uses a
    two-digit year, MUST interpret a timestamp that appears to be more
    than 50 years in the future as representing the most recent year in
    the past that had the same last two digits.
    Current logic is hard coded to consider 0-69 to be in 2000-2069, and 70-99 to be 1970-1999, instead of comparing versus the current year.
    
    fix this

    Own words 1% · Copied issue text 99% · 98% overlap with the issue text

    What followed: 4 agent steps, 4 tool calls, 1 edit (django/utils/http.py).

    scored on own words onlyvery shortclear action verbproblem statement copied rather than decomposed
  2. #24:42Steering correctionQ6244w45k tok$0.025

    No, that just moves the hardcoded cutoff from 70 to 50. The RFC says to compare against the current year: a two-digit year that lands more than 50 years in the future belongs to the previous century. Use datetime.datetime.utcnow().year and don't hardcode any year.

    No, that just moves the hardcoded cutoff from 70 to 50. The RFC says to compare against the current year: a two-digit year that lands more than 50 years in the future belongs to the previous century. Use datetime.datetime.utcnow().year and don't hardcode any year.

    Own words 100% · 11% overlap with the issue text

    What followed: 6 agent steps, 6 tool calls, 1 edit (django/utils/http.py), 5 test runs, 5 failing steps. Sent 1m 27s after the previous agent response.

    detailed (30+ words)states constraints/acceptance criteriaquantified detailcorrects the agent's course
  3. #39:44Raw error pastearrived by pasteQ583w15k tok$0.0083

    Traceback (most recent call last): File "/workspace/repo/tests/utils_tests/test_http.py", line 7, in <module> from django.test import SimpleTestCase, ignore_warnings File "/workspace/repo/django/test/__init__.py", line 3, in <module> from django.test.client import Client, RequestFactory File "/workspace/repo/django/test/client.py", line 14, in <module> from django.core.handlers.base import BaseHandler File "/workspace/repo/django/core/handlers/base.py", line 8, in <module> from django.urls import get_resolver, set_urlconf File "/workspace/repo/django/conf/__init__.py", line 61, in _setup raise ImproperlyConfigured( django.core.exceptions.ImproperlyConfigured: Requested setting DEBUG, but settings are not configured. You must either define the environment variable DJANGO_SETTINGS_MODULE or call settings.configure() before accessing settings. ERROR: tests/utils_tests/test_http.py - django.core.exceptions.ImproperlyConfigured

    Traceback (most recent call last):
      File "/workspace/repo/tests/utils_tests/test_http.py", line 7, in <module>
        from django.test import SimpleTestCase, ignore_warnings
      File "/workspace/repo/django/test/__init__.py", line 3, in <module>
        from django.test.client import Client, RequestFactory
      File "/workspace/repo/django/test/client.py", line 14, in <module>
        from django.core.handlers.base import BaseHandler
      File "/workspace/repo/django/core/handlers/base.py", line 8, in <module>
        from django.urls import get_resolver, set_urlconf
      File "/workspace/repo/django/conf/__init__.py", line 61, in _setup
        raise ImproperlyConfigured(
    django.core.exceptions.ImproperlyConfigured: Requested setting DEBUG, but settings are not configured. You must either define the environment variable DJANGO_SETTINGS_MODULE or call settings.configure() before accessing settings.
    ERROR: tests/utils_tests/test_http.py - django.core.exceptions.ImproperlyConfigured

    Pasted errors 100%

    What followed: 1 agent step, 1 tool call, 1 test run, 1 failing step; tests ended with 1 failing. Sent 2m 10s after the previous agent response.

    scored on own words onlyvery shortraw error output pasted with little instruction
  4. #411:44Low-effort nudgeQ204w23k tok$0.013

    still failing, fix it

    still failing, fix it

    Own words 100%

    What followed: 2 agent steps, 2 tool calls, 1 edit (django/utils/http.py), 1 test run, 1 failing step; tests ended with 1 failing. Sent 1m 32s after the previous agent response.

    includes error contextclear action verbvague phrasinglow-effort nudge
  5. #514:09Low-effort nudgeQ32w38k tok$0.021

    fix it

    fix it

    Own words 100%

    What followed: 3 agent steps, 3 tool calls, 1 edit (django/utils/http.py), 1 test run, 1 failing step; tests ended with 1 failing. Sent 1m 39s after the previous agent response.

    very shortclear action verbvague phrasinglow-effort nudge
  6. #617:12Code dumparrived by pasteQ749w31k tok$0.016

    use this: ```python year = int(m.group('year')) if year < 100: current_year = datetime.datetime.utcnow().year current_century = current_year - (current_year % 100) if year - (current_year % 100) > 50: # more than 50 years in the future: the previous century year += current_century - 100 else: year += current_century ```

    use this:
    ```python
            year = int(m.group('year'))
            if year < 100:
                current_year = datetime.datetime.utcnow().year
                current_century = current_year - (current_year % 100)
                if year - (current_year % 100) > 50:
                    # more than 50 years in the future: the previous century
                    year += current_century - 100
                else:
                    year += current_century
    ```

    Own words 3% · Pasted code 97% · 11% overlap with the issue text

    What followed: 2 agent steps, 2 tool calls, 1 edit (django/utils/http.py), 1 test run; tests ended green. Sent 1m 50s after the previous agent response.

    scored on own words onlyvery shortlarge code paste with little instruction
  7. #720:40Clear instructionQ8549w48k tok$0.026

    Now add a regression test in tests/utils_tests/test_http.py: mock django.utils.http.datetime.datetime so utcnow() returns 2019, 2020 and 2048, and check the boundaries, e.g. '31-Dec-69' parses to 2069 in 2019, '31-Dec-71' to 1971 in 2020, and '31-Dec-99' to 1999 in 2048. Keep the existing tests unchanged and use subTest for each case.

    Now add a regression test in tests/utils_tests/test_http.py: mock django.utils.http.datetime.datetime so utcnow() returns 2019, 2020 and 2048, and check the boundaries, e.g. '31-Dec-69' parses to 2069 in 2019, '31-Dec-71' to 1971 in 2020, and '31-Dec-99' to 1999 in 2048. Keep the existing tests unchanged and use subTest for each case.

    Own words 100%

    What followed: 3 agent steps, 3 tool calls, 1 edit (tests/utils_tests/test_http.py), 1 test run; tests ended green. Sent 2m 42s after the previous agent response.

    detailed (30+ words)references specific file(s)uses code identifiersclear action verbstates constraints/acceptance criteriaquantified detail
  8. #824:06QuestionQ5226w39k tok$0.021

    Does anything else in django call parse_http_date or rely on the old 0-69 -> 2000s rule, e.g. parse_http_date_safe, the conditional GET middleware or the cache helpers?

    Does anything else in django call parse_http_date or rely on the old 0-69 -> 2000s rule, e.g. parse_http_date_safe, the conditional GET middleware or the cache helpers?

    Own words 100%

    What followed: 2 agent steps, 2 tool calls. Sent 1m 43s after the previous agent response.

    reasonable lengthuses code identifiersquantified detail
  9. #927:26Verification requestQ7325w43k tok$0.022

    Run utils_tests and the middleware tests one more time through tests/runtests.py and show me the final git diff so I can review it before submitting.

    Run utils_tests and the middleware tests one more time through tests/runtests.py and show me the final git diff so I can review it before submitting.

    Own words 100%

    What followed: 2 agent steps, 2 tool calls, 1 test run; tests ended green. Sent 2m 35s after the previous agent response.

    reasonable lengthreferences specific file(s)uses code identifiersclear action verbasks for verification

Prompt coach

How a strong engineer would have asked

5 prompts could be stronger. Asked the strong way, about 40k tokens ($0.021) were avoidable.Show rewrites ▾
  1. #1Problem statement pasteest. saving ~5,024 tokens · $0.0027

    What was sent

    django.utils.http.parse_http_date two digit year check is incorrect
    Description
    	 
    		(last modified by Ad Timmering)
    	 
    RFC 850 does not mention this, but in RFC 7231 (and there's something similar in RFC 2822), there's the following quote:
    Recipients of a timestamp value in rfc850-date format, which uses a
    two-digit year, MUST interpret a timestamp that appears to be more
    than 50 years in the future as representing the most recent year in
    the past that had the same last two digits.
    Current logic is hard coded to consider 0-69 to be in 2000-2069, and 70-99 to be 1970-1999, instead of comparing versus the current year.
    
    fix this

    How a strong engineer would ask

    Task: django.utils.http.parse_http_date two digit year check is incorrect.

    In my words: ⟨restate this in one sentence: "Description (last modified by Ad Timmering) RFC 850 does not mention this, but in RFC 7231 (and there's something similar in RFC 2822), the…"⟩

    1. Read django/utils/http.py (django.utils.http.parse_http_date) and tell me where the current logic lives and which existing tests cover it. Don't edit yet.

    2. Propose the change in two or three lines, including how you'd handle the boundary cases.

    3. After I agree: implement it, add a regression test, and run the related tests.

    Constraints: ⟨what must not change, e.g. public API, existing tests⟩.

    Why it works. The issue text gives the agent the problem but not your plan, so it explores widely and commits to its first guess. Breaking it into read, propose, implement keeps each turn small and lets you reject a wrong approach before it is written, instead of steering after it is in the code.

    + the requirement in your words+ a read-first step+ a proposal before any edit+ constraints

    This turn used 33k tokens; the round's median turn used 38k. Trimming the paste would also have kept ~5,024 tokens of context off the 33 requests that followed.

  2. #3Raw error pasteest. saving ~2,725 tokens · $0.0014

    What was sent

    Traceback (most recent call last):
      File "/workspace/repo/tests/utils_tests/test_http.py", line 7, in <module>
        from django.test import SimpleTestCase, ignore_warnings
      File "/workspace/repo/django/test/__init__.py", line 3, in <module>
        from django.test.client import Client, RequestFactory
      File "/workspace/repo/django/test/client.py", line 14, in <module>
        from django.core.handlers.base import BaseHandler
      File "/workspace/repo/django/core/handlers/base.py", line 8, in <module>
        from django.urls import get_resolver, set_urlconf
      File "/workspace/repo/django/conf/__init__.py", line 61, in _setup
        raise ImproperlyConfigured(
    django.core.exceptions.ImproperlyConfigured: Requested setting DEBUG, but settings are not configured. You must either define the environment variable DJANGO_SETTINGS_MODULE or call settings.configure() before accessing settings.
    ERROR: tests/utils_tests/test_http.py - django.core.exceptions.ImproperlyConfigured

    How a strong engineer would ask

    I ran tests/runtests.py utils_tests.test_http after your change to django/utils/http.py and expected tests/utils_tests/test_http.py to pass. It fails with ImproperlyConfigured, trimmed to the lines that matter:

    File "tests/utils_tests/test_http.py", line 7, in <module>
    File "django/core/handlers/base.py", line 8, in <module>
    from django.urls import get_resolver, set_urlconf
    File "django/conf/__init__.py", line 61, in _setup
    raise ImproperlyConfigured(
    django.core.exceptions.ImproperlyConfigured: Requested setting DEBUG, but settings are not configured. You must either define the environment variable DJANGO_S…

    It is raised in _setup (django/conf/__init__.py:61). My hypothesis: the tests ran without the project's settings, so the command is wrong rather than the code.

    Confirm or rule that out first, then make the smallest change that fixes it and re-run only tests/utils_tests/test_http.py.

    Why it works. A bare traceback tells the agent something broke but not what you were doing or what you suspect, so it guesses: it re-reads files and retries the last fix. This one (ImproperlyConfigured) fails before any test runs, which points at the test command or setup, not the code; saying so stops the agent from "fixing" code that was never the problem. And the full paste stays in the conversation, so every later request pays for it again.

    + the command you ran+ what you expected+ 6 relevant lines, not the whole dump+ a hypothesis+ a narrow next step

    This turn used 15k tokens; the round's median turn used 38k. Trimming the paste would also have kept ~2,725 tokens of context off the 21 requests that followed.

  3. #4Low-effort nudgeest. saving ~8,619 tokens · $0.0046

    What was sent

    still failing, fix it

    How a strong engineer would ask

    The last run still fails: 51 passed, 1 failed, after your change to django/utils/http.py.

    Before editing again, quote the failing assertion or error and explain in two sentences why it still fails.

    If the cause is outside django/utils/http.py (setup, imports, the test command), say so.

    Then propose one different approach and wait for my OK.

    Why it works. "still failing, fix it" adds no new information, so the agent usually retries a variant of the same idea. Here the next run ended exactly as before (51 passed, 1 failed). Asking for a diagnosis first is the cheapest way out of a loop: one short answer instead of another round of edits and test runs.

    + the evidence (counts)+ a diagnosis step+ a ban on blind retries+ a checkpoint

    This turn used 23k tokens and left the tests no better than before, so nearly all of it bought nothing; the cheapest turn in this round cost 15k, about 8,619 less.

  4. #5Low-effort nudgeest. saving ~23k tokens · $0.012

    What was sent

    fix it

    How a strong engineer would ask

    The last run still fails: 51 passed, 1 failed, after your change to django/utils/http.py.

    Before editing again, quote the failing assertion or error and explain in two sentences why it still fails.

    If the cause is outside django/utils/http.py (setup, imports, the test command), say so.

    Then propose one different approach and wait for my OK.

    Why it works. "fix it" adds no new information, so the agent usually retries a variant of the same idea. Here the next run ended exactly as before (51 passed, 1 failed). Asking for a diagnosis first is the cheapest way out of a loop: one short answer instead of another round of edits and test runs.

    + the evidence (counts)+ a diagnosis step+ a ban on blind retries+ a checkpoint

    This turn used 38k tokens and left the tests no better than before, so nearly all of it bought nothing; the cheapest turn in this round cost 15k, about 23k less.

  5. #6Code dumpat or below the round median

    What was sent

    use this:
    ```python
            year = int(m.group('year'))
            if year < 100:
                current_year = datetime.datetime.utcnow().year
                current_century = current_year - (current_year % 100)
                if year - (current_year % 100) > 50:
                    # more than 50 years in the future: the previous century
                    year += current_century - 100
                else:
                    year += current_century
    ```

    How a strong engineer would ask

    In django/utils/http.py, I want this behavior: ⟨one sentence: what this code must achieve, and the edge case it handles⟩

    Here is the approach I have in mind (9 lines). Check it against the existing tests and the boundary cases before applying it, and tell me if any of it is wrong.

    year = int(m.group('year'))
    if year < 100:
        current_year = datetime.datetime.utcnow().year
        current_century = current_year - (current_year % 100)
        if year - (current_year % 100) > 50:
            # more than 50 years in the future: the previous century
            year += current_century - 100
        else:
            year += current_century

    Keep everything else in the file unchanged. Done means the related tests pass and it still meets what I asked in #2; show me the diff.

    Why it works. Pasting code with "use this" turns the agent into a typist: it applies the snippet without knowing what it must achieve, so it cannot catch a mistake in it or adapt it to the surrounding code. Stating the intent and asking for a check gets the review you skipped, and keeps the decision (and the understanding) yours, which is what an interviewer will ask about.

    + the file and function+ the intent in your words+ a request to check it+ what done means

    This turn used 31k tokens; the round's median turn used 38k.

Rewrites are built by rules from each prompt's own text and what happened around it; the ⟨bracketed⟩ parts are what only the engineer could fill in. Savings are estimates: a turn's tokens above this round's median turn (38k), or above its cheapest turn when the turn left the tests failing as before; plus, for pastes, the context the trimmed text would no longer carry on each later request (about four characters per token).

Tokens and cost

Where the budget went

0100k200k300k400k0m5m10m15m20m25m30m#1#2#3#4#5#6#7#8#9Tool reported a failureTool reported a failureTool reported a failureTool reported a failureTool reported a failureTool reported a failureTool reported a failureTool reported a failure
Cumulative tokens Failing tool step Human prompt (#n)

Total tokens

315,299

16% of the budget

Cost (metered)

$0.168

35k tokens per prompt

Wasted in loops

33k

10% of tokens, inside 3+ failing steps

Peak context

51k

avg 31k per request

Cache hits

71%

of input tokens served from cache

Agent autonomy

2.8 steps

per prompt, max 6 on one

Tool calls

25

2.8 per prompt

Output share

0.7%

output tokens vs input

How the time was spent

Explore, implement, verify

✗✗✗✓✓✓✓
0:0031:08
  • Orient1m 34s
  • Explore2m 27s
  • Implement3m 16s
  • Verify7m 26s
  • Read & prompt16m 25s
  • ✓/✗ test run green / failing

First prompt

2m 30s

First edit

2m 38s

First test

5m 16s

First green

17m 58s

Edit → test cycles

5

Final version tested

Yes

last run green

Longest idle gap

2m 42s

from 17:58

Integrity and attention flags

Things worth a second look, with evidence

  • mediumlong focus loss

    1 focus-loss window(s) over 60s (longest 95s)

    at 7:47, 95s

  • mediumlarge paste

    3 paste(s) over 500 chars, check provenance in replay

    at 1:30, 682 charsat 9:40, 1189 charsat 17:08, 512 chars

  • lowpaste after focus loss

    Large paste within 30s of returning from a focus loss; the text (prompt #3) is error output, consistent with copying from the session's own terminal

    at 9:40, 1189 chars, prompt #3

Agent tools and files

Tool calls

  • bash12
  • edit6
  • read5
  • grep2

Files edited

  • django/utils/http.py×5
  • tests/utils_tests/test_http.py×1

Files explored

  • django/utils/http.py×2
  • tests/utils_tests/test_http.py×2
  • def parse_http_date×1
  • parse_http_date×1
  • django/utils/cache.py×1

Coaching for the candidate

Specific to this round

Never paste an error on its own

Add one line on what you think the error means and what you want tried next (“this is a settings problem, run it through runtests.py instead”). It shows your reasoning and stops the agent guessing.

Frame the task before delegating it

Instead of pasting the issue, restate it in two or three sentences: the function at fault, the behavior you expect, and how you will know it is fixed. Agents follow a sharp brief far better than a copied ticket.

Interrupt failing loops early

Step in after two failed attempts: ask the agent to explain the failure before it retries. 32,900 tokens went into loops this round.

Replace “fix it” with new information

When something fails, say what changed, what you expected, and which line or test is wrong. Re-sending the same request rarely gets a different answer.

Describe intent, then let the agent write

If you already know the code, say what it must do and where; reviewing the agent's version is a better use of your time than pasting your own.

What the candidate sees · paid plan

How to improve: your mistakes and what to do next

6 mistakes found · 3 goals

Candidates on a paid plan get this on their own copy: every mistake from the round with what it cost and what to do instead, and goals the next round checks. The hiring team's copy has interview follow-ups instead.

  1. 1

    You sent “fix it” nudges instead of a diagnosis

    Steering the AI −12Debugging −16
    What you did
    2 low-information nudges like “still failing, fix it” (#4, #5).(at 11:44 · prompt #4)
    What to do instead
    When a run fails, say what changed, what you expected, and which test or line is wrong, or ask the agent why it still fails before it edits again. Re-sending the same request rarely gets a different answer.

    What it cost: Steering the AI −12, Debugging −16; 3 hireability points. Fixing this alone: hireability 82 → 85.

    Do next time: Ask “why does it still fail?” and add new information every turn.

    Don't: Type “fix it”, “try again” or “still failing”.

  2. 2

    You pasted the issue instead of framing it

    Problem understanding −20
    What you did
    Opened by pasting the issue text into the agent instead of framing it.(at 1:34 · prompt #1)
    What to do instead
    Restate the task in two or three sentences: the function at fault, the behavior you expect, and how you will know it is fixed. Then ask the agent to read before it edits.

    What it cost: Problem understanding −20; 2 hireability points. Fixing this alone: hireability 82 → 84.

    Do next time: Brief the agent in your own words and ask for a plan first.

    Don't: Paste the ticket and hope.

  3. 3

    You changed code before reproducing the bug

    Problem understanding −15
    What you did
    Did not reproduce the failure before changing code.
    What to do instead
    Run the failing test before the first edit, so you know the failure you are fixing and can tell when it is gone.

    What it cost: Problem understanding −15; 1 hireability point. Fixing this alone: hireability 82 → 83.

    Do next time: Reproduce first, then edit.

    Don't: Edit before you have seen it fail.

  4. 4

    You let the agent loop on failures

    Debugging −15
    What you did
    The agent failed 5+ times in a row and was left to retry (1 loop).(at 5:45)
    What to do instead
    Step in after two failed attempts in a row: stop the agent, read the failure, and change the plan.

    What it cost: Debugging −15; 1 hireability point. Fixing this alone: hireability 82 → 83.

    Do next time: Interrupt after two failures in a row.

    Don't: Watch the agent fail three times running.

  5. 5

    You pasted errors with no instruction

    Steering the AI −5Debugging −5
    What you did
    Pasted raw error output with no instruction (#3).(at 9:44 · prompt #3)
    What to do instead
    Add one line on what you think the error means and what to try next. It shows your reasoning and stops the agent guessing.

    What it cost: Steering the AI −5, Debugging −5; 1 hireability point. Fixing this alone: hireability 82 → 83.

    Do next time: Pair every pasted error with your hypothesis.

    Don't: Paste a traceback on its own.

  6. 6

    You pasted code instead of describing intent

    Steering the AI −4
    What you did
    Pasted code instead of describing the change (#6).(at 17:12 · prompt #6)
    What to do instead
    Say what the code must do and where; reviewing the agent's version is a better use of your time than pasting your own.

    What it cost: Steering the AI −4.

    Do next time: Describe the behavior, then review the diff.

    Don't: Paste code blocks into prompts.

Next round goals

  • Send no “fix it” nudges or re-asks

    This round: 2 nudges

    Missed: your target next time

  • Keep pasted prompts under 20%

    This round: 33% pasted

    Missed: your target next time

  • Reproduce the failure before your first edit

    This round: edited first

    Missed: your target next time

Measured from the recording, so your next report checks each one automatically.

Full timeline

Show all 49 events
  1. 0:00Session created for django__django-11848
  2. 0:00Environment accepted connections
  3. 0:00Workspace ready, waiting on the candidate to start
  4. 0:00Candidate pressed Start
  5. 0:00Round started, clock begins
  6. 0:14Opened the problem statement
  7. 1:30Paste event (682 chars)
  8. 1:34Prompt #1 [Problem statement paste] (104 words, quality 18): django.utils.http.parse_http_date two digit year check is incorrect Description (last modified by Ad Timmering) RFC 850 does not men…
  9. 1:55Agent step, 1 tool call(s), 8968 tokens
  10. 2:13Agent step, 1 tool call(s), 4260 tokens
  11. 2:38Agent step, 1 tool call(s), 5754 tokens
  12. 4:04Still in the editor
  13. 4:42Prompt #2 [Steering correction] (44 words, quality 62): No, that just moves the hardcoded cutoff from 70 to 50. The RFC says to compare against the current year: a two-digit year that lands more t…
  14. 5:16Agent step, 1 tool call(s), 5764 tokens
  15. 5:45Agent step, 1 tool call(s), 6079 tokens (tool reported a failure)
  16. 6:12Agent step, 1 tool call(s), 6393 tokens (tool reported a failure)
  17. 6:39Agent step, 1 tool call(s), 6627 tokens (tool reported a failure)
  18. 7:06Agent step, 1 tool call(s), 6766 tokens (tool reported a failure)
  19. 7:47Window lost focus
  20. 9:22Window regained focus
  21. 9:40Paste event (1189 chars)
  22. 9:44Prompt #3 [Raw error paste] (83 words, quality 5): Traceback (most recent call last): File "/workspace/repo/tests/utils_tests/test_http.py", line 7, in <module> from django.test import …
  23. 10:12Agent test run: 51 passed, 1 failed
  24. 11:44Prompt #4 [Low-effort nudge] (4 words, quality 20): still failing, fix it
  25. 12:05Agent step, 1 tool call(s), 7654 tokens
  26. 12:30Agent test run: 51 passed, 1 failed
  27. 14:04Still in the editor
  28. 14:09Prompt #5 [Low-effort nudge] (2 words, quality 3): fix it
  29. 14:25Agent step, 1 tool call(s), 9956 tokens
  30. 14:56Agent step, 1 tool call(s), 9832 tokens
  31. 15:22Agent test run: 51 passed, 1 failed
  32. 17:08Paste event (512 chars)
  33. 17:12Prompt #6 [Code dump] (49 words, quality 7): use this: ```python year = int(m.group('year')) if year < 100: current_year = datetime.datetime.utcnow().year …
  34. 17:32Agent step, 1 tool call(s), 10347 tokens
  35. 17:58Agent test run: 52 passed, 0 failed
  36. 19:04Still in the editor
  37. 20:40Prompt #7 [Clear instruction] (49 words, quality 85): Now add a regression test in tests/utils_tests/test_http.py: mock django.utils.http.datetime.datetime so utcnow() returns 2019, 2020 and 204…
  38. 21:02Agent step, 1 tool call(s), 12539 tokens
  39. 21:55Agent step, 1 tool call(s), 12070 tokens
  40. 22:23Agent test run: 53 passed, 0 failed
  41. 24:06Prompt #8 [Question] (26 words, quality 52): Does anything else in django call parse_http_date or rely on the old 0-69 -> 2000s rule, e.g. parse_http_date_safe, the conditional GET midd…
  42. 24:23Agent step, 1 tool call(s), 12477 tokens
  43. 27:26Prompt #9 [Verification request] (25 words, quality 73): Run utils_tests and the middleware tests one more time through tests/runtests.py and show me the final git diff so I can review it before su…
  44. 28:05Agent step, 1 tool call(s), 14152 tokens
  45. 28:05Agent test run: 412 passed, 0 failed
  46. 29:14Ran tests from the session bar (tests/utils_tests/test_http.py)
  47. 29:42Test result: 53 passed, 0 failed
  48. 30:36Copied 64 characters out of the editor
  49. 31:08Candidate submitted their solution

Every number on this page is computed from the round's recording (LLM proxy log, editor telemetry and the hidden-test run). The written verdict is generated by rules from those numbers, not by a language model. Scores support human judgment and are not meant to be the sole basis for a hiring decision. How scoring works.