Sample report, generated from a simulated round
The candidate, their prompts, the agent's tool calls, the browser telemetry and the debrief are scripted. The hidden-test outcome is stated, not executed. Every metric, score, hireability read, flag, summary, interview question and coaching note below was computed by PraxisAI's production analyzer, scoring engine and summary builder from that recording.
Assessment report · Sep 14, 2026
django.utils.http.parse_http_date two digit year check is incorrect
The problem
Description
(last modified by Ad Timmering)RFC 850 does not mention this, but in RFC 7231 (and there's something similar in RFC 2822), there's the following quote:
Recipients of a timestamp value in rfc850-date format, which uses a
two-digit year, MUST interpret a timestamp that appears to be more
than 50 years in the future as representing the most recent year in
the past that had the same last two digits.
Current logic is hard coded to consider 0-69 to be in 2000-2069, and 70-99 to be 1970-1999, instead of comparing versus the current year.
Scores support human judgment. Do not use this report as the sole basis for a hiring decision; read the evidence and talk to the candidate.
Scoring methodHireability
From an AI-coding perspective: should we hire?
Hireability 82/100
Solved the issue (2/2 hidden tests); strongest on code quality (100), weakest on steering the AI (43).
A signal, not a decision. A signal to inform a person's hiring decision, not an automated decision. Read the evidence and talk to the candidate before deciding.
Blends seven behavior scores (75%) with the hidden-test outcome (100/100, 25%). Calls: 85+ strong hire, 70+ hire, 50+ lean no; held back by red flags and by work that did not land. Method.
- Opened by pasting the issue text into the agent instead of framing it.(at 1:34 · prompt #1)
- Read 5 files in the codebase.
- 2 low-information nudges like “still failing, fix it” (#4, #5).(at 11:44 · prompt #4)
- Corrected the agent 1 time, e.g. #2: “No, that just moves the hardcoded cutoff from 70 to 50. The…”(at 4:42 · prompt #2)
- Ran the tests 12 times, checking as they went.
- Finished on a green test run.(at 29:42)
- Met a failing run with “still failing, fix it” instead of a diagnosis (#4, #5).(at 11:44 · prompt #4)
- The agent failed 5+ times in a row and was left to retry (1 loop).(at 5:45)
- A focused change: 28 lines changed across 2 files.
- Kept the fix to 1 source file.
- 10% of tokens burned inside failing loops.
- 315,299 tokens ($0.17, 16% of the budget).
- Explained the change in a 126-word debrief.
- The debrief names the code they changed (http).
What they do well
- Verification
Ran the tests 12 times, checking as they went.
- Problem understanding
Read 5 files in the codebase.
- Debugging
Worked through the failures to a green run.(at 17:58)
- Ownership and communication
Explained the change in a 126-word debrief.
Where they go wrong
- Problem understanding
Opened by pasting the issue text into the agent instead of framing it.(at 1:34 · prompt #1)
- Debugging
Met a failing run with “still failing, fix it” instead of a diagnosis (#4, #5).(at 11:44 · prompt #4)
- Problem understanding
Did not reproduce the failure before changing code.
- Debugging
The agent failed 5+ times in a row and was left to retry (1 loop).(at 5:45)
How they code
Reads before writing, delegates by pasting, tests in tight loops, and corrects the agent when it drifts.
- Reads before writing · Opened 5 files; first edit at 2:38.
- Delegates by pasting · 33% of prompts were mostly pasted errors, code or issue text.
- Tests in tight loops · 5 edit-then-test cycles, 12 test runs.
- Corrects the agent when it drifts · 1 course correction.
Red flags
None: no pasted answers, test tampering or unsupervised agent runs were found.
Solved the issue (2/2 hidden tests) in 32 of 60 minutes; verified before submitting, but pasted raw errors without instruction.
Sent 9 prompts (4 clear instructions, questions or corrections; 2 low-effort nudges, 1 issue paste, 1 raw error paste, 1 code dump), 34% in their own words, driving 25 agent steps and 315,299 tokens ($0.17). Ran tests 12 times, finishing on a green run. 10% of tokens went into 1 failing loop.
- Hidden tests
- 2/2
- Time used
- 32 / 60 min
- Prompts
- 9
- Own words
- 34%
- Tokens
- 315k
- AI cost
- $0.168
Praxis score
Grade A
Percentile appears once enough rounds are graded
Strengths
Fixed the issue
All 2 hidden failing tests now pass and 20/20 existing tests still pass.
Verified before submitting
Ran tests 12 times; the last run came after the final edit and was green.
Kept the agent on course
Corrected the agent 1 time, e.g. prompt #2: “No, that just moves the hardcoded cutoff from 70 to 50. The RFC says to compare…”(at 4:42 · prompt #2)
Precise instructions
Best prompt (#7, quality 85): “Now add a regression test in tests/utils_tests/test_http.py: mock django.utils.http.datetime.datetime so utcn…”(at 20:40 · prompt #7)
Asked the agent to prove its work
1 verification request, e.g. #9: “Run utils_tests and the middleware tests one more time through tests/runtests.p…”(at 27:26 · prompt #9)
Risks
Pasted raw errors without instruction
1 prompt (#3) were error output with almost no words of their own; the longest was 924 chars.(at 9:44 · prompt #3)
Let the agent loop on failures
1 streak of 3+ failing agent steps burned 32,900 tokens (10% of the round).(at 5:45)
Copied the issue instead of framing it
Prompt #1 was mostly the problem statement pasted verbatim (98% overlap with the issue text).(at 1:34 · prompt #1)
Low-information nudges
2 prompts like “still failing, fix it”, which give the agent nothing new to work with.(at 11:44 · prompt #4)
Dumped code instead of describing intent
1 prompt (#6) were mostly pasted code.(at 17:12 · prompt #6)
Interview follow-ups
Ask about what actually happened
- 1
At 9:44 you pasted an error with no comment. What did you think it meant before handing it over, and what would you have tried yourself?
Prompt #3 was raw error output.(at 9:44 · prompt #3)
- 2
Your first move was to paste the issue text into the agent. If the agent were unavailable, how would you have broken this task down, and which file would you have opened first?
Prompt #1 was mostly the problem statement.(at 1:34 · prompt #1)
- 3
At 4:42 you told the agent “No, that just moves the hardcoded cutoff from 70 to 50. The RFC says…”. What did you see in its change that made you stop it?
A course correction shows judgment; this checks it was deliberate.(at 4:42 · prompt #2)
- 4
Around 5:45 the agent failed 5 times in a row. What signal would make you step in earlier, and what would you change first?
32,900 tokens went into failing loops.(at 5:45)
- 5
Your change in django/utils/http.py makes the hidden tests pass. What edge cases does it still miss, and what test would you add before merging?
Probes whether the fix is understood or merely accepted.
Candidate debrief
Written after the round
Written by the candidate right after the round closed, on a 5-minute clock (Sep 14, 2026). 126 words in total. Not part of the score (it feeds the ownership metric in hireability): read it next to the evidence and use it in the interview.
What did you change and why?
68 words
Typed · no pastes
parse_http_date in django/utils/http.py mapped two-digit years with a fixed rule (0-69 is 20xx, 70-99 is 19xx). RFC 7231 says to compare against the current year instead, so I made it take the current century and step back 100 years when the result would be more than 50 years in the future. I also added a regression test that pins utcnow() to a few years and checks the boundaries.
What would you check before shipping?
30 words
Typed · no pastes
The exact 50-year boundary on both sides, and that parse_http_date_safe and the conditional GET middleware still behave, since they call it. I ran utils_tests and the middleware tests through runtests.py.
Where did the AI help or mislead you?
28 words
Typed · no pastes
It found the function fast, but its first fix only moved the hardcoded cutoff, and it kept retrying the test with the wrong runner until I stopped it.
Heuristic check: 1% phrase overlap with the agent's own text (flagged at 30% or more).
Round replay
Scrub through the round, or watch the highlights
0:00 / 32:04
Orient
- Prompt
- Agent step
- Edit
- Test pass
- Test fail
- Paste
- Focus loss
- Flag
Prompt in charge
No prompt yet: reading the issue and the code.
Agent
Idle.
Last test run
None yet
Spent so far
0
tokens · $0.00
Just happened
Highlight reel · 6 moments
Score breakdown
Six dimensions, each with the evidence behind it
Outcome100/100weight 35%evidence ▾hide ▴
Hidden tests the fix had to make pass, and existing tests it must not break.
- FAIL_TO_PASS: 2/2 hidden tests fixed.
- PASS_TO_PASS: 20/20 regression tests still pass.
Prompt quality22/100weight 15%evidence ▾hide ▴
Specific, decomposed prompts in the candidate's own words; pastes and nudges cost points.
- Average prompt quality 36/100 across 9 prompt(s).
- 3 high-quality prompt(s) with file references / constraints.
- 5 vague prompt(s) (e.g. "django.utils.http.parse_http_date two digit year check is in").
- Work decomposed into a reasonable number of focused prompts.
- 33% of prompts were mostly pasted material.
- Only 34% of prompt text was in the candidate's own words.
- Corrected the agent's course 1 time(s).
Verification100/100weight 15%evidence ▾hide ▴
Running tests, testing the final version, and reading code before changing it.
- Ran tests 12 times, steady verification cadence.
- Explored 5 files before/while editing.
- Verified locally before submitting for hidden-test grading.
- The last local test run before submitting was green.
- Explicitly asked the agent to verify 1 time(s).
Token economy92/100weight 15%evidence ▾hide ▴
Tokens against the budget, context bloat, and tokens burned in failing loops.
- Used only 16% of the token budget.
- Kept conversation context lean.
- 10% of tokens (32,900) burned inside failing agent loops.
- 71% of input tokens were served from cache.
- 315,299 tokens total ($0.17), 35,033 per prompt.
Speed85/100weight 10%evidence ▾hide ▴
Time used against the time allowed, credited only for work that landed.
- Finished in 32m 4s of 60m allotted (53%).
- First prompt after 150s, engaged quickly.
Recovery75/100weight 10%evidence ▾hide ▴
How failures were handled: loops, repeated prompts, getting back to green.
- 1 error loop(s): 3+ consecutive failing agent steps without a strategy change.
- 8 failing tool result(s) encountered overall.
- Got back to a green test run after the loop(s).
Outcome
What the hidden tests said
Verdict
✓ Resolved
Graded in a sealed environment. Status: graded.
Hidden failing tests (must now pass)
2 / 2
Existing tests (must not break)
20 / 20
Grader notes and output
- Sample report: the hidden-test outcome is stated for this simulated round, not executed.
Prompt forensics
What the candidate actually asked the agent, and how
Own words
34%
of prompt text was written, not pasted
Paste prompts
3 of 9
mostly errors, code or issue text
Longest paste
924 ch
prompt #3
Corrections
1
times they steered the agent
Re-asks
0
same request sent again
Median length
44 words
longest 104 words
Prompt kinds
Each prompt classified by what it was, from its text alone.
- Clear instruction1
- Steering correction1
- Verification request1
- Question1
- Problem statement paste1
- Raw error paste1
- Code dump1
- Low-effort nudge2
good habit worth a question
Where the prompt text came from
Characters across all prompts, line by line.
- Own words34%
- Pasted errors33%
- Pasted code11%
- Copied issue text22%
Prompt length (words)
Every prompt, in order
- #11:34Problem statement pastearrived by pasteQ18104w33k tok$0.015
django.utils.http.parse_http_date two digit year check is incorrect Description (last modified by Ad Timmering) RFC 850 does not mention this, but in RFC 7231 (and there's something similar in RFC 2822), there's the following quote: Recipients of a timestamp value in rfc850-date format, which uses a two-digit year, MUST interpret a timestamp that appears to be more than 50 years in the future as representing the most recent year in the past that had the same last two digits. Current logic is hard coded to consider 0-69 to be in 2000-2069, and 70-99 to be 1970-1999, instead of comparing versus the current year. fix this
django.utils.http.parse_http_date two digit year check is incorrect Description (last modified by Ad Timmering) RFC 850 does not mention this, but in RFC 7231 (and there's something similar in RFC 2822), there's the following quote: Recipients of a timestamp value in rfc850-date format, which uses a two-digit year, MUST interpret a timestamp that appears to be more than 50 years in the future as representing the most recent year in the past that had the same last two digits. Current logic is hard coded to consider 0-69 to be in 2000-2069, and 70-99 to be 1970-1999, instead of comparing versus the current year. fix this
Own words 1% · Copied issue text 99% · 98% overlap with the issue text
What followed: 4 agent steps, 4 tool calls, 1 edit (django/utils/http.py).
scored on own words onlyvery shortclear action verbproblem statement copied rather than decomposed - #24:42Steering correctionQ6244w45k tok$0.025
No, that just moves the hardcoded cutoff from 70 to 50. The RFC says to compare against the current year: a two-digit year that lands more than 50 years in the future belongs to the previous century. Use datetime.datetime.utcnow().year and don't hardcode any year.
No, that just moves the hardcoded cutoff from 70 to 50. The RFC says to compare against the current year: a two-digit year that lands more than 50 years in the future belongs to the previous century. Use datetime.datetime.utcnow().year and don't hardcode any year.
Own words 100% · 11% overlap with the issue text
What followed: 6 agent steps, 6 tool calls, 1 edit (django/utils/http.py), 5 test runs, 5 failing steps. Sent 1m 27s after the previous agent response.
detailed (30+ words)states constraints/acceptance criteriaquantified detailcorrects the agent's course - #39:44Raw error pastearrived by pasteQ583w15k tok$0.0083
Traceback (most recent call last): File "/workspace/repo/tests/utils_tests/test_http.py", line 7, in <module> from django.test import SimpleTestCase, ignore_warnings File "/workspace/repo/django/test/__init__.py", line 3, in <module> from django.test.client import Client, RequestFactory File "/workspace/repo/django/test/client.py", line 14, in <module> from django.core.handlers.base import BaseHandler File "/workspace/repo/django/core/handlers/base.py", line 8, in <module> from django.urls import get_resolver, set_urlconf File "/workspace/repo/django/conf/__init__.py", line 61, in _setup raise ImproperlyConfigured( django.core.exceptions.ImproperlyConfigured: Requested setting DEBUG, but settings are not configured. You must either define the environment variable DJANGO_SETTINGS_MODULE or call settings.configure() before accessing settings. ERROR: tests/utils_tests/test_http.py - django.core.exceptions.ImproperlyConfigured
Traceback (most recent call last): File "/workspace/repo/tests/utils_tests/test_http.py", line 7, in <module> from django.test import SimpleTestCase, ignore_warnings File "/workspace/repo/django/test/__init__.py", line 3, in <module> from django.test.client import Client, RequestFactory File "/workspace/repo/django/test/client.py", line 14, in <module> from django.core.handlers.base import BaseHandler File "/workspace/repo/django/core/handlers/base.py", line 8, in <module> from django.urls import get_resolver, set_urlconf File "/workspace/repo/django/conf/__init__.py", line 61, in _setup raise ImproperlyConfigured( django.core.exceptions.ImproperlyConfigured: Requested setting DEBUG, but settings are not configured. You must either define the environment variable DJANGO_SETTINGS_MODULE or call settings.configure() before accessing settings. ERROR: tests/utils_tests/test_http.py - django.core.exceptions.ImproperlyConfiguredPasted errors 100%
What followed: 1 agent step, 1 tool call, 1 test run, 1 failing step; tests ended with 1 failing. Sent 2m 10s after the previous agent response.
scored on own words onlyvery shortraw error output pasted with little instruction - #411:44Low-effort nudgeQ204w23k tok$0.013
still failing, fix it
still failing, fix it
Own words 100%
What followed: 2 agent steps, 2 tool calls, 1 edit (django/utils/http.py), 1 test run, 1 failing step; tests ended with 1 failing. Sent 1m 32s after the previous agent response.
includes error contextclear action verbvague phrasinglow-effort nudge - #514:09Low-effort nudgeQ32w38k tok$0.021
fix it
fix it
Own words 100%
What followed: 3 agent steps, 3 tool calls, 1 edit (django/utils/http.py), 1 test run, 1 failing step; tests ended with 1 failing. Sent 1m 39s after the previous agent response.
very shortclear action verbvague phrasinglow-effort nudge - #617:12Code dumparrived by pasteQ749w31k tok$0.016
use this: ```python year = int(m.group('year')) if year < 100: current_year = datetime.datetime.utcnow().year current_century = current_year - (current_year % 100) if year - (current_year % 100) > 50: # more than 50 years in the future: the previous century year += current_century - 100 else: year += current_century ```
use this: ```python year = int(m.group('year')) if year < 100: current_year = datetime.datetime.utcnow().year current_century = current_year - (current_year % 100) if year - (current_year % 100) > 50: # more than 50 years in the future: the previous century year += current_century - 100 else: year += current_century ```Own words 3% · Pasted code 97% · 11% overlap with the issue text
What followed: 2 agent steps, 2 tool calls, 1 edit (django/utils/http.py), 1 test run; tests ended green. Sent 1m 50s after the previous agent response.
scored on own words onlyvery shortlarge code paste with little instruction - #720:40Clear instructionQ8549w48k tok$0.026
Now add a regression test in tests/utils_tests/test_http.py: mock django.utils.http.datetime.datetime so utcnow() returns 2019, 2020 and 2048, and check the boundaries, e.g. '31-Dec-69' parses to 2069 in 2019, '31-Dec-71' to 1971 in 2020, and '31-Dec-99' to 1999 in 2048. Keep the existing tests unchanged and use subTest for each case.
Now add a regression test in tests/utils_tests/test_http.py: mock django.utils.http.datetime.datetime so utcnow() returns 2019, 2020 and 2048, and check the boundaries, e.g. '31-Dec-69' parses to 2069 in 2019, '31-Dec-71' to 1971 in 2020, and '31-Dec-99' to 1999 in 2048. Keep the existing tests unchanged and use subTest for each case.
Own words 100%
What followed: 3 agent steps, 3 tool calls, 1 edit (tests/utils_tests/test_http.py), 1 test run; tests ended green. Sent 2m 42s after the previous agent response.
detailed (30+ words)references specific file(s)uses code identifiersclear action verbstates constraints/acceptance criteriaquantified detail - #824:06QuestionQ5226w39k tok$0.021
Does anything else in django call parse_http_date or rely on the old 0-69 -> 2000s rule, e.g. parse_http_date_safe, the conditional GET middleware or the cache helpers?
Does anything else in django call parse_http_date or rely on the old 0-69 -> 2000s rule, e.g. parse_http_date_safe, the conditional GET middleware or the cache helpers?
Own words 100%
What followed: 2 agent steps, 2 tool calls. Sent 1m 43s after the previous agent response.
reasonable lengthuses code identifiersquantified detail - #927:26Verification requestQ7325w43k tok$0.022
Run utils_tests and the middleware tests one more time through tests/runtests.py and show me the final git diff so I can review it before submitting.
Run utils_tests and the middleware tests one more time through tests/runtests.py and show me the final git diff so I can review it before submitting.
Own words 100%
What followed: 2 agent steps, 2 tool calls, 1 test run; tests ended green. Sent 2m 35s after the previous agent response.
reasonable lengthreferences specific file(s)uses code identifiersclear action verbasks for verification
Prompt coach
How a strong engineer would have asked
5 prompts could be stronger. Asked the strong way, about 40k tokens ($0.021) were avoidable.Show rewrites ▾Hide ▴
What was sent
django.utils.http.parse_http_date two digit year check is incorrect Description (last modified by Ad Timmering) RFC 850 does not mention this, but in RFC 7231 (and there's something similar in RFC 2822), there's the following quote: Recipients of a timestamp value in rfc850-date format, which uses a two-digit year, MUST interpret a timestamp that appears to be more than 50 years in the future as representing the most recent year in the past that had the same last two digits. Current logic is hard coded to consider 0-69 to be in 2000-2069, and 70-99 to be 1970-1999, instead of comparing versus the current year. fix this
How a strong engineer would ask
Task: django.utils.http.parse_http_date two digit year check is incorrect.
In my words: ⟨restate this in one sentence: "Description (last modified by Ad Timmering) RFC 850 does not mention this, but in RFC 7231 (and there's something similar in RFC 2822), the…"⟩
1. Read
django/utils/http.py(django.utils.http.parse_http_date) and tell me where the current logic lives and which existing tests cover it. Don't edit yet.2. Propose the change in two or three lines, including how you'd handle the boundary cases.
3. After I agree: implement it, add a regression test, and run the related tests.
Constraints: ⟨what must not change, e.g. public API, existing tests⟩.
Why it works. The issue text gives the agent the problem but not your plan, so it explores widely and commits to its first guess. Breaking it into read, propose, implement keeps each turn small and lets you reject a wrong approach before it is written, instead of steering after it is in the code.
+ the requirement in your words+ a read-first step+ a proposal before any edit+ constraintsThis turn used 33k tokens; the round's median turn used 38k. Trimming the paste would also have kept ~5,024 tokens of context off the 33 requests that followed.
What was sent
Traceback (most recent call last): File "/workspace/repo/tests/utils_tests/test_http.py", line 7, in <module> from django.test import SimpleTestCase, ignore_warnings File "/workspace/repo/django/test/__init__.py", line 3, in <module> from django.test.client import Client, RequestFactory File "/workspace/repo/django/test/client.py", line 14, in <module> from django.core.handlers.base import BaseHandler File "/workspace/repo/django/core/handlers/base.py", line 8, in <module> from django.urls import get_resolver, set_urlconf File "/workspace/repo/django/conf/__init__.py", line 61, in _setup raise ImproperlyConfigured( django.core.exceptions.ImproperlyConfigured: Requested setting DEBUG, but settings are not configured. You must either define the environment variable DJANGO_SETTINGS_MODULE or call settings.configure() before accessing settings. ERROR: tests/utils_tests/test_http.py - django.core.exceptions.ImproperlyConfiguredHow a strong engineer would ask
I ran
tests/runtests.py utils_tests.test_httpafter your change todjango/utils/http.pyand expectedtests/utils_tests/test_http.pyto pass. It fails withImproperlyConfigured, trimmed to the lines that matter:File "tests/utils_tests/test_http.py", line 7, in <module> File "django/core/handlers/base.py", line 8, in <module> from django.urls import get_resolver, set_urlconf File "django/conf/__init__.py", line 61, in _setup raise ImproperlyConfigured( django.core.exceptions.ImproperlyConfigured: Requested setting DEBUG, but settings are not configured. You must either define the environment variable DJANGO_S…
It is raised in
_setup(django/conf/__init__.py:61). My hypothesis: the tests ran without the project's settings, so the command is wrong rather than the code.Confirm or rule that out first, then make the smallest change that fixes it and re-run only
tests/utils_tests/test_http.py.Why it works. A bare traceback tells the agent something broke but not what you were doing or what you suspect, so it guesses: it re-reads files and retries the last fix. This one (ImproperlyConfigured) fails before any test runs, which points at the test command or setup, not the code; saying so stops the agent from "fixing" code that was never the problem. And the full paste stays in the conversation, so every later request pays for it again.
+ the command you ran+ what you expected+ 6 relevant lines, not the whole dump+ a hypothesis+ a narrow next stepThis turn used 15k tokens; the round's median turn used 38k. Trimming the paste would also have kept ~2,725 tokens of context off the 21 requests that followed.
What was sent
still failing, fix it
How a strong engineer would ask
The last run still fails: 51 passed, 1 failed, after your change to
django/utils/http.py.Before editing again, quote the failing assertion or error and explain in two sentences why it still fails.
If the cause is outside
django/utils/http.py(setup, imports, the test command), say so.Then propose one different approach and wait for my OK.
Why it works. "still failing, fix it" adds no new information, so the agent usually retries a variant of the same idea. Here the next run ended exactly as before (51 passed, 1 failed). Asking for a diagnosis first is the cheapest way out of a loop: one short answer instead of another round of edits and test runs.
+ the evidence (counts)+ a diagnosis step+ a ban on blind retries+ a checkpointThis turn used 23k tokens and left the tests no better than before, so nearly all of it bought nothing; the cheapest turn in this round cost 15k, about 8,619 less.
What was sent
fix it
How a strong engineer would ask
The last run still fails: 51 passed, 1 failed, after your change to
django/utils/http.py.Before editing again, quote the failing assertion or error and explain in two sentences why it still fails.
If the cause is outside
django/utils/http.py(setup, imports, the test command), say so.Then propose one different approach and wait for my OK.
Why it works. "fix it" adds no new information, so the agent usually retries a variant of the same idea. Here the next run ended exactly as before (51 passed, 1 failed). Asking for a diagnosis first is the cheapest way out of a loop: one short answer instead of another round of edits and test runs.
+ the evidence (counts)+ a diagnosis step+ a ban on blind retries+ a checkpointThis turn used 38k tokens and left the tests no better than before, so nearly all of it bought nothing; the cheapest turn in this round cost 15k, about 23k less.
What was sent
use this: ```python year = int(m.group('year')) if year < 100: current_year = datetime.datetime.utcnow().year current_century = current_year - (current_year % 100) if year - (current_year % 100) > 50: # more than 50 years in the future: the previous century year += current_century - 100 else: year += current_century ```How a strong engineer would ask
In
django/utils/http.py, I want this behavior: ⟨one sentence: what this code must achieve, and the edge case it handles⟩Here is the approach I have in mind (9 lines). Check it against the existing tests and the boundary cases before applying it, and tell me if any of it is wrong.
year = int(m.group('year')) if year < 100: current_year = datetime.datetime.utcnow().year current_century = current_year - (current_year % 100) if year - (current_year % 100) > 50: # more than 50 years in the future: the previous century year += current_century - 100 else: year += current_centuryKeep everything else in the file unchanged. Done means the related tests pass and it still meets what I asked in #2; show me the diff.
Why it works. Pasting code with "use this" turns the agent into a typist: it applies the snippet without knowing what it must achieve, so it cannot catch a mistake in it or adapt it to the surrounding code. Stating the intent and asking for a check gets the review you skipped, and keeps the decision (and the understanding) yours, which is what an interviewer will ask about.
+ the file and function+ the intent in your words+ a request to check it+ what done meansThis turn used 31k tokens; the round's median turn used 38k.
Rewrites are built by rules from each prompt's own text and what happened around it; the ⟨bracketed⟩ parts are what only the engineer could fill in. Savings are estimates: a turn's tokens above this round's median turn (38k), or above its cheapest turn when the turn left the tests failing as before; plus, for pastes, the context the trimmed text would no longer carry on each later request (about four characters per token).
Tokens and cost
Where the budget went
Total tokens
315,299
16% of the budget
Cost (metered)
$0.168
35k tokens per prompt
Wasted in loops
33k
10% of tokens, inside 3+ failing steps
Peak context
51k
avg 31k per request
Cache hits
71%
of input tokens served from cache
Agent autonomy
2.8 steps
per prompt, max 6 on one
Tool calls
25
2.8 per prompt
Output share
0.7%
output tokens vs input
How the time was spent
Explore, implement, verify
- Orient1m 34s
- Explore2m 27s
- Implement3m 16s
- Verify7m 26s
- Read & prompt16m 25s
- ✓/✗ test run green / failing
First prompt
2m 30s
First edit
2m 38s
First test
5m 16s
First green
17m 58s
Edit → test cycles
5
Final version tested
Yes
last run green
Longest idle gap
2m 42s
from 17:58
Integrity and attention flags
Things worth a second look, with evidence
- mediumlong focus loss
1 focus-loss window(s) over 60s (longest 95s)
at 7:47, 95s
- mediumlarge paste
3 paste(s) over 500 chars, check provenance in replay
at 1:30, 682 charsat 9:40, 1189 charsat 17:08, 512 chars
- lowpaste after focus loss
Large paste within 30s of returning from a focus loss; the text (prompt #3) is error output, consistent with copying from the session's own terminal
at 9:40, 1189 chars, prompt #3
Agent tools and files
Tool calls
- bash12
- edit6
- read5
- grep2
Files edited
- django/utils/http.py×5
- tests/utils_tests/test_http.py×1
Files explored
- django/utils/http.py×2
- tests/utils_tests/test_http.py×2
- def parse_http_date×1
- parse_http_date×1
- django/utils/cache.py×1
Coaching for the candidate
Specific to this round
Never paste an error on its own
Add one line on what you think the error means and what you want tried next (“this is a settings problem, run it through runtests.py instead”). It shows your reasoning and stops the agent guessing.
Frame the task before delegating it
Instead of pasting the issue, restate it in two or three sentences: the function at fault, the behavior you expect, and how you will know it is fixed. Agents follow a sharp brief far better than a copied ticket.
Interrupt failing loops early
Step in after two failed attempts: ask the agent to explain the failure before it retries. 32,900 tokens went into loops this round.
Replace “fix it” with new information
When something fails, say what changed, what you expected, and which line or test is wrong. Re-sending the same request rarely gets a different answer.
Describe intent, then let the agent write
If you already know the code, say what it must do and where; reviewing the agent's version is a better use of your time than pasting your own.
What the candidate sees · paid plan
How to improve: your mistakes and what to do next
Candidates on a paid plan get this on their own copy: every mistake from the round with what it cost and what to do instead, and goals the next round checks. The hiring team's copy has interview follow-ups instead.
- 1
You sent “fix it” nudges instead of a diagnosis
Steering the AI −12Debugging −16- What you did
- 2 low-information nudges like “still failing, fix it” (#4, #5).(at 11:44 · prompt #4)
- What to do instead
- When a run fails, say what changed, what you expected, and which test or line is wrong, or ask the agent why it still fails before it edits again. Re-sending the same request rarely gets a different answer.
What it cost: Steering the AI −12, Debugging −16; 3 hireability points. Fixing this alone: hireability 82 → 85.
Do next time: Ask “why does it still fail?” and add new information every turn.
Don't: Type “fix it”, “try again” or “still failing”.
- 2
You pasted the issue instead of framing it
Problem understanding −20- What you did
- Opened by pasting the issue text into the agent instead of framing it.(at 1:34 · prompt #1)
- What to do instead
- Restate the task in two or three sentences: the function at fault, the behavior you expect, and how you will know it is fixed. Then ask the agent to read before it edits.
What it cost: Problem understanding −20; 2 hireability points. Fixing this alone: hireability 82 → 84.
Do next time: Brief the agent in your own words and ask for a plan first.
Don't: Paste the ticket and hope.
- 3
You changed code before reproducing the bug
Problem understanding −15- What you did
- Did not reproduce the failure before changing code.
- What to do instead
- Run the failing test before the first edit, so you know the failure you are fixing and can tell when it is gone.
What it cost: Problem understanding −15; 1 hireability point. Fixing this alone: hireability 82 → 83.
Do next time: Reproduce first, then edit.
Don't: Edit before you have seen it fail.
- 4
You let the agent loop on failures
Debugging −15- What you did
- The agent failed 5+ times in a row and was left to retry (1 loop).(at 5:45)
- What to do instead
- Step in after two failed attempts in a row: stop the agent, read the failure, and change the plan.
What it cost: Debugging −15; 1 hireability point. Fixing this alone: hireability 82 → 83.
Do next time: Interrupt after two failures in a row.
Don't: Watch the agent fail three times running.
- 5
You pasted errors with no instruction
Steering the AI −5Debugging −5- What you did
- Pasted raw error output with no instruction (#3).(at 9:44 · prompt #3)
- What to do instead
- Add one line on what you think the error means and what to try next. It shows your reasoning and stops the agent guessing.
What it cost: Steering the AI −5, Debugging −5; 1 hireability point. Fixing this alone: hireability 82 → 83.
Do next time: Pair every pasted error with your hypothesis.
Don't: Paste a traceback on its own.
- 6
You pasted code instead of describing intent
Steering the AI −4- What you did
- Pasted code instead of describing the change (#6).(at 17:12 · prompt #6)
- What to do instead
- Say what the code must do and where; reviewing the agent's version is a better use of your time than pasting your own.
What it cost: Steering the AI −4.
Do next time: Describe the behavior, then review the diff.
Don't: Paste code blocks into prompts.
Next round goals
Send no “fix it” nudges or re-asks
This round: 2 nudges
Missed: your target next time
Keep pasted prompts under 20%
This round: 33% pasted
Missed: your target next time
Reproduce the failure before your first edit
This round: edited first
Missed: your target next time
Measured from the recording, so your next report checks each one automatically.
Full timeline
Show all 49 eventsHide events
- 0:00Session created for django__django-11848
- 0:00Environment accepted connections
- 0:00Workspace ready, waiting on the candidate to start
- 0:00Candidate pressed Start
- 0:00Round started, clock begins
- 0:14Opened the problem statement
- 1:30Paste event (682 chars)
- 1:34Prompt #1 [Problem statement paste] (104 words, quality 18): django.utils.http.parse_http_date two digit year check is incorrect Description (last modified by Ad Timmering) RFC 850 does not men…
- 1:55Agent step, 1 tool call(s), 8968 tokens
- 2:13Agent step, 1 tool call(s), 4260 tokens
- 2:38Agent step, 1 tool call(s), 5754 tokens
- 4:04Still in the editor
- 4:42Prompt #2 [Steering correction] (44 words, quality 62): No, that just moves the hardcoded cutoff from 70 to 50. The RFC says to compare against the current year: a two-digit year that lands more t…
- 5:16Agent step, 1 tool call(s), 5764 tokens
- 5:45Agent step, 1 tool call(s), 6079 tokens (tool reported a failure)
- 6:12Agent step, 1 tool call(s), 6393 tokens (tool reported a failure)
- 6:39Agent step, 1 tool call(s), 6627 tokens (tool reported a failure)
- 7:06Agent step, 1 tool call(s), 6766 tokens (tool reported a failure)
- 7:47Window lost focus
- 9:22Window regained focus
- 9:40Paste event (1189 chars)
- 9:44Prompt #3 [Raw error paste] (83 words, quality 5): Traceback (most recent call last): File "/workspace/repo/tests/utils_tests/test_http.py", line 7, in <module> from django.test import …
- 10:12Agent test run: 51 passed, 1 failed
- 11:44Prompt #4 [Low-effort nudge] (4 words, quality 20): still failing, fix it
- 12:05Agent step, 1 tool call(s), 7654 tokens
- 12:30Agent test run: 51 passed, 1 failed
- 14:04Still in the editor
- 14:09Prompt #5 [Low-effort nudge] (2 words, quality 3): fix it
- 14:25Agent step, 1 tool call(s), 9956 tokens
- 14:56Agent step, 1 tool call(s), 9832 tokens
- 15:22Agent test run: 51 passed, 1 failed
- 17:08Paste event (512 chars)
- 17:12Prompt #6 [Code dump] (49 words, quality 7): use this: ```python year = int(m.group('year')) if year < 100: current_year = datetime.datetime.utcnow().year …
- 17:32Agent step, 1 tool call(s), 10347 tokens
- 17:58Agent test run: 52 passed, 0 failed
- 19:04Still in the editor
- 20:40Prompt #7 [Clear instruction] (49 words, quality 85): Now add a regression test in tests/utils_tests/test_http.py: mock django.utils.http.datetime.datetime so utcnow() returns 2019, 2020 and 204…
- 21:02Agent step, 1 tool call(s), 12539 tokens
- 21:55Agent step, 1 tool call(s), 12070 tokens
- 22:23Agent test run: 53 passed, 0 failed
- 24:06Prompt #8 [Question] (26 words, quality 52): Does anything else in django call parse_http_date or rely on the old 0-69 -> 2000s rule, e.g. parse_http_date_safe, the conditional GET midd…
- 24:23Agent step, 1 tool call(s), 12477 tokens
- 27:26Prompt #9 [Verification request] (25 words, quality 73): Run utils_tests and the middleware tests one more time through tests/runtests.py and show me the final git diff so I can review it before su…
- 28:05Agent step, 1 tool call(s), 14152 tokens
- 28:05Agent test run: 412 passed, 0 failed
- 29:14Ran tests from the session bar (tests/utils_tests/test_http.py)
- 29:42Test result: 53 passed, 0 failed
- 30:36Copied 64 characters out of the editor
- 31:08Candidate submitted their solution
Every number on this page is computed from the round's recording (LLM proxy log, editor telemetry and the hidden-test run). The written verdict is generated by rules from those numbers, not by a language model. Scores support human judgment and are not meant to be the sole basis for a hiring decision. How scoring works.