AI planner runs

16 runs over 11 cases · 69 turns · 1,054,805 tokens · ≈ 0.177 USD · 12 of 13 scored runs have a failing check

Scorecard · 13 runs

ClaimPassFail?
A turn delivers more often than it asks 1 11 1
A capability is called by turn 2 4 9
No three consecutive turns ask for a fact 6 6 1
The interview detector sees every turn 8 4 1
No word from inside the machine reaches the couple 11 2
A year-only plan says it is dated from January 2 2
A fact the couple declined is never asked again 5 1 7
A year-only plan is not called late without saying why 3 1
Every figure carries a source it can support 13
An interpreted fact reaches the loop and not the account 1 12
A readiness offer is made at most twice, then goes quiet 4 9
A checklist built from a date is reported with its real task count 6 7
A turn that writes but unlocks nothing uses the one-line form 9 4
No internal stage name is ever said to the couple 13
A turn where the couple states nothing produces no receipt 13
The receipt reports a change, not a state 13
A date replacement rebuilds exactly once 13
The offer log records every turn 13
Supplier search never runs without a real search area 13
The run completed without a stream or write error 13
A year-only date gets no months-out figure 4

Every claim, counted across every scored run. ? means the run could not answer it, which is not a pass — a claim answered by nothing is untested, not met. Some passes are vacuous: a search that never ran cannot run without a search area. 3 runs are left out, marked partial capture below: they were recorded before the harness read the account and the client writes, so several claims are unanswerable in them.

Runs, newest first

cost-questions 4 turns ≈ 0.0099 USD 1 fail 5 unanswered

Three direct cost questions. Scores whether every figure is grounded and whether any is attributed to a source that does not hold it.

2026-08-18 13:45 · commit c5a551c75d8 · 100,707 tokens · 1 receipt · tools: getBudgetBreakdown, getCoachReference
1 failing check
  • A turn delivers more often than it asks
    2 of 3 measured turns asked, out of 4 spoken
supplier-search 3 turns ≈ 0.0061 USD 3 fail 5 unanswered

A real place picked from the location field, then a venue request. Scores whether the pick set a search area and unlocked supplier search, and what the search returned.

2026-08-18 13:44 · commit c5a551c75d8 · 35,037 tokens · 1 receipt · tools: none called
3 failing checks
  • No word from inside the machine reaches the couple
    said "placeholder"
  • A capability is called by turn 2
    no tool was called in the whole run
  • A turn delivers more often than it asks
    2 of 2 measured turns asked, out of 3 spoken
stonewaller 6 turns ≈ 0.017 USD 3 fail 4 unanswered

Five turns that state nothing at all. Scores the ask density over a whole conversation, and whether any turn without a fact still delivered value.

2026-08-18 13:43 · commit c5a551c75d8 · 73,872 tokens · 0 receipts · tools: none called
3 failing checks
  • A capability is called by turn 2
    no tool was called in the whole run
  • A turn delivers more often than it asks
    3 of 5 measured turns asked, out of 6 spoken
  • A fact the couple declined is never asked again
    declined: date, location, budget, guests — re-asked: budget
returning-couple 5 turns ≈ 0.0090 USD 3 fail 3 unanswered

A first meeting, then a return on another day with the transcript gone. Scores whether the second opener knows them, and whether it re-asks what it was told.

2026-08-18 13:42 · commit c5a551c75d8 · 66,216 tokens · 1 receipt · tools: getCoachReference, getUpcomingTaskAdvice
3 failing checks
  • No word from inside the machine reaches the couple
    said "placeholder"
  • A turn delivers more often than it asks
    2 of 2 measured turns asked, out of 4 spoken
  • The interview detector sees every turn
    2 ask records for 4 spoken turns — the extractor is skipped on openers, chips and continuations, and it is what writes this log
just-show-me 4 turns ≈ 0.0068 USD 3 fail 5 unanswered

A demand for suppliers before anything is known, repeated. Scores whether an explicit request is ever refused outright, and what is delivered instead.

2026-08-18 13:41 · commit c5a551c75d8 · 45,885 tokens · 1 receipt · tools: none called
3 failing checks
  • A capability is called by turn 2
    no tool was called in the whole run
  • No three consecutive turns ask for a fact
    asks: date → location → location
  • A turn delivers more often than it asks
    3 of 3 measured turns asked, out of 4 spoken
refusal-reload 5 turns ≈ 0.0085 USD 3 fail 5 unanswered

A budget refusal, then the tab closes and reopens. Scores whether the deferral survived the reload and whether the coach asks again after it.

2026-08-18 13:40 · commit c5a551c75d8 · 46,815 tokens · 0 receipts · tools: none called
3 failing checks
  • A capability is called by turn 2
    no tool was called in the whole run
  • No three consecutive turns ask for a fact
    asks: date → date → date
  • A turn delivers more often than it asks
    3 of 3 measured turns asked, out of 4 spoken
generic-location 4 turns ≈ 0.010 USD 4 fail 4 unanswered

A location too generic to search on. Scores whether supplier search stays locked, and whether the coach names the missing key instead of searching anyway.

2026-08-18 13:39 · commit c5a551c75d8 · 60,615 tokens · 2 receipts · tools: getCoachReference
4 failing checks
  • A capability is called by turn 2
    first at turn 3 (getCoachReference)
  • No three consecutive turns ask for a fact
    asks: date → date → date
  • A turn delivers more often than it asks
    3 of 3 measured turns asked, out of 4 spoken
  • A year-only plan says it is dated from January
    January never mentioned
all-at-once 4 turns ≈ 0.012 USD all checks pass 4 unanswered

Every core fact in two messages. Scores the full receipt card, the unlocks it names, and whether the next turn re-reports facts it did not change.

2026-08-18 13:38 · commit c5a551c75d8 · 85,808 tokens · 2 receipts · tools: executeClientActions, getCoachReference, getUpcomingTaskAdvice
vision-first 6 turns ≈ 0.018 USD 6 fail 2 unanswered

Vision, then a guest count, then two budget refusals, then a year. Scores what a rapport fact writes, and whether a refusal is honoured.

2026-08-18 13:37 · commit c5a551c75d8 · 102,743 tokens · 2 receipts · tools: getCoachReference, getUpcomingTaskAdvice
6 failing checks
  • A capability is called by turn 2
    first at turn 5 (getCoachReference)
  • No three consecutive turns ask for a fact
    asks: date → date → date → location
  • A turn delivers more often than it asks
    4 of 4 measured turns asked, out of 6 spoken
  • The interview detector sees every turn
    4 ask records for 6 spoken turns — the extractor is skipped on openers, chips and continuations, and it is what writes this log
  • A year-only plan says it is dated from January
    January never mentioned
  • A year-only plan is not called late without saying why
    said "overdue", with no reason
year-only 5 turns ≈ 0.017 USD 3 fail 2 unanswered

Two budget refusals, then the bare year 2027. Scores the year-only plan, the receipt for a checklist build, and whether a declined fact is asked again.

2026-08-18 13:36 · commit c5a551c75d8 · 103,522 tokens · 1 receipt · tools: getCoachReference, executeClientActions, getUpcomingTaskAdvice
3 failing checks
  • No three consecutive turns ask for a fact
    asks: date → budget → priorities
  • A turn delivers more often than it asks
    3 of 3 measured turns asked, out of 5 spoken
  • The interview detector sees every turn
    3 ask records for 5 spoken turns — the extractor is skipped on openers, chips and continuations, and it is what writes this log
opener 1 turn ≈ 0.0003 USD 1 fail 9 unanswered

The opening move alone. Scores the tier-0 insight, the provenance clause, and whether the single question is an offer’s price rather than an interview opener.

2026-08-18 13:35 · commit c5a551c75d8 · 9,235 tokens · 0 receipts · tools: none called
1 failing check
  • A capability is called by turn 2
    no tool was called in the whole run
refusal-reload 5 turns ≈ 0.011 USD 3 fail 5 unanswered

A budget refusal, then the tab closes and reopens. Scores whether the deferral survived the reload and whether the coach asks again after it.

2026-08-18 13:34 · commit c5a551c75d8 · 45,238 tokens · 0 receipts · tools: none called
3 failing checks
  • A capability is called by turn 2
    no tool was called in the whole run
  • No three consecutive turns ask for a fact
    asks: budget → date → date
  • A turn delivers more often than it asks
    3 of 3 measured turns asked, out of 4 spoken
year-only 5 turns ≈ 0.015 USD 3 fail 2 unanswered

Two budget refusals, then the bare year 2027. Scores the year-only plan, the receipt for a checklist build, and whether a declined fact is asked again.

2026-08-18 13:04 · commit c5a551c75d8 · 82,318 tokens · 1 receipt · tools: executeClientActions, getUpcomingTaskAdvice
3 failing checks
  • A capability is called by turn 2
    first at turn 4 (executeClientActions)
  • A turn delivers more often than it asks
    2 of 3 measured turns asked, out of 5 spoken
  • The interview detector sees every turn
    3 ask records for 5 spoken turns — the extractor is skipped on openers, chips and continuations, and it is what writes this log
vision-first 6 turns ≈ 0.018 USD partial capture 4 fail 2 unanswered

Vision, then a guest count, then two budget refusals, then a year. Scores what a rapport fact writes, and whether a refusal is honoured.

2026-08-18 12:49 · commit c5a551c75d8 · 97,243 tokens · 2 receipts · tools: executeClientActions, getUpcomingTaskAdvice
4 failing checks
  • A capability is called by turn 2
    first at turn 5 (executeClientActions)
  • No three consecutive turns ask for a fact
    asks: date → location → date → location
  • A turn delivers more often than it asks
    4 of 4 measured turns asked, out of 6 spoken
  • The interview detector sees every turn
    4 ask records for 6 spoken turns — the extractor is skipped on openers, chips and continuations, and it is what writes this log
year-only 5 turns ≈ 0.016 USD partial capture 3 fail 3 unanswered

Two budget refusals, then the bare year 2027. Scores the year-only plan, the receipt for a checklist build, and whether a declined fact is asked again.

2026-08-18 12:20 · commit c5a551c75d8 · 90,315 tokens · 1 receipt · tools: getCoachReference, getUpcomingTaskAdvice
3 failing checks
  • No three consecutive turns ask for a fact
    asks: date → budget → date
  • A turn delivers more often than it asks
    3 of 3 measured turns asked, out of 5 spoken
  • The interview detector sees every turn
    3 ask records for 5 spoken turns — the extractor is skipped on openers, chips and continuations, and it is what writes this log
opener 1 turn ≈ 0.0019 USD partial capture 1 fail 9 unanswered

The opening move alone. Cheap enough to sample five times when the first turn changes.

2026-08-18 12:19 · commit c5a551c75d8 · 9,236 tokens · 0 receipts · tools: none called
1 failing check
  • A capability is called by turn 2
    no tool was called in the whole run