Scorecard · 13 runs
| Claim | Pass | Fail | ? | |
|---|---|---|---|---|
| A turn delivers more often than it asks | 1 | 11 | 1 | |
| A capability is called by turn 2 | 4 | 9 | ||
| No three consecutive turns ask for a fact | 6 | 6 | 1 | |
| The interview detector sees every turn | 8 | 4 | 1 | |
| No word from inside the machine reaches the couple | 11 | 2 | ||
| A year-only plan says it is dated from January | 2 | 2 | ||
| A fact the couple declined is never asked again | 5 | 1 | 7 | |
| A year-only plan is not called late without saying why | 3 | 1 | ||
| Every figure carries a source it can support | 13 | |||
| An interpreted fact reaches the loop and not the account | 1 | 12 | ||
| A readiness offer is made at most twice, then goes quiet | 4 | 9 | ||
| A checklist built from a date is reported with its real task count | 6 | 7 | ||
| A turn that writes but unlocks nothing uses the one-line form | 9 | 4 | ||
| No internal stage name is ever said to the couple | 13 | |||
| A turn where the couple states nothing produces no receipt | 13 | |||
| The receipt reports a change, not a state | 13 | |||
| A date replacement rebuilds exactly once | 13 | |||
| The offer log records every turn | 13 | |||
| Supplier search never runs without a real search area | 13 | |||
| The run completed without a stream or write error | 13 | |||
| A year-only date gets no months-out figure | 4 |
Every claim, counted across every scored run. ? means the run could not answer it, which is not a pass — a claim answered by nothing is untested, not met. Some passes are vacuous: a search that never ran cannot run without a search area. 3 runs are left out, marked partial capture below: they were recorded before the harness read the account and the client writes, so several claims are unanswerable in them.
Runs, newest first
Three direct cost questions. Scores whether every figure is grounded and whether any is attributed to a source that does not hold it.
c5a551c75d8 ·
100,707 tokens ·
1 receipt ·
tools: getBudgetBreakdown, getCoachReference
1 failing check
- A turn delivers more often than it asks
2 of 3 measured turns asked, out of 4 spoken
A real place picked from the location field, then a venue request. Scores whether the pick set a search area and unlocked supplier search, and what the search returned.
c5a551c75d8 ·
35,037 tokens ·
1 receipt ·
tools: none called
3 failing checks
- No word from inside the machine reaches the couple
said "placeholder" - A capability is called by turn 2
no tool was called in the whole run - A turn delivers more often than it asks
2 of 2 measured turns asked, out of 3 spoken
Five turns that state nothing at all. Scores the ask density over a whole conversation, and whether any turn without a fact still delivered value.
c5a551c75d8 ·
73,872 tokens ·
0 receipts ·
tools: none called
3 failing checks
- A capability is called by turn 2
no tool was called in the whole run - A turn delivers more often than it asks
3 of 5 measured turns asked, out of 6 spoken - A fact the couple declined is never asked again
declined: date, location, budget, guests — re-asked: budget
A first meeting, then a return on another day with the transcript gone. Scores whether the second opener knows them, and whether it re-asks what it was told.
c5a551c75d8 ·
66,216 tokens ·
1 receipt ·
tools: getCoachReference, getUpcomingTaskAdvice
3 failing checks
- No word from inside the machine reaches the couple
said "placeholder" - A turn delivers more often than it asks
2 of 2 measured turns asked, out of 4 spoken - The interview detector sees every turn
2 ask records for 4 spoken turns — the extractor is skipped on openers, chips and continuations, and it is what writes this log
A demand for suppliers before anything is known, repeated. Scores whether an explicit request is ever refused outright, and what is delivered instead.
c5a551c75d8 ·
45,885 tokens ·
1 receipt ·
tools: none called
3 failing checks
- A capability is called by turn 2
no tool was called in the whole run - No three consecutive turns ask for a fact
asks: date → location → location - A turn delivers more often than it asks
3 of 3 measured turns asked, out of 4 spoken
A budget refusal, then the tab closes and reopens. Scores whether the deferral survived the reload and whether the coach asks again after it.
c5a551c75d8 ·
46,815 tokens ·
0 receipts ·
tools: none called
3 failing checks
- A capability is called by turn 2
no tool was called in the whole run - No three consecutive turns ask for a fact
asks: date → date → date - A turn delivers more often than it asks
3 of 3 measured turns asked, out of 4 spoken
A location too generic to search on. Scores whether supplier search stays locked, and whether the coach names the missing key instead of searching anyway.
c5a551c75d8 ·
60,615 tokens ·
2 receipts ·
tools: getCoachReference
4 failing checks
- A capability is called by turn 2
first at turn 3 (getCoachReference) - No three consecutive turns ask for a fact
asks: date → date → date - A turn delivers more often than it asks
3 of 3 measured turns asked, out of 4 spoken - A year-only plan says it is dated from January
January never mentioned
Every core fact in two messages. Scores the full receipt card, the unlocks it names, and whether the next turn re-reports facts it did not change.
c5a551c75d8 ·
85,808 tokens ·
2 receipts ·
tools: executeClientActions, getCoachReference, getUpcomingTaskAdvice
Vision, then a guest count, then two budget refusals, then a year. Scores what a rapport fact writes, and whether a refusal is honoured.
c5a551c75d8 ·
102,743 tokens ·
2 receipts ·
tools: getCoachReference, getUpcomingTaskAdvice
6 failing checks
- A capability is called by turn 2
first at turn 5 (getCoachReference) - No three consecutive turns ask for a fact
asks: date → date → date → location - A turn delivers more often than it asks
4 of 4 measured turns asked, out of 6 spoken - The interview detector sees every turn
4 ask records for 6 spoken turns — the extractor is skipped on openers, chips and continuations, and it is what writes this log - A year-only plan says it is dated from January
January never mentioned - A year-only plan is not called late without saying why
said "overdue", with no reason
Two budget refusals, then the bare year 2027. Scores the year-only plan, the receipt for a checklist build, and whether a declined fact is asked again.
c5a551c75d8 ·
103,522 tokens ·
1 receipt ·
tools: getCoachReference, executeClientActions, getUpcomingTaskAdvice
3 failing checks
- No three consecutive turns ask for a fact
asks: date → budget → priorities - A turn delivers more often than it asks
3 of 3 measured turns asked, out of 5 spoken - The interview detector sees every turn
3 ask records for 5 spoken turns — the extractor is skipped on openers, chips and continuations, and it is what writes this log
The opening move alone. Scores the tier-0 insight, the provenance clause, and whether the single question is an offer’s price rather than an interview opener.
c5a551c75d8 ·
9,235 tokens ·
0 receipts ·
tools: none called
1 failing check
- A capability is called by turn 2
no tool was called in the whole run
A budget refusal, then the tab closes and reopens. Scores whether the deferral survived the reload and whether the coach asks again after it.
c5a551c75d8 ·
45,238 tokens ·
0 receipts ·
tools: none called
3 failing checks
- A capability is called by turn 2
no tool was called in the whole run - No three consecutive turns ask for a fact
asks: budget → date → date - A turn delivers more often than it asks
3 of 3 measured turns asked, out of 4 spoken
Two budget refusals, then the bare year 2027. Scores the year-only plan, the receipt for a checklist build, and whether a declined fact is asked again.
c5a551c75d8 ·
82,318 tokens ·
1 receipt ·
tools: executeClientActions, getUpcomingTaskAdvice
3 failing checks
- A capability is called by turn 2
first at turn 4 (executeClientActions) - A turn delivers more often than it asks
2 of 3 measured turns asked, out of 5 spoken - The interview detector sees every turn
3 ask records for 5 spoken turns — the extractor is skipped on openers, chips and continuations, and it is what writes this log
Vision, then a guest count, then two budget refusals, then a year. Scores what a rapport fact writes, and whether a refusal is honoured.
c5a551c75d8 ·
97,243 tokens ·
2 receipts ·
tools: executeClientActions, getUpcomingTaskAdvice
4 failing checks
- A capability is called by turn 2
first at turn 5 (executeClientActions) - No three consecutive turns ask for a fact
asks: date → location → date → location - A turn delivers more often than it asks
4 of 4 measured turns asked, out of 6 spoken - The interview detector sees every turn
4 ask records for 6 spoken turns — the extractor is skipped on openers, chips and continuations, and it is what writes this log
Two budget refusals, then the bare year 2027. Scores the year-only plan, the receipt for a checklist build, and whether a declined fact is asked again.
c5a551c75d8 ·
90,315 tokens ·
1 receipt ·
tools: getCoachReference, getUpcomingTaskAdvice
3 failing checks
- No three consecutive turns ask for a fact
asks: date → budget → date - A turn delivers more often than it asks
3 of 3 measured turns asked, out of 5 spoken - The interview detector sees every turn
3 ask records for 5 spoken turns — the extractor is skipped on openers, chips and continuations, and it is what writes this log
The opening move alone. Cheap enough to sample five times when the first turn changes.
c5a551c75d8 ·
9,236 tokens ·
0 receipts ·
tools: none called
1 failing check
- A capability is called by turn 2
no tool was called in the whole run