Relationships
#2778 gatorwalk-factory: record_usage cannot take the token count Claude Code reports, so no dispatch gets usage
Opened by skunk-ape · 9/30/2026· Shipped 9/30/2026
The cue trial (work item cue-er7koww5) finished with "Tokens (attested): 0 in, 0 out over 0 dispatch(es); 14 without usage" (summary, 23:26:07 UTC). record_usage was never called. Six of the 14 dispatches were reviewer subagents (dispatches 2, 4, 6, 8, 10, 13). The round-1 prompt told the reviewer to end with "tokens: unknown (or the counts if you know them)", and it answered "tokens: unknown". A subagent cannot see its own token count.
Claude Code does report usage, but not where the skill says to look and not in the shape record_usage accepts. The reviewers ran as background agents. Each one's completion <task-notification> in the driver's context carried <usage><subagent_tokens>N</subagent_tokens><tool_uses>..</tool_uses><duration_ms>..</duration_ms></usage>: 65,155, 80,523, 79,635, 86,217, 78,701 and 49,659 tokens (jsonl lines 142, 213, 266, 339, 380, 503). That is one total per subagent, with no input/output split. record_usage requires both inputTokens and outputTokens (work_item.ts:185-193, run_record.ts:65), so an honest driver can't record the total without making up a split. The skill points the driver the wrong way: SKILL.md:60-61 says "record_usage for a dispatch whose tokens you know", and driving.md:227 says "When the work used tokens you can count (a subagent reports them)". The per-turn split exists only in the subagent's own transcript (subagents/agent-<id>.jsonl), and the launch result tells the driver not to read that.
The eight interactive dispatches (the driver's own planning, implementing, checking and landing) have no usage the driver can see at all.
Fix direction:
- Engine: let
record_usagetaketotalTokensalone, or together with the split.metrics.ts:435-449andsummary.ts:128-152should sum and show totals, and keep "without usage" for dispatches that have nothing. Optionally also taketoolUsesanddurationMs, which the notification reports too. Decision for the implementer: keepinputTokens/outputTokensrequired whentotalTokensis absent, or make every field optional with at least one present. - Skill: say where the number comes from. It is the harness's report of the subagent (Claude Code: the
<usage>block in the task notification, or the Agent tool result when the call is synchronous), not the subagent's own text. Record it right after the hand-back, keyed by that subagent's dispatch id. Takemodelfrom what the Agent call resolved to. Drop "(a subagent reports them)", and stop asking reviewers to report tokens. State plainly that interactive dispatches normally have no usage, so the summary's "without usage" count for them is expected.
Acceptance: in a trial with dispatch stages, every subagent dispatch has usage and the summary's attested total equals the sum of the harness-reported counts. driving.md's record_usage command, with totalTokens, is covered by skill_test.ts.
Shipped
Click a lifecycle step above to view its details.
Sign in to post a ripple.