Skip to main content
← Back to list
01Issue
BugShippedExtensionsPublic
Assigneesskunk-ape

Relationships

#2778 gatorwalk-factory: record_usage cannot take the token count Claude Code reports, so no dispatch gets usage

Opened by skunk-ape · 9/30/2026· Shipped 9/30/2026

The cue trial (work item cue-er7koww5) finished with "Tokens (attested): 0 in, 0 out over 0 dispatch(es); 14 without usage" (summary, 23:26:07 UTC). record_usage was never called. Six of the 14 dispatches were reviewer subagents (dispatches 2, 4, 6, 8, 10, 13). The round-1 prompt told the reviewer to end with "tokens: unknown (or the counts if you know them)", and it answered "tokens: unknown". A subagent cannot see its own token count.

Claude Code does report usage, but not where the skill says to look and not in the shape record_usage accepts. The reviewers ran as background agents. Each one's completion <task-notification> in the driver's context carried <usage><subagent_tokens>N</subagent_tokens><tool_uses>..</tool_uses><duration_ms>..</duration_ms></usage>: 65,155, 80,523, 79,635, 86,217, 78,701 and 49,659 tokens (jsonl lines 142, 213, 266, 339, 380, 503). That is one total per subagent, with no input/output split. record_usage requires both inputTokens and outputTokens (work_item.ts:185-193, run_record.ts:65), so an honest driver can't record the total without making up a split. The skill points the driver the wrong way: SKILL.md:60-61 says "record_usage for a dispatch whose tokens you know", and driving.md:227 says "When the work used tokens you can count (a subagent reports them)". The per-turn split exists only in the subagent's own transcript (subagents/agent-<id>.jsonl), and the launch result tells the driver not to read that.

The eight interactive dispatches (the driver's own planning, implementing, checking and landing) have no usage the driver can see at all.

Fix direction:

  • Engine: let record_usage take totalTokens alone, or together with the split. metrics.ts:435-449 and summary.ts:128-152 should sum and show totals, and keep "without usage" for dispatches that have nothing. Optionally also take toolUses and durationMs, which the notification reports too. Decision for the implementer: keep inputTokens/outputTokens required when totalTokens is absent, or make every field optional with at least one present.
  • Skill: say where the number comes from. It is the harness's report of the subagent (Claude Code: the <usage> block in the task notification, or the Agent tool result when the call is synchronous), not the subagent's own text. Record it right after the hand-back, keyed by that subagent's dispatch id. Take model from what the Agent call resolved to. Drop "(a subagent reports them)", and stop asking reviewers to report tokens. State plainly that interactive dispatches normally have no usage, so the summary's "without usage" count for them is expected.

Acceptance: in a trial with dispatch stages, every subagent dispatch has usage and the summary's attested total equals the sum of the harness-reported counts. driving.md's record_usage command, with totalTokens, is covered by skill_test.ts.

02Bog Flow
✓OPEN✓TRIAGED✓IN PROGRESS✓SHIPPED+ 1 MOREASSIGNED+ 5 MOREREVIEW+ 10 MOREPR_LINKED+ 2 MORESESSION_SUMMARIZED

Shipped

9/30/2026, 9:43:20 PM

Click a lifecycle step above to view its details.

03Sludge Pulse
skunk-ape assigned skunk-ape9/30/2026, 7:24:28 PM

Sign in to post a ripple.