Skip to main content
← Back to list
01Issue
BugClosedExtensionsPublic
AssigneesNone

Relationships

≡ duplicate of #2776

#2777 gatorwalk-factory: subagent findings are retyped into record_artifact instead of recorded as returned

Opened by skunk-ape · 9/30/2026

In the cue trial (work item cue-er7koww5), every reviewer returned a valid {"findings":[...]} object in its hand-back. The driver recorded none of them as returned. For each of the six reviews it retyped the findings into a heredoc (findings.json, f2.json to f5.json, cr.json; jsonl lines 140, 211, 264, 337, 378, 501) and recorded that. Comparing each hand-back with the stored artifact-plan-review versions 1-5 and artifact-code-review version 1: ids and severities survived. 30 of the 31 descriptions were rewritten, and together they shrank to 48% of the reviewers' text (23,314 to 11,130 characters). Most file:line citations the prompt asked for were cut. PR-2 lost a whole sentence ("Without the chat_live.ex change, that test also fails, because the read_timer assign does not exist at HEAD"). PR4-1 went from 1,295 to 544 characters. In round 1 the driver started a review.py to read the task's output file, and the file still holds a dead ... if False else None line (22:09:02, line 140). It gave that up and typed the findings by hand.

Rule 6 ("Record only what happened") and driving.md:204 ("Record what they return") set the goal but give no mechanical way to meet it. The hand-back arrives as text in the driver's context. The only documented ways to record are payload='<json>' or an --input-file YAML the driver writes (driving.md:239-260), and both mean the driver types the payload out again. A model that retypes 1,000-character findings will summarise them. The driver is also the plan's author, so recording the reviews in its own words is exactly what rule 6 guards against.

swamp can already do this without retyping. record_artifact accepts payload as a JSON string (work_item_ops.ts:154-157), and swamp reads key=@path inputs from a file. I checked this against swamp 20260929.202912.0: record_artifact t1 --input name=plan --input payload=@p.json --input expectedStage=... recorded the file's contents. --input-file and --input also merge, with the file as the base (swamp src/cli/input_parser.ts:246-289).

Fix direction (skill, driving.md "Do the stage's work" and "Record products"): for a dispatch stage, the driver picks a result path per subagent, for example <scratch>/<key>-d<dispatchId>-<n>.json. It tells the subagent to write its product payload there as JSON and nothing else. This is the one write a read-only reviewer is allowed. The driver then records with --input payload=@<path> plus name and expectation. It may read the file to show the person, but never edits it. If the file is missing or invalid, the driver sends the subagent back to fix it (SendMessage) instead of repairing it itself. For several subagents on one findings artifact, the implementer must decide how they merge: concatenate findings arrays mechanically (a documented jq -s line), or have the engine accept one record per subagent. The packet could also name the result path (see the dispatch-packet schema issue), so the driver doesn't choose it.

Acceptance: driving.md documents the result-file flow, and skill_test.ts checks the payload=@<path> command. In a trial, each recorded findings artifact is byte-identical (after JSON canonicalisation) to what its reviewer wrote, and the driver's transcript has no findings heredocs.

02Bog Flow
✓OPEN○TRIAGED○IN PROGRESS◉CLOSED

Closed

9/30/2026, 4:48:39 PM

No activity in this phase yet.

03Sludge Pulse
skunk-ape linked duplicate of #27769/30/2026, 4:48:38 PM
Editable. Press Enter to edit.

skunk-ape commented 9/30/2026, 4:48:39 PM

Merged into #2776 (dispatch fidelity), which now carries this issue's full text as part 2.

Sign in to post a ripple.