Skip to main content
← Back to list
01Issue
BugShippedExtensionsPublic
Assigneesskunk-ape

Relationships

⊘ blocks #2781≡ duplicated by #2777≡ duplicated by #2779

#2776 gatorwalk-factory: dispatch fidelity: subagents get the rendered prompt and product contract, and their results are recorded as returned

Opened by skunk-ape · 9/30/2026· Shipped 9/30/2026

This issue now covers #2777 and #2779 as well; both are closed as duplicates of it. All three are one problem: in a dispatch stage, what the subagent receives and what gets recorded are both whatever the driver retypes. The fix is the same for each: the packet carries a complete, ready-to-send contract (the rendered prompt, the products to produce with their schemas, and where to write the result), and the driver records the subagent's own output file.

Parts:

  1. The rendered prompt is sent verbatim. Any addendum is fixed and documented, or recorded on the dispatch.
  2. The packet names the products and their schemas (products in DispatchPacket).
  3. Results are recorded as returned. The subagent writes its payload JSON to a result path, and the driver records it with payload=@<path> without editing it.

Decide in planning:

  • a ready-to-send subagentPrompt versus a recorded addendum;
  • inline schemas versus a fetch command;
  • who names the result path (the packet, preferably);
  • how several subagents' findings merge.

Done when all three acceptance sections below hold.

Part 1: dispatch prompts rewritten (originally this issue)

In the cue trial (work item cue-er7koww5, session a20250a4), none of the six reviewer subagents got the prompt dispatch rendered. The driver wrote its own prompt each time. It added a file path for the plan, a list of files to read, "READ ONLY", a JSON-only return format and each earlier round's findings. By round 5 it had also added "That is a decided risk acceptance, not a gap for you to re-litigate" and "Calibrate severity to the change's size: critical/high only for something that would ship a real defect" (Agent call 22:37:14 UTC, jsonl line 361). In round 3 (Agent call 22:28:15, line 241) the reviewer got the literal text $(summary is in the plan file) where the plan summary should have been.

The engine did not cause the round-3 leak. The run record's dispatch 6 holds the correctly rendered prompt, with the real summary, so renderTemplate (_lib/template.ts) worked. The driver never read that prompt. It ran dispatch ... --log 2>&1 | grep -E "dispatch [0-9]|rror|runaway" | head -3 (line 232, 22:28:02), which dropped everything after the first log line. Then it rebuilt the prompt from the lifecycle template it had read with swamp data get <key> lifecycle at 22:06:44. {{planSummary}} became a shell-style placeholder the model wrote itself. Nothing expands text in an Agent prompt. The result is that dispatches 2, 4, 6, 8, 10 and 13 each record a prompt the subagent never saw, and the replay record the packet exists for (dispatch.ts:24-33, "recordDispatch stores its inputs and prompt for replay") is wrong.

The skill makes this likely. driving.md:202-204 says to "hand the prompt to subagents subagents" and "give each the products named in inject". It never says to send the prompt verbatim. It also doesn't say how to deliver injected products or what the subagent should return, so the driver has to write that part itself. The packet (dispatch.ts:35-55) has prompt and inject names but no output contract. Once the driver is writing instructions anyway, it also changes the reviewer's scope and severity bar. That is the plan's author tuning its own review.

Fix direction:

  • Skill: send packet.prompt unchanged, taken from the dispatch output and not retyped. The driver may add only a fixed, documented addendum: where each injected product can be read (swamp data get <key> artifact-<name> --json), and how to return the result (see the record-without-retyping issue). The driver must never change scope, severity guidance or the reviewer's stance. A decision the person made, such as "the JS is hand-checked after landing", belongs in the product being reviewed (the plan says it) or in an approval note. It does not go into a rewritten prompt.
  • Engine (decision for the implementer): either (a) the packet carries a ready-to-send subagentPrompt that already includes the inject read commands and the output contract, so the driver adds nothing, or (b) dispatch takes an optional addendum input that is recorded beside the rendered prompt, so replay shows exactly what was sent. (a) keeps drivers honest by construction. (b) is smaller and still auditable.

Acceptance: in a trial with a dispatch stage, every subagent prompt starts with the recorded dispatch prompt byte for byte. Anything after it is either the documented addendum or recorded on the dispatch. driving.md says this in the dispatch-mode bullet, and skill_test.ts still passes.

Part 2: findings retyped instead of recorded (was #2777)

In the cue trial (work item cue-er7koww5), every reviewer returned a valid {"findings":[...]} object in its hand-back. The driver recorded none of them as returned. For each of the six reviews it retyped the findings into a heredoc (findings.json, f2.json to f5.json, cr.json; jsonl lines 140, 211, 264, 337, 378, 501) and recorded that. Comparing each hand-back with the stored artifact-plan-review versions 1-5 and artifact-code-review version 1: ids and severities survived. 30 of the 31 descriptions were rewritten, and together they shrank to 48% of the reviewers' text (23,314 to 11,130 characters). Most file:line citations the prompt asked for were cut. PR-2 lost a whole sentence ("Without the chat_live.ex change, that test also fails, because the read_timer assign does not exist at HEAD"). PR4-1 went from 1,295 to 544 characters. In round 1 the driver started a review.py to read the task's output file, and the file still holds a dead ... if False else None line (22:09:02, line 140). It gave that up and typed the findings by hand.

Rule 6 ("Record only what happened") and driving.md:204 ("Record what they return") set the goal but give no mechanical way to meet it. The hand-back arrives as text in the driver's context. The only documented ways to record are payload='<json>' or an --input-file YAML the driver writes (driving.md:239-260), and both mean the driver types the payload out again. A model that retypes 1,000-character findings will summarise them. The driver is also the plan's author, so recording the reviews in its own words is exactly what rule 6 guards against.

swamp can already do this without retyping. record_artifact accepts payload as a JSON string (work_item_ops.ts:154-157), and swamp reads key=@path inputs from a file. I checked this against swamp 20260929.202912.0: record_artifact t1 --input name=plan --input payload=@p.json --input expectedStage=... recorded the file's contents. --input-file and --input also merge, with the file as the base (swamp src/cli/input_parser.ts:246-289).

Fix direction (skill, driving.md "Do the stage's work" and "Record products"): for a dispatch stage, the driver picks a result path per subagent, for example <scratch>/<key>-d<dispatchId>-<n>.json. It tells the subagent to write its product payload there as JSON and nothing else. This is the one write a read-only reviewer is allowed. The driver then records with --input payload=@<path> plus name and expectation. It may read the file to show the person, but never edits it. If the file is missing or invalid, the driver sends the subagent back to fix it (SendMessage) instead of repairing it itself. For several subagents on one findings artifact, the implementer must decide how they merge: concatenate findings arrays mechanically (a documented jq -s line), or have the engine accept one record per subagent. The packet could also name the result path (see the dispatch-packet schema issue), so the driver doesn't choose it.

Acceptance: driving.md documents the result-file flow, and skill_test.ts checks the payload=@<path> command. In a trial, each recorded findings artifact is byte-identical (after JSON canonicalisation) to what its reviewer wrote, and the driver's transcript has no findings heredocs.

Part 3: the packet does not name products or schemas (was #2779)

In the cue trial (work item cue-er7koww5), the first dispatch (plan, 22:06:40 UTC, jsonl line 55) printed the stage prompt and a packet of stage, cycle, mode, skills, subagents, values, inject, problems and ready. It said nothing about what the stage has to record. The driver said "I need the plan artifact's schema too" (22:06:43) and ran swamp data get cue-er7koww5 lifecycle --json | jq '.content // .' | head -300 (line 64) to find plan's required summary, steps[{description, files}] (with additionalProperties: false) and testingStrategy. For the review stages it copied the findings shape from driving.md:263-264 into every reviewer prompt, because the packet doesn't give a subagent the output contract either.

buildDispatch (_lib/dispatch.ts:92-110) builds the packet from stage.work alone. The stage's artifacts, evidence and resultEvidence are never read, and neither are the built-in contracts that apply to them: FINDINGS_SCHEMA for kind: findings (payload_schema.ts:561-582) and OUTCOME_SCHEMA for result evidence. status names a missing product only through a gate failure ("artifact 'plan' has not been recorded"), and only for products a gate reads. So the only full statement of what to record is the pinned lifecycle document, which the skill never tells the driver to read.

Fix direction: add products to DispatchPacket (dispatch.ts:35-55). It holds one entry per artifact and evidence the stage declares: kind (artifact or evidence), name, reviews when set, and schema, the effective contract used to validate a payload (the declared schema combined with the findings or outcome contract, the same pair validateArtifactPayload checks at payload_schema.ts:604-617). The dispatch log already prints the packet as JSON (work_item_ops.ts:1098-1104), so the driver sees it with no other change. driving.md "Do the stage's work" then says to record each entry in products, and to pass a product's schema to any subagent that produces it. Decision for the implementer: inline full schemas (simple, but a stage like cue's check has a 40-line evidence schema printed on every dispatch), or print name and kind plus one swamp data get command that shows the schemas. This goes together with the dispatch-prompt and record-without-retyping issues: if the packet also names a result path per product, the subagent's output contract is complete without the driver writing any of it.

Acceptance: dispatch on a stage with a declared artifact prints that artifact's name and effective schema. A findings stage shows the findings contract. A workflow stage shows the outcome contract for its result evidence. A test in dispatch_test.ts covers all three. In a trial, the driver records products without reading the pinned lifecycle.

02Bog Flow
✓OPEN✓TRIAGED✓IN PROGRESS✓SHIPPED+ 1 MOREASSIGNED+ 5 MOREREVIEW+ 13 MOREPR_MERGED+ 2 MORESESSION_SUMMARIZED

Shipped

9/30/2026, 7:16:23 PM

Click a lifecycle step above to view its details.

03Sludge Pulse
skunk-ape assigned skunk-ape9/30/2026, 6:22:50 PM
skunk-ape linked duplicated by #27779/30/2026, 4:48:38 PM
skunk-ape linked duplicated by #27799/30/2026, 4:48:40 PM
skunk-ape linked blocks #27819/30/2026, 6:17:27 PM

Sign in to post a ripple.