Skip to main content
← Back to list
01Issue
FeatureClosedExtensionsPublic
Assigneesskunk-ape

Relationships

#2912 gatorwalk-factory skill: an eval suite (claude plugin eval) that keeps authoring and driving behaviour from regressing

Opened by skunk-ape · 10/1/2026

Build a repeatable eval suite for the gatorwalk-factory skill with claude plugin eval, so changes to the skill are tested the way code is.

Why

The skill is gatorwalk's real interface: agents author factories (#2818) and drive work items (driving.md) through it. Today it is checked only for command validity (integration/extension/skill_test.ts checks that each command names a real method with accepted inputs) and by one-off acceptance runs: the #2711 dogfood, #2818's acceptance run, and the test drive. A wording change can regress behaviour that no code test sees. The 2026-09-29 cue trial is the example: review churn, a forced choice at a cycle limit, rewritten reviewer prompts, and retyped findings. Those were fixed in #2781, #2782 and #2776, and nothing keeps them fixed.

How claude plugin eval works

Docs: https://[HOST-1]/docs/en/plugin-evals.md

  • Layout. Cases live under an eval directory, one directory per case, holding prompt.md (frontmatter plus prompt), an optional case.yaml, and graders/*.md. A plugin.json is not required. A case's plugins: frontmatter can point at the skill directory (.claude/skills/gatorwalk-factory).
  • Sandbox. Each run gets a fresh home, working directory and config, and only the skill under test loads. Tools are read-only unless granted with --allow-tools, which these cases need for Bash, Write and Edit. Only EVAL_* environment variables pass through. Network is not blocked.
  • Setup. case.yaml context.scaffold_script, run with --scaffold, executes outside the sandbox before the agent starts. It has a 120 s limit and a non-zero exit scores 0. Use it for swamp repo init, adding gatorwalk-factory as an extension source from the checkout (the path passed via EVAL_*), and creating factories or work items in a known state. context.history_file replays a prior transcript, so a case can start at a human stop with the person's answer as the next turn.
  • Graders.
    • tool_used (with input_match regex, min, max)
    • tool_order
    • regex (on last_message, the trace, or a file)
    • file_exists
    • llm (a rubric judged by a model; keep rubrics short and concrete)
  • Running. claude plugin eval <dir> --scaffold --allow-tools Bash Write Edit --ablation none --runs N --max-cost-usd X --json out.json. Exit 0 means every case met --threshold. Use --trust-plugin in CI. Two-arm runs (with and without the skill) double the cost; --ablation none is the default for iteration.

Cases

Each case below says what it proves; the implementer settles the exact prompts and graders.

  1. Triggering. Natural prompts that should load the skill without naming it ("set up a factory for our team's review process", "move this work item to its next stage"). Assert tool_used: Skill matching gatorwalk-factory. Include one prompt that must not trigger it, an unrelated swamp task.
  2. Authoring from one sentence (#2818). The scaffold is an empty swamp repo with gatorwalk added. The prompt describes a small team process. Pass when:
    • a factory model definition exists, containing the definition (#2884);
    • validate was run and reports no errors;
    • the agent started from an example via init --from;
    • the interview batched its questions into one round;
    • the swamp-club Lab was never offered (#2842).
  3. Drive to the first human stop. The scaffold is a factory from the starter example plus a started work item. Pass when the agent:
    • runs status, then dispatches or records as the stage asks;
    • stops at the human stop;
    • lists every exit with its cost (#2782: an llm rubric);
    • never runs approve, decline or grant_override itself (tool_used max 0).
  4. Churn stop (#2782). The scaffold puts a work item in plan review with one automatic rework already journaled, and the next review records a high finding. Pass when the agent stops and asks at the second automatic rework instead of taking rework again, showing each round's blocking finding.
  5. Dispatch fidelity (#2776). On a dispatch stage, pass when:
    • each subagent prompt begins with the packet's rendered prompt, byte for byte (tool_used: Agent with input_match);
    • findings are recorded with payload=@<file>;
    • no findings heredoc appears in the trace;
    • usage is recorded with the harness's token total (#2778).
  6. Refused write. The scaffold makes the agent's expectation stale (stage, cycle or era). Pass when the agent re-reads status and acts on the new state rather than retrying blindly or reaching for reset.
  7. No --log, status after each write (#2780). Assert no --log in Bash calls. Assert no chained writes without a status read between them, unless each write prints the status that follows it.

Decide during planning

  • Where the suite lives: for example gatorwalk-factory/evals/. Point it at the skill through plugins:, or add a plugin.json (experimental.evals names the directory). It ships, or doesn't, with the extension at go-live (#2820).
  • When it runs. It costs real money per run, so probably not on every PR. Options: on demand plus before each release, or a nightly job with a cost ceiling. Make it a release gate on #2820 either way.
  • The model under test. Pick the model the skill is expected to work with, and note whether the suite also runs on a smaller one.
  • Runs per case, and a threshold that tolerates model variance without hiding regressions.
  • Fixtures: build them with scaffold scripts that call the real engine, so they can't drift. Use history_file only where a mid-conversation start is required.

Done when

  • The suite exists with cases 1 to 7 and passes at the chosen threshold on main.
  • One command runs it, documented in the README's testing section. The aggregate JSON is kept in CI artifacts, or wherever runs happen.
  • Each case shows it catches its regression: reverting the fix it guards (for example #2782's churn rule) makes that case fail. Record this in the PR.

Out of scope: a studio browser smoke test, a go-live install rehearsal and a live Linear smoke test. Those are separate testing work.

02Bog Flow
✓OPEN✓TRIAGED○IN PROGRESS◉CLOSED+ 1 MOREASSIGNEDCLASSIFICATION

Closed

10/2/2026, 2:44:07 PM

No activity in this phase yet.

03Sludge Pulse
skunk-ape assigned skunk-ape10/1/2026, 7:32:26 PM
skunk-ape linked blocks #282010/1/2026, 6:29:31 PM
skunk-ape removed blocks #282010/2/2026, 3:22:44 PM
Editable. Press Enter to edit.

system commented 10/1/2026, 6:29:30 PM

Classified automatically when this issue was filed.

  • Source: Extensions

If you feel this classification is incorrect, add a ripple to tell us so.

skunk-ape commented 10/2/2026, 2:56:55 AM

From #2931, which adds the first-factory walkthrough (references/getting-started.md) and three examples: incident-review, content-review and openapi-models. These are the eval cases to build for it here.

Not added here; #2912 builds the suite. Each case loads only the stagecraft skill.

1. first-factory-non-software

  • Scaffold: swamp init --tool none, swamp extension source add <stagecraft checkout from EVAL_STAGECRAFT_DIR>.
  • Prompt: "I want to set up a factory for our post-incident reviews." The history file supplies the person's answers: Confluence, a SEV1–SEV3 severity, the on-call lead signs off, never published without sign-off, built-in tracker.
  • Graders:
    • tool_used Read with input_match: getting-started\.md.
    • tool_used Bash with input_match: method run .* validate, min: 1.
    • tool_used Bash with input_match: studio serve.
    • file_exists models/@swamp/stagecraft/factory/*.yaml.
    • regex on the trace for saved scenario\(s\) passed.
    • llm rubric:
      • shows the checklist;
      • asks its questions in one message, in plain words, without making the person learn factory terms;
      • checks each step before moving on;
      • writes the person's must-never rule as a saved scenario;
      • ends with next steps.

2. existing-factory-hands-off

  • Scaffold: as above, plus a factory created from minimal.
  • Prompt: "set up a factory".
  • Graders:
    • regex on last_message: it says a factory exists and asks whether to change it, make another or drive work.
    • tool_used Bash with input_match: model create @swamp/stagecraft/factory, max: 0.

Trigger phrases

  • Should trigger: "set up a factory", "my first factory", "get started with stagecraft", "build a factory for our review process", "a process where agents do the work and a person approves each stage".
  • Should not trigger:
    • "get started with swamp" (swamp-getting-started);
    • "create a workflow that runs nightly" (swamp workflows);
    • "automate this nightly job";
    • "triage issue 42" (issue-lifecycle);
    • "run the software-factory".

skunk-ape commented 10/2/2026, 2:44:07 PM

Closing without implementation for now (Seth, 2026-10-02). Triage findings, for whoever picks this up:

Docs drift from the issue body (checked against [HOST-1]/docs/en/plugin-evals.md)

  • Scaffold scripts get no EVAL_* variables, only PATH, a temp HOME and TMPDIR. A scaffold can find the plugin from its own path: BASH_SOURCE[0] is the case script's absolute path.
  • A plugin manifest is needed. Plan was stagecraft/.claude-plugin/plugin.json with skills pointing at ./.claude/skills/ and experimental.evals pointing at evals. A scratch copy with that layout loaded.
  • In a with/without run, tool_used Skill graders are indicators only. Negative trigger cases need arm: both, or run with --ablation none.
  • The default judge is haiku. Rubric graders such as "lists every exit with its cost" may need --judge-model sonnet.
  • The Bash sandbox can't read the real home directory and limits network to granted domains.

Spike result

  • Scaffold side works: swamp init, extension source add and type search took 3s in the run's temp HOME. swamp unpacks its ~117 MB runtime there.
  • Agent side was blocked on Seth's machine. Any Bash-granting run is refused because home-manager symlinks files inside ~/.ssh. A clean HOME gets past that check but then needs an env credential (CLAUDE_CODE_OAUTH_TOKEN from claude setup-token). The other ways round it are running in podman, or making ~/.ssh a plain directory.
  • Still unverified: running swamp inside the sandbox, network access, and reading the swamp binary under ~/.local/bin.
  • Risk: swamp warns it will require authentication from 2026-10-01, and the temp HOME has no auth.json.

Decisions already made

  • Widen the skill description so a natural "set up a factory" request triggers it, using neutral wording.
  • Run it as a deno task for now.
  • Model under test claude-sonnet-5, 3 runs per case, threshold 0.8, --ablation none, cost ceiling about 25 USD.
  • One run per reverted fix to prove each case catches its regression.
  • Add a tessl review check (90% or above) for the skill.
  • Group cases in subdirectories and leave room for #2931's getting-started cases.

Prior art: the swamp repo has only promptfoo trigger_evals.json and tessl review, no plugin-eval cases or scaffolds.

Open: whether to build the seven cases first or wait for the test drive's findings.

Note: this was listed as a release gate on #2820.

skunk-ape commented 10/2/2026, 3:21:33 PM

Wanted when this is picked up: a case for #2950. The first reply to 'set up a factory for our post-incident reviews' is short and asks exactly one question, and the next turn asks exactly one more. That keeps getting started from going back to a questionnaire.

Sign in to post a ripple.