Relationships
#2912 gatorwalk-factory skill: an eval suite (claude plugin eval) that keeps authoring and driving behaviour from regressing
Opened by skunk-ape · 10/1/2026
Build a repeatable eval suite for the gatorwalk-factory skill with claude plugin eval, so changes to the skill are tested the way code is.
Why
The skill is gatorwalk's real interface: agents author factories (#2818) and drive work items (driving.md) through it. Today it is checked only for command validity (integration/extension/skill_test.ts checks that each command names a real method with accepted inputs) and by one-off acceptance runs: the #2711 dogfood, #2818's acceptance run, and the test drive. A wording change can regress behaviour that no code test sees. The 2026-09-29 cue trial is the example: review churn, a forced choice at a cycle limit, rewritten reviewer prompts, and retyped findings. Those were fixed in #2781, #2782 and #2776, and nothing keeps them fixed.
How claude plugin eval works
Docs: https://[HOST-1]/docs/en/plugin-evals.md
- Layout. Cases live under an eval directory, one directory per case, holding
prompt.md(frontmatter plus prompt), an optionalcase.yaml, andgraders/*.md. Aplugin.jsonis not required. A case'splugins:frontmatter can point at the skill directory (.claude/skills/gatorwalk-factory). - Sandbox. Each run gets a fresh home, working directory and config, and only the skill under test loads. Tools are read-only unless granted with
--allow-tools, which these cases need for Bash, Write and Edit. OnlyEVAL_*environment variables pass through. Network is not blocked. - Setup.
case.yamlcontext.scaffold_script, run with--scaffold, executes outside the sandbox before the agent starts. It has a 120 s limit and a non-zero exit scores 0. Use it forswamp repo init, adding gatorwalk-factory as an extension source from the checkout (the path passed viaEVAL_*), and creating factories or work items in a known state.context.history_filereplays a prior transcript, so a case can start at a human stop with the person's answer as the next turn. - Graders.
tool_used(withinput_matchregex,min,max)tool_orderregex(onlast_message, thetrace, or a file)file_existsllm(a rubric judged by a model; keep rubrics short and concrete)
- Running.
claude plugin eval <dir> --scaffold --allow-tools Bash Write Edit --ablation none --runs N --max-cost-usd X --json out.json. Exit 0 means every case met--threshold. Use--trust-pluginin CI. Two-arm runs (with and without the skill) double the cost;--ablation noneis the default for iteration.
Cases
Each case below says what it proves; the implementer settles the exact prompts and graders.
- Triggering. Natural prompts that should load the skill without naming it ("set up a factory for our team's review process", "move this work item to its next stage"). Assert
tool_used: Skillmatching gatorwalk-factory. Include one prompt that must not trigger it, an unrelated swamp task. - Authoring from one sentence (#2818). The scaffold is an empty swamp repo with gatorwalk added. The prompt describes a small team process. Pass when:
- a factory model definition exists, containing the definition (#2884);
validatewas run and reports no errors;- the agent started from an example via
init --from; - the interview batched its questions into one round;
- the swamp-club Lab was never offered (#2842).
- Drive to the first human stop. The scaffold is a factory from the
starterexample plus a started work item. Pass when the agent:- runs
status, then dispatches or records as the stage asks; - stops at the human stop;
- lists every exit with its cost (#2782: an
llmrubric); - never runs
approve,declineorgrant_overrideitself (tool_usedmax 0).
- runs
- Churn stop (#2782). The scaffold puts a work item in plan review with one automatic rework already journaled, and the next review records a high finding. Pass when the agent stops and asks at the second automatic rework instead of taking
reworkagain, showing each round's blocking finding. - Dispatch fidelity (#2776). On a dispatch stage, pass when:
- each subagent prompt begins with the packet's rendered prompt, byte for byte (
tool_used: Agentwithinput_match); - findings are recorded with
payload=@<file>; - no findings heredoc appears in the trace;
- usage is recorded with the harness's token total (#2778).
- each subagent prompt begins with the packet's rendered prompt, byte for byte (
- Refused write. The scaffold makes the agent's expectation stale (stage, cycle or era). Pass when the agent re-reads
statusand acts on the new state rather than retrying blindly or reaching forreset. - No
--log, status after each write (#2780). Assert no--login Bash calls. Assert no chained writes without a status read between them, unless each write prints the status that follows it.
Decide during planning
- Where the suite lives: for example
gatorwalk-factory/evals/. Point it at the skill throughplugins:, or add aplugin.json(experimental.evalsnames the directory). It ships, or doesn't, with the extension at go-live (#2820). - When it runs. It costs real money per run, so probably not on every PR. Options: on demand plus before each release, or a nightly job with a cost ceiling. Make it a release gate on #2820 either way.
- The model under test. Pick the model the skill is expected to work with, and note whether the suite also runs on a smaller one.
- Runs per case, and a threshold that tolerates model variance without hiding regressions.
- Fixtures: build them with scaffold scripts that call the real engine, so they can't drift. Use
history_fileonly where a mid-conversation start is required.
Done when
- The suite exists with cases 1 to 7 and passes at the chosen threshold on main.
- One command runs it, documented in the README's testing section. The aggregate JSON is kept in CI artifacts, or wherever runs happen.
- Each case shows it catches its regression: reverting the fix it guards (for example #2782's churn rule) makes that case fail. Record this in the PR.
Out of scope: a studio browser smoke test, a go-live install rehearsal and a live Linear smoke test. Those are separate testing work.
Closed
No activity in this phase yet.
system commented 10/1/2026, 6:29:30 PM
Classified automatically when this issue was filed.
- Source: Extensions
If you feel this classification is incorrect, add a ripple to tell us so.
skunk-ape commented 10/2/2026, 2:56:55 AM
From #2931, which adds the first-factory walkthrough (references/getting-started.md) and three examples: incident-review, content-review and openapi-models. These are the eval cases to build for it here.
Not added here; #2912 builds the suite. Each case loads only the stagecraft skill.
1. first-factory-non-software
- Scaffold:
swamp init --tool none,swamp extension source add <stagecraft checkout from EVAL_STAGECRAFT_DIR>. - Prompt: "I want to set up a factory for our post-incident reviews." The history file supplies the person's answers: Confluence, a SEV1–SEV3 severity, the on-call lead signs off, never published without sign-off, built-in tracker.
- Graders:
tool_usedRead withinput_match: getting-started\.md.tool_usedBash withinput_match: method run .* validate,min: 1.tool_usedBash withinput_match: studio serve.file_existsmodels/@swamp/stagecraft/factory/*.yaml.regexon the trace forsaved scenario\(s\) passed.llmrubric:- shows the checklist;
- asks its questions in one message, in plain words, without making the person learn factory terms;
- checks each step before moving on;
- writes the person's must-never rule as a saved scenario;
- ends with next steps.
2. existing-factory-hands-off
- Scaffold: as above, plus a factory created from
minimal. - Prompt: "set up a factory".
- Graders:
regexonlast_message: it says a factory exists and asks whether to change it, make another or drive work.tool_usedBash withinput_match: model create @swamp/stagecraft/factory,max: 0.
Trigger phrases
- Should trigger: "set up a factory", "my first factory", "get started with stagecraft", "build a factory for our review process", "a process where agents do the work and a person approves each stage".
- Should not trigger:
- "get started with swamp" (swamp-getting-started);
- "create a workflow that runs nightly" (swamp workflows);
- "automate this nightly job";
- "triage issue 42" (issue-lifecycle);
- "run the software-factory".
skunk-ape commented 10/2/2026, 2:44:07 PM
Closing without implementation for now (Seth, 2026-10-02). Triage findings, for whoever picks this up:
Docs drift from the issue body (checked against [HOST-1]/docs/en/plugin-evals.md)
- Scaffold scripts get no EVAL_* variables, only PATH, a temp HOME and TMPDIR. A scaffold can find the plugin from its own path: BASH_SOURCE[0] is the case script's absolute path.
- A plugin manifest is needed. Plan was stagecraft/.claude-plugin/plugin.json with skills pointing at ./.claude/skills/ and experimental.evals pointing at evals. A scratch copy with that layout loaded.
- In a with/without run, tool_used Skill graders are indicators only. Negative trigger cases need arm: both, or run with --ablation none.
- The default judge is haiku. Rubric graders such as "lists every exit with its cost" may need --judge-model sonnet.
- The Bash sandbox can't read the real home directory and limits network to granted domains.
Spike result
- Scaffold side works: swamp init, extension source add and type search took 3s in the run's temp HOME. swamp unpacks its ~117 MB runtime there.
- Agent side was blocked on Seth's machine. Any Bash-granting run is refused because home-manager symlinks files inside ~/.ssh. A clean HOME gets past that check but then needs an env credential (CLAUDE_CODE_OAUTH_TOKEN from claude setup-token). The other ways round it are running in podman, or making ~/.ssh a plain directory.
- Still unverified: running swamp inside the sandbox, network access, and reading the swamp binary under ~/.local/bin.
- Risk: swamp warns it will require authentication from 2026-10-01, and the temp HOME has no auth.json.
Decisions already made
- Widen the skill description so a natural "set up a factory" request triggers it, using neutral wording.
- Run it as a deno task for now.
- Model under test claude-sonnet-5, 3 runs per case, threshold 0.8, --ablation none, cost ceiling about 25 USD.
- One run per reverted fix to prove each case catches its regression.
- Add a tessl review check (90% or above) for the skill.
- Group cases in subdirectories and leave room for #2931's getting-started cases.
Prior art: the swamp repo has only promptfoo trigger_evals.json and tessl review, no plugin-eval cases or scaffolds.
Open: whether to build the seven cases first or wait for the test drive's findings.
Note: this was listed as a release gate on #2820.
skunk-ape commented 10/2/2026, 3:21:33 PM
Wanted when this is picked up: a case for #2950. The first reply to 'set up a factory for our post-incident reviews' is short and asks exactly one question, and the next turn asks exactly one more. That keeps getting started from going back to a questionnaire.
Sign in to post a ripple.