Relationships
#1735 knowledge-base retrieve persists query output into the state resource, with lifetime infinite and the query text in the instance name
Opened by esteban · 8/19/2026· Shipped 10/6/2026
Summary
@swamp/aws/bedrock/knowledge-base.retrieve (added in 2026.08.18.1, from feature request #1571 — thank you for shipping it) calls the Bedrock data plane correctly, but persists its result into the model's state resource. state is the knowledge base's own configuration state: declared lifetime: "infinite", and its schema requires KnowledgeBaseId while describing KnowledgeBaseConfiguration, StorageConfiguration, RoleArn, Status. A retrieval result does not fit that shape, is retained forever, and lands under an instance name built from every argument value — including the natural-language query.
All line references below are from the pulled source of 2026.08.19.1, re-verified today.
The code
models/knowledge_base.ts:1253-1269:
const result = await retrieve(mergedArgs, credentials);
const argKeys = Object.keys(args).filter((k) => args[k] !== undefined);
const suffix = argKeys.length > 0
? "-" + argKeys.map((k) => String(args[k])).join("-")
: "";
const instanceName = ("retrieve" + suffix).replace(/[\/\\]/g, "_")
.replace(/\.\./g, "_").replace(/\0/g, "");
const handle = await context.writeResource("state", instanceName, result);The target resource, models/knowledge_base.ts:989-995:
resources: {
state: {
description: "Bedrock KnowledgeBase resource state",
schema: StateSchema,
lifetime: "infinite",
garbageCollection: 10,
},
},Three problems, in order of severity
| # | Problem | Why it matters |
|---|---|---|
| 1 | Writes into state |
StateSchema (:764-795) has KnowledgeBaseId: z.string() required, plus optional KnowledgeBaseConfiguration, StorageConfiguration, RoleArn, Status, CreatedAt... The retrieve payload is { results, resultCount, nextToken? } (:696-742) and carries no KnowledgeBaseId at all, so it cannot satisfy the schema. It also occupies the slot that genuine resource state belongs in — a retrieve call and a sync call now write the same resource with unrelated shapes |
| 2 | lifetime: "infinite" |
Correct for configuration state, wrong for query output. Retrieval results never expire |
| 3 | Instance name embeds every argument value | retrieve-<query text>-<kbId>-<numberOfResults>-..., path-sanitized. The full natural-language query becomes part of a data instance name: unbounded cardinality, unreadable in data list, and a brand-new instance for every distinct question asked |
Note that 2 and 3 compound. garbageCollection: 10 caps versions per instance name, so it would bound things if retrievals reused one name — but because every distinct query produces a new name, the cap never engages. The result is an unbounded set of instances, each retained forever.
Honest scope note: problem 1 is read from the two schemas rather than from a failed run — we renamed our own method to coexist (see below) and have been using that, so I have not observed whether the mismatch throws at validation or is stored unvalidated. Either way the resource choice looks unintended.
Suggested fix
Declare a dedicated resource for retrieval rather than reusing state:
- a
retrievalresource whose schema describes chunks —text,score,metadata,location,sourceUri; - a short
lifetimesuch as7dwithgarbageCollectionset, since query output is derived data; - an instance name from something stable and low-cardinality — the knowledge-base ID is the natural key — instead of the query text.
A working reference implementation
This repo has been doing exactly that since before the first-party method existed (it is what prompted #1571): dedicated retrieval resource, lifetime: "7d", garbageCollection: 10, instance keyed by knowledgeBaseId. Live-verified against Bedrock: 3 chunks with real relevance scores in 1875ms, persisting to exactly one instance named after the KB ID at v13 — no per-query instance sprawl. Happy to offer it as a reference or a patch if that is useful.
Why we still carry ours
Our extension is named retrieve_chunks, not retrieve, because a method-name collision between a base type and an extending extension silently drops the whole extension file — filed separately as swamp Lab #1734. Once the persistence here is fixed, our extension becomes redundant and we delete it outright; that is the outcome we are after.
Environment
@swamp/aws/bedrock2026.08.19.1(identical model source to2026.08.18.1— that release changed only the version string inmanifest.yaml)- swamp
20260809.004828.0-sha.b61c9de2, Deno 2.8.3, Linux x86_64 - Related: #1571 (the request that added
retrieve), #1734 (silent collision, why we cannot simply drop ours)
Upstream repository: https://github.com/swamp-club/swamp-extensions
Environment
- Extension:
@swamp/aws/bedrock@2026.08.19.1 - swamp:
20260819.011806.0-sha.a9c7ee9a - OS:
linux(x86_64) - Deno:
2.8.3 - Shell:
/bin/zsh
Shipped
Click a lifecycle step above to view its details.
stack72 commented 10/6/2026, 6:27:46 PM
Fix is up in https://git.swamp-club.com/swamp-club/swamp-extensions/pulls/462 (thanks for the detailed report and the reference implementation).
What changes once it ships:
@swamp/aws/bedrock/knowledge-base.retrievewrites to a dedicatedretrievalresource (lifetime 7d, garbageCollection 10), one instance per knowledge base (retrieval/<knowledgeBaseId>). The query text is in the payload, not the instance name. Each new query becomes the latest version; earlier ones remain as prior versions.- The same codegen template also drove
@swamp/aws/events/event-bus.put_eventsand the four@swamp/aws/cloudformation/stack-setmethods, so they move too:putEventsResult/<bus Name>,stackInstance/<Account>-<Region>,operation/<OperationId>,driftDetection/<StackSetName>. - Nothing is migrated. Old
stateinstances written by these methods (names startingretrieve-,put_events-,describeOperation-,detectDrift-, and numeric names on stack-set models) stay until you delete them withswamp data. Anything that read these method results fromstateneeds to point at the new resource names.
Your retrieve_chunks extension should be redundant once this is released. The method-name collision you hit is tracked separately in #1734.
stack72 commented 10/6/2026, 6:32:08 PM
Thanks @esteban for reporting this! We shipped: Stop AWS enrichment custom methods from writing their output into the CloudControl state resource. Each custom method declares a dedicated output resource (schema, lifetime, garbageCollection, stable instance key) in its enrichment config; the generator validates and emits those resources and writes to them. Applies to Bedrock retrieve (the reported bug), Events put_events and the four CloudFormation StackSet methods, which share the same template.. The fix has been merged and a release is on its way. We appreciate your contribution to swamp.
Sign in to post a ripple.