Brief Sanitize
Sanitizer for untrusted text (email bodies, calendar invites, scraped pages, third-party payloads) before it is fed to an LLM prompt. Strips HTML to plain text via a real parser (sanitize-html), decodes entities, strips Unicode obfuscation characters (NFKC-normalized, \p{Cf} + variation selectors), defangs http(s)/javascript/data/protocol-relative URLs, truncates, and wraps each item behind a random per-call nonce delimiter (Microsoft "spotlighting", arXiv:2403.14720) so a downstream prompt can safely quote it as data rather than instructions. Defense-in-depth against prompt injection: no LLM, no network, no secrets.
Resources
Stable: prompt-injection sanitizer, 3 adversarial review rounds survived, 88 tests, 14/14; verified in the daily-brief pipeline.
- Has README or module doc2/2earned
- README has a code example1/1earned
- README is substantive1/1earned
- Most symbols documented1/1earned
- No slow types (deprecated)1/1earned
- Dependencies pass trust audit2/2earned
- Has description1/1earned
- Platform support declared (or universal)2/2earned
- License declared1/1earned
- Verified public repository2/2earned