Jev Reliability
@vcjdeboer/jev-reliability · v2026.09.19.2
Is this Jev question safe to build on? A pre-flight check for TypeSafe System One questions. Runs a fully crossed item x framing x repeat grid against the API and reports what you need settled before putting a model's number behind an `if`: repeatability (does an identical request give an identical answer), framing sensitivity (does rewording move the answer, and by how much more than plain noise), resolution (can the question tell your items apart, or does the scale saturate into unrankable ties), threshold stability (does a fixed cut-off flap on identical input), and answerability (is it confidently rating things that have nothing to rate). Works on all three primitives — score for rating, noul for gates, choice for routers. The headline is a decision flip rate: if you reran this, how often would the decision your code acts on come out differently? The report adds a nested variance decomposition that separates repeat noise from framing sensitivity from real differences between items, fitted by a Gibbs/slice sampler in TypeScript so it needs no Stan, R or Python, and it withholds the posterior when the convergence gate fails. It does not take your paraphrases on trust: it hashes every request body and refuses to count a perturbation that produced a byte-identical request.