Early access
Methodology & Limitations
We tell you exactly what ClearRun does and does not do, so a score is never misread. Here it is, plainly.
What ClearRun measures
ClearRun measures honesty, which we define narrowly and specifically as confidence-justification alignment: whether the confidence an answer projects is actually earned by the reasoning and evidence it provides. A calm, well-hedged answer that supports its claims scores well. A confident answer that asserts more than it justifies scores poorly. The output is a 0 to 100 honesty score, a set of named failure categories (for example: unsupported claim, overconfidence, false precision, internal inconsistency), and, on request, a signed verdict.
What ClearRun does NOT do
- It does not fact-check. ClearRun does not verify claims against the world. A well-justified answer can still be factually wrong, and an honest answer can express appropriate uncertainty. A high score is not a claim of truth.
- It is not a correctness or safety guarantee. Do not treat a score as permission to act without your own judgment. It is one signal, not a warranty.
- It does not censor or block. ClearRun evaluates output; it never modifies or withholds it.
Properties of a verdict
- Deterministic. The same input produces the same verdict, every time. This is what makes a verdict reproducible and reviewable, unlike a probabilistic “LLM-as-judge” that varies run to run.
- Signed and verifiable. Each verdict is signed (Ed25519). Anyone can confirm a verdict came from ClearRun, unaltered, using the public key at
/api/verify/{id}. No account required to verify.
Validation status (in progress, stated honestly)
ClearRun is in early access. Its reproducibility and signature are true and demonstrable today. Its accuracy (how well the honesty score predicts real unreliability) is being validated now, and we will not claim it before we can show it. Our validation is pre-registered: we commit the datasets, metrics, and thresholds in advance, then run and publish the results, including where ClearRun underperforms. The study uses public, human-labeled benchmarks (such as HaluEval, RAGTruth, and TruthfulQA) and reports standard measures (ranking performance and calibration) against real baselines. Until that study is published, please read scores as an honesty-alignment signal, not a validated accuracy metric.
ClearRun is an Adepoco product. Questions: contact@adepoco.com.