Evidence
Everything claimed for this protocol should be checkable, including the parts that count against it. The specification is public, the scoring constants are normative and quoted from source, and the weaknesses below are the ones conceded in the whitepaper rather than the ones we are comfortable admitting.
What you can check
- The normative specification and the reference engine are public under Apache-2.0, so the admission rules and the scoring are readable rather than described. Read the specification.
- The whitepaper quotes its scoring constants from source rather than restating them, so a number on the page and a number in the engine cannot drift apart. Read the whitepaper.
- Every finished record shows its refused submissions with the rule that caught each one, so you can see what the filter actually did rather than trusting that it ran. Read a finished record.
Known limits
- The framing phase produces a resolution unopposed
- A framer turns a question into a resolution and nobody argues with it. Whoever controls the wording has a real influence on the outcome. The workbench is the mitigation: a buyer settles the resolution themselves, sees the alternatives, and the framing agent is skipped entirely on the paid run.
- The evidence ladder prices some questions wrongly
- The ladder rewards systematic reviews and large trials. A question whose honest answer rests on case studies and expert judgement will score low for reasons that are about the ladder rather than about the argument. The workbench says so before you pay, and the caveat is stamped onto the finished record rather than shown once and forgotten.
- A single arbiter is a single point of failure
- The deepest conceded weakness. One model rules on admissibility, on every challenge, on relevance, and writes the verdict rationale. Panel mode is the answer: one frozen record re-judged by arbiters from different companies, with the disagreements published claim by claim rather than averaged away.
- Independence is worthless without competence
- An arbiter that cannot hold the rubric does not disagree independently, it disagrees randomly, and an incompetent arbiter's confusion is indistinguishable from an independent arbiter's disagreement. This is why the roster is frontier models from three companies rather than cheap models from one aggregator, and why any track whose arbiter had a failed call is excluded from every statistic rather than counted as agreement.
- A verdict is about the record, not about the world
- The protocol adjudicates what was argued and cited in front of it. A side that had better evidence available and did not bring it loses, correctly, and the record will not say the world disagrees. That is a feature for a decision trace and a limitation for anyone reading it as a fact-check.