What the study found
The calibration run produced a false-REFUTE operating characteristic of 0.046, which was consistent with the registered prediction of about 0.05. The abstract describes this as a formal clearance record for the test apparatus rather than a substantive empirical claim.
Why the authors say this matters
The authors say the record shows the instrument onus has been discharged for the test itself, not the underlying empirical claim. They also state that the load-bearing result is reproducible bit-for-bit by re-running the adjudication on the registered data, without trusting the multi-node pipeline.
What the researchers tested
The study reports a pre-registered Bayesian equivalence study of LLM factual-commitment retention on the retrieval axis, with an operating-characteristic calibration run. The apparatus was checked through a two-node adversarial pass on independent hardware, and the record includes four hash-pinned modules, a sealed golden configuration, an adversarial clearance ledger, three documentation conditions, and a stratified refusal report.
What worked and what didn't
The calibration run succeeded in producing a false-REFUTE operating characteristic of 0.046, with a Wilson 95% confidence interval of [0.0373, 0.0567] and a half-width of 0.0097, based on 1,803 converged replicates and a 1.21% refusal rate. The abstract says PIN-B refused-path count-identity and resume-equivalence under a mid-replicate crash were each verified, rather than assumed.
What to keep in mind
The abstract says this is not a substantive empirical claim and should be retired only by the registered confirmatory run on real data plus independent reproduction. It also presents verification provenance as method rather than evidence, and does not describe other limitations beyond that scope.
Key points
- The calibration run reported a false-REFUTE operating characteristic of 0.046.
- That value was consistent with the registered prediction of about 0.05.
- The result was based on 1,803 converged replicates with a 1.21% refusal rate.
- Two checks were verified on independent hardware: refused-path count-identity and resume-equivalence under a mid-replicate crash.
- The abstract says this is a clearance record, not a substantive empirical claim.
Disclosure
- Research title:
- Calibration run matches registered false-REFUTE rate
- Authors:
- John Richard Smith, SHAI / HATI
- Institutions:
- Symbiom (Czechia)
- Publication date:
- 2026-06-29
- OpenAlex record:
- View
Get the weekly research newsletter
Stay current with scholarly research without reading academic papers — one filtered digest, every Friday.