Research update: evidence checks and TriForge

ExoAI-S completed a small offline evaluation of an experimental evidence checker. All 12 synthetic test cases produced the expected results in two repeated rounds, including correctly reporting uncertainty when completion could not be established. Both rounds used the same cases, not independent trials.

TriForge brings together bounded task preparation, two-assistant review, and recorded evidence in a human-supervised research workflow. We are evaluating its usefulness and have not yet demonstrated improved coding efficiency.

A two-assistant review recommended testing separately authored evidence formats next: can the checker recognize valid work without accepting unsupported claims?

These are early engineering results—not proof of real-world security, a breakthrough, or improved efficiency. Further testing and comparison with established tools remain necessary.

Previous
Previous

A little light. A world of possibility.

Next
Next

ExoAI-S reports bounded progress on Guardian Receipts research