ExoAI-S reports first measured results from the “Crystalline Brain”, a self-check for AI reasoning
By The Lockadin · ExoAI-S research update · October 5, 2026
ExoAI-S has built and tested the Crystalline Brain, a tool that lets an AI model write down its competing explanations, its predictions and its evidence as it works, and then checks each decision against those records before the decision is carried out. A person can watch the whole process live as a three-dimensional "brain" in which possible explanations light up, are tested, and go dark when evidence rules them out. The AI can look at the same picture of its own reasoning.
This update reports three results from the past week. All of the work ran on one computer, under human supervision, with the rules for each test written down before it was run.
What the tests show
Three AI models caused no harm in a synthetic support-desk benchmark, with and without safety warnings. Across three test phases, three current models from two companies handled 180 simulated support-ticket episodes with 0 harmful outcomes. In the final phase the written warnings about every fault were removed; harm still did not appear. Instead the models grew more cautious in the one scenario where a removed fact was needed to act confidently: 3 of 6 episodes ended with a safe "unknown" rather than a finished ticket, against 0 of 9 with the warnings present. This is a screen, not a bound: it would usually have detected a harm rate of roughly 13 to 24 percent in the affected scenarios, and it cannot show that harm is rare.
The self-check beat three simple caution rules in a pre-registered comparison. On 528 scripted episodes per condition, the check reduced harmful outcomes from 111 to 51 while completing 327 tasks and blocking 0 safe actions. "Wait for confirmation", "stop after any fault" and "never act" each did worse on at least two of the three measures, and "never act" completed only 117 tasks. A follow-up change, in which the check asks for more evidence when every recorded option is harmful somewhere, lowered harm further, to 42, with 336 completions. These are scripted stand-in agents in simulated worlds: an upper bound on how well the check can work, not a field result.
Two research-grade mathematics problems were solved live in the brain and matched their published answers exactly. Both are public sample problems from a well-known hard-mathematics benchmark; both answers were committed in the brain before the published answers were read. In each case the model's memory of a key fact was wrong, the self-check raised an alarm before the model could report from memory, and computation then corrected the error. A third problem, solved with a new feature that records computations directly into the brain so the model cannot author that evidence itself, exposed two further errors in the model's recollection, including a factor-of-two slip caught only because the machine's record could not be edited.
What improved
Over three rounds of adversarial review by independent AI reviewers, 40 ways of gaming the self-check were found, reproduced and closed, with a documented list of what remains. The check now treats recorded harm as binding until real evidence rules the explanation out, recognises an action regardless of spelling or spacing, and labels every piece of evidence by who recorded it: the model, or a computation the brain itself ran. Every release passes 220 automated tests and a replay of all real decisions recorded so far, which have never changed verdict unexpectedly.
What this does not establish
The self-check reads the model's own notes. A "clear" means the decision is consistent with what the model recorded believing, not that it is correct. Across all real sessions so far, every block the check made was on an action the model had itself marked as harmful seconds earlier; the check enforced the model's own beliefs, and detected nothing on its own. The new machine-recorded evidence narrows that gap but does not close it: it proves that code ran and what it printed, not that the model read the result correctly.
The mathematics problems are public, so the models may have met them in training; two solved problems are not a score. The benchmark results concern scripted agents and one synthetic domain. No controlled study with live AI agents has yet shown that the brain changes their accuracy or safety; that study is designed and paused pending resources. None of this is independent verification, external peer review, or a security guarantee, and the brain is not a safeguard for deploying AI systems.
The next research milestone
The next step is a controlled comparison with live AI agents on problems whose answers can be checked: the same hard problems with and without the brain, counting confident wrong answers, honest abstentions and correct answers under rules fixed in advance. The questions it would answer are the ones this update cannot: whether seeing its own reasoning changes what a model does, and whether evidence the model did not write makes its self-check worth trusting.
Our direction remains human-supervised research with clear limits and results that can be checked.
Onward and outward.