I measure what AI systems lose between two processing steps.
Independent researcher on structural stability in human–AI coupling and in agentic processing chains. Originator of the +1 Principle. Twenty-five years in banking and finance, most of them in risk controlling and model validation — which is where the question comes from: how do you validate a model that hands its output to the next model?
Across thousands of documented exchanges one pattern held: only ethically guided feedback sustained coherence over time, while unethical reinforcement collapsed it. The +1 Principle states that a system stays stable when each step adds a corrective increment instead of amplifying the previous one — and that this is measurable rather than merely desirable.
Naturalistic, single-subject observation — not curated, not staged. This is the empirical origin of the +1 Principle and the calibration base of the stability coefficient.
A controlled measurement campaign across four AI models from three families, five scenarios, twelve steps and five repetitions each. Every transition between two consecutive steps was measured against the chain's point of origin.
In one scenario a step invented a lending value of 250,000 euros and two steps later carried it forward as the loan amount applied for — 600,000 became 250,000 euros, and nobody noticed before the decision. That is the point: the object under review is the transition, not the answer.
Design, control chains, sensitivity analyses and the limits of what can be concluded are documented openly → full report.
There are two distinct questions, and they need two distinct instruments. Both are deterministic: no second AI model, no access to the system being measured, byte-identical results for identical input.
Measures one answer against the prompt it came from. What is examined is how the answer is built — hallucination, sycophancy, false authority and further markers of structural instability, combined into one value on a scale from 0 to 90.
Blind to anything already lost before this step. A perfectly built answer to an input corrupted earlier still scores well — correctly so.
Measures the transition between two consecutive structured work products against the chain's point of origin, and returns chain-level indicators — tipping point, volatility, end-to-end traceability. Four classes of structural deviation:
Blind to whether an individual step, in itself, overstates its certainty. A chain can be perfectly traceable and still built of confident guesses.
The two fail in opposite directions, which is why neither replaces the other. An answer can be flawless in itself while carrying a constraint that disappeared three steps earlier — that is what the campaign of 2026 documents. And a chain can be faultlessly traceable while every step quietly overstates what it knows — that is what the observation of 2025 documents. Covering both layers is the point of this research.
How the markers and classes are computed, weighted and thresholded is not published. That is deliberate — and it is also why the method can be verified without being disclosed: send logs of your own chains twice, and the curve comes back identical to the decimal.
10.5281/zenodo.1716696310.5281/zenodo.17166964A peer-reviewed paper on the 2026 campaign is in preparation.
Deployers must ensure effective human oversight. Effectiveness presupposes that it can be established. That gap is what this research addresses.
Contributed technical comments to the public consultation in 2026, including one verified error in the draft's OSCAL referencing.
The German two-pillar framework for AI test tools and test data. sReact is built along its governance requirements: separation from AI development, black-box access, versioning, tamper protection by checksum.
I think it could be a valuable insight.
The direction of this work was encouraged and influenced by James A. Yorke. He answered a stranger without an institution behind him, took the idea seriously, and pointed it toward something better: collecting real cases where AI feedback led to bad outcomes. That suggestion became the case register this research now rests on. I am more grateful for that openness than these few lines can carry.
Early encouragement also came from Dirk Helbing, who wrote that this perspective would be “particularly interesting for many people in the AI business” and worth substantiating “by a business and banking insider”.
If you work on evaluation, model risk, oversight or agentic reliability — in research, in supervision, or inside an organisation running such chains — I am glad to hear from you. Sending logs of your own chains is the fastest way to see what the method does.