Consider two reviews with the same verdict and confidence figure. One is based on technical evidence. The other only has a short project description. The number looks comparable, but the underlying evidence is not.
Confidence should describe the basis of a judgement. In AI-generated analysis it can also reflect how firmly the answer is written. The outcome alone does not reveal that difference.
Why confidence needs context
Language models are not hopeless at judging their own certainty. Kadavath and colleagues showed in 2022 that large models are reasonably calibrated on multiple-choice and true/false questions, provided those are put to them in the right format.¹ The question “is this product ready to go live” is not a multiple-choice question. There is no answer key, no fixed range and no moment inside the run where the answer turns out to be right or wrong.
Outside closed question formats, the relationship between certainty and correctness becomes harder to establish. A polished analysis based on little material can sound more convincing than a cautious analysis based on much more evidence.
No finding is not the same as no risk
Altman and Bland made the distinction explicit in 1995: absence of evidence is not evidence of absence.² Their note concerned clinical research, but the methodological point also applies to product reviews.
A reviewer without access to code cannot establish whether a specific injection flaw is present. A report with no finding on that point is therefore incomplete, not automatically clean.
A control nobody performed does not belong in a report as passed. It belongs there as not established, and keeping those two apart is the system’s job, not the reader’s.
Recording the scope of the evidence
Forenta records which kinds of material were available for an organizational review. That may include the project description, decisions from earlier stages and context from a source the organization explicitly connected. Missing material remains visible as missing.
This distinction is derived from the material supplied to the review, not from a model's claim that it inspected something. A sentence about having reviewed an architecture is not itself evidence that the architecture was available.
When a judgement depends mainly on written context, the result should say so and express more uncertainty than a review supported by direct technical evidence. The public explanation stops there deliberately: exact internal decision rules are implementation details, not proof of quality.
Derived from the supplied review context, not from the model's account of what it claims to have read.
Comparable controls across reviews
Findings written as free text are hard to follow over time. Last month’s “authorisation needs attention” and this month’s “access control is thin in places” may be the same problem or two different ones, and nobody can tell from the words.
Stable control identifiers make successive reviews easier to compare. Each result can distinguish established, unresolved and not applicable points, so a later review can show what changed without relying on similar-sounding prose.
A positive finding without matching evidence is not allowed to read as established. The result keeps the original claim traceable and records why its status was adjusted.
Not applicable never blocks. Fail, partial and not established all do.
Readiness is constrained by unresolved blockers
An average is the wrong operation for readiness. Nine controls in order and one missing key rotation does not make a project ninety percent ready. It makes it not ready, with one thing to arrange.
Forenta therefore treats unresolved blocking controls as constraints on readiness. A point that has not been established does not count as approved. A point that genuinely does not apply does not block progress.
The result names the blockers. That gives the team a concrete next action instead of a percentage with no owner.
| Reading | Result | What you do with it |
|---|---|---|
| Average of the statuses | 80% ready | Nothing specific. The number has no owner. |
| Constrained by blockers | Prototype. Blocked by one control. | Establish that control, then run again. |
What this does not solve
None of this makes a judgement correct. It makes visible what the judgement rests on. A review with full evidence scope can still be wrong, and a capped review can be right.
Forge, the public analysis product, works with text supplied by the user. It does not fetch a URL or clone a repository. In a controlled Forenta for Business pilot, an organization may explicitly connect a bounded source as read-only context. Findings still remain AI-generated analysis, not direct observations by a human auditor.
The method has not yet been validated against a large dataset of project outcomes. Its reliability therefore has to be assessed through documented cases, human review and transparent limitations rather than through the interface alone.
A number is only useful when you know where it came from
Structured expert elicitation reached this conclusion long before anyone asked a model for a verdict. Cooke’s work on the subject is built around calibrating experts against known quantities and being explicit about where genuine uncertainty remains, rather than forcing a clean answer.³ The NIST guidance for generative AI systems points the same way, at documenting the basis of a claim rather than the claim alone.⁴ Its broader AI Risk Management Framework also treats measurement as part of an ongoing governance process, not as a self-sufficient score.⁵
The practical version is smaller than the literature. Ask any system that gives you a verdict what it was able to look at. If it cannot answer, the verdict is a sentence with a number after it.
What is evidence scope?
A per-review record of which kinds of evidence were actually available and which were missing. It is derived from the submitted material, not from the model’s description of what it claims to have read.
Why does not established block just as hard as fail?
Because not checked is not approved. Treating an unverified control as passed is how a readiness claim ends up resting on nothing, which is worse than no claim because it looks like one.
Does Forenta read my repository?
Forge does not read a repository. It analyses text that you submit. Within a controlled Forenta for Business pilot, an organization may explicitly connect a bounded source as read-only context. The available scope is shown in that environment.
References
- 1.Kadavath, S., et al. (2022). Language Models (Mostly) Know What They Know. arXiv:2207.05221.
- 2.Altman, D. G., & Bland, J. M. (1995). Statistics notes: Absence of evidence is not evidence of absence. BMJ, 311, 485.
- 3.Cooke, R. M. (1991). Experts in Uncertainty: Opinion and Subjective Probability in Science. Oxford University Press.
- 4.NIST (2024). Artificial Intelligence Risk Management Framework: Generative AI Profile (NIST AI 600-1).
- 5.NIST (2023). Artificial Intelligence Risk Management Framework (AI RMF 1.0).