Engine31 July 20265 min read

How Evidence Shapes the Reliability of AI Product Reviews

By Forenta Team · updated 7 August 2026

In this article

Consider two reviews with the same verdict and confidence figure. One is based on technical evidence. The other only has a short project description. The number looks comparable, but the underlying evidence is not.

Confidence should describe the basis of a judgement. In AI-generated analysis it can also reflect how firmly the answer is written. The outcome alone does not reveal that difference.

Why confidence needs context

Language models are not hopeless at judging their own certainty. Kadavath and colleagues showed in 2022 that large models are reasonably calibrated on multiple-choice and true/false questions, provided those are put to them in the right format.¹ The question “is this product ready to go live” is not a multiple-choice question. There is no answer key, no fixed range and no moment inside the run where the answer turns out to be right or wrong.

Outside closed question formats, the relationship between certainty and correctness becomes harder to establish. A polished analysis based on little material can sound more convincing than a cautious analysis based on much more evidence.

No finding is not the same as no risk

Altman and Bland made the distinction explicit in 1995: absence of evidence is not evidence of absence.² Their note concerned clinical research, but the methodological point also applies to product reviews.

A reviewer without access to code cannot establish whether a specific injection flaw is present. A report with no finding on that point is therefore incomplete, not automatically clean.

A control nobody performed does not belong in a report as passed. It belongs there as not established, and keeping those two apart is the system’s job, not the reader’s.

Recording the scope of the evidence

Forenta records which kinds of material were available for an organizational review. That may include the project description, decisions from earlier stages and context from a source the organization explicitly connected. Missing material remains visible as missing.

This distinction is derived from the material supplied to the review, not from a model's claim that it inspected something. A sentence about having reviewed an architecture is not itself evidence that the architecture was available.

When a judgement depends mainly on written context, the result should say so and express more uncertainty than a review supported by direct technical evidence. The public explanation stops there deliberately: exact internal decision rules are implementation details, not proof of quality.

What the review could inspect
Project descriptionWhat you wrote downinspected
Earlier decisionsVerdicts from previous stagesinspected
Assessment criteriaThe checks relevant to the stageinspected
Connected sourcesBounded context connected by the organizationunavailable
Confidence adjusted: technical findings remain inference until matching evidence is available.

Derived from the supplied review context, not from the model's account of what it claims to have read.

Comparable controls across reviews

Findings written as free text are hard to follow over time. Last month’s “authorisation needs attention” and this month’s “access control is thin in places” may be the same problem or two different ones, and nobody can tell from the words.

Stable control identifiers make successive reviews easier to compare. Each result can distinguish established, unresolved and not applicable points, so a later review can show what changed without relying on similar-sounding prose.

A positive finding without matching evidence is not allowed to read as established. The result keeps the original claim traceable and records why its status was adjusted.

What happens when supporting evidence is missing
Control marked as passed
The result uses a stable control ID
Forenta checks the evidence
Was the relevant source inspected?
Status changes to not established
The reason is recorded

Not applicable never blocks. Fail, partial and not established all do.

Readiness is constrained by unresolved blockers

An average is the wrong operation for readiness. Nine controls in order and one missing key rotation does not make a project ninety percent ready. It makes it not ready, with one thing to arrange.

Forenta therefore treats unresolved blocking controls as constraints on readiness. A point that has not been established does not count as approved. A point that genuinely does not apply does not block progress.

The result names the blockers. That gives the team a concrete next action instead of a percentage with no owner.

ReadingResultWhat you do with it
Average of the statuses80% readyNothing specific. The number has no owner.
Constrained by blockersPrototype. Blocked by one control.Establish that control, then run again.
The same five controls, read two ways. Averaging hides the one that matters; constraining surfaces it.

What this does not solve

None of this makes a judgement correct. It makes visible what the judgement rests on. A review with full evidence scope can still be wrong, and a capped review can be right.

Forge, the public analysis product, works with text supplied by the user. It does not fetch a URL or clone a repository. In a controlled Forenta for Business pilot, an organization may explicitly connect a bounded source as read-only context. Findings still remain AI-generated analysis, not direct observations by a human auditor.

The method has not yet been validated against a large dataset of project outcomes. Its reliability therefore has to be assessed through documented cases, human review and transparent limitations rather than through the interface alone.

A number is only useful when you know where it came from

Structured expert elicitation reached this conclusion long before anyone asked a model for a verdict. Cooke’s work on the subject is built around calibrating experts against known quantities and being explicit about where genuine uncertainty remains, rather than forcing a clean answer.³ The NIST guidance for generative AI systems points the same way, at documenting the basis of a claim rather than the claim alone. Its broader AI Risk Management Framework also treats measurement as part of an ongoing governance process, not as a self-sufficient score.

The practical version is smaller than the literature. Ask any system that gives you a verdict what it was able to look at. If it cannot answer, the verdict is a sentence with a number after it.

What is evidence scope?

A per-review record of which kinds of evidence were actually available and which were missing. It is derived from the submitted material, not from the model’s description of what it claims to have read.

Why does not established block just as hard as fail?

Because not checked is not approved. Treating an unverified control as passed is how a readiness claim ends up resting on nothing, which is worse than no claim because it looks like one.

Does Forenta read my repository?

Forge does not read a repository. It analyses text that you submit. Within a controlled Forenta for Business pilot, an organization may explicitly connect a bounded source as read-only context. The available scope is shown in that environment.

Back to Journal