Engine13 May 20267 min read

Independent Review in AI Product Development

By Forenta Team · updated 7 August 2026

In this article

Poor decisions do not always result from missing information. A team can have enough information to feel certain without testing whether that certainty is justified.

As building gets cheaper, that becomes a more expensive mistake. When a team can go from idea to working prototype in days, the constraint moves. The question stops being whether something can be built quickly and becomes whether the reasoning behind it has been challenged by anything other than its author.

The persuasive force of a single analysis

A single author organizes the material into a coherent narrative. That can make the analysis appear more certain than the evidence warrants. Uncertainties omitted during drafting are no longer visible to the reader. This applies whether the author is a colleague, consultant or language model.

Tversky and Kahneman documented the mechanism underneath it in 1974: when people estimate, they anchor on an initial value and adjust too little, even when the anchor is arbitrary.¹ Once the first analysis in a discussion sets a frame, later thinking adjusts around it instead of starting again.

At group level the effect compounds. Janis studied a set of policy failures, Pearl Harbor and the Bay of Pigs among them, and named the pattern groupthink: cohesive teams under time pressure suppress dissent, converge early and stop testing assumptions.² The driver is social, not intellectual.

One analysis or several independent analyses
One analysis sets the frame
The first analysis defines the frame
Later judgements adjust to it
Blind spots remain
Independent analyses first
Each person assesses separately
Each analysis finds different gaps
Differences become visible

Concept: Tversky and Kahneman (1974), Science.

What research on group judgement shows

The findings are narrower than popular summaries suggest, but they still offer useful guidance.

Galton reported the well-known case in 1907. At a fatstock exhibition in Plymouth the year before, visitors paid to guess the weight of an ox. Of the cards, 787 were usable, and the middlemost estimate came to 1207 pounds against a true dressed weight of 1198, within roughly one percent.³ Galton's point was not that crowds beat experts, and he did not claim that. His point was about the shape of the distribution and the condition that produces it: the guesses were made independently, by people who could not see each other's cards.

That condition is the whole mechanism. Independent errors cancel. Correlated errors accumulate, and the group's apparent agreement becomes evidence that people read each other rather than evidence about the ox.

Why independent estimates reveal more
Discuss first
Estimate independently first

The jade dot with a ring is the true value. Concept: Galton (1907), Nature.

Woolley and colleagues reported in 2010 that group performance across varied tasks loads onto a general factor, correlated more with even participation and social perception than with the members' individual intelligence. Worth carrying the caveat that the finding has a contested replication record, with later work arguing that member intelligence predicts group performance more strongly than the original paper suggested. Treat it as suggestive rather than settled.

The Delphi method and early group influence

In the early 1950s RAND sought forecasts from expert groups without distortion from rank, dominance or premature consensus. Dalkey and Helmer published the resulting Delphi method in 1963. Experts first respond independently and anonymously, then see the distribution without names and may revise their judgement.

Cooke later formalised the harder half of the problem. Reliable elicitation means calibrating experts against quantities whose answers are known, weighting by track record and stating explicitly where uncertainty remains instead of forcing a clean answer. The output is not an average of opinions. It is a summary of what the evidence supports, including where it stops.

The Delphi method in four steps
Answer independently
Nobody sees anyone else's answer
Compare the distribution
Show the spread, not the names
Revise using evidence
Use the evidence, not the speaker
Report the result
Include remaining uncertainty

Source: Dalkey and Helmer (1963), Management Science.

Klein's premortem approaches the problem differently. Participants assume that a plan has failed and identify possible causes. Mitchell, Russo and Pennington described how treating an outcome as settled changes the explanations people generate. Percentages commonly attached to this effect are not adequately supported by the paper and are omitted here. The method remains useful because it makes criticism easier to express.

The evidence for structured disagreement

Modestly, and with caveats worth stating.

Schwenk's 1990 meta-analysis found advantages for devil's advocacy and dialectical inquiry over consensus approaches on decision quality, drawn largely from laboratory studies with student samples, and inconclusive between the two techniques. Nemeth and colleagues found in 2001 that authentic minority dissent stimulated more divergent thinking than one assigned devil's-advocate condition, with some evidence that role-played dissent hardened the majority instead.¹⁰ Neither result licenses a universal claim. Both point the same way.

Blinding does measurable work. Tomkins, Zhang and Heavlin ran a controlled experiment on the reviewing for a 2017 conference and found that reviewers who could see author identity favoured papers from prestigious institutions and famous authors, a preference absent from the double-blind arm.¹¹

More opinions do not automatically produce better decisions. The evidence supports independent first judgements, visible disagreement and summaries that preserve uncertainty instead of averaging it away.

How Forenta applies these principles

The implementation differs between the two product environments. Forge analyses text supplied by an individual user and returns a single structured assessment. It should not be read as an expert panel or an independent audit.

Forenta for Business supports a controlled organizational process. Relevant perspectives are kept separate before they are brought together, material disagreement remains visible and the resulting advice is attached to a decision point. The public explanation is intentionally limited to these design principles. The exact orchestration is part of the internal system.

How a structured review runs
Separate perspectives
Reduce early anchoring
Bring the evidence together
Keep sources and uncertainty visible
Dissent recorded
Do not flatten disagreement
Controls outside the model
The model cannot bypass them
Human decision
An override requires a reason

The depth of the review follows the project stage and risk class. Exact orchestration remains internal.

Safeguards with operational consequences run outside the language model. An unresolved blocking risk cannot disappear because a synthesis sounds confident. A person may still override a recommendation, but the reason and the original advice remain in the record.

DisciplineIn the literatureApplication in Forenta
Independent first passDelphi, blind submissionRelevant perspectives begin separately
Summary without averagingCooke, calibrated elicitationThe synthesis keeps evidence and uncertainty visible
Disagreement survivesSchwenk, NemethMaterial dissent remains attached to the decision
Human authorityGovernance and accountabilityA person decides and must explain an override
How Forenta currently applies the principles described in the research, including the stated limitations.

Where this falls short

A structured process can reduce anchoring and premature consensus, but it cannot guarantee an accurate decision. AI-generated perspectives can still share blind spots, especially when they rely on similar models and the same source material.

Independent human judgement on real decisions remains the relevant test. Forenta has not yet published enough outcome data to claim that its method improves project success rates.

Forge does not look up missing facts. It reasons about the text the user supplies. A Forenta for Business pilot may use explicitly connected read-only context, but the analysis remains limited to the evidence made available.

Judgement is a process, not a moment

The finding is narrow and it holds. Independent analysis, collected before group influence, structured so disagreement survives, and summarised without averaging, produces better calibrated judgement than consensus-first thinking.

The practical version for anyone building something: the moment before execution is the cheapest one to discover a weak assumption. Every week of work built on an unexamined premise raises the cost of the correction. That is not overhead. It is the only thing that stops a bad premise from compounding.

Does Forge provide an independent expert panel?

No. Forge provides an AI-generated structured analysis of the text you submit. It supports a decision but does not replace independent expert review.

Can the model talk its way past a negative finding?

A blocking finding cannot disappear solely because the synthesis sounds confident. Operational safeguards run outside the language model. A person may override the advice, but the reason and the original recommendation remain visible.

Do multiple AI perspectives provide real independence?

Only partially. Separating roles can produce different reasoning, but shared models and source material can still create correlated errors. Material decisions therefore remain subject to human review.

References

  1. 1.Tversky, A., & Kahneman, D. (1974). Judgment under uncertainty: Heuristics and biases. Science, 185(4157), 1124–1131.
  2. 2.Janis, I. L. (1982). Groupthink: Psychological Studies of Policy Decisions and Fiascoes (2nd ed.). Houghton Mifflin.
  3. 3.Galton, F. (1907). Vox populi. Nature, 75, 450–451.
  4. 4.Woolley, A. W., Chabris, C. F., Pentland, A., Hashmi, N., & Malone, T. W. (2010). Evidence for a collective intelligence factor in the performance of human groups. Science, 330(6004), 686–688.
  5. 5.Dalkey, N. C., & Helmer, O. (1963). An experimental application of the Delphi method to the use of experts. Management Science, 9(3), 458–467.
  6. 6.Cooke, R. M. (1991). Experts in Uncertainty: Opinion and Subjective Probability in Science. Oxford University Press.
  7. 7.Klein, G. (2007). Performing a project premortem. Harvard Business Review, 85(9), 18–19.
  8. 8.Mitchell, D. J., Russo, J. E., & Pennington, N. (1989). Back to the future: Temporal perspective in the explanation of events. Journal of Behavioral Decision Making, 2(1), 25–38.
  9. 9.Schwenk, C. R. (1990). Effects of devil's advocacy and dialectical inquiry on decision making: A meta-analysis. Organizational Behavior and Human Decision Processes, 47(1), 161–176.
  10. 10.Nemeth, C. J., Brown, K., & Rogers, J. (2001). Devil's advocate versus authentic dissent: Stimulating quantity and quality. European Journal of Social Psychology, 31(6), 707–720.
  11. 11.Tomkins, A., Zhang, M., & Heavlin, W. D. (2017). Reviewer bias in single- versus double-blind peer review. Proceedings of the National Academy of Sciences, 114(48), 12708–12713.
Back to Journal