Building24 June 20265 min read

Using Voice Input to Develop Early Product Ideas

By Forenta Team · updated 7 August 2026

In this article

An idea can feel clear while walking and become difficult to formulate when an input field opens. Writing requires the user to formulate, spell and type at the same time.

Research on working memory helps explain that difference.

The cognitive cost of writing and speaking

Bourdin and Fayol directly compared speaking and writing in 1994. Participants produced the same content orally and in writing.¹ Written production required more capacity, partly because spelling and the motor act of producing words consumed working memory.

Kellogg's model of writing describes the same squeeze from the inside: formulating, executing and monitoring all draw on the same limited pool, and translating an idea into language carries the heaviest load.² Worth being precise about, because Kellogg studied writing. He did not study speaking, and a model of writing cannot on its own establish that speaking is better.

Flower and Hayes described writing as a recursive process of planning, translating and reviewing rather than a simple transfer of a finished thought. Olive's later review connects those processes explicitly to working-memory limits. Together, they support an interface that separates capture from revision, without proving that speech is always the better capture method.

Why writing requires more working memory
Writing
Decide what you want to say
Formulate it at the same time
Both tasks use working memory
Details are easily lost
Speaking
Decide what you want to say
Say it directly
Keep more attention on the content
Structure the text afterwards

Source: Bourdin and Fayol (1994), International Journal of Psychology, who ran this comparison directly. Kellogg (1996) describes the load inside writing; he did not study speaking.

The limits of thinking aloud

Here is the part that usually gets left out of articles like this one. Schooler, Ohlsson and Brooks found in 1993 that verbalising while you work impairs performance on insight problems, the ones where the answer arrives whole rather than by steps, while leaving analytic problems unaffected.³ Putting a half-formed intuition into words can push out the non-verbal version that was about to produce the answer.

So the honest claim is narrower than the popular one. Speaking helps when you already half-know what you mean and need it out of your head. It does not help when the answer still has to arrive.

For a project description that boundary is comfortable. Describing a product you have been thinking about for weeks is articulation, not insight. Deciding whether to build it at all is a different task, and one you should probably not do while dictating.

How Forenta handles voice input

The feature contains three separate steps, each with a different owner.

The transcription is handled through the browser's speech-recognition interface rather than by an audio service operated by Forenta. Quality, availability and processing therefore depend on the browser and its configuration. Browser implementations may use an external recognition service.

Nothing is rewritten while the user speaks. The transcript appears first. The user may then ask Flux to structure the text by pressing a button.

The user retains control of the final text. The result screen shows the structured version next to the original transcript and asks which version should be used. Flux therefore cannot silently replace the source text or remove a material detail without review.

From a spoken idea to an assessable brief
Say the rough version
No structure needed
Transcribed
Speech becomes text
You tighten it
Cut, name, make concrete
Ready for assessment
Specific enough for a substantive judgement

The last step stays yours. The engine assesses what you submit, not what you meant.

Why the wording of the description matters

Forge reads what you submit, not what you meant. It does not fetch a URL and it does not clone a repository, so the text is the whole input. A description that says “a compliance tool for mid-market SaaS, and we have SOC 2 in scope for Q4” gives it something to work with. “A B2B platform for compliance” does not, and no amount of engine quality repairs that.

Speech can help capture concrete details before editing begins. The user must still revise the transcript for completeness and precision.

What this does not fix

Dictation needs a room where you can talk, and a tolerance for a first version that reads badly. It is poor at precise technical wording, exact names and anything where the spelling matters, because that is the part your browser is guessing at.

No cited study establishes that spoken product descriptions are more specific than typed descriptions. That remains a plausible mechanism rather than a measured product effect, so Forenta does not present it as proven.

A first version enables revision

What voice input changes is the starting point. A messy spoken description is something you can cut, sharpen and argue with. An empty field that stays empty is not.

Is speaking always better than typing for early ideas?

No. It helps when you already roughly know what you want to say and the cost is getting it out. Schooler, Ohlsson and Brooks showed that verbalising can impair insight problems, where the answer has not arrived yet. Use speech for articulation, not for the moment of figuring it out.

Where does the transcription happen?

Through the browser's speech-recognition interface. Behaviour and availability differ by browser, and the browser implementation may use an external recognition service. Forenta does not operate the audio-recognition service in this flow.

Does Flux rewrite what I said without asking?

No. Structuring is a button you press. In the app you then see the shaped version next to your raw transcript and choose which one is used.

Back to Journal