# Why a model thinks before it answers
> Modern models write pages of working-out before answering. That 'thinking' is computation bought with tokens: what it buys, what it costs, when it's wasted.
Source: https://rfriedmann.de/blog/why-models-think-before-answering/
Published: 2026-09-14 · Track: learn · Level: Beginner
Ask a current model something hard and you often see the same thing: a collapsed
grey block labelled "thinking", filling up with the model talking to itself, before
a single word of the real answer appears. Let me set this up. Wait, that's not
right. Try the other way.
It looks like the model deliberating. It is something more literal than that, and
once you see the mechanism the whole feature stops being mysterious, including the
part where it is sometimes a waste of your money.
## A model gets one pass per token, and no scratchpad
Here is the constraint everything else follows from. An LLM produces
[one token at a time](/blog/what-is-an-llm/), and each token costs one trip through
the network: a fixed amount of arithmetic, the same for "the" as for the hardest
step of a proof. There is no place to work something out privately. There is no
variable it can hold a partial result in between tokens.
So a model asked for a hard answer immediately has to start emitting the answer,
with only the fixed compute of a single pass behind each word. Anything it has not
worked out by then, it has to guess, and a model that guesses fluently is exactly
how you get [confident nonsense](/blog/why-models-make-things-up/).
Writing the working-out changes this in two ways at once. It buys more passes: a
thousand tokens of reasoning is a thousand more trips through the network aimed at
your problem. And it makes the intermediate results *readable*, because everything
the model writes lands in [the context window](/blog/the-context-window/) and is
re-read on every following token. The scratchpad the architecture does not provide,
the model builds out of its own output.
The working-out is the scratchpad
[diagram omitted — see the page for the chart]
Same weights, same fixed compute per token in both rows. The bottom row simply spends more tokens before committing, and each one can see the ones before it. That is the entire trick: computation you can only buy by writing.
## Nobody prompts it into this any more
The old version of this was a prompting trick: append "let's think step by step"
and watch accuracy on maths problems jump. It worked, and it was fragile, and it
depended on you remembering to ask.
Current reasoning models have the habit trained in, and trained in a specific way.
After the [usual stages](/blog/pretraining-finetuning-rlhf/), the model is put
through reinforcement learning on problems where the answer can be *checked
automatically*: maths with a known result, code that either passes the tests or
does not, puzzles with a verifier. Generate many attempts per problem, reward the
ones that end up right, repeat.
Nobody tells the model what good reasoning looks like. The reward only looks at the
final answer, so whatever style of working-out tends to arrive there gets
reinforced. What emerges, across labs and model families, looks remarkably
consistent: restating the problem, trying an approach, noticing the approach is
wrong, backing up, sanity-checking the result at the end. Those behaviours survived
because they *pay*, on problems where correctness is machine-checkable.
That last clause is the important caveat, and it explains most of what follows.
## What it costs you
Thinking tokens are tokens. They are not free, not fast, and not invisible to the
rest of the system.
- **You pay for them.** Providers bill them as output tokens, the expensive kind,
and a hard question can run to thousands before the answer starts.
- **They eat the context.** The working-out occupies the same
[window](/blog/the-context-window/) as your documents and the conversation so
far. In an [agent loop](/blog/how-an-agent-uses-tools-and-memory/), where the
transcript is already the memory, this is the difference between a task
completing and a task running out of room.
- **They delay the first word.** This one is easy to feel locally. A 30B model on a
single card decodes [around 300 tokens a
second](/blog/serving-30b-models-rtx-5090/), so 2,000 tokens of thinking is about
seven seconds of silence before the answer begins. On a slower setup, or with a
model that likes to deliberate, it is much worse.
Which is why every serious implementation exposes a knob: a thinking budget, an
effort setting, or a hard on/off switch per request. Open models such as the Qwen3
family ship a mode flag you can flip per message. The knob is not a nicety, it is
how you keep the cost proportionate to the question.
## When it earns its keep, and when it doesn't
The pattern follows straight from how the habit was trained. Reasoning pays where
there is a chain of steps that can go wrong, and where a later step can catch an
earlier mistake.
Worth the tokens
Maths, and anything with units or arithmetic
Code with real constraints to satisfy
Multi-step planning before acting
Puzzles, deductions, tricky edge cases
Just latency
Recall: facts the model either knows or doesn't
Rewriting, translating, changing tone
Summarising a document you supplied
Extraction and formatting
The right-hand column is worth dwelling on, because it is where the money goes. No
amount of deliberation retrieves a fact the weights never held. Thinking about a
summary of a text sitting right there in the context adds nothing the first pass
did not already have. And on trivially easy questions, extended deliberation can
actively hurt: the model talks itself out of a correct first instinct, which is a
failure mode you will recognise from humans.
## The working-out is not a log
One last thing, and it is the one people get wrong most often.
The thinking block reads like an explanation of how the model reached its answer.
It isn't. It is text, generated by the same next-token process as everything else,
and there is no mechanism forcing it to correspond to whatever actually drove the
final output. Research into this keeps finding the same thing: models can be
influenced by a hint in the prompt and then produce reasoning that never mentions
it, constructing a tidy justification after the fact.
So read it as useful working-out, which it genuinely is. Do not read it as an audit
trail. If your process needs to know *why* a decision was made, the thinking block
is evidence of the same rank as the answer itself, which is to say it needs
[checking](/blog/when-not-to-use-ai/) too.
None of this makes reasoning models less impressive. It makes them legible: a model
with no scratchpad, given the ability to buy computation by the token, and trained
by grading only the destination. Spend the tokens where the problem has steps.
Switch it off where it doesn't, and [ask
clearly](/blog/how-to-prompt/) either way.