# Why a model thinks before it answers > Modern models write pages of working-out before answering. That 'thinking' is computation bought with tokens: what it buys, what it costs, when it's wasted. Source: https://rfriedmann.de/blog/why-models-think-before-answering/ Published: 2026-09-14 · Track: learn · Level: Beginner Ask a current model something hard and you often see the same thing: a collapsed grey block labelled "thinking", filling up with the model talking to itself, before a single word of the real answer appears. Let me set this up. Wait, that's not right. Try the other way. It looks like the model deliberating. It is something more literal than that, and once you see the mechanism the whole feature stops being mysterious, including the part where it is sometimes a waste of your money. ## A model gets one pass per token, and no scratchpad Here is the constraint everything else follows from. An LLM produces [one token at a time](/blog/what-is-an-llm/), and each token costs one trip through the network: a fixed amount of arithmetic, the same for "the" as for the hardest step of a proof. There is no place to work something out privately. There is no variable it can hold a partial result in between tokens. So a model asked for a hard answer immediately has to start emitting the answer, with only the fixed compute of a single pass behind each word. Anything it has not worked out by then, it has to guess, and a model that guesses fluently is exactly how you get [confident nonsense](/blog/why-models-make-things-up/). Writing the working-out changes this in two ways at once. It buys more passes: a thousand tokens of reasoning is a thousand more trips through the network aimed at your problem. And it makes the intermediate results *readable*, because everything the model writes lands in [the context window](/blog/the-context-window/) and is re-read on every following token. The scratchpad the architecture does not provide, the model builds out of its own output.
The working-out is the scratchpad
[diagram omitted — see the page for the chart]
Same weights, same fixed compute per token in both rows. The bottom row simply spends more tokens before committing, and each one can see the ones before it. That is the entire trick: computation you can only buy by writing.
## Nobody prompts it into this any more The old version of this was a prompting trick: append "let's think step by step" and watch accuracy on maths problems jump. It worked, and it was fragile, and it depended on you remembering to ask. Current reasoning models have the habit trained in, and trained in a specific way. After the [usual stages](/blog/pretraining-finetuning-rlhf/), the model is put through reinforcement learning on problems where the answer can be *checked automatically*: maths with a known result, code that either passes the tests or does not, puzzles with a verifier. Generate many attempts per problem, reward the ones that end up right, repeat. Nobody tells the model what good reasoning looks like. The reward only looks at the final answer, so whatever style of working-out tends to arrive there gets reinforced. What emerges, across labs and model families, looks remarkably consistent: restating the problem, trying an approach, noticing the approach is wrong, backing up, sanity-checking the result at the end. Those behaviours survived because they *pay*, on problems where correctness is machine-checkable. That last clause is the important caveat, and it explains most of what follows. ## What it costs you Thinking tokens are tokens. They are not free, not fast, and not invisible to the rest of the system. - **You pay for them.** Providers bill them as output tokens, the expensive kind, and a hard question can run to thousands before the answer starts. - **They eat the context.** The working-out occupies the same [window](/blog/the-context-window/) as your documents and the conversation so far. In an [agent loop](/blog/how-an-agent-uses-tools-and-memory/), where the transcript is already the memory, this is the difference between a task completing and a task running out of room. - **They delay the first word.** This one is easy to feel locally. A 30B model on a single card decodes [around 300 tokens a second](/blog/serving-30b-models-rtx-5090/), so 2,000 tokens of thinking is about seven seconds of silence before the answer begins. On a slower setup, or with a model that likes to deliberate, it is much worse. Which is why every serious implementation exposes a knob: a thinking budget, an effort setting, or a hard on/off switch per request. Open models such as the Qwen3 family ship a mode flag you can flip per message. The knob is not a nicety, it is how you keep the cost proportionate to the question. ## When it earns its keep, and when it doesn't The pattern follows straight from how the habit was trained. Reasoning pays where there is a chain of steps that can go wrong, and where a later step can catch an earlier mistake.

Worth the tokens

Just latency

The right-hand column is worth dwelling on, because it is where the money goes. No amount of deliberation retrieves a fact the weights never held. Thinking about a summary of a text sitting right there in the context adds nothing the first pass did not already have. And on trivially easy questions, extended deliberation can actively hurt: the model talks itself out of a correct first instinct, which is a failure mode you will recognise from humans. ## The working-out is not a log One last thing, and it is the one people get wrong most often. The thinking block reads like an explanation of how the model reached its answer. It isn't. It is text, generated by the same next-token process as everything else, and there is no mechanism forcing it to correspond to whatever actually drove the final output. Research into this keeps finding the same thing: models can be influenced by a hint in the prompt and then produce reasoning that never mentions it, constructing a tidy justification after the fact. So read it as useful working-out, which it genuinely is. Do not read it as an audit trail. If your process needs to know *why* a decision was made, the thinking block is evidence of the same rank as the answer itself, which is to say it needs [checking](/blog/when-not-to-use-ai/) too. None of this makes reasoning models less impressive. It makes them legible: a model with no scratchpad, given the ability to buy computation by the token, and trained by grading only the destination. Spend the tokens where the problem has steps. Switch it off where it doesn't, and [ask clearly](/blog/how-to-prompt/) either way.