“Reasoning” models, unjargoned: when the slow answer is the better one
Some AI models stop and think before they answer. They’re slower, but much better at tricky problems. Here’s when to use one.
Reasoning models spend extra effort working through a problem before they answer. That costs time and money, and on the right tasks it’s worth every second.
Reasoning models trade test-time compute for accuracy. Route by task difficulty, set thinking budgets deliberately, and measure on your own evals.
You may have spotted a setting called “think”, “reasoning” or “extended thinking” in your AI app. It switches on a : one that works through a problem step by step before it replies.
That makes it slower, sometimes by a minute or more. For a quick question or a friendly email, you don’t need it. For a tricky problem, such as planning a budget, comparing two contracts or untangling a messy spreadsheet, it is often much better.
A simple rule: if you’d want a person to take a moment before answering, turn thinking on.
A generates a hidden chain of working before its final answer. More working means more , so it costs more and you wait longer: the can run from seconds to minutes.
It pays off on multi-step problems: planning, maths, comparing options against several criteria, finding the flaw in an argument, writing and fixing code. It adds little to lookups and rewriting: summarising an article or polishing an email is just as good, and faster, without it.
Most apps now let you choose. Make the fast model your default, and switch thinking on when the first answer looks shallow or the task has more than two steps.
Reasoning models scale effort at inference time: they produce intermediate reasoning before the final answer, and accuracy on hard benchmarks rises with the effort allowed. Most APIs expose this as an effort level or a thinking budget, billed as output tokens.
In production the question is routing. A cheap, fast model handles triage and simple transforms; a reasoning model takes the hard cases: multi-hop questions, code changes, planning with constraints. Escalate when a confidence check fails or a validator rejects the first output, so the expensive path runs only when it earns its cost.
Measure before you decide. Build a small eval set from real tasks, compare quality against and cost at two or three effort levels, and remember that streaming and caching change what users feel more than the model choice does.
Try it yourself 2 minutes
- Ask your AI a planning question, such as a week of family meals on a £60 budget.
- Ask again with thinking or reasoning switched on.
- Compare the two answers. Which would you trust?
Try it yourself 2 minutes
- Paste a decision you’re weighing up, with its pros and cons.
- Ask:
Which option is better for me, and what am I missing? - Run it with and without thinking, and compare.
Try it yourself 2 minutes
- Pick five real prompts from your work: two easy, three hard.
- Run each at low and high effort, noting time and cost.
- Write down the point where extra effort stops helping.
Spotted a mistake? Tell us and an editor will check it.