Claude has several models, and for each you can set how much it should think. Instead of going by impressions, we measured it: we gave the same work with mail and notes to seven different settings and scored the results against a key written beforehand. Here is what came out.
Karel DerflIn short
- Sonnet is enough for ordinary office work. Mail, notes, drafts, questions about your notes.
- The harder the task, the more a smarter model pays off. Finding connections, contradictions and things that don't add up. But a smarter model uses up the limit faster.
- Don't switch models in the middle of long work. After switching, the whole conversation has to be loaded again. It's cheapest after
/compactor in a new conversation.
What we tested
One task of the kind Miládka does every day:
- Go through five emails. Sort them, see whose move it is, and match them with tasks.
- Write a draft reply. The dictation in the brief deliberately contained a sentence that was no longer true. A good model notices.
- Answer five questions from the notes. One of them had no answer, because it wasn't in the vault. A good model says "I don't know" rather than making up a number.
- Find connections between mail, tasks and the calendar.
Each run got points out of ten. We scored against a key that was written before the test.
Results
| Model | Thinking | Quality (out of 10) | Time |
|---|---|---|---|
| Haiku | - | 4.5 | 2 min |
| Sonnet | normal | 8 | 4 min |
| Sonnet | a lot | 9 | 12 min |
| Opus | normal | 9 | 5 min |
| Opus | a lot | 9.5 | 11 min |
| Fable | normal | 9 | 5 min |
| Fable | a lot | 9.5 | 6 min |
How fast they use up the limit, from least to most:
- Haiku takes the least.
- Sonnet with normal thinking is the middle ground you can work with all day.
- Opus with normal thinking takes noticeably more than Sonnet.
- Sonnet and Opus with a lot of thinking take more than Opus with normal thinking. They take more steps, and at each one they read the whole conversation again.
- Fable takes the most, whether it thinks a lot or a little.
What follows from it
Everyone except Haiku got the facts right. Looking up in the notes when someone was last in touch or until when a link is valid is something every one of the bigger models can do. The difference only came where thinking was needed.
The difference is in judgement. It showed best on the draft with the untrue sentence. Opus and Fable both noticed the contradiction and wrote the draft according to the actual state of things. Sonnet found the contradiction too, but kept the version with the mistake as the main text and offered the correct one only as an alternative. Haiku didn't notice the contradiction at all and put the mistake into the email.
More thinking doesn't pay off for Sonnet and Opus on ordinary work. It adds half a point to a point, but the work takes two to three times longer and uses up the limit much faster. What it found on top were deeper things: an outdated profile of a person, or a rule that contradicts another one. Useful when tidying up, not for everyday mail.
Haiku isn't enough for work with Miládka. It's the fastest and uses up the least of the limit, but it only works with what it sees in the email and doesn't connect it with the notes. It made up three small details.
Fable is the smartest, but it uses up the limit the fastest. With normal thinking it's as good as Opus.
Which one to pick
- Sonnet for an ordinary day. Mail, notes, drafts, questions. The mistakes it makes are in judgement, not in facts, and you read drafts before they go out anyway.
- Opus when the result matters. An important email, a decision, finding connections across projects. It uses up the limit faster, but you get better judgement.
- More thinking for tidying up. When you want Miládka to go through your notes and find what doesn't add up or has gone stale. For everyday work, leave it on normal.
- Not Haiku. For work with mail and notes it's too little.
When to switch models
The model and the level of thinking are changed at the bottom of the window, next to where you type. Click the name of the model and choose another one.

Right next to it is the level of thinking (Effort). With the slider you set it from faster to smarter.

In the terminal, the /model command does the same.
After switching to another model, the whole conversation is loaded again. Claude keeps recently read text in a cache, but every model has its own. The new model has to read the conversation from the start, and that is the most demanding part of the work. We measured it in a conversation of about 126 thousand tokens: the first reply after switching had to load almost a hundred thousand tokens again. An ordinary reply in the same conversation loads a few hundred.
You pay for it only once. The following replies on the new model run normally. And if you go back within an hour to the model you were on, its cache is still there and going back costs almost nothing.
So:
- Switch to another model after
/compactor in a new conversation, when the conversation is short and loading it costs little. - Longer demanding work (tidying up notes, a big document) is better started on the smarter model right away.
- Change the level of thinking any time. The conversation isn't loaded again because of it, we measured that too.
What to keep in mind
- It was one test on one vault. The results show a direction, not an exact ranking. For you it may come out differently depending on the work you do.
- The order in using up the limit is an estimate. Anthropic doesn't publish how much each model takes from the subscription limit. We go by how much text each model read and wrote in the test, and by how expensive each model is outside the subscription.
- There is one limit for the whole account. All conversations draw on it, not just Miládka. On some subscriptions Fable also has its own weekly limit.
Why the limit runs out at all and what gets read in a conversation is in the article Why the limit runs out.