The most common question we get: why is my limit running out when I've hardly done anything? The answer usually isn't in what you ask Miládka to do, but in how AI works and how long the conversation you talk to her in is. We have spent dozens of hours on making her economical. Here is what we measured and what we do about it.
Karel DerflClaude reads everything every time
Claude doesn't remember anything between replies by itself. To answer you, it reads the whole conversation again from the start every time: Miládka's rules, everything you have written to each other, and everything it has read in files and mail along the way.
The subscription limit is counted by how much text Claude reads and writes. It's measured in tokens, a token is roughly a piece of a word.
Text Claude read recently is set aside and reading it a second time is much cheaper. But cheaper doesn't mean free, and when the conversation is long, a lot gets read.
Where you see it
Next to the model picker at the bottom of the window there is a small circle. When you hover over it, it shows how long the conversation is and how much of your limits you have used.

- Context window is the length of the conversation. In the screenshot almost 200 thousand tokens out of 300 thousand, at which the conversation condenses itself.
- 5-hour limit renews after five hours.
- Weekly is the weekly limit. That's the one that runs out sooner.
We build Miládka to be economical
Most of what Miládka does could be built more simply: let the AI watch and read everything by itself. It works, but then the limit disappears even when nothing is happening. We took the longer road. We built and measured every part again and again until the AI only did what AI is really needed for, and we spent dozens of hours on it. Where an ordinary program is enough, we leave the AI out on purpose.
- Watching mail without AI. A small program on your computer looks at your mail, every few minutes if you like. It wakes Miládka only when something arrives. When nobody writes, it takes nothing from the limit. If the AI checked the mail itself, every empty check in a long conversation would cost over a million tokens. For us that used to be twenty-two times a day.
- WhatsApp the same. A program watches the messages, Miládka wakes up only for a real message.
- What's done isn't read again. Mail Miládka has gone through gets a label, and next time the program leaves it out. So Miládka doesn't keep reading the whole inbox, only what is new.
- One step instead of forty. When Miládka labels forty emails after a pass, she does it in one go. Every extra step would mean reading the whole conversation again.
- What a program can do, a program does. A program adds the signature to an email, and a program makes the backup too. The AI doesn't touch it, so it takes nothing from the limit.
- Watching only when it makes sense. Mail can be watched only during working hours, for example.
- Short rules. The rules are read with every reply, so every extra sentence is paid for again and again. That's why we keep them as short as possible, and what isn't needed all the time, such as how to write an email, is loaded only when Miládka is writing one.
- She condenses the conversation herself. When the conversation grows to about a third of what Claude can hold, Miládka summarises it. What matters stays in the notes, the rest gets shortened.
The result is that Miládka uses up the limit only when she is really working. When nothing is happening, she doesn't.
What takes the most from the limit
We gave the same work with mail and notes to seven different models and measured how much each read and wrote. How they did in quality is in the article Which model for Miládka. The usage split like this:
- The start of a conversation. The rules and the descriptions of tools that are loaded before you write your first word. For us 80 to 130 thousand tokens, and for short work more than half of the whole usage. In a long conversation it spreads out.
- The number of steps. Every opening of an email or a file is a step, and every step reads the whole conversation again. A model that took thirty steps instead of twelve read more than three times as much.
- Writing replies. Noticeable, but the smallest of the three.
More thinking takes more mainly because of the steps. Claude then doesn't just write more, it mainly checks and opens more things. That's why you don't need it for ordinary work.
Switching models loads the conversation again
Each model sets text aside separately. When you switch to another one in a conversation that's under way, it has to load the whole thing again. We measured it in a conversation of about 126 thousand tokens: the first reply after switching loaded almost a hundred thousand tokens afresh, an ordinary reply loads a few hundred.
You pay for it once, the following replies run normally. Going back to the original model within an hour costs almost nothing, and neither does changing the level of thinking.
So change the model after /compact or in a new conversation, when the conversation is short.
There is one limit for the whole account
All your conversations with Claude draw on it, not just Miládka. When you work on something else next to her, the limit goes down for both.
A smarter model uses up the limit faster, even when it does the same work. Exactly how much, Anthropic doesn't publish. But the order is clear: Haiku the least, then Sonnet, Opus, and Fable the most. On some subscriptions Fable also has its own weekly limit, in the screenshot above that's the "Weekly · Fable" line.
What you can do
/compactafter a big job. Type it into the chat and the conversation condenses right away. Handy after a long tidy-up, processing documents or a big search.- Don't start a new conversation because of the limit. Condensing is enough. A new conversation loads the rules again and briefly stops watching mail and WhatsApp, which starts again with your first message.
- Change the model at the start, not in the middle. Best right after
/compact. - Normal thinking for ordinary work. Leave more thinking for tidying up and complicated things.
- Don't give Miládka everything at once. A huge file or hundreds of emails at once stay in the conversation. Better in parts.
- Watching only where you need it. Every wake-up reads the whole conversation. Waking Miládka for every message from every group isn't worth it.
Why does my limit run out even though I'm hardly doing anything?
/compact helps.