Every answer you get costs something. Most people find that out the week they run out. This is how the meter works, the handful of habits that stretch it furthest, and how to pick the right engine for the job instead of running the big one all day.
What you are actually spending, where to see it, and the one mechanic that explains almost every surprise bill.
Six habits that do the real work. None of them depend on which AI you use.
Three engines. When to reach for each one, and what it costs you to guess wrong.
The control next to the model name that almost nobody touches.
How a team seat works, and what to do the day you run out anyway.
Parts 1 and 2 work anywhere. Parts 3 and 4 are the Claude-specific settings, because that is what we all have.
What you are spending, and why it adds up faster than it feels like it should
A token is about three quarters of a word. You already know that. What nobody covers is that tokens are the unit your plan is metered in, and four things move the meter:
Those are the four dials. The rest of this session is which ones are worth turning.
Bottom right corner. That little circle fills as you use it. Click it and it tells you the truth: how much of the window you have left, and exactly when it resets.
Check it on Monday, not on Thursday when it is already blinking at you.
People think they have one budget. They have two, and they run out of the short one first.
Use it hard for an afternoon and you can hit this by three o'clock. It refills a few hours later on its own. Annoying, not fatal.
This is the one that ruins a week. It resets at a fixed time assigned to your account, not on Monday morning because you want it to.
Find your reset time once and write it down. Planning the heavy work right after it is free money.
Every time you send a message, the AI re-reads the entire conversation from the top before it answers. Not the last thing you said. All of it.
It has no memory between turns. The only way it knows what you were talking about is that the whole transcript gets handed back to it, every single time.
This is not a Claude thing. Every chat assistant on the market works this way.
Same question, same model, different point in the conversation:
Which means the single most expensive habit in the building is typing "thanks, that's perfect" at the bottom of a sixty message thread. That one word re-read the whole thing.
Second half of the same mechanic. Every document you attach and every file it opens stays in the conversation for the rest of the session. Not just for the answer you wanted it for.
So the cost of opening the wrong thing is not paid once. It is paid again on every turn that follows it.
Six things to change. None of them care which AI you use.
This is the whole session in four words. If you change nothing else, change this.
People keep one giant chat open for weeks because it feels like a relationship. It is not. It is a receipt that keeps getting longer.
Before you close a long one, ask for the handoff. Then paste it into the new chat as your first message.
Two hundred words of summary replaces sixty messages of history. You keep the thinking and drop the freight.
Five separate questions means five full re-reads of everything above them. One message with five questions in it means one.
Same result. One turn instead of five. Think it through before you type, the way you would before you called somebody.
If you paste the same price sheet into a new chat every morning, you are paying for that price sheet every morning.
This is the official advice too, and it is the one people skip because setting up a project takes four minutes.
Not all files cost the same. Two are much worse than people expect:
A PDF gets read as text and as a picture of every page. You pay for both. A long scanned document can eat a serious chunk of a window just sitting there.
Images are not cheap. A screenshot of a spreadsheet costs far more than pasting the eight numbers you actually care about.
Copy the paragraph. Paste the rows. Text is the cheapest thing you can hand it, and usually the clearest.
Before you attach a fifty page document, ask whether you need page 31 or the whole binder. Usually it is page 31.
Left alone, it writes long. Long answers cost you twice: once to produce, and again on every future turn when it re-reads them.
The fastest saving in this deck is a six word instruction at the end of your prompt.
Connected tools are not free. Every tool switched on gets described to the model before it answers, whether you use it that day or not.
Official guidance calls tools token intensive. Treat them like shop lights. On when you are in the bay, off when you leave.
These are not equal. Ranked by what they actually save:
| 1 | New job, new chat | Biggest saving available to you, by a distance |
| 2 | Right model for the job | Several times the cost for the same question, Part 3 |
| 3 | Batch your questions | Turns five re-reads into one |
| 4 | Projects instead of re-pasting | Stops you paying for the same file daily |
| 5 | Ask for shorter answers | Cheap to do, compounds quietly |
| 6 | Turn off unused tools | Small per message, real over a week |
| 7 | Do not switch engines mid chat | Throws away the cache and reprocesses everything, Part 4 |
One and two are most of it. The rest is tidying.
Three engines. You would not run the big one to go get coffee.
Every AI company sells a range. Claude's is three named tiers, and the naming never tells you what they are for. So:
Quick lookups, tidying a list, pulling fields out of text, simple rewrites. When you want it now and the job is not hard.
Where most work belongs. Drafting, editing, summarizing, everyday analysis. For a lot of people this is the only one they ever need.
Genuinely difficult reasoning, dense analysis, long multi-step work, and anything where being wrong is expensive.
Start on Sonnet. Move down for the easy stuff, up for the hard stuff. That is the whole strategy.
The same conversation drains your week at very different speeds depending on which engine you left it on.
Treat those bars as direction, not gospel. The exact ratios move as models change. The point that does not move: an identical chat can cost several times more on the top tier, and most chats do not need it.
| Reformat this list | Haiku |
| Pull the dates out of these emails | Haiku |
| Draft the customer follow up | Sonnet |
| Summarize this meeting | Sonnet |
| Build the weekly scoreboard | Sonnet |
| Work out why the numbers disagree | Opus |
| Plan a change that affects the whole team | Opus |
Pick before you start, and check what it is set to: the model and the effort both carry over from whatever you used last time. Plenty of people run the top tier all week without meaning to.
The control sitting right next to the model name that almost nobody touches
Separate from which model you picked, you can set how much work it puts in. Click the model name, then Effort. Five settings:
High is the default on most models, and it is lit above because it is where you already are. The useful move is usually stepping down, not up.
Routine work. Reformatting, quick questions, tidying, anything where you already know what good looks like. Official guidance says these stretch your usage further, and for this kind of job you will not see the difference.
The default, and the best balance of quality and speed. If you are not sure, you are fine here. Most people should leave it alone and spend their attention on Part 2.
Genuinely hard reasoning, long multi step work, the analysis you are going to make a decision on. Deliberate choices for a specific job, not a setting you leave on.
Set it at the top of the chat, the same as the model. The next slide is why flipping it halfway through is not the free move it looks like.
Continuing a conversation is normally cheap, because the history is cached. It does not get read from scratch every time, it gets read from a shortcut.
Change the model or the effort partway through and that shortcut is gone. The whole conversation gets processed again at full price.
So decide at the top of the chat. And if you genuinely need a different engine halfway through, that is usually a sign it became a different job, which means a new chat anyway.
You find these settings, you think "well I want the best answers," and you put the top model on max effort and leave it there.
Higher is not better. Higher is more thorough, slower, and it spends your week faster. For a task that did not need it, you paid a premium for an answer you could not tell apart.
Spend the top of the range where being wrong is expensive. Spend the bottom everywhere else. That is the entire skill.
Do not get so careful that you start doing the work by hand to save tokens. That is not thrift, that is just your afternoon.
The goal is not using less. It is not wasting the budget on the easy things so it is there for the hard ones.
Anthropic publishes their own version of this for people running long working sessions. It is the same four things:
| 1 | Keep them short | Every turn re-sends everything before it. The longer the session, the more each turn costs. |
| 2 | Limit your context | Everything it reads stays in the conversation. Only let in what the job needs. |
| 3 | Match the model to the task | Model and effort multiply everything else, and both carry over from last time. |
| 4 | Preserve your cache | Switching model or effort mid conversation reprocesses the whole thing from scratch. |
If you remember the card and nothing else from today, you are most of the way there.
How a team seat works, and what to do the day you run out anyway
Worth saying plainly, because people worry about the wrong thing:
So there is no reason to be shy with it. There is a reason not to waste it.
It happens, usually in a week that earned it. Your options, in the order worth trying:
Tell somebody if you are hitting it every week. That is useful information about how the work is going, not a confession.
Click the circle in the bottom right. Note how much is left and what time your week resets. Put that time in your calendar.
Find your longest running chat. Ask it for the handoff summary, paste that into a fresh chat, and close the old one.
Pick one routine job you do most days. Run it on the small model at low effort. If you cannot tell the difference, that is your new default for that job.
Nothing here needs permission or setup. All three are settings you already have.
New job, new chat. Then match the engine to the job instead of running the big one all day.
Everything else in this session is a refinement on those two. Get them right and you will stop thinking about limits at all, which is the actual goal. You are not trying to use less. You are trying to have it there on the day it matters.