Why the app uses more of your Claude plan than the terminal does - and what actually spends it

Modified on Tue, 4 Aug at 12:33 PM

You ran one morning sweep and Claude answered with usage limit reached. Or you have watched the numbers climb all week and written to us in the words several owners have used: insane token use, the app is obviously bleeding tokens, and that costs actual money to us. Or you ran the same work in a plain terminal session, saw it cost a fraction of what the app costs, and asked the fair question: why does the CLI use less? This article answers all of that straight, including the part where you are right and we are not going to argue with you.

Before anything else: are you paying per token, or paying a flat monthly fee?

This one question decides whether your problem is money or time, and the two have completely different answers. Answer it first, because most of what follows changes meaning depending on which one you are.

  • Signed in with a claude.ai plan (Claude Pro, about $20 a month, or Claude Max): nothing bills per token. That monthly fee is the entire spend. There is no meter running underneath it and no invoice arriving later. What you hit instead is your plan's usage window: heavy runs eat the allowance, and Claude pauses you until the window resets, roughly every 5 hours. Infuriating when you have work to do, but it is lost time, not lost dollars.
  • Running on an Anthropic API key (a key from console.anthropic.com active on your computer): you pay per token, in real money, for every run. Here the words that costs actual money to us are literally true.

If you are on an API key, switching to a claude.ai plan is the single biggest lever there is, and nothing else on this page comes close to it. You give up nothing by moving: while a key is the credential in use, the claude.ai account connections (Gmail, Google Calendar and Drive) are switched off anyway, so a plan hands function back rather than taking any away. The article Which AI runs Booked Solid? Your own Claude account (API keys are the exception) explains how a key quietly takes over and how to stop it.

Yes, the app costs more than a bare terminal session

One owner put it to us like this: this doesn't explain why the same processes run in the CLI take so many less tokens. He was right. Here is the real reason rather than a defence.

A bare terminal session starts almost empty. Every run inside the app carries real overhead before it does a single useful thing:

  • the operating instructions that make it behave like a booking department instead of a chatbot,
  • your business Brain: your leads, their history, and what has actually worked for you before,
  • the connector tools for mail, calendars and the rest.

That overhead is exactly what buys a draft that sounds like you and a follow-up that knows each lead's story. It is also paid for out of your usage window, and on Claude Pro you feel it quickly. The two heaviest surfaces are the daily sweep and chatting with Otto, because both carry the whole machinery on every single message.

What spends tokens, and what costs nothing at all

These cost nothing

  • Emailing support, and reading your own ticket. If you are wondering does emailing support use my tokens, the answer is no, not one. Your support tickets and Otto are two completely separate systems: the email thread runs on the Kivi Media help desk, and nothing in it touches your Claude plan. One owner typed even this is using up tokens! into a reply, and it was not. Write in as often as you like.
  • The dataset side of School Radar and Library Radar. The bundled data and the ranking that orders it are deterministic and offline. No tokens, no network calls, nothing contacts anyone. Tokens get spent only when Otto goes out for live web research and writes the drafts, and that happens only when you ask.
  • Two of the scheduled jobs. Back up my data and Weekly brain check run inside the app itself rather than through Claude, so they cost nothing to leave switched on.
  • Moving around the app. Opening the board, reading a lead, checking Today, looking at a past run. Those screens read your own files on your own computer.

These spend tokens

  • A Play you launch.
  • An Autopilot pass, or a scheduled task that runs through Otto.
  • Every chat message you send to Otto.
  • Live web research.
  • Writing drafts.

Drafting is much heavier than a plain chat, and this surprises people. To write one reply, Otto reads the gig, your notes and the whole thread, then thinks through several passes before a word appears, carrying your business Brain through all of it. That is why a morning that felt light can still reach the end of a Claude Pro window.

The levers, in the order that saves the most

  1. Move the daily sweep, or pause it and run it on demand. It is the single biggest consumer of the day. In Settings, find the section headed Runs on its own (described as "Things I do for you at a set time while the app is open"). Move it to a time you are not working, so the window has recovered before you sit down. Or switch it off there and press Run now on the days you actually want it.
  2. Start a fresh conversation for each topic. A long thread re-reads its whole history on every message, so it burns the window faster the longer it runs. Use Start fresh in Chat when you change subject. Talking things out with Otto all day is a completely legitimate way to work; it costs far less in short conversations.
  3. Do pure thinking in plain Claude. For talking something through that needs none of your leads and produces no drafts, use the Claude you already have, in the terminal or in the Claude app. Same plan, same window, but without the app's overhead it burns far slower. Keep the app runs for work that needs the machinery.
  4. Keep Thoroughness on Balanced or Fastest for routine work. Click Advanced on the right of the top bar, open the Thoroughness dropdown, and lift it to Most thorough only for pricing, outreach and locking in a booking. A fresh install already starts on Balanced, so you are not silently burning the most expensive option, but it is worth confirming. Detail is in Most thorough, Balanced, or Fastest: what the thoroughness setting changes.
  5. Ask for one thing at a time, and box the ask in. A broad instruction makes Otto read widely to work out where to start, and that reading is what costs. When you already know the target, name it and forbid the wandering: "For the [gig name] gig only, write the thank-you draft into that gig's folder. Do not read anything else first." That finishes in a fraction of what a broad launch costs. And do not ask Otto to re-review work that is already correct: a repeat pass costs about as much as fresh work.
  6. Archive leads that are done. When a lead is closed, lost, or a past one-off, archive it so future runs are not paying to re-read old history.
  7. Only after all of the above: if you run a daily sweep plus all-day chat, you are a heavy user, and Claude Pro's window will keep cutting you off however well the app is trimmed. Claude Max, about $100 a month, has a much larger allowance. You upgrade inside your Claude account at claude.ai and nothing changes in Booked Solid. That is the only time this article will suggest it; if you have decided against it, that is a fair call and the rest of this page still applies.

See where it actually went

You do not have to guess. Click Advanced in the top bar and open Token history. It shows what each run spent and where inside a long job the spending happened, run by run, labelled with the play or the task. If one play is eating far more than everything else, that is where to trim. From the same screen, Flag beside any run saves a support bundle with that one run's full token ledger attached, ready to send to us. It carries numbers and task names, unless you tick the box that adds that run's conversation, and the file is yours to open and read before anything leaves your computer.

When a run stops at the limit

If a job stops partway because the window ran out, the app now says so plainly at the moment it happens: your Claude plan's usage window is used up for now, that is the plan refilling and not a fault in the run. Nothing is lost. Autopilot and long jobs keep an append-only progress log as they work, so when your limit resets you run the same job again and it picks up from where it left off instead of starting over.

One thing to separate out: a pause whose own message says it stopped to protect your Claude plan and quotes a token number is Booked Solid's own safety brake rather than your plan running out. It is a different thing with a different fix, and the help center has a dedicated article on that pause and its Keep going button.

Running everything through Pearl in the terminal is not a downgrade

Some of you came from the terminal, watched the Facebook group say the app eats through tokens, and concluded there were things to fix before I want to run my business with it. That is a reasonable position, and here is the part worth keeping: the lean terminal path is the product, not a consolation prize.

The app does not replace your terminal, it wraps it. It opens your existing Booking HQ, puts you on the Terminal view (the tab labelled Power view) with your own setup and past sessions right where you left them, and keeps Pearl exactly as you taught her. The board, the chat and Otto's Autopilot sit on top as extras, never as a swap. Your Booking HQ, your past work and everything you have taught Pearl stay yours either way, so working the lean way costs you nothing extra and loses you nothing.

What has already shipped, and what is still running

An owner told us that until the token use was under control, the app is not usable even for testing at this point. That was fair, and it is why this is treated as work rather than a talking point.

  • 1.7.6 and later carry prompt caching (repeated context is no longer re-billed at full price), local file search (the app stops re-reading your files through the model), and tighter per-request context.
  • 1.11.37 fixed a real background drain. A pass that runs every few minutes to keep track of how your notes relate to each other used to ask the same failed question again on every pass, for as long as the app stayed open, which on one account meant hundreds of paid calls a day with nobody touching the keyboard. It now backs off after each attempt that gets nowhere and recovers on its own. If you have seen your usage climb with the app open and nothing running, update from booked.kivimedia.co/download before anything else.
  • Ongoing: a dedicated measurement programme that compares exactly what the app loads into every run against a bare terminal session and strips whatever does not earn its keep. It reaches you free in a normal update, with nothing to do on your end. There is no ship date to give you and no before-and-after numbers yet, and we would rather say that than invent either.

If it still does not add up

If your usage drains far quicker than the work seems to justify, even on Balanced with the sweep moved and short conversations, that is worth a look at our end. Take a screenshot of Token history, generate a support bundle (see Generate a support bundle and turn on support logging), and open a ticket at booked.kivimedia.co/support. Bundles carry numbers and task names only: no message text and no client details. The one exception is the box you can tick when you flag a run, which adds that run's own conversation, and you read that file before it leaves your computer. And to repeat the one thing worth repeating: writing to us costs you nothing on your Claude plan, so send as much detail as you like.

Was this article helpful?

That’s Great!

Thank you for your feedback

Sorry! We couldn't be helpful

Thank you for your feedback

Let us know how can we improve this article!

Select at least one of the reasons
CAPTCHA verification is required.

Feedback sent

We appreciate your effort and will try to fix the article