You have probably had this moment. You are deep into something with an AI assistant, the work is going well, and then you glance at your usage. It has climbed somewhere you did not expect.
So you go looking for advice, and everyone gives you the same useless sentence. Use fewer tokens.
Right. Where?
Here is what almost nobody explains. There are only four places your tokens go. Four. And over the past year, four separate tools have been built to attack them, one each. They have peculiar names: RTK, Headroom, Caveman, and Ponytail. A router called 9Router bundles all four together, which is how most people who have heard of them heard of them.
Most people have not heard of them.
That is a shame, because the ideas underneath are genuinely clever, and three of the four are useful to you even if you never install anything. Two of them will change how you use AI tomorrow morning. One of them is quietly costing people money right now.
Let me show you the map first, because once you see it you cannot unsee it.
The Four Leaks
Every token you are billed for is doing one of two jobs. It is either something the AI reads, or something the AI writes. There is no third category.
Now split each of those into narrow and broad, and you have the whole map. Four boxes. Four techniques, one per box.
Read that bottom right box again. Makes the AI build less. Not say less, not read less. Actually produce a smaller thing.
Hold that thought, because it turns out to be the one that works.
Why this map is worth memorising: the next time someone sells you a way to cut your AI costs, you can ask which box it is in. If it cannot answer, it is not a technique. It is a vibe.
The Detour That Explains Everything
Before we get to the four, there is one piece of plumbing you need. It takes three minutes and it is the reason most people's attempts to cut their AI bill do nothing at all.
AI models have no memory between messages. None. So every time you send a message, the entire conversation gets re-sent from the beginning. Your fiftieth message drags all forty-nine previous ones along with it.
That would be ruinously expensive, so the providers invented caching. The repeated part gets stored, and every time it is read back you pay one tenth of the normal price.
Here is what that does to everything you are about to read.
Suppose a tool proudly trims 10,000 tokens out of something the AI was about to read. Big number. Feels great.
But if those tokens were going to be cache reads, you did not save 10,000 tokens' worth of money. You saved about 1,000. The dashboard shows you the first number. Your invoice reflects the second.
One user posted real telemetry from 26 days and around 1,300 prompts, and it puts the whole thing in perspective. Cache reads were about 97 percent of his tokens. The visible replies, the part you actually read, were 0.6 percent.
Burn this into your brain: tokens saved and money saved are different numbers, and the gap between them is roughly ten times. Every claim in the rest of this article gets measured against that gap.
RTK: The Wall of Text Problem
Picture an AI assistant running a test suite on your code. The tests run. And 1,842 lines of output come back, exactly as a human would see them in a terminal.
The AI has to read all 1,842 lines. You pay for all 1,842 lines. And of those, maybe three actually mattered: how many passed, how many failed, which ones broke.
That is the leak RTK was built for. It stands for Rust Token Killer, it is free and open source, and the idea is almost embarrassingly sensible. Sit between the command and the AI. Recognise what kind of output this is. Rewrite it into the short version before the AI ever sees it.
Those 1,842 lines become three. A file listing becomes a tidy tree with counts instead of a wall of permissions and timestamps. And if a filter cannot make sense of something, it passes the original through untouched rather than mangling it.
It is a good idea, cleanly built. So here is the uncomfortable part.
It made things more expensive
The team at JetBrains ran the test properly. Not a demo, not a dashboard: 425 billed trials across 86 tasks, around 320 dollars of real compute, comparing actual invoices with the tool on and off.
| What they measured | Result |
|---|---|
| Advertised saving | 60 to 90% |
| Actual cost change | 7.6% more expensive |
| Quality | No difference |
Three reasons, and they are worth knowing because they apply to this whole category.
It only sees part of the traffic. Coding assistants have their own built-in ways of reading files, and those skip RTK entirely. It was catching about a fifth of what the AI read, which caps the best possible saving at roughly three percent before you start.
It takes credit for savings that were never going to be billed. Assistants already cut off absurdly long output on their own. The benchmark found 190 giant file reads where RTK logged half a million saved tokens each, on text the AI was never going to receive anyway.
And it prices those savings at full rate when most of them were cache reads at a tenth. That is the gap from Section 02, showing up exactly where you would expect.
Now, credit where it is due, and there is real credit here. The RTK team publicly accepted the findings and conceded a realistic ceiling of three to four percent. Their own README says plainly that cutting 90 percent of what the AI reads is not the same as cutting your bill by 90 percent. That is a more honest response than most projects manage.
The hidden cost of compression: throwing text away means sometimes throwing the wrong text away. One user found a file listing quietly capped at 200 of roughly 850 results, and the AI concluded whole folders did not exist. Another found that compressing test failures stripped out precisely the detail needed to fix them, turning a twenty minute job into an hour.
What you can do today
Nothing, if you use AI through a chat window. RTK compresses the output of commands, and if you are not running commands there is nothing for it to compress. That is worth saying plainly, because it is marketed as a token saver and it will do nothing at all for most people reading this.
But the instinct is portable, and it is one of the highest-value habits you can build. Stop pasting raw dumps into a chat. Paste the part that matters.
If you have a spreadsheet with 4,000 rows and a question about a pattern, do not upload the spreadsheet. Upload thirty representative rows and describe the rest. If you have a five-page error log, paste the error, not the log. You are doing RTK's job by hand, for free, and better, because you know which part matters and it has to guess.
RTK only ever sees a fraction of what flows into the AI. Which raises an obvious question. What if something caught everything?
Headroom: The Clever One
Headroom is the answer to that question, and it is by some distance the most sophisticated thing in this lineup.
Instead of recognising specific commands, it looks at whatever is flowing toward the AI and asks what kind of content this actually is. Then it picks a technique to match.
And it does two things the others do not. Nothing is actually destroyed, because originals stay in a local store and the AI can ask for the full version back. And everything runs on your own machine, so your work is not shipped off somewhere to be summarised.
One independent reviewer checked its accounting against a real provider bill and found it matched to within four ten-thousandths of a percent. Whatever else is true, this thing does exactly what it says to the bytes.
Which makes what happened next genuinely interesting.
The trap nobody saw coming
Go back to caching for a second. It only works while the front of your request stays byte for byte identical. Change one character near the beginning and the discount evaporates.
Now think about what a context compressor does for a living. It rewrites things. Including, in early versions, things that were already cached.
So Headroom was handing back tokens billed at a tenth of a cent, in exchange for sending fewer at full price. Its own repository contains an audit of six bugs filed under the heading "cache-killer smoking guns." In May, the maintainer wrote plainly that rewriting cached content was losing more than the compression saved, and that some users' bills went up.
That has since been rebuilt, and the safe default now leaves cached content alone. But read what the changelog says about that mode: with the cached part left untouched, tokens saved is near-always zero by design. The safe setting and the saving setting turned out to be different settings.
One engineer ran 25 tasks deliberately chosen to flatter the tool. Prompt tokens fell 39 percent. The cache discount he forfeited came to $16.37. The compression saved him $15.41.
He paid 96 cents for the privilege.
The strangest finding in this whole article: a carefully controlled test found the wrapper command flips a setting in Claude Code as a side effect, and that setting trims the tool descriptions by about 29 percent all on its own. Turning it on with Headroom not installed saved 20.5 percent. Headroom on top of it added 3.2 percent, which is noise. The savings were mostly not coming from the tool.
Here is what I want to be fair about, because this team has behaved well. They published the cache-killer audit themselves. The maintainer filed the pull requests correcting his own inflated numbers, including one showing his dashboard reported 97 percent saved where reality was around 13. They publish fleet data showing a 4.8 percent median against a marketing page claiming far more. They quietly revised their own headline from 92 percent down to 21.
That is more honesty than this category usually manages. It also tells you what to expect.
What you can do today
Headroom needs Python, a terminal, and a background process. There is no chat-window version. But its core insight is the single most useful habit in this article, and it is free.
Headroom works by keeping things lean at the edge instead of letting them pile up. You can do that by hand:
- Start a fresh conversation every 10 to 20 turns. Remember, the whole history rides along on every message. Long chats get expensive in a curve, not a line.
- Carry a summary, not the transcript. Ask for a short recap of what matters, open a new chat, paste it in as the first message. This is real context compression, and unlike Headroom it works on your conversation.
- Turn off connectors you are not using. Every connected tool adds its description to every single message whether you touch it or not. That is the same 29 percent effect from the callout above, and you can have it for nothing.
- Send text, not pictures of text. One high-resolution screenshot can cost 1,500 to 2,000 tokens. The words in it cost a fraction of that.
So that is both halves of what the AI reads. Two clever tools, two disappointing invoices, and a handful of free habits that beat them.
Now for the other side of the map, where things get considerably more fun.
Caveman: Why Use Many Token When Few Token Do Trick
In April, a developer named Julius Brussee had a thought that was mostly a joke. What if you just told the AI to talk like a caveman?
No articles. No "I'd be happy to help with that." No "it's worth noting that there are several approaches." Drop the filler, drop the hedging, drop the pleasantries, and answer in fragments.
He posted it. It went to the top of Reddit. It hit a hundred thousand stars on GitHub. Companies deployed it. And the number everyone repeated was that it cut output by 65 percent.
Now here is the part that makes this my favourite story in the article. The author himself, in public, said it was very much intended to be a joke rather than research, and that the fair criticism is that his number came from preliminary testing rather than a real benchmark.
The project has since formally retracted it. Its own honest-numbers documentation lists the output reduction as "not published."
It is smarter than the joke suggests
Which is a pity, because the actual design is thoughtful. The preservation list is strict: code unchanged, error messages quoted exactly, numbers and technical terms exact. It will never drop a "not" or "only," because flipping a meaning is worse than any token saved.
It bans fake caveman padding, too. You cannot add words to sound more primitive, and invented abbreviations and arrows are forbidden on the grounds that they cost the same tokens as the real word while being harder to read.
Best of all, it has an automatic off switch. Security warnings, confirmations before anything irreversible, and multi-step instructions all revert to normal English. So does anything you are going to send to another person.
So does it save money?
A user ran the cleanest test I found: 20 prompts, five times each, three runs. Three hundred runs a side.
| Measurement | Normal | Caveman | Change |
|---|---|---|---|
| Output tokens | 369,825 | 206,894 | 44% fewer |
| Cached context | 879,808 | 1,339,463 | 52% more |
| Total cost | $26.01 | $26.02 | +0.06% |
Look at that for a second. The compression is real. Forty-four percent fewer words, repeatably. And the bill did not move, because the rulebook that produced those shorter answers is itself about 1,500 tokens riding along on every single message.
You are paying for the instruction that saves you money. On short exchanges, the instruction costs more than the answer.
Then there is the finding that made me laugh. An independent benchmark put five approaches head to head across 24 prompts, with a separate model scoring the quality of all 120 answers.
| Approach | Quality | Tokens |
|---|---|---|
| No instruction at all | 0.985 | 636 |
| Just "be brief" | 0.985 | 419 |
| Caveman lite | 0.976 | 401 |
| Caveman full | 0.975 | 404 |
| Caveman ultra | 0.970 | 449 |
Two words matched a hundred-thousand-star project on tokens, and beat it slightly on quality, without the 1,500-token rulebook attached.
It may not even stick any more. Newer versions of Claude's coding tool ship their own instruction saying readable beats concise, and specifically not to compress writing into fragments or arrow chains. That directly contradicts Caveman. Users report it holding for a turn or two and then quietly reverting.
What you can do today
This is the one. Of the four, this is the technique that is genuinely aimed at you.
Coding sessions bury the prose under code that cannot be compressed, which is why the measurements are so grim. But in ordinary conversation, the reply is the entire output. Research, drafting, planning, answering questions. Compress the reply and you have compressed everything.
You do not need the plugin. You need about five lines in your custom instructions.
That exclusion list matters more than the brevity line. Without it, dropped words genuinely do flip meanings. One reported case turned "make sure the file exists" into "make file exists," which is not a shorter way of saying the same thing. It is a different instruction.
Caveman makes the AI say less. Which leaves one box on the map, and one question nobody else was asking.
Ponytail: The One That Worked
Every technique so far takes something that already exists and squeezes it. Smaller command output. Smaller context. Smaller sentences.
Dietrich Gebert asked a different question. What if the AI just made a smaller thing in the first place?
So he gave it the personality of a senior developer who has been around long enough to find unnecessary work genuinely irritating. Someone who, handed a request, instinctively looks for the laziest thing that actually solves it. The kind of person who has seen enough clever code at 3am to prefer boring code.
The mechanism is a ladder. Before building anything, the AI climbs it and stops at the first rung that holds.
Worth a note here, because almost every write-up of this gets it wrong, including 9Router's own. Most summaries list five rungs. The real ladder has seven, and the two that usually get dropped are the first two. Which is unfortunate, because "does this need to exist" and "do we already have this" are where the big wins live.
And it has a list of things it refuses to be lazy about. It will never skip input validation at trust boundaries, error handling that prevents data loss, security, accessibility basics, or anything you explicitly asked for. If you insist on the full version, it builds the full version and does not argue.
It is also forbidden from being lazy about understanding. The rule is that the ladder shortens the solution, never the reading, because laziness that skips comprehension dresses up as efficiency and ships a confident wrong answer.
The results
Same independent team, same careful method, 80 tasks.
| Measurement | Advertised | Measured |
|---|---|---|
| Code written | 54% less | 15.4% less |
| Cost | 20% less | 10.3% less |
| Quality | not stated | No difference |
Roughly half the advertised figure. And statistically solid, repeatable, in the right direction, with no quality cost.
After a tool that made things more expensive, a tool whose savings turned out to be an environment variable, and a tool that moved the bill by six hundredths of a percent, a real ten percent is a genuine result.
But the number that actually matters is what happened when they sorted the tasks by size.
There it is. The whole article in four boxes.
Ponytail does not save you a fixed percentage. It reclaims whatever room the AI had to over-build. Give it room and it saves a third. Give it none and it does nothing, because there was nothing there to save.
That is true of all four techniques, and of every token-saving idea you will ever be sold. They pay out in proportion to how much waste you already had.
Two footnotes worth your time. The project withdrew its own earlier claim of 80 to 94 percent after a consultancy CTO showed the comparison was unfair. And in that same critique, adding just "follow YAGNI principles, and one-liner solutions" to a prompt produced even less code than the skill did. But the rebuilt benchmark found that seven-word version was erratic, and it was the only approach that quietly deleted a security check. The skill wrote three more lines. Those three lines were the check.
What you can do today
The author is explicit that this is for building software and not for general questions or prose. But the instinct travels, because over-building is not a programming problem. It is a human one, and AI has it badly.
Use it on a spreadsheet, a landing page, an onboarding sequence, a content calendar. The saving is not really about tokens. It is about not receiving a fourteen-tab system when you needed one tab.
Which Ones Should You Actually Run?
Because each technique works on a different box, they genuinely do stack. What they do not do is add up. Running all four does not give you 40 plus 65 plus 20 percent. Each one is taking a slice of what the last one already shrank, and every slice gets divided by ten on the way to your invoice.
One replay analysis across 614 million tokens from real sessions put all three compressors together at 3.7 percent of total spend.
One pairing to avoid: RTK and Headroom are both squeezing the same box. They are not complementary, they are competing. Run one, alone, and measure it.
| If you are... | Do this |
|---|---|
| A chat user ChatGPT, Claude, Gemini in a browser or app |
The brevity instruction from Section 05 and the scoping ladder from Section 06, plus fresh chats every 10 to 20 turns. No installs exist for you and you do not need any. |
| Building with AI in an editor Cursor, Copilot, Claude Code |
Ponytail. It is the one with a solid independent result. Scope your prompts on top of it. |
| Running agents heavily Long sessions, real spend |
Ponytail, plus trimming your connected tools. Measure two comparable weeks before you add anything else. |
| Building on the API Your own product or automation |
Set up caching properly first. It is a 90 percent discount and it dwarfs everything in this article. |
The Free Levers That Beat All Four
Here is the uncomfortable pattern across every independent test in this article. The tools produced single digits. The habits produce double digits, cost nothing, and require trusting no one.
Match the model to the job
Running a top-tier reasoning model to reformat a list or fix a typo is the most common and most expensive habit in AI use. Smaller models are five to twenty times cheaper and perfectly capable of ordinary work. This one change beats every tool in this article combined.
Scope the request
"Build a login page" invites the AI to invent a password reset flow, a remember-me checkbox, and three abstractions you never asked for. "Build a login page with email and password only, nothing else" does not. Most of what Ponytail does is compensate for the first kind of prompt.
Batch, and edit surgically
Three separate messages load the whole conversation three separate times. One message with three questions loads it once. And when something needs fixing, say "redo section three only, leave the rest alone" rather than "redo the report."
Start fresh conversations
The longer a thread runs, the more every message costs, because the whole history rides along. Every ten to twenty turns, ask for a summary, open a new chat, paste it in. For ordinary users this is the highest-value habit on the list.
Judge by the invoice, never the dashboard
This is the real lesson. Every tool here reports savings against its own guess about what would otherwise have happened. RTK's counter reported 96.2 million tokens saved in trials where the bill went up. Headroom's maintainer measured 13 percent where his own dashboard said 97.
Want to know if something works for you? Run a week with it and a comparable week without, and compare what you were charged. As the JetBrains team put it, your invoice is the only part of the benchmark that bills you.
The Thing Worth Remembering
None of these tools are scams. The engineering is real, the people building them are serious, and several of them responded to hostile testing by publishing corrections against their own marketing. That is rarer than it should be.
But look at what separated the one that worked from the three that did not.
RTK compressed. Headroom compressed. Caveman compressed. Ponytail did not compress anything. It reduced the work.
That is the whole lesson, and it survives long after these four tools have been replaced by four others. Compression fights the symptom. The tokens were already generated, already sent, already paid for at some rate. Reduction means they never existed.
So the cheapest token is the one nobody had a reason to produce. Ask for less. Ask more precisely. Use a smaller model when a smaller model will do. Start a new chat before the old one gets heavy.
None of that requires installing anything, and all of it will beat any plugin you could add.
And then, if you are still spending more than you would like, install one thing, run it for a week, and go look at what you were actually charged. Not what the dashboard says you saved. What the invoice says you paid.