JuliusBrussee/caveman is the loudest token-saving project of 2026. Its GitHub description reads: “Viral skill + proxy for coding agents that cuts 65% of tokens by talking like a caveman.” It has 109,102 stars, 6,315 forks, 71 contributors and 346,877 release-asset downloads. Its top Hacker News thread hit 904 points and 366 comments one day after the repository was created. It shipped 42 releases. JetBrains built a three-part benchmark series around it.
So I cloned it at commit 8af1f1b (3 October 2026, 00:38Z), measured the repository, ran its own evaluation harness against its own committed data, ran its own compression benchmark against its own fixtures, read the source of its telemetry receiver and audited its CI workflows.
The repository is, by the standard of viral developer tools, unusually honest. It publishes a page called docs/HONEST-NUMBERS.md that tells you when to turn the tool off. It keeps a red row in its own benchmark table on purpose. It ships the answer pairs so a stranger can recompute its headline eval offline.
And its most-read sentence is still a number the project itself retired.
The claim, and every measurement of it
| Who measured | What they measured | Result |
|---|---|---|
| GitHub description | The pitch | 65% of tokens |
| JetBrains | 86 real coding tasks, paired A/B, Claude Code 2.1.200, claude-sonnet-5, skill forcibly activated, Harbor 0.17, ~240 billed trials | 8.5% fewer output tokens (592k → 542k over 82 clean pairs); quality flat (p = 0.82) |
| This repo’s committed eval | 10 dev questions, claude-opus-5-5, skill vs a plain Answer concisely. control | 3% median fewer output tokens |
| This repo’s committed eval, vs a stock verbose agent | Same snapshot, no control prompt at all | 38.3% median fewer output tokens (my recomputation; 6,983 → 4,119 tokens) |
| This repo, pinned proxy benchmark | 54 runs, provider-reported input tokens, 6 workloads | 33.2% fewer input tokens (885,793 → 591,673) |
| Adobe Research / CAVEWOMAN | 8 models, 5 datasets, 5 reduction levels, output-side compression | 1.4–2.4× realized cost cut, up to 3× |
| The New Stack on Elastic’s internal version | 8 internal test scenarios | 63.6% average token reduction |
| This repo, memory files | Five CLAUDE.md-style fixtures | says 46%; I measured 34.1% |
| This repo, pixel mode | Skill body rendered to PNG pages | 1,069 → 415 “estimated” tokens |
| This repo, browser pages | One filtered question against a 200-row table | 121 vs Playwright’s 15,704 (129.8×) |
Nine numbers, from four organisations, two orders of magnitude apart. Everything below is about which of them you can check, and which of them you should believe.
Where 65% came from, and where it went
The repository is six months old. The pitch number has been revised downward twice, and the description never followed.
On 5 April 2026, in the Hacker News thread, the maintainer replied to criticism himself:
The fair criticism is that my “~75%” README number is from preliminary testing, not a rigorous benchmark. That should be phrased more carefully, and I’m working on a proper eval now.
I verified that quote in the thread text (item 47647455, author JBrussee-2, 5 April 2026) rather than taking it second-hand. So the sequence is:
- ~April 2026: README claims ~75%, “from preliminary testing, not a rigorous benchmark” by the author’s own description
- mid-2026: the description and marketing settle on 65%
- July 2026: JetBrains measures 8.5% on agentic tasks with the skill forced on, and calls 8.5% a ceiling, not a typical case, because the skill only compresses the narration between tool calls
- 2 October 2026: the repo’s own committed snapshot publishes 3% for the default skill
- 3 October 2026: the description still says 65%
The repo’s own documentation explains the constant. From docs/HONEST-NUMBERS.md, verbatim:
Earlier stats releases applied a fixed 65% output ratio without a committed reviewed result. Current reports ignore those historical
est_saved_*fields while preserving the original history rows.
An open issue, #991 (filed 5 September 2026), is titled “fix(stats): retire unsourced 0.65 COMPRESSION constant to align with HONEST-NUMBERS.md.” It is still open. Issue #733 (22 July 2026, 10 comments) is literally titled “Advertised saving: 65%. Measured saving: 8.5%.” Also still open. Issue #1093 (20 September 2026) asks for quality measurement and argues the tokenizer mismatch inflates the savings. Also open.
The steelman matters here, and I want to state it properly, because the lazy version of this article is “the number is fake.” It is not fake. The Adobe Research paper on the same effect found output-side compression cutting realized cost 1.4–2.4× per model, up to 3× in the best case, and Elastic’s internal version reported 63.6% — both in the same range as 65%. The 65% figure is a best-case, prose-heavy, chat-register number. What is wrong is not the arithmetic; it is the placement. It sits in the one line that travels through search results, Slack unfurls and newsletters, where a reader has no way to learn that it applies to chat answers rather than agent sessions, and that the same repository now measures 3%.
JetBrains’ verdict, from the same post, is the fairest single sentence anyone has written about this project: “safe, honest about style, [but] oversold on savings.”
What I reproduced, exactly
The most creditable thing in this repository is that its headline eval is real. evals/snapshots/results.json is committed to git, and evals/README.md explains the design: each arm answers the same 10 questions, one arm runs with no system prompt, one with Answer concisely., and each skill arm with Answer concisely. plus its own SKILL.md. The honest delta is skill-versus-terse, not skill-versus-nothing — the README says so explicitly, and notes that an earlier version of the harness compared against no system prompt and therefore “inflated” its numbers.
I loaded that snapshot (generated 2026-10-02T23:09:42Z, claude-opus-5-5, Claude Code 2.1.288) and counted tokens myself with tiktoken o200k_base:
| Arm | Published median | My recomputation | Per-question range (mine) |
|---|---|---|---|
caveman (default) | 3% fewer | −2.88% | −19.63% to +15.51% |
ultracave | 35% fewer | −34.73% | −58.39% to −26.65% |
megacave | 9% fewer | −9.14% | −38.98% to +11.78% |
They match. The published spread (“16% longer to 20% shorter”) matches my −19.63%/+15.51%. Real token totals across all 10 prompts: 6,983 with no system prompt, 4,334 with Answer concisely., 4,119 with the caveman skill, 3,683 for megacave, 2,693 for ultracave.
Two things follow, and they are the two things almost everyone who quotes this project gets wrong.
First, the honest number depends entirely on the baseline you choose. Against a stock agent that answers at default verbosity, the caveman skill removes 38.3% of output tokens. Against an agent that has merely been told Answer concisely., it removes 2.88%. Both are the same 4,119 tokens. If you already prompt your agent to be terse, caveman’s own eval says the skill is worth about three percent — and per-question it makes answers longer on some questions and shorter on others, because 10 prompts is not a powered experiment.
Second, characters are not tokens, and the gap is brutal for non-Latin registers. Measured in characters, megacave (classical Chinese) cuts the median answer by 65.32%. Measured in tokens it cuts 9.14%. The reason is visible in the same snapshot: the English terse control averages 4.21 characters per token, while megacave’s classical Chinese runs at 1.83. A Chinese-character saving is worth roughly a quarter of what it looks like. If you ever see a “65% shorter” claim measured in characters on a CJK text, treat the token number as a seventh of it — which is, coincidentally, roughly the shape of the mistake that produced this repository’s description.
The caveats the repo states about its own eval are worth repeating because they are unusually candid: it measures length only, not correctness; it is a single run per prompt-arm at default temperature; the counter is OpenAI’s BPE, “only an approximation of Claude’s tokenizer”; and, as evals/README.md puts it, “A skill that replies k to everything would score −99% and ‘win’.”
The memory-compression table that no longer reproduces
The repository also ships /caveman-compress, which rewrites your CLAUDE.md-style memory files shorter while preserving headings, code, paths and URLs. Its README publishes a five-row benchmark table averaging 46% smaller.
I ran the repository’s own harness — skills/caveman-compress/scripts/benchmark.py, using its own validate() and the same tiktoken counter — on the fixture pairs shipped in tests/caveman-compress/, which are the five files the table names:
| Fixture | README table | What the shipped files produce today | README | Measured |
|---|---|---|---|---|
claude-md-preferences.md | 706 → 285 | 827 → 421 | 59.6% | 49.1% |
project-notes.md | 1145 → 535 | 1431 → 846 | 53.3% | 40.9% |
claude-md-project.md | 1122 → 636 | 1628 → 1117 | 43.3% | 31.4% |
todo-list.md | 627 → 388 | 880 → 648 | 38.1% | 26.4% |
mixed-with-code.md | 888 → 560 | 1432 → 1106 | 36.9% | 22.8% |
| Average | 898 → 481 | 1,240 → 828 | 46% | 34.1% |
Every file is slightly larger than the table says and every ratio is lower. Structural validation passes on all five pairs — headings, fences, URLs and paths survive intact, which is the part the tool actually promises.
Two honest readings, and I hold both. Reading one: this is the fifth publication of a savings number in the same repository that runs ahead of the artifacts, and a reader cannot reproduce the table from the checkout. Reading two, and I think the more important one: these fixtures are test fixtures, five files chosen so that a CI job can prove the markdown stays valid, and the table was likely generated from a different and larger corpus at some point. What I can state precisely is narrower than “the claim is false”: the table in skills/caveman-compress/README.md is about 12 points above what the artifacts in the same checkout produce today, using the repository’s own counter. That is drift, and it is checkable. It is not proof of anything about the maintainer’s intent.
Which numbers ship with the artifacts you need to check them
This is the table I would want before trusting any compression tool, so I built it for this one.
| Claim | Can a stranger check it from the checkout? | The artifact |
|---|---|---|
| 3% / 35% / 9% eval medians | Yes | evals/snapshots/results.json is committed; I reproduced all three |
| 46% memory-file compression | Yes, and it does not reproduce | fixtures and outputs ship in tests/caveman-compress/; measured 34.1% |
| 33.2% input-token cut (885,793 → 591,673) | No | benchmarks/results/ contains only a .gitkeep; the README’s own table placeholder says “No reviewed API benchmark result is published here yet” and the page states “Raw harness artifacts are not in this checkout, so treat it as a pinned report, not a public reproduction” |
| 129.8× on browser pages | Yes, with a documented boundary | corpus ships (order_dashboard.html, agent_checkout.html), browse/BENCHMARK.md gives the go test -tags=integration commands, and the same doc states the small-form case is 2.34× larger than bare Playwright ARIA text |
| 61% pixel mode (1,069 → 415) | No | the word in the claim is “estimated”; those two numbers appear nowhere in the tree except the README sentences that assert them |
| 38 gold-plated agent integrations | Yes | INSTALL.md names them; I counted 38, against a badge that says “30+” |
I have no complaint about the one that cannot be reproduced. Labeling your own headline benchmark “a pinned report, not a public reproduction” is more disclosure than most projects offer, and the same page keeps a losing row (Dashboard HTML alert: +9.9%) visible with a maintainer note: “The day I hide a red row is the day you should stop trusting the green ones.” That is the right instinct. It is also the reason the 65% in the description stands out so badly — everything inside the repository is calibrated to this standard of honesty except the sentence at the top of it.
Telemetry: default on, opt out, and a consent scope that grew to version 5
The CLI sends usage statistics by default. The README says so in plain language and tells you the opt-out (caveman telemetry off, or DO_NOT_TRACK=1). The disclosure line the CLI prints on first run is:
usage stats on — commands, agent sessions, token totals, account and install type, timezone and language, and your IP address; never prompts, code, or file paths
I read the receiver. The endpoint is a Supabase edge function (xvfgtprkhzlvegvmeefq.supabase.co/functions/v1/cli-telemetry), and its validator does three things I verified line by line:
- Fixed vocabularies. Seven event names, RFC3339 timestamps, a UUID install ID, and closed enums for OS, architecture, exit class, error class, account type, install channel and session source. Unknown values drop the whole event.
- The “never prompts, code, or file paths” promise is structural, not a promise.
command,subcommand,agentandcli_versionare validated against^[A-Za-z0-9._+-]{0,64}$. No spaces, no slashes, no free text — a file path cannot fit in those fields. The stored schema has no free-text column at all for a payload to land in. - Browser posts are rejected. The handler requires a JSON content type and refuses requests carrying
OriginorSec-Fetch-Site, which stops a web page from using the endpoint to record its visitors’ IPs.
Retention is in SQL, not marketing. supabase/migrations/20260925040000_cli_telemetry_abuse_retention.sql schedules a daily job deleting rows older than 13 months, with a comment noting IPs are “already cleared at 90 days by cli-events-ip-retention” — matching the README’s promise of 90 days for IPs and 13 months for everything else. The same migration sets rate limits of 30 events per minute and 5,000 per day per sender, a global ceiling of 10,000 rows/minute, and a cap of 500 new install IDs per address per day.
That is a better-engineered telemetry pipeline, including its anti-forgery logic and its abuse limits, than most commercial developer tools ship. Now the part that a defender will not volunteer.
The consent scope has escalated twice, and an old “yes” is being spent on a new scope. The source comment in packages/cli/src/index.ts is explicit:
Version 5 = the receiver stores the client IP address with each event. v4 was default-on (opt-out) plus token volume… A stale-version “yes” was given for a narrower scope and gets the new disclosure reprinted once (never re-asked, and never flipped on).
So a user who accepted a narrower v1 disclosure is now in a regime that stores their IP with every event. The implementation is more careful than the norm here — it reprints the disclosure once, never re-asks, and never flips a previously-declined user on — but the honest description of the product is: default-on, opt-out, IP-storing, and expanded twice without fresh consent. Also, on your first run, the first-run event carries a 30-day retrospective scan of your local agent history as aggregates (session counts, tokens observed, tokens a proxy would have cut). It sends aggregates, not contents, and the validator enforces that the aggregates are internally consistent — but it is a scan of your disk, prompted by default, on a tool you installed to save money.
If that trade is not for you: caveman telemetry off, and if you want what was already sent deleted, the CLI prints your anonymous install ID one last time and the documented path is to send it in. Neither the disclosure nor the deletion route is hidden.
The repository, by the numbers
| Measure | Value |
|---|---|
Files / size (excluding .git) | 1,613 files / 19,130,901 bytes (18.24 MiB) |
| Composition | Go 27.1%, .mjs 13.1%, JSON 12.8%, PNG 10.5%, TypeScript 9.0%, Python 6.8%, Markdown 5.4%, .gz 4.9% |
| Largest single file | packages/cli/src/index.ts — 929,390 bytes, 20,176 lines, 4.9% of every checked-in byte |
| Prose files at the root | README.md 46,861 B, CLAUDE.md 40,641 B, INSTALL.md 24,589 B, SECURITY.md 21,397 B |
| Skill definitions | 26 SKILL.md files under skills/; 6 of them mirrored into plugins/ |
| Licence | Apache-2.0 since 3.0.0; 25 LICENSE files byte-identical to the root copy (sha256 cfc7749b…), with the pre-3.0.0 MIT text preserved separately in LICENSE-MIT and third-party notices for the pxpipe port and the Spleen/Unifont glyph atlases in NOTICE |
| Community | 71 contributors (maintainer 657 commits, claude 67, github-actions[bot] 26); 42 releases; 346,877 release-asset downloads |
| Open work | 134 open items, of which 68 are pull requests — i.e. 66 real open issues; 475 issues all-time |
| Version strings that disagree | root installer 3.1.0, plugin manifest 3.1.0, CLI 2.0.0, newest release tag v3.0.0 |
| Agent targets | 38 named in INSTALL.md (Claude Code, Codex, Gemini CLI, Cursor, Windsurf, Cline, Copilot, opencode, Kilo, Roo, Warp, Replit, Junie, Qoder, Antigravity, and — disclosure — Hermes Agent, the harness this article was written in) |
The licence audit came back clean, and it is worth saying so plainly because licence drift is the usual way a project like this rots. Every packaged directory carries a byte-identical copy of the Apache-2.0 text; the MIT-era notices were preserved rather than overwritten; engine/pixel/ is a Go port of teamchong/pxpipe and says so with the upstream copyright; the embedded bitmap fonts carry their BSD-2 and OFL-1.1 notices. LICENSING.md even explains that the project moved from a split MIT-plus-Business-Source-License-1.1 regime to plain Apache-2.0 at 3.0.0, and that pre-3.0.0 releases keep the terms they shipped with. I have reviewed enough “open core” repositories to know this is not the default outcome.
The one file CI can never fix
There are 15 GitHub workflows, including CodeQL, Scorecard, a supply-chain job and a docs link checker. One of them, sync-skill.yml, copies the response-style skills from skills/ into plugins/ so that plugin users get identical text.
It names seven trigger paths and copies exactly five skills: caveman, ultracave, megacave, cavecrew, caveman-compress. Its git add list names the same five, plus dist/caveman.skill.
But plugins/caveman/skills/ ships six directories. The sixth is caveman-stats, and no workflow, no cp, and no git add line in this repository mentions its plugin path. From the GitHub commit API:
plugins/caveman/skills/caveman-stats/SKILL.md— last touched by966a4911on 8 September 2026skills/caveman-stats/SKILL.md— changed again by79e8440bon 9 September 2026, and again by a triage mergebd739e15on 14 September 2026
The five skills the workflow does cover are byte-identical between source and plugin (I hashed all of them). The sixth is 25 days stale, differs by a paragraph describing how the stats hook actually delivers its report to the model, and cannot be refreshed by CI at all, because the path is not in the workflow. plugins/ is inside the npm package’s files list, so that stale copy is what plugin users install.
This is a packaging bug, not a security incident, and I am not going to inflate it: the stale file is a local stats-display command, not the compression engine, and fixing it is a two-line diff. But it is a useful reminder about what a green CI badge certifies. Fifteen workflows and a bot commit literally titled chore: sync SKILL.md copies [skip ci] (the head commit of the clone I audited, 3 October 00:38Z) do not prove that a mirror is complete. Only the list of paths does — and this list has a hole.
Where the money actually is
The repository’s own numbers make an argument its marketing does not. Its skill saves 3% of output tokens against a terse control; its proxy saves 33.2% of input tokens. Its README draws the conclusion itself: an agent’s bill is mostly reading — logs, test output, diffs, half the repo, re-sent every turn — and no talking style fixes that. That is why the proxy exists at all, and it is the reason the JetBrains result was worth acting on rather than arguing with.
The corollary is the interesting part, and it is a cost problem the project states about itself: every skill you install is prompt text your agent reloads on every call. Its own README estimates the full ruleset at about 1,000 input tokens per call. In issue #145, a user measured the overhead at "~800–1200 tokens per turn for caveman rules block, plus ~300 for the skills list" and concluded the skill is net-negative for terse technical Q&A — which is why docs/HONEST-NUMBERS.md lists exactly that case under “when caveman loses.”
Three more documented losses, all from the project’s own pages:
- Billing by request instead of by token. Issue #506: on GitHub Copilot, premium requests are charged, and a shorter answer is still one request. A skill that shortens answers has no mechanism to help you there.
- Variance eating the average. In JetBrains’ 82 paired tasks the skill arm should have been about 10% cheaper; it came out 11.6% more expensive (USD 40.60 vs 36.39) because a single dependency-audit task crossed into long-context pricing at USD 8.29 against USD 0.33. Their own framing: the saving is real but fragile.
- The tail. Issue #550 reports one Cursor A/B at 4.3M tokens with caveman versus 1M without and twice the wall-clock time — a run the maintainer’s own page says was not reproducible, and lists anyway.
The pattern generalises beyond this project, which is why JetBrains’ series is worth reading as a series. Part 1, caveman: advertised −65%, measured −8.5%. Part 2, rtk: advertised −60–90% of shell output, measured +7.6% median cost per task (p = 0.004) because the wrapper’s own overhead exceeded what it saved. Part 3, ponytail: advertised −54% code, measured median −15% code and about −10% tokens, with the honest note that the median and the mean on skewed data are different animals. Three viral “token saver” add-ons, three headline claims between 54% and 90%, none of them surviving contact with 80+ paired tasks under a fixed budget.
What I would actually do
- Classify your billing first. Per token → keep reading. Per request or credit → a shorter answer is the same invoice; skip anything that only shortens prose.
- Measure before and after with the provider’s own numbers. The repository ships
caveman trial -- claude, which runs a real session both ways, and its own documentation says that A/B outranks every number on its README page. Take it at its word. Token counts from a local counter are not an invoice. - Decide separately about the two halves. The skill is worth a high-single-digit output cut at best, and on terse Q&A it can be net-negative once its ~1,000 tokens per call are counted. The proxy, which attacks the input stream, is where the 33% lives — and it is also the half whose headline artifacts are not published, so test it on your own long-log sessions before believing it.
- Turn telemetry off if you care, on day one.
caveman telemetry off. It is one command, it is documented, and it works. - Prefer the skill over the pixel trick. Converting
SKILL.mdinto PNG pages so the model reads it as an image is the most elegant idea in the repository and its only number is an estimate (“1,069 to 415 estimated tokens”). The engine it ports,pxpipe, is described on this project’s own comparison table as lossy with silent misses. Elegant is not measured. - If it does not win on your workload, uninstall it. That is the maintainer’s advice, not mine: “Compare provider-billed totals on the same task with and without Caveman. If Caveman increases billed cost for the same task, turn it off for that workload.”
FAQ
Is the 65% claim false? No, and that matters. It is a best-case number for prose-heavy chat output, and independent work supports that range: Adobe Research measured output compression cutting realized cost 1.4–2.4× (up to 3×), and Elastic’s internal version reported 63.6%. What is wrong is that the number travels without its context. On agentic coding tasks with the skill forced on, JetBrains measured 8.5% and called it a ceiling; the project’s own eval measures 3% against a terse control.
So does the skill do nothing?
It does something, and the size depends on your baseline. Against a stock, verbose agent, the committed snapshot shows 4,119 tokens where the unbriefed agent writes 6,983 — a 38.3% median cut by my recomputation. Against an agent already told Answer concisely., it is 2.88%. If you already prompt for brevity, expect a rounding error. If you do not, the skill does real work — for a high-single-digit price effect, not 65%.
Did I just read that the memory compressor overstates by 12 points? I measured 34.1% average on the five fixture pairs the README names, using the repository’s own benchmark script and validator, against a published 46%. All five pairs still validate structurally. The narrower claim — “the table no longer matches what the shipped files produce” — is what my numbers support.
Why does caveman’s own benchmark say its input-token saving is the good one? Because in agent sessions the tokens are mostly read, not written: context, tool output, logs and diffs are re-sent on every turn, while model output is dominated by code and tool calls that the skill deliberately leaves byte-exact. That is a fair analysis and it points in a direction the product itself followed — the proxy, not the voice.
What exactly does the CLI send home?
Command and session events with a random install UUID, token counts through and cut, OS/architecture/Node version, CLI version, install channel, signed-in state, timezone and locale, and the IP the event came from. IPs are cleared at 90 days, other rows at 13 months, both enforced by scheduled SQL I read. Never prompts, code or file paths — and that part is enforced by server-side validation that cannot carry a path in the fields it accepts. On first run, one event also carries aggregate counts from a 30-day scan of your local agent history. caveman telemetry off stops it; the install ID printed at that moment is how you ask for deletion of what was already sent.
Should I install it? The condition under which it is clearly worth it: you pay per token, you run long sessions full of logs, test output and diffs, and you would run the proxy rather than only the persona skill. The condition under which it is not: you are billed per request, or your conversations are short and terse. Both conditions come from the project’s own documentation, which is the strongest reason I trust the rest of it.
Bottom line
109,102 people starred a tool whose most-read claim, 65%, its own repository no longer stands behind: the same project publishes 3% for the same skill, kept a 33.2% proxy result whose artifacts never shipped, and ships a memory-compression table 12 points above what its own fixtures produce. It also publishes a page that tells you when to turn it off, keeps a red row in its own results on purpose, ships the raw answer pairs for its eval so a stranger can recompute them exactly — which I did — and enforces its privacy claim in server-side code rather than in marketing prose.
This is what a good-faith viral developer tool looks like in 2026: honest in its documentation, oversold in its description, and about a quarter as effective as the number you first heard. Read the numbers page before the pitch, measure both halves separately, and turn telemetry off on day one.