Auditing the Caveman Skill: 109,102 Stars, a 65% Headline, and the 3% Its Own Eval Publishes

Caveman promises to cut 65% of your agent's tokens. I cloned it at commit 8af1f1b (1,613 files, 19,130,901 bytes), reproduced its committed eval with tiktoken myself and got the 3%/35%/9% medians it publishes, then ran its own memory-compression harness on its own shipped fixtures and got 34.1% where its README table says 46%. Also: the 65% traces to a fixed constant its own docs retired, JetBrains measured 8.5% against that 65% while its default skill beat a stock verbose agent by 38.3% and the words “Answer concisely” by only 2.88%, the CLI ships default-on telemetry whose consent scope escalated to storing your IP in version 5, and a plugin skill copy has been stale since 9 September because the sync workflow never names its path.

JuliusBrussee/caveman is the loudest token-saving project of 2026. Its GitHub description reads: “Viral skill + proxy for coding agents that cuts 65% of tokens by talking like a caveman.” It has 109,102 stars, 6,315 forks, 71 contributors and 346,877 release-asset downloads. Its top Hacker News thread hit 904 points and 366 comments one day after the repository was created. It shipped 42 releases. JetBrains built a three-part benchmark series around it.

So I cloned it at commit 8af1f1b (3 October 2026, 00:38Z), measured the repository, ran its own evaluation harness against its own committed data, ran its own compression benchmark against its own fixtures, read the source of its telemetry receiver and audited its CI workflows.

The repository is, by the standard of viral developer tools, unusually honest. It publishes a page called docs/HONEST-NUMBERS.md that tells you when to turn the tool off. It keeps a red row in its own benchmark table on purpose. It ships the answer pairs so a stranger can recompute its headline eval offline.

And its most-read sentence is still a number the project itself retired.

The claim, and every measurement of it

Who measuredWhat they measuredResult
GitHub descriptionThe pitch65% of tokens
JetBrains86 real coding tasks, paired A/B, Claude Code 2.1.200, claude-sonnet-5, skill forcibly activated, Harbor 0.17, ~240 billed trials8.5% fewer output tokens (592k → 542k over 82 clean pairs); quality flat (p = 0.82)
This repo’s committed eval10 dev questions, claude-opus-5-5, skill vs a plain Answer concisely. control3% median fewer output tokens
This repo’s committed eval, vs a stock verbose agentSame snapshot, no control prompt at all38.3% median fewer output tokens (my recomputation; 6,983 → 4,119 tokens)
This repo, pinned proxy benchmark54 runs, provider-reported input tokens, 6 workloads33.2% fewer input tokens (885,793 → 591,673)
Adobe Research / CAVEWOMAN8 models, 5 datasets, 5 reduction levels, output-side compression1.4–2.4× realized cost cut, up to 3×
The New Stack on Elastic’s internal version8 internal test scenarios63.6% average token reduction
This repo, memory filesFive CLAUDE.md-style fixturessays 46%; I measured 34.1%
This repo, pixel modeSkill body rendered to PNG pages1,069 → 415 “estimated” tokens
This repo, browser pagesOne filtered question against a 200-row table121 vs Playwright’s 15,704 (129.8×)

Nine numbers, from four organisations, two orders of magnitude apart. Everything below is about which of them you can check, and which of them you should believe.

Where 65% came from, and where it went

The repository is six months old. The pitch number has been revised downward twice, and the description never followed.

On 5 April 2026, in the Hacker News thread, the maintainer replied to criticism himself:

The fair criticism is that my “~75%” README number is from preliminary testing, not a rigorous benchmark. That should be phrased more carefully, and I’m working on a proper eval now.

I verified that quote in the thread text (item 47647455, author JBrussee-2, 5 April 2026) rather than taking it second-hand. So the sequence is:

  • ~April 2026: README claims ~75%, “from preliminary testing, not a rigorous benchmark” by the author’s own description
  • mid-2026: the description and marketing settle on 65%
  • July 2026: JetBrains measures 8.5% on agentic tasks with the skill forced on, and calls 8.5% a ceiling, not a typical case, because the skill only compresses the narration between tool calls
  • 2 October 2026: the repo’s own committed snapshot publishes 3% for the default skill
  • 3 October 2026: the description still says 65%

The repo’s own documentation explains the constant. From docs/HONEST-NUMBERS.md, verbatim:

Earlier stats releases applied a fixed 65% output ratio without a committed reviewed result. Current reports ignore those historical est_saved_* fields while preserving the original history rows.

An open issue, #991 (filed 5 September 2026), is titled “fix(stats): retire unsourced 0.65 COMPRESSION constant to align with HONEST-NUMBERS.md.” It is still open. Issue #733 (22 July 2026, 10 comments) is literally titled “Advertised saving: 65%. Measured saving: 8.5%.” Also still open. Issue #1093 (20 September 2026) asks for quality measurement and argues the tokenizer mismatch inflates the savings. Also open.

The steelman matters here, and I want to state it properly, because the lazy version of this article is “the number is fake.” It is not fake. The Adobe Research paper on the same effect found output-side compression cutting realized cost 1.4–2.4× per model, up to 3× in the best case, and Elastic’s internal version reported 63.6% — both in the same range as 65%. The 65% figure is a best-case, prose-heavy, chat-register number. What is wrong is not the arithmetic; it is the placement. It sits in the one line that travels through search results, Slack unfurls and newsletters, where a reader has no way to learn that it applies to chat answers rather than agent sessions, and that the same repository now measures 3%.

JetBrains’ verdict, from the same post, is the fairest single sentence anyone has written about this project: “safe, honest about style, [but] oversold on savings.”

What I reproduced, exactly

The most creditable thing in this repository is that its headline eval is real. evals/snapshots/results.json is committed to git, and evals/README.md explains the design: each arm answers the same 10 questions, one arm runs with no system prompt, one with Answer concisely., and each skill arm with Answer concisely. plus its own SKILL.md. The honest delta is skill-versus-terse, not skill-versus-nothing — the README says so explicitly, and notes that an earlier version of the harness compared against no system prompt and therefore “inflated” its numbers.

I loaded that snapshot (generated 2026-10-02T23:09:42Z, claude-opus-5-5, Claude Code 2.1.288) and counted tokens myself with tiktoken o200k_base:

ArmPublished medianMy recomputationPer-question range (mine)
caveman (default)3% fewer−2.88%−19.63% to +15.51%
ultracave35% fewer−34.73%−58.39% to −26.65%
megacave9% fewer−9.14%−38.98% to +11.78%

They match. The published spread (“16% longer to 20% shorter”) matches my −19.63%/+15.51%. Real token totals across all 10 prompts: 6,983 with no system prompt, 4,334 with Answer concisely., 4,119 with the caveman skill, 3,683 for megacave, 2,693 for ultracave.

Two things follow, and they are the two things almost everyone who quotes this project gets wrong.

First, the honest number depends entirely on the baseline you choose. Against a stock agent that answers at default verbosity, the caveman skill removes 38.3% of output tokens. Against an agent that has merely been told Answer concisely., it removes 2.88%. Both are the same 4,119 tokens. If you already prompt your agent to be terse, caveman’s own eval says the skill is worth about three percent — and per-question it makes answers longer on some questions and shorter on others, because 10 prompts is not a powered experiment.

Second, characters are not tokens, and the gap is brutal for non-Latin registers. Measured in characters, megacave (classical Chinese) cuts the median answer by 65.32%. Measured in tokens it cuts 9.14%. The reason is visible in the same snapshot: the English terse control averages 4.21 characters per token, while megacave’s classical Chinese runs at 1.83. A Chinese-character saving is worth roughly a quarter of what it looks like. If you ever see a “65% shorter” claim measured in characters on a CJK text, treat the token number as a seventh of it — which is, coincidentally, roughly the shape of the mistake that produced this repository’s description.

The caveats the repo states about its own eval are worth repeating because they are unusually candid: it measures length only, not correctness; it is a single run per prompt-arm at default temperature; the counter is OpenAI’s BPE, “only an approximation of Claude’s tokenizer”; and, as evals/README.md puts it, “A skill that replies k to everything would score −99% and ‘win’.”

The memory-compression table that no longer reproduces

The repository also ships /caveman-compress, which rewrites your CLAUDE.md-style memory files shorter while preserving headings, code, paths and URLs. Its README publishes a five-row benchmark table averaging 46% smaller.

I ran the repository’s own harness — skills/caveman-compress/scripts/benchmark.py, using its own validate() and the same tiktoken counter — on the fixture pairs shipped in tests/caveman-compress/, which are the five files the table names:

FixtureREADME tableWhat the shipped files produce todayREADMEMeasured
claude-md-preferences.md706 → 285827 → 42159.6%49.1%
project-notes.md1145 → 5351431 → 84653.3%40.9%
claude-md-project.md1122 → 6361628 → 111743.3%31.4%
todo-list.md627 → 388880 → 64838.1%26.4%
mixed-with-code.md888 → 5601432 → 110636.9%22.8%
Average898 → 4811,240 → 82846%34.1%

Every file is slightly larger than the table says and every ratio is lower. Structural validation passes on all five pairs — headings, fences, URLs and paths survive intact, which is the part the tool actually promises.

Two honest readings, and I hold both. Reading one: this is the fifth publication of a savings number in the same repository that runs ahead of the artifacts, and a reader cannot reproduce the table from the checkout. Reading two, and I think the more important one: these fixtures are test fixtures, five files chosen so that a CI job can prove the markdown stays valid, and the table was likely generated from a different and larger corpus at some point. What I can state precisely is narrower than “the claim is false”: the table in skills/caveman-compress/README.md is about 12 points above what the artifacts in the same checkout produce today, using the repository’s own counter. That is drift, and it is checkable. It is not proof of anything about the maintainer’s intent.

Which numbers ship with the artifacts you need to check them

This is the table I would want before trusting any compression tool, so I built it for this one.

ClaimCan a stranger check it from the checkout?The artifact
3% / 35% / 9% eval mediansYesevals/snapshots/results.json is committed; I reproduced all three
46% memory-file compressionYes, and it does not reproducefixtures and outputs ship in tests/caveman-compress/; measured 34.1%
33.2% input-token cut (885,793 → 591,673)Nobenchmarks/results/ contains only a .gitkeep; the README’s own table placeholder says “No reviewed API benchmark result is published here yet” and the page states “Raw harness artifacts are not in this checkout, so treat it as a pinned report, not a public reproduction”
129.8× on browser pagesYes, with a documented boundarycorpus ships (order_dashboard.html, agent_checkout.html), browse/BENCHMARK.md gives the go test -tags=integration commands, and the same doc states the small-form case is 2.34× larger than bare Playwright ARIA text
61% pixel mode (1,069 → 415)Nothe word in the claim is “estimated”; those two numbers appear nowhere in the tree except the README sentences that assert them
38 gold-plated agent integrationsYesINSTALL.md names them; I counted 38, against a badge that says “30+”

I have no complaint about the one that cannot be reproduced. Labeling your own headline benchmark “a pinned report, not a public reproduction” is more disclosure than most projects offer, and the same page keeps a losing row (Dashboard HTML alert: +9.9%) visible with a maintainer note: “The day I hide a red row is the day you should stop trusting the green ones.” That is the right instinct. It is also the reason the 65% in the description stands out so badly — everything inside the repository is calibrated to this standard of honesty except the sentence at the top of it.

The CLI sends usage statistics by default. The README says so in plain language and tells you the opt-out (caveman telemetry off, or DO_NOT_TRACK=1). The disclosure line the CLI prints on first run is:

usage stats on — commands, agent sessions, token totals, account and install type, timezone and language, and your IP address; never prompts, code, or file paths

I read the receiver. The endpoint is a Supabase edge function (xvfgtprkhzlvegvmeefq.supabase.co/functions/v1/cli-telemetry), and its validator does three things I verified line by line:

  • Fixed vocabularies. Seven event names, RFC3339 timestamps, a UUID install ID, and closed enums for OS, architecture, exit class, error class, account type, install channel and session source. Unknown values drop the whole event.
  • The “never prompts, code, or file paths” promise is structural, not a promise. command, subcommand, agent and cli_version are validated against ^[A-Za-z0-9._+-]{0,64}$. No spaces, no slashes, no free text — a file path cannot fit in those fields. The stored schema has no free-text column at all for a payload to land in.
  • Browser posts are rejected. The handler requires a JSON content type and refuses requests carrying Origin or Sec-Fetch-Site, which stops a web page from using the endpoint to record its visitors’ IPs.

Retention is in SQL, not marketing. supabase/migrations/20260925040000_cli_telemetry_abuse_retention.sql schedules a daily job deleting rows older than 13 months, with a comment noting IPs are “already cleared at 90 days by cli-events-ip-retention” — matching the README’s promise of 90 days for IPs and 13 months for everything else. The same migration sets rate limits of 30 events per minute and 5,000 per day per sender, a global ceiling of 10,000 rows/minute, and a cap of 500 new install IDs per address per day.

That is a better-engineered telemetry pipeline, including its anti-forgery logic and its abuse limits, than most commercial developer tools ship. Now the part that a defender will not volunteer.

The consent scope has escalated twice, and an old “yes” is being spent on a new scope. The source comment in packages/cli/src/index.ts is explicit:

Version 5 = the receiver stores the client IP address with each event. v4 was default-on (opt-out) plus token volume… A stale-version “yes” was given for a narrower scope and gets the new disclosure reprinted once (never re-asked, and never flipped on).

So a user who accepted a narrower v1 disclosure is now in a regime that stores their IP with every event. The implementation is more careful than the norm here — it reprints the disclosure once, never re-asks, and never flips a previously-declined user on — but the honest description of the product is: default-on, opt-out, IP-storing, and expanded twice without fresh consent. Also, on your first run, the first-run event carries a 30-day retrospective scan of your local agent history as aggregates (session counts, tokens observed, tokens a proxy would have cut). It sends aggregates, not contents, and the validator enforces that the aggregates are internally consistent — but it is a scan of your disk, prompted by default, on a tool you installed to save money.

If that trade is not for you: caveman telemetry off, and if you want what was already sent deleted, the CLI prints your anonymous install ID one last time and the documented path is to send it in. Neither the disclosure nor the deletion route is hidden.

The repository, by the numbers

MeasureValue
Files / size (excluding .git)1,613 files / 19,130,901 bytes (18.24 MiB)
CompositionGo 27.1%, .mjs 13.1%, JSON 12.8%, PNG 10.5%, TypeScript 9.0%, Python 6.8%, Markdown 5.4%, .gz 4.9%
Largest single filepackages/cli/src/index.ts — 929,390 bytes, 20,176 lines, 4.9% of every checked-in byte
Prose files at the rootREADME.md 46,861 B, CLAUDE.md 40,641 B, INSTALL.md 24,589 B, SECURITY.md 21,397 B
Skill definitions26 SKILL.md files under skills/; 6 of them mirrored into plugins/
LicenceApache-2.0 since 3.0.0; 25 LICENSE files byte-identical to the root copy (sha256 cfc7749b…), with the pre-3.0.0 MIT text preserved separately in LICENSE-MIT and third-party notices for the pxpipe port and the Spleen/Unifont glyph atlases in NOTICE
Community71 contributors (maintainer 657 commits, claude 67, github-actions[bot] 26); 42 releases; 346,877 release-asset downloads
Open work134 open items, of which 68 are pull requests — i.e. 66 real open issues; 475 issues all-time
Version strings that disagreeroot installer 3.1.0, plugin manifest 3.1.0, CLI 2.0.0, newest release tag v3.0.0
Agent targets38 named in INSTALL.md (Claude Code, Codex, Gemini CLI, Cursor, Windsurf, Cline, Copilot, opencode, Kilo, Roo, Warp, Replit, Junie, Qoder, Antigravity, and — disclosure — Hermes Agent, the harness this article was written in)

The licence audit came back clean, and it is worth saying so plainly because licence drift is the usual way a project like this rots. Every packaged directory carries a byte-identical copy of the Apache-2.0 text; the MIT-era notices were preserved rather than overwritten; engine/pixel/ is a Go port of teamchong/pxpipe and says so with the upstream copyright; the embedded bitmap fonts carry their BSD-2 and OFL-1.1 notices. LICENSING.md even explains that the project moved from a split MIT-plus-Business-Source-License-1.1 regime to plain Apache-2.0 at 3.0.0, and that pre-3.0.0 releases keep the terms they shipped with. I have reviewed enough “open core” repositories to know this is not the default outcome.

The one file CI can never fix

There are 15 GitHub workflows, including CodeQL, Scorecard, a supply-chain job and a docs link checker. One of them, sync-skill.yml, copies the response-style skills from skills/ into plugins/ so that plugin users get identical text.

It names seven trigger paths and copies exactly five skills: caveman, ultracave, megacave, cavecrew, caveman-compress. Its git add list names the same five, plus dist/caveman.skill.

But plugins/caveman/skills/ ships six directories. The sixth is caveman-stats, and no workflow, no cp, and no git add line in this repository mentions its plugin path. From the GitHub commit API:

  • plugins/caveman/skills/caveman-stats/SKILL.md — last touched by 966a4911 on 8 September 2026
  • skills/caveman-stats/SKILL.md — changed again by 79e8440b on 9 September 2026, and again by a triage merge bd739e15 on 14 September 2026

The five skills the workflow does cover are byte-identical between source and plugin (I hashed all of them). The sixth is 25 days stale, differs by a paragraph describing how the stats hook actually delivers its report to the model, and cannot be refreshed by CI at all, because the path is not in the workflow. plugins/ is inside the npm package’s files list, so that stale copy is what plugin users install.

This is a packaging bug, not a security incident, and I am not going to inflate it: the stale file is a local stats-display command, not the compression engine, and fixing it is a two-line diff. But it is a useful reminder about what a green CI badge certifies. Fifteen workflows and a bot commit literally titled chore: sync SKILL.md copies [skip ci] (the head commit of the clone I audited, 3 October 00:38Z) do not prove that a mirror is complete. Only the list of paths does — and this list has a hole.

Where the money actually is

The repository’s own numbers make an argument its marketing does not. Its skill saves 3% of output tokens against a terse control; its proxy saves 33.2% of input tokens. Its README draws the conclusion itself: an agent’s bill is mostly reading — logs, test output, diffs, half the repo, re-sent every turn — and no talking style fixes that. That is why the proxy exists at all, and it is the reason the JetBrains result was worth acting on rather than arguing with.

The corollary is the interesting part, and it is a cost problem the project states about itself: every skill you install is prompt text your agent reloads on every call. Its own README estimates the full ruleset at about 1,000 input tokens per call. In issue #145, a user measured the overhead at "~800–1200 tokens per turn for caveman rules block, plus ~300 for the skills list" and concluded the skill is net-negative for terse technical Q&A — which is why docs/HONEST-NUMBERS.md lists exactly that case under “when caveman loses.”

Three more documented losses, all from the project’s own pages:

  • Billing by request instead of by token. Issue #506: on GitHub Copilot, premium requests are charged, and a shorter answer is still one request. A skill that shortens answers has no mechanism to help you there.
  • Variance eating the average. In JetBrains’ 82 paired tasks the skill arm should have been about 10% cheaper; it came out 11.6% more expensive (USD 40.60 vs 36.39) because a single dependency-audit task crossed into long-context pricing at USD 8.29 against USD 0.33. Their own framing: the saving is real but fragile.
  • The tail. Issue #550 reports one Cursor A/B at 4.3M tokens with caveman versus 1M without and twice the wall-clock time — a run the maintainer’s own page says was not reproducible, and lists anyway.

The pattern generalises beyond this project, which is why JetBrains’ series is worth reading as a series. Part 1, caveman: advertised −65%, measured −8.5%. Part 2, rtk: advertised −60–90% of shell output, measured +7.6% median cost per task (p = 0.004) because the wrapper’s own overhead exceeded what it saved. Part 3, ponytail: advertised −54% code, measured median −15% code and about −10% tokens, with the honest note that the median and the mean on skewed data are different animals. Three viral “token saver” add-ons, three headline claims between 54% and 90%, none of them surviving contact with 80+ paired tasks under a fixed budget.

What I would actually do

  1. Classify your billing first. Per token → keep reading. Per request or credit → a shorter answer is the same invoice; skip anything that only shortens prose.
  2. Measure before and after with the provider’s own numbers. The repository ships caveman trial -- claude, which runs a real session both ways, and its own documentation says that A/B outranks every number on its README page. Take it at its word. Token counts from a local counter are not an invoice.
  3. Decide separately about the two halves. The skill is worth a high-single-digit output cut at best, and on terse Q&A it can be net-negative once its ~1,000 tokens per call are counted. The proxy, which attacks the input stream, is where the 33% lives — and it is also the half whose headline artifacts are not published, so test it on your own long-log sessions before believing it.
  4. Turn telemetry off if you care, on day one. caveman telemetry off. It is one command, it is documented, and it works.
  5. Prefer the skill over the pixel trick. Converting SKILL.md into PNG pages so the model reads it as an image is the most elegant idea in the repository and its only number is an estimate (“1,069 to 415 estimated tokens”). The engine it ports, pxpipe, is described on this project’s own comparison table as lossy with silent misses. Elegant is not measured.
  6. If it does not win on your workload, uninstall it. That is the maintainer’s advice, not mine: “Compare provider-billed totals on the same task with and without Caveman. If Caveman increases billed cost for the same task, turn it off for that workload.”

FAQ

Is the 65% claim false? No, and that matters. It is a best-case number for prose-heavy chat output, and independent work supports that range: Adobe Research measured output compression cutting realized cost 1.4–2.4× (up to 3×), and Elastic’s internal version reported 63.6%. What is wrong is that the number travels without its context. On agentic coding tasks with the skill forced on, JetBrains measured 8.5% and called it a ceiling; the project’s own eval measures 3% against a terse control.

So does the skill do nothing? It does something, and the size depends on your baseline. Against a stock, verbose agent, the committed snapshot shows 4,119 tokens where the unbriefed agent writes 6,983 — a 38.3% median cut by my recomputation. Against an agent already told Answer concisely., it is 2.88%. If you already prompt for brevity, expect a rounding error. If you do not, the skill does real work — for a high-single-digit price effect, not 65%.

Did I just read that the memory compressor overstates by 12 points? I measured 34.1% average on the five fixture pairs the README names, using the repository’s own benchmark script and validator, against a published 46%. All five pairs still validate structurally. The narrower claim — “the table no longer matches what the shipped files produce” — is what my numbers support.

Why does caveman’s own benchmark say its input-token saving is the good one? Because in agent sessions the tokens are mostly read, not written: context, tool output, logs and diffs are re-sent on every turn, while model output is dominated by code and tool calls that the skill deliberately leaves byte-exact. That is a fair analysis and it points in a direction the product itself followed — the proxy, not the voice.

What exactly does the CLI send home? Command and session events with a random install UUID, token counts through and cut, OS/architecture/Node version, CLI version, install channel, signed-in state, timezone and locale, and the IP the event came from. IPs are cleared at 90 days, other rows at 13 months, both enforced by scheduled SQL I read. Never prompts, code or file paths — and that part is enforced by server-side validation that cannot carry a path in the fields it accepts. On first run, one event also carries aggregate counts from a 30-day scan of your local agent history. caveman telemetry off stops it; the install ID printed at that moment is how you ask for deletion of what was already sent.

Should I install it? The condition under which it is clearly worth it: you pay per token, you run long sessions full of logs, test output and diffs, and you would run the proxy rather than only the persona skill. The condition under which it is not: you are billed per request, or your conversations are short and terse. Both conditions come from the project’s own documentation, which is the strongest reason I trust the rest of it.

Bottom line

109,102 people starred a tool whose most-read claim, 65%, its own repository no longer stands behind: the same project publishes 3% for the same skill, kept a 33.2% proxy result whose artifacts never shipped, and ships a memory-compression table 12 points above what its own fixtures produce. It also publishes a page that tells you when to turn it off, keeps a red row in its own results on purpose, ships the raw answer pairs for its eval so a stranger can recompute them exactly — which I did — and enforces its privacy claim in server-side code rather than in marketing prose.

This is what a good-faith viral developer tool looks like in 2026: honest in its documentation, oversold in its description, and about a quarter as effective as the number you first heard. Read the numbers page before the pitch, measure both halves separately, and turn telemetry off on day one.