“Give your agent CAD superpowers” is the entire pitch: 17,978 stars say the pitch landed, and the top-half GitHub Trending list — plus a Hacker News thread that hit 186 points on 2026-05-01 — says people believe it. So I installed it, built a real part with it, and measured everything the README does not put a number on.
The short version: this is one of the more disciplined agent-tooling repositories currently shipping (76 releases in 129 days, none without release notes, MIT, byte-determinism as a documented promise), and the interesting part is precisely where its guarantees stop. The project is fanatically rigorous about the layer a machine can check — source into bytes, geometry into verdicts — and silent about the layer where the ambiguity actually lives, which is your sentence about the part you want.
The numbers that reproduce
Everything below came out of the GitHub API, the PyPI JSON API and pypistats, not the README.
| Metric | Value |
|---|---|
| Stars / forks / watchers | 17,978 / 1,804 / 93 |
| Watcher ratio / fork ratio | 0.52% / 10.0% |
| Releases (2026-05-30 → 2026-10-06) | 76, latest v0.7.15 |
| Release notes | 76 of 76 non-empty, median gap 16.5 hours, 24 in the last 30 days |
| Contributors | 18 accounts, 1,391 commits |
| Commit concentration | earthtojake 1,110 (79.8%), github-actions[bot] 102 (7.3%), claude 87 (6.3%) |
| Repository size | 2,902 files, 34.5 MB of blobs, 1,436 Python files |
| cadgen versions on PyPI | 53 since 2026-08-11 |
| cadgen downloads | 42,213 in 56 days; 23,603 in the last 30 (≈787/day) |
Two of those rows deserve a second look.
93 watchers against 17,978 stars. Watchers are the people who asked to be notified about change; a 0.52% ratio is low even by the standards of viral repositories. Forks, on the other hand, are 10.0% — high, which fits a project whose install instructions are a git clone for one of the five supported agent apps.
claude with 87 commits, plus 102 from github-actions[bot]. The GitHub interface shows commits authored by “earthtojake and claude”, and the bot account is doing release plumbing. Counting the two, roughly 13.6% of commits are machine-authored; the maintainer is 79.8%. That is a bus factor of one with an automated tail — worth knowing before you build a production pipeline on it.
What the agent actually receives
This is where the marketing vocabulary and the architecture diverge. “Plugin” and “MCP server” suggest a protocol surface of CAD operations. I spoke the protocol to it directly — a JSON-RPC initialize followed by tools/list over stdio against cadgen mcp — and got this:
| Layer | What it exposes |
|---|---|
| MCP tools | 2: cad_show (one argument: path), cad_analytics (one argument: action) |
| MCP resources | 1: ui://cad/…/app.html, text/html;profile=mcp-app |
| Skills (prompt-level) | 12, matching the repository tree exactly |
| CLI | 30+ verbs (step build, stl build, step snapshot, dxf snapshot, urdf validate, store why, …) |
The MCP server is a viewer, not a modelling API. Its job is to render a file the agent points at (viewer cards in a chat host, a sidebar tab in Codex) and to switch anonymous analytics on and off. The CAD work itself happens in twelve skill packages — CAD, step.parts, engineering-drawing, DXF, URDF, SRDF, SDF, SendCutSend, DfAM check, DFM, G-code, Bambu Labs — which are markdown instructions that tell the agent to write a Python file and shell out to cadgen.
That architecture is defensible (the CLI is the same door for every agent, and the 12 skills are versioned with the wheel), but it has a consequence you should price in: the agent’s CAD ability is a prompt contract, not a typed tool surface. Nothing validates the geometry the agent claims to have produced except the checks the skill text asks it to run.
The runtime bill, measured
cadgen-0.7.15-py3-none-any.whl is 11.01 MB. Here is what installing it produced in a clean uv virtualenv on this machine:
| Item | Measured |
|---|---|
| Wheel | 11.01 MB |
| Installed site-packages | 822 MB (75× the wheel) |
OCP (OpenCascade binding) | 158 MB |
cadquery_ocp_novtk.libs | 95 MB |
playwright (pinned ==1.63.0) | 137 MB |
scipy + scikit-learn + numpy + sympy + matplotlib | ≈195 MB |
| Install wall time | 12.5 s (warm package cache) |
| Bundled runtime inside the wheel | 22 MB — viewer client 14 MB, browser render bundle 1.4 MB, Node 0.4 MB |
The dependency list is short and mostly unavoidable: build123d>=0.11.1,<0.12, cadquery-ocp-novtk>=7.9,<8, ezdxf, shapely, matplotlib, pillow, and that exact-pinned Playwright. Note the two deliberate choices visible in it: the -novtk OCP build exists to skip VTK, and Playwright is pinned exactly rather than ranged. Rendering is genuinely a browser problem here — snapshot_core.py describes “the headless browser driver” and loads render.html plus snapshot-render.js, with tessellation floors mirrored between the Python renderer and packages/core/src/common/source.js and covered by a parity test.
So the honest first-run figure is: 822 MB of Python environment, plus a headless Chromium on the first snapshot (261 MB for the chromium_headless_shell build in this machine’s Playwright cache; the snapshot itself ran fine because a compatible browser was already present). The README says “Its first start downloads CAD’s runtime” without a size. It is a laptop-class install, not a curl | sh.
The daemon you did not ask for
After one python src/mount_plate.py — a 60 × 40 × 5 mm plate with four M3 clearance holes, built by eleven lines of build123d — the build itself took 1.55 seconds wall clock, and 0.19 seconds on the second run, which correctly answered “STEP/mount_plate.step is current; not rebuilt”.
Then I looked at the process table:
| Process | RSS |
|---|---|
python -P -m cadgen.daemon | 284 MB |
worker ×3 (cadgen.daemon.worker) | 499 MB, 477 MB, 471 MB |
| Total resident | 1.65 GiB |
Four processes, 1.65 GiB, still resident a minute and a half after the script exited. The daemon’s own log states the terms: idle timeout 3600s. Reading the source gives the design constants:
DEFAULT_SPARES = 2— spare workers that “have finished importing build123d and are bound to nothing”, so a new model’s first build pays no import cost.WORKER_SEED_BYTES = 512 MiB— what one worker is charged at admission until the pool has measured a real idle worker.- Memory admission defaults to 70% of physical memory (
CADGEN_MEMORY_MBoverrides it,0disables admission). DEFAULT_IDLE_UNBIND_SECONDS = 600.0— a bound worker unbound after ten idle minutes;DEFAULT_RECYCLE_AFTER = 1000jobs per worker as a leak hedge.
The stated rule in the pool’s docstring is “nothing waits on another build”. That is a real engineering answer to a real problem — an agent making ten edits in a row should not pay ten kernel imports — and it is also an hour of ~1.7 GiB on a laptop for a part that took a second to build. On the 8 GB Linux box I measured on, that was 21% of physical RAM committed by a tool you invoked once, silently, as a side effect of running a Python file. The store it maintains under ~/.cache/cadgen has a 20 GB cap and eight index categories (model, document, output, component, surface, bounds, mesh, drawing); the housekeeping that keeps it under the cap runs when the daemon is idle.
The 19 laws, and the one they cannot cover
The most unusual thing about this project is on its PyPI page, not its README. The package description is 28,118 characters of specification titled “The design laws” — 19 numbered laws, with pressure-tests, that the code is expected to obey. Three of them are worth quoting because they predict how the tool behaves when you push on it:
8. No backwards compatibility. Hard cutovers only. Every retired surface fails loudly with a teaching error naming its replacement — never an alias, never a shim.
10. Loud failure or correct output, nothing between. The cardinal sin is plausible-wrong output at exit 0. No silent fallbacks, no globs, no guessing; a failed render leaves NO file at the requested path.
19. Nothing waits on the network. Once installed, cadgen works offline: every command, build, snapshot, server and page.
Law 10 is the one to hold the project’s own benchmarks against. In the 2026-05 Hacker News thread, a commenter who goes by voidUpdate did exactly that, and their three observations have not been answered in the README since:
“In the benchmarks, there is a strange lack of measurements that I’d expect in a CAD process (EG in benchmark 1, the positions of the 4 holes are not specified at all). I’m assuming that’s why the gussets in benchmark 3 overlap the holes and make a part that cannot be used.”
A part whose gussets overlap its holes is plausible-wrong output. It satisfies law 10 at the layer cadgen controls — the source is deterministic, the STEP is valid, the snapshot renders — and violates the intent of law 10 at the layer it does not: mapping an underspecified sentence onto geometry. That is the honest boundary of this design. Enforceable contracts stop where natural language starts, and no amount of byte determinism reaches across it.
The same thread contains the cleanest statement of the other half of the problem, from randusername:
“I don’t feel like text-to-CAD is a viable workflow for me because of the ’language barrier’. I would need, like, a visual dictionary of terms.”
alnwlsn, arguing from drafting practice, put the counter-position: “What a nightmare to describe all this in text! when the language of drafting is able to describe it perfectly, wordlessly and unambiguous, in a single drawing sheet.” Eisenstein reported the sobering version from real use: the trick with LLM CAD “is to be an expert in the field already”.
One original finding: the SURF container
The repository ships a geometry format that appears in no document I could find: 19 files ending in .surf, described nowhere in the README. I decoded one. sun_gear.surf (332,302 bytes) begins:
SURF \x02\x00\x00\x00 \x8a\xc8\x04\x00 {"version":2,"shapes":[{"ord":1,"kind":"solid","volume":12491.658278013321}],
"faces":[{"ord":1,"shape":1,"reversed":true,"uv":[...],"surfaceType":"plane","area":48.371601946768976,
"center":[...],"bbox":[...],"surface":...}]}
A four-byte magic, a version, a length-prefixed JSON index carrying per-face B-rep metadata (surface type, uv, area, centre, bounding box) and per-solid volume, followed by a raw float64 payload. It is the viewer’s own compact container for the docs hero and fixture assets — the same data STEP carries, without the text overhead. Undocumented, MIT-licensed, and a reasonable sign of how far this project will go for its render path.
Where adoption actually shows
Star counts measure attention. Downloads measure use, so I checked both distribution channels:
| Channel | Number |
|---|---|
GitHub release assets (cad-openai-plugin-0.7.15.zip) | 45 downloads |
| Same asset, previous eight releases | 56, 2, 4, 9, 16, 3, 2, 2 |
PyPI cadgen downloads, last 30 days | 23,603 (peak 3,184 on 2026-10-05) |
PyPI cadgen downloads, first 56 days | 42,213 |
The GitHub release page is a rounding error — nobody installs a plugin from a release asset here — while PyPI moves about 787 installs a day, up from a first-day 876. The download counts include CI and uvx re-resolutions (each uvx --from cadgen==0.7.15 run may re-resolve), so treat 23,603/month as an upper bound on humans. Even so: 17,978 stars, 93 watchers, and roughly 800 package downloads a day is a coherent picture of a tool that many people tried and relatively few run continuously.
Practical guidance
Install it if: you are doing parametric, machine-checkable work (brackets, fixtures, adapters, enclosures) with dimensions you actually know; you want STEP and DXF plus DFM and printability checks inside the same agent session; you are on macOS or Linux; you have disk and RAM to spare.
Skip it if: you are on a fresh Windows 11 machine with Smart App Control on — the README is refreshingly blunt that the unsigned native OCP module is blocked, every cadgen command fails with ImportError: DLL load failed while importing OCP, the Event Viewer logs Event ID 3077, and your only options are turning Smart App Control off (irreversible without reinstalling Windows) or running under WSL. Skip it too if your parts are artistic rather than dimensional, or if you need to run it inside a sealed container where a 1.65 GiB idle daemon and a browser dependency are unacceptable.
Audit the store before you trust the cache: cadgen store why <model>.py prints a five-check gate (record, source closure, children, tree hash, outputs) and explains an unexpected rebuild in one command. That is better observability than most build tools of this vintage ship.
Two operational notes from the source that are not in the README: the daemon’s memory budget defaults to 70% of your physical RAM and there are two spare workers pre-imported by default, so both numbers scale with your machine rather than with your part; and the analytics are off until answered, send a one-way code plus each file’s format (never names, paths, or prompts), and stop entirely with DO_NOT_TRACK=1. The daily update check talks to api.texttocad.dev/v1/versions at most once and can be disabled with CADGEN_UPDATE_CHECK=0.
FAQ
Is text-to-cad the same as cadgen?
No. The repository is the plugin plus twelve skills; cadgen is the PyPI package (11 MB wheel, 225 Python modules in 0.7.15) that does the work. The skills pin an exact version, and the plugin and the CLI share one installation.
Does it generate STEP files directly from an LLM?
No. The LLM writes Python that uses build123d; cadgen runs it and exports STEP, STL, 3MF, GLB or DXF through OpenCascade’s OCP binding. Version-to-version, the CLI is generated from the public function signatures, and law 6 makes the decorator, the function and the CLI one named surface.
Why does a CAD tool depend on Playwright?
Snapshots. The render path drives a headless browser over render.html and snapshot-render.js, which is how it can produce solid, render, xray, hidden-line and wireframe views, sections, and .mp4/.gif video of animation clips. In my run the first snapshot cost 9.7 seconds and produced a 120 KB PNG; render mode cost 27.7 seconds for 1.6 MB.
Is the daemon optional?
Not via a documented flag. Idle timeout is one hour, spare workers default to two, and the pool’s stated rule is designed for interactive agents rather than one-shot batch use. CADGEN_DAEMON_SPARES, CADGEN_DAEMON_IDLE_TIMEOUT, CADGEN_DAEMON_IDLE_UNBIND and CADGEN_MEMORY_MB exist as environment variables and are the knobs to use.
What did the benchmarks get wrong?
The most specific public criticism, from voidUpdate on Hacker News, is that the headline benchmark prompts do not pin hole positions, that the L-bracket specification asks for gussets that overlap its holes, and that benchmark 7’s through-hole does not visibly pass through. If you copy those numbers into a procurement document, add the missing dimensions yourself.
Does it work offline? Per law 19, yes once installed: builds, snapshots, the viewer and the MCP server all work with every network request refused. Only two requests exist — anonymous analytics (opt-in) and the daily version check — and neither blocks a command.
The verdict
The engineering here is more serious than the star count prepares you for: content-addressed build caching with an explainable gate, byte-determinism as a promise with the kernel explicitly excluded from it, closed vocabularies and teaching errors instead of aliases, 76 releases with notes in four months, and a PyPI page that publishes its own constitution. The bill is real too, and it is invisible in a one-line install instruction: 822 MB of environment, a browser, up to 1.65 GiB of idle daemon, and an MCP surface of exactly two tools because the capability was never in the protocol to begin with.
The thing to internalise is the boundary. This tool will faithfully, reproducibly, verifiably build the part you specified. It will not tell you that you failed to specify the hole positions — and law 10, however loudly it is written, only governs the half of the pipeline where a machine is checking.