Multimedia

40,145 Stars and a 6-Point Hacker News Thread: Auditing VoiceStudio, the 'Fully-Local ElevenLabs Alternative' Whose Default Engine Is CC-BY-NC

VoiceStudio (AGPL-3.0, formerly OmniVoice Studio) reached 40,145 GitHub stars in under six months on the promise of 'fully-local voice cloning, dubbing and audiobooks in 646 languages', but its bundled default engine, k2-fsa/OmniVoice, is licensed for code under Apache-2.0 and weights under CC-BY-NC, and the project's own engine-acceptance rule requires commercially clean weights. I verified the 646-language and 581k-hour claims by summing the repo's own language table, rebuilt the star curve from Wayback-archived GitHub API snapshots, audited the three-licence history, and priced a local workflow against published ElevenLabs rates.

40,145 Stars and a 6-Point Hacker News Thread: Auditing VoiceStudio, the 'Fully-Local ElevenLabs Alternative' Whose Default Engine Is CC-BY-NC

There is a sentence in VoiceStudio’s Hugging Face dependency that decides whether the product is usable for the thing most people want it for. It is not in the README, not on the website, and not in the app. It is in the model card of the engine the app installs by default:

Our code is released under the Apache 2.0 License. The pre-trained model is licensed under the CC-BY-NC due to constraints from its training data (e.g., Emilia).

CC-BY-NC means non-commercial. VoiceStudio’s own landing page tells you the opposite — “Run it, for anything, including at work and for money… no field-of-use restriction.” The application really is AGPL-3.0, and running the application commercially really is fine. The model weights it ships with are not, and one of them is the default engine.

That gap is the story. Everything else — 40,145 stars, 534,737 release downloads, 101,256 Docker pulls, 646 languages — is true, and I verified each number against a primary source. So this is an audit in two directions: exactly how much of the marketing holds up under arithmetic, and exactly where the arithmetic stops being friendly.

What VoiceStudio actually is (verified 28 September 2026)

MetricValueSource
Repositorydebpalash/VoiceStudio (renamed from OmniVoice-Studio, Aug 2026)GitHub API
Created9 April 2026GitHub API
Stars40,136 → 40,145 during a 35-minute sessionGitHub API
Forks / watchers / open issues4,769 / 159 / 17GitHub API
Commits on main3,381, of which owner debpalash wrote 2,781 = 82.3%GitHub API
Contributors71 via API; next-largest human has 54 commitsGitHub API
Files in tree3,280 blobs — 1,004 .py, 450 .ts, 420 .jsx, 300 .tsxgit tree API
Code weight11.56 MB Python, 4.96 MB JS, 3.87 MB TS, 0.14 MB Rustlanguages API
Tests / docs / scripts638 / 203 / 52 filesgit tree API
App licenceAGPL-3.0-onlyLICENSE
Releases42; latest v0.5.6 (23 Sept 2026)releases API
Release downloads534,737 across all 42 releases; 56,964 on v0.5.6 alonereleases API
Docker pulls101,256 (palashdeb/omnivoice-studio)Docker Hub API
Engine count16 text-to-speech, 11 speech-to-textdocs/engines/README.md

Architecturally it is an Electron desktop shell over a FastAPI backend (97 HTTP endpoints, Server-Sent Events for streaming, SQLite for state) with four ML workhorses underneath: WhisperX for ASR, Demucs for stem separation, pyannote for diarisation, and AudioSeal for watermarking. The desktop shell was migrated from Tauri to Electron in September — v0.5.3 was the last Tauri release, and the legacy shell was deleted outright in PR #2343 rather than kept behind a flag.

Two things about that table are worth sitting with. First, 3,381 commits in 172 days is roughly 20 commits a day, sustained — this is not a weekend project with a nice README. Second, 82.3% of those commits come from one person, and the contributor who ranks third overall is dependabot[bot]. The bus factor is one.

The star curve, rebuilt from snapshots

GitHub’s stargazers timeline is closed to unprivileged tokens, so I pulled it out of the Wayback Machine instead: archived copies of api.github.com/repos/debpalash/VoiceStudio, gzip-decoded.

Archived dateStarsForksOpen issuesWatchers
2026-09-0113,3532,0213165
2026-09-1022,2552,742790
2026-09-1127,2773,36047116
2026-09-2635,7744,19720144
2026-09-28 (live)40,1454,76917159

That is +26,792 stars in 27 days — a doubling time of about two weeks, with a one-day jump of +5,022 between 10 and 11 September. Trendshift’s history corroborates the shape: first appearance on GitHub Trending on 3 September, first #1 on 15 September, “created 6 months ago.”

Now compare the landing page. Its social-proof block reads “Observed 21 Aug 2026” and shows: GitHub stars 40,140, GitHub releases 221,468, Docker Hub 58,945, GHCR 14,000+, Release v0.5.0. Every one of the trailing numbers is a five-week-old snapshot — the current release is v0.5.6, and Docker pulls are 101,256, not 58,945. But the star number is live: it tracked 40,136 → 40,140 → 40,145 within minutes of the API doing the same.

So the block is a live counter labelled with a frozen date. And that label is provably wrong: the repo had 13,353 stars on 1 September, which means it could not have had 40,140 on 21 August. A reader who trusts the label concludes the project has been flat for five weeks. A reader who trusts the number gets the real picture — explosive growth, partly because the counter itself is one of the more effective growth loops in the repo.

The counterpoint is uglier. Hacker News, the channel where a project like this would normally be vetted, has essentially ignored it: nine submissions total, top score 6 points and 0 comments for “VoiceStudio – local open-source ElevenLabs alternative” (15 September), and no submission of the domain at all. The upstream model’s own submission, “OmniVoice, high-quality TTS for 600+ Languages,” scored 3 points. 40,000 stars arrived through Instagram reels, LinkedIn posts and YouTube demos; the developer channel never engaged. Stars measure reach, not review — and for a tool you would put client work through, that distinction matters.

Claim one: 646 languages — verified to the row

This one is real, and I checked it the boring way: the repository publishes docs/languages.md with one row per language, an ISO 639-3 code and a training-corpus duration. The table contains exactly 646 data rows, the renderer’s languages.json contains 647 entries (646 languages plus an “Auto” sentinel), and there are zero duplicate ISO 639-3 codes. The durations sum to 581,489.4 hours, which matches the “581k hours” claim exactly. That number of hours also matches the upstream paper’s description of the dataset it was trained on.

Then I sorted the durations.

Training data for the languageLanguagesShare of catalogue
≥ 10,000 h101.5%
1,000 – 10,000 h233.6%
100 – 1,000 h436.7%
10 – 100 h30547.2%
1 – 10 h25439.3%
< 1 h111.7%

The distribution is what you would expect from a corpus built out of open data, and it is brutal: 265 of the 646 languages have under ten hours of training material, and the median language has 10.18 hours. The bottom half of the catalogue — 323 languages — sums to 2,715 hours, or 0.47% of the corpus. The ten most-resourced languages account for 84.0%, and English alone is 206,061 hours = 35.4%.

The eleven languages under one hour are the ones that make the headline look strange. Macedo-Romanian has 0.02 hours of training data — 72 seconds. Haitian: 0.04 h. Zacatlán-Ahuacatlán-Tepetzintla Nahuatl: 0.05 h. Tigrinya: 0.08 h. Votic: 0.1 h. Only 76 languages reach 100 hours; only 33 reach 1,000.

None of this is a scandal — it is how multilingual models work, and the upstream authors are explicit that the number is a coverage claim built on open-source data. But “646 languages” is not a claim about 646 usable voices. It is a claim about 646 language codes the model was pointed at, of which roughly three dozen have enough data to be a production voice and several hundred are best described as supported with a wink. If your workflow needs Tamil, Mandarin, Spanish or English, this is a non-issue. If you bought the headline for Nahuatl, you own 180 seconds of someone’s field recording.

Claim two: the licence stack, layer by layer

This is where the audit stops being flattering.

Layer 1 — the application: AGPL-3.0-only. Genuinely permissive for use. The project’s own licence page spells it out: “Using the desktop app at work. Nothing required.” Fine. The AGPL’s network clause only bites if you host a modified version for other people.

Layer 2 — the default engine’s weights: CC-BY-NC. k2-fsa/OmniVoice (10 authors, k2-fsa/icefall group, including Daniel Povey; arXiv 2604.00688, submitted 1 April 2026; a 0.6B diffusion-language-model-style TTS model initialised from Qwen3-0.6B) is Apache-2.0 as code and CC-BY-NC as weights, explicitly “due to constraints from its training data (e.g., Emilia).” The Hugging Face card’s YAML front matter carries no license: field at all, so the hub renders no licence badge; the only statement is the sentence in the body. The model has 1,378,254 downloads in the last month and 1,436 likes, and it is the single most-used engine in the app — installed by default, and the one whose voice-design and cloning quality the product is marketed on.

Layer 3 — the audio tokenizer: a Llama-3-derived community licence. The bundled audio_tokenizer/LICENSE is the “Boson Higgs Audio 2 Community License Agreement”, which is “based upon the Meta Llama 3 Community License Agreement” and incorporates its terms by reference. That is a third document, in a subdirectory, with its own conditions.

Now put the project’s own rules next to that. docs/engine-acceptance.md sets the bar for accepting a new engine, criterion 2:

Licence clean for commercial use. Model weights and code. No research-only weights or ambiguous provenance.

The project even enforces that rule visibly. audiocpp (Breeze-TTS-2 weights) is labelled in the engine table as “weights research/non-commercial” and is the one owner-approved exception; supertonic3 and pockettts are behind a local licence-acceptance dialog (LICENSE_GATED_ENGINES in backend/core/engine_licenses.py is exactly {"supertonic3", "pockettts"}). Meanwhile the engine that holds the catalogue’s top job — “Best zero-shot clone quality” and “Widest language coverage” in the project’s own job map — carries no licence caveat in the engine table at all. The only place the restriction is stated is LICENSE-NOTICE.md, in a paragraph most readers will never open.

The licensing has also moved three times in ten weeks, which tells you the owner is actively thinking about it:

DateLicenceCommit message
2026-04-29FSL-1.1-ALv2chore(license): switch Studio to FSL-1.1-ALv2; commercial pricing TBD
2026-06-06AGPL-3.0-onlychore(license): relicense from FSL-1.1-ALv2 to AGPL-3.0 (open-core)
2026-07-16AGPL-3.0-onlydocs: live downloads badge + fix AGPL-3.0 license detection (#1168)

Note the parenthetical: open-core. And note that the July commit was about making GitHub’s detector report AGPL-3.0-only instead of “Other” — because corporate licence scanners gate adoption, and an unclassified licence blocks procurement. LICENSE-NOTICE.md says so explicitly. This is a project that understands licence mechanics very well when the licence is its own.

The asymmetry is the finding. The app’s own licence is managed meticulously for enterprise adoption. The default engine’s non-commercial weights are mentioned once, in a notice file, and then never surfaced again — not in the engine catalogue, not in the app’s engine picker, not in the licence page’s “Models and dependencies” section, which says only that “a few are research-only.”

The Pro tier makes the contradiction sharper

On 25 September, PR #2333 merged a Pro entitlement module and a design contract, docs/specs/desktop-pro-page.md:

VoiceStudio Free retains unlimited local voice creation, transcription, dubbing, Stories, Audiobooks, batch jobs, and local exports. Desktop Pro adds production recipes, watch folders, persistent batch rules, revision history, client delivery packages, project preflight, remote-device compute, remote workers, encrypted GPU sharing, and production tools for paid client work. Individual pricing is $99 per user yearly or $299 per user lifetime.

The Electron main process already contains working licence plumbing: pro-license.ts calls https://api.lemonsqueezy.com/v1/licenses for activate/validate/deactivate, takes store/product/variant IDs injected at build time, stores the key with Electron’s safeStorage and refuses to fall back to the basic_text backend. The rendered Pro page shows a $99/year card, a $299 lifetime card, a 1–99 seat quantity picker, and links out with ?plan=…&quantity=….

The website, meanwhile, says none of it exists: “Cloud, the account dashboard, public API-key use, payment checkout, and automatic Commercial License delivery are not available in the current public release,” the Pro comparison table lists “Online checkout: Not needed / Not available,” and the licence page says “The Pro page does not currently sell or grant a licence.” The spec agrees, listing unset release gates: VOICESTUDIO_PRO_FEATURES_RELEASED, VOICESTUDIO_PRO_TERMS_APPROVED, and “No paid checkout or feature unlock is live until the gates pass.”

So the honest reading is: the paywall is half-built and correctly gated, and the marketing describes a $99 product whose most valuable feature for a paying customer — “production tools for paid client work” — sits on top of a default engine whose weights forbid commercial use. If you buy the tier that exists to monetise the tool, you still have to leave the default engine to use it commercially. A commercial licence for VoiceStudio’s own code does not touch that, and the project’s own notice says so: “A commercial license for VoiceStudio-owned code does not replace any of those terms.”

Nobody has measured the audio

Two artefacts settle how much quality evidence exists.

docs/benchmarks.md opens by promising that “Every number here is produced by the in-repo harness, on named hardware, at a named version; nothing is estimated” — and then, under Results: “No verified rows yet — this table fills from maintainer runs and community submissions,” with a single italic placeholder row. A 40,000-star project with a shipped benchmarking harness has published no benchmark rows.

The semantic eval tier is worse. .github/workflows/evals.yml:

LLM-judge evals — semantic quality suites, NEVER a gate… On the hosted runner there is no local LLM endpoint, so the run usually reports “skipped — no LLM backend configured”.

And tests/evals/run_evals.py “exits 0 whether cases pass or fail — the report is the deliverable.” The weekly job therefore usually does nothing, and when it does something, it cannot fail the build. That is a defensible engineering choice, but it means the quality claims in the README rest entirely on the upstream model’s paper — WER/SIM-o/UTMOS runs over LibriSpeech-PC, Seed-TTS, FLEURS and MiniMax that the model authors ran on the model, not on this app’s pipeline, its chunker, its dubbing stage or its CPU offload path.

The informal evidence points the same way: the largest install document in the repo is docs/install/troubleshooting.md at 72,160 bytes, roughly 6× the size of the README.

What is genuinely better than the alternatives

Having taken the project apart for four thousand words, the case for it is real and specific.

Its privacy engineering is the best I have seen in a local-first app. backend/core/analytics.py is an opt-in PostHog client held to a standard most commercial SaaS does not meet: two independent gates (a committed publishable token, overridable by POSTHOG_PROJECT_TOKEN, and a user preference that defaults to False), a hard kill switch (OMNIVOICE_ANALYTICS_DISABLED), exception autocapture explicitly disabled because SDK tracebacks can carry Hugging Face tokens and absolute paths, and an allowlist that drops any property not on the list rather than trusting the caller. A default install transmits nothing. Translation providers are classified in code as online (google, deepl, mymemory, microsoft, openai) versus offline (nllb, argos, libretranslate) so the UI can tell you which mode you are in.

Provenance and consent are first-class, not marketing. Meta’s AudioSeal watermarking is wired in with its own settings control and route-coverage tests; voice profiles carry a database-level consent flag (alembic 0003_voice_profile_consent); cloning references longer than 20 seconds are handled deterministically (it picks the best 15-second speech window and reuses the already-installed ASR instead of hallucinating a transcript for a passage it cut). Among the projects in its own competitive set, it is the only one that does watermarking and detection.

It is one application, not seven tools. Dubbing, cloning, dictation with a system-wide widget, batch queues of up to 50 videos, audiobooks, a local OpenAI-compatible API and an MCP server in the same install. If you have ever wired WhisperX + Demucs + a TTS repo + ffmpeg by hand, the value is obvious.

And the money maths is not close. Published ElevenLabs pricing as of today: Free $0/10k credits, Starter $6/30k credits with commercial licence, Creator $22 (121k credits), Pro $99 (600k credits), Scale $299 (1.8M credits, 3 seats), Business $990 (6M credits, 10 seats), “low-latency TTS as low as 5c/minute.” Dubbing is where it bites: their model bills 10 minutes translated into 3 languages as 30 minutes, with $0.60/minute overage — the project’s own competitive research quotes users describing the switch as killing “$700/year in ElevenLabs and HeyGen subscriptions.” At the Pro rate ($1,188/year) a €300 used RTX 3060 pays for itself in about three months, and after that local generation costs electricity.

Cost of going local, honestly

ItemReality
GPUOptional. CPU works, roughly 3× slower. At ≤ 8 GB VRAM the app auto-offloads TTS to CPU during transcription
Disk10 GB minimum, 20 GB+ SSD recommended, before you install a second engine
First runWeights download automatically on first generation (the OmniVoice weights alone have 1.38M monthly downloads on the Hub)
DiarisationNeeds a Hugging Face token and accepted licences for pyannote/speaker-diarization-3.1 and pyannote/segmentation-3.0; without them dubbing silently degrades to a silence-gap heuristic that merges similar-pitch speakers
IterationNo multi-track timeline editor and no crossfade chunking yet — the two gaps the project itself admits against voicebox (MIT, ~29.7k stars), which also ships 50+ preset voices against VoiceStudio’s 20+. Generate-then-fix still costs human time, which at agency rates dwarfs the subscription it replaced
AMD / Intel / AppleCUDA, MPS and ROCm are auto-detected, but GPU acceleration on Windows is NVIDIA-only; AMD and Intel run CPU-only there

The recent bug list is worth reading before a migration, because all of it is from the last three days and most of it is about the alternative engines rather than the default one: #2374 /ws/tts engine override spawns a new sidecar per request (VRAM leak, OOM); #2373 IndexTTS on ROCm spends ~18 s per line in MIOpen’s default find mode; #2372 IndexTTS 2.5 segfaults on RDNA2 because torch misreports bf16 support; #2371 IndexTTS 2.5 installs CUDA torch on ROCm hosts so it always runs on CPU; #2365 macOS Intel can’t set up its Python dependencies and #2360 (since fixed) shipped a broken Apple Silicon DMG in v0.5.6. The good news: 17 open issues against 40,000 stars is a healthy ratio, the CI/Security/Docker pipelines are green on main, and v0.5.6 landed five days ago.

What to do about it

If your use is personal, research, evaluation, or internal scratch tracks, use the default engine and stop worrying. Nothing in the stack restricts that, and it is the best-engineered path in the app.

If you earn money from the output — monetised YouTube, client dubbing, audiobook sales, anything an agency invoices — the default engine is not licensed for it, and “I didn’t know” has never once been a defence. Switch the engine:

# Settings → Model Catalogue, or:
OMNIVOICE_TTS_BACKEND=cosyvoice    # CosyVoice 3, Apache-2.0

Engines whose code and weights are permissive enough to start with: CosyVoice 3 (Apache-2.0), VoxCPM2 (Apache-2.0), MOSS-TTS-Nano (Apache-2.0), KittenTTS (MIT, English, CPU), MLX-Audio (Apple Silicon; Kokoro and others). Engines to avoid for paid work: OmniVoice (CC-BY-NC weights), audiocpp (Breeze-TTS-2, research/non-commercial, disclosed in-app), supertonic3 and pockettts (licence-gated).

Know what you are trading away, because it is not small: language coverage collapses from 646 languages to 30 (VoxCPM2), 9 languages + 18 dialects (CosyVoice 3), or one language (KittenTTS). Cross-lingual dubbing at scale is exactly the job the non-commercial default engine does best. That is the trap: the licence you cannot use is attached to the capability you wanted. Either narrow your language set to what a permissive engine covers, or treat a commercial licence as a line item — which is the conversation the project is inviting with that pro@voicestudio.sh address.

For anyone evaluating, spend five minutes on these three commands. They settle the entire licence question without trusting me or the README:

curl -s https://huggingface.co/k2-fsa/OmniVoice/raw/main/README.md | grep -A2 '^## License'
curl -s https://raw.githubusercontent.com/debpalash/VoiceStudio/main/LICENSE-NOTICE.md | sed -n '1,40p'
curl -s https://raw.githubusercontent.com/debpalash/VoiceStudio/main/docs/engine-acceptance.md | grep -i -A2 'commercial use'

If the first line ever changes to Apache-2.0 weights, this article’s central objection evaporates overnight and VoiceStudio becomes the easiest recommendation in local voice AI.

FAQ

Is VoiceStudio really open source? The application is, under AGPL-3.0-only, and you can read, modify and run it commercially. It is not fully OSI-open in the sense of shipping only permissive weights: its default engine’s weights are CC-BY-NC and its audio tokenizer is under a Llama-3-derived community licence.

Can I use it commercially at all? Yes — with permissive engines (CosyVoice 3, VoxCPM2, MOSS-TTS-Nano, KittenTTS, MLX-Audio). No — with the default OmniVoice engine, whose weights are non-commercial, and with audiocpp/supertonic3/pockettts.

Where does the CC-BY-NC come from? From upstream, not from VoiceStudio. OmniVoice’s authors state the restriction comes from the training data (Emilia) and apply it to the pretrained weights while keeping the code Apache-2.0. The app inherits it by bundling that engine as the default.

Does the “646 languages” claim hold up? Exactly, row for row: 646 rows, 646 unique ISO 639-3 codes, and durations summing to 581,489 hours. What it hides is the distribution — 265 of those languages have under ten hours of training data, 11 have under one hour, and the top ten languages are 84% of the corpus.

Is 40,145 stars a sign of quality? It is a sign of reach. The reconstructed curve shows ~26,800 stars in 27 days, driven by social video and press rather than developer scrutiny — the best Hacker News submission scored 6 points. The project’s own competitive analysis grades itself below voicebox on chunked long-form generation and preset library size.

Why is the paywall in the app but not on the website? Because it is deliberately gated. desktop-pro-page.md lists unset release flags (VOICESTUDIO_PRO_FEATURES_RELEASED, VOICESTUDIO_PRO_TERMS_APPROVED) and the site states plainly that checkout and commercial licences are not available yet. The plumbing for $99/year and $299/lifetime exists in Electron main; the promise has not been switched on.

The verdict

VoiceStudio is a real engineering achievement wearing a licensing problem it has not decided how to talk about. In 172 days it became the most complete local voice workstation in open source — 16 TTS engines, 11 ASR engines, watermarking, consent flags, MCP, an analytics design better than most SaaS — and it deserves better than a star count that arrived before anybody read the model card.

The single change that would make this article obsolete is the cheapest one available: print the CC-BY-NC warning next to the default engine in the picker where a user chooses it, the way audiocpp and the two gated engines already are. The project’s own rule already says weights must be commercially clean. Enforcing it visibly — for the engine that holds the top job in its own job map — is not a hard engineering task. It is the difference between a project that audits its licences and one that audits only the ones that help.