<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>Audio Generation on SoloSoft</title><link>https://www.solosoft.dev/tags/audio-generation/</link><description>Recent content in Audio Generation on SoloSoft</description><generator>Hugo</generator><language>en-us</language><lastBuildDate>Fri, 01 May 2026 00:00:00 +0000</lastBuildDate><atom:link href="https://www.solosoft.dev/tags/audio-generation/index.xml" rel="self" type="application/rss+xml"/><item><title>AudioCraft: Meta's Open-Source AI Audio Generation Toolkit</title><link>https://www.solosoft.dev/post/audiocraft-musicgen-2026/</link><pubDate>Fri, 01 May 2026 00:00:00 +0000</pubDate><guid>https://www.solosoft.dev/post/audiocraft-musicgen-2026/</guid><description>&lt;p&gt;The ability to generate high-quality audio from text descriptions has long been a holy grail of artificial intelligence. &lt;strong&gt;AudioCraft&lt;/strong&gt;, Meta&amp;rsquo;s open-source PyTorch library, brings this capability to the broader AI community with a comprehensive suite of audio generation models that cover music, sound effects, and neural audio compression.&lt;/p&gt;
&lt;p&gt;AudioCraft unifies three distinct audio generation capabilities under a single codebase: MusicGen for generating music from text prompts, AudioGen for creating sound effects and environmental audio, and EnCodec for neural audio compression. Each component is state-of-the-art in its domain, and together they form one of the most powerful open-source audio AI toolkits available.&lt;/p&gt;</description></item><item><title>Higgs Audio: Boson AI's Open-Source Text-Audio Foundation Model</title><link>https://www.solosoft.dev/post/higgs-audio-generation-2026/</link><pubDate>Fri, 01 May 2026 00:00:00 +0000</pubDate><guid>https://www.solosoft.dev/post/higgs-audio-generation-2026/</guid><description>&lt;p&gt;Text-to-speech technology has advanced dramatically in recent years, transitioning from robotic, monotone synthesis to remarkably natural voice generation. &lt;strong&gt;Higgs Audio&lt;/strong&gt; by Boson AI represents the state of the art in open-source audio generation, offering a text-to-audio foundation model that produces speech indistinguishable from human recordings across multiple voices, languages, and emotional registers.&lt;/p&gt;
&lt;p&gt;What distinguishes Higgs Audio from previous TTS systems is its scale and architecture. Pretrained on over 10 million hours of diverse audio data &amp;ndash; far more than any prior open-source TTS model &amp;ndash; Higgs Audio has learned the full richness and variety of human speech. It can generate expressive speech with appropriate emotion, emphasis, and pacing, clone a voice from just a few seconds of audio, produce multi-speaker dialogues with distinct voices, and even transfer speaking styles between voices.&lt;/p&gt;</description></item></channel></rss>