<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>Multilingual on SoloSoft</title><link>https://www.solosoft.dev/tags/multilingual/</link><description>Recent content in Multilingual on SoloSoft</description><generator>Hugo</generator><language>en-us</language><lastBuildDate>Fri, 01 May 2026 00:00:00 +0000</lastBuildDate><atom:link href="https://www.solosoft.dev/tags/multilingual/index.xml" rel="self" type="application/rss+xml"/><item><title>Surya: Open-Source Multilingual OCR and Document Understanding</title><link>https://www.solosoft.dev/post/surya-ocr-2026/</link><pubDate>Fri, 01 May 2026 00:00:00 +0000</pubDate><guid>https://www.solosoft.dev/post/surya-ocr-2026/</guid><description>&lt;p&gt;Optical Character Recognition is one of the oldest applications of computer vision, but traditional OCR engines have struggled to keep pace with modern demands. Documents today are more diverse in layout, multilingual in content, and variable in quality than ever before. &lt;strong&gt;Surya&lt;/strong&gt; represents a modern approach to OCR, built on deep learning architectures that handle the complexity of real-world documents with accuracy that traditional engines cannot match.&lt;/p&gt;
&lt;p&gt;Developed by the datalab-to team (the same group behind Marker), Surya is designed as both a standalone OCR system and a component for larger document processing pipelines. It provides three core capabilities: text detection (finding where text is on a page), text recognition (reading what it says), and layout analysis (understanding the document structure). The unified architecture means that a single model handles text across dozens of scripts and languages.&lt;/p&gt;</description></item><item><title>VoxCPM2: OpenBMB's Tokenizer-Free TTS for Multilingual Speech Generation</title><link>https://www.solosoft.dev/post/voxcpm-tts-2026/</link><pubDate>Fri, 01 May 2026 00:00:00 +0000</pubDate><guid>https://www.solosoft.dev/post/voxcpm-tts-2026/</guid><description>&lt;p&gt;VoxCPM2 is a tokenizer-free text-to-speech (TTS) model developed by &lt;a href="https://www.openbmb.cn/"&gt;OpenBMB&lt;/a&gt;, an open-source AI research community affiliated with Tsinghua University and the Beijing Academy of Artificial Intelligence (BAAI). With 2 billion parameters, VoxCPM2 represents a paradigm shift in speech synthesis by operating directly on continuous speech representations, eliminating the need for discrete audio tokenizers that typically degrade voice quality.&lt;/p&gt;
&lt;p&gt;The model supports over 30 languages with capabilities spanning zero-shot voice cloning, voice design (creating entirely new voices from text descriptions), and real-time streaming inference. VoxCPM2 has quickly become one of the most talked-about open-source TTS models of 2026, competing directly with commercial offerings like ElevenLabs and OpenAI&amp;rsquo;s TTS while remaining freely available under the Apache 2.0 license.&lt;/p&gt;</description></item></channel></rss>