<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>MoE on SoloSoft</title><link>https://www.solosoft.dev/tags/moe/</link><description>Recent content in MoE on SoloSoft</description><generator>Hugo</generator><language>en-us</language><atom:link href="https://www.solosoft.dev/tags/moe/index.xml" rel="self" type="application/rss+xml"/><item><title>Seed1.5-VL: ByteDance's Vision-Language Foundation Model Achieving 38 SOTA Benchmarks</title><link>https://www.solosoft.dev/post/seed15-vl-vision-language-2026/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://www.solosoft.dev/post/seed15-vl-vision-language-2026/</guid><description>&lt;p&gt;In the rapidly advancing field of vision-language models, a new heavyweight has emerged from an unexpected corner. &lt;strong&gt;Seed1.5-VL&lt;/strong&gt;, developed by ByteDance&amp;rsquo;s Seed team, has achieved state-of-the-art results on an astonishing 38 out of 60 public benchmarks, spanning image understanding, video comprehension, document parsing, and multi-image reasoning.&lt;/p&gt;
&lt;p&gt;Built on a 20-billion parameter Mixture-of-Experts (MoE) architecture with approximately 2 billion activated parameters per token, Seed1.5-VL represents a careful balancing act between raw capability and computational efficiency. It outperforms models with far larger parameter counts while maintaining inference speeds suitable for real-world applications.&lt;/p&gt;
&lt;p&gt;The model&amp;rsquo;s benchmark sweep is remarkable not just for the number of wins, but for the breadth of categories it dominates. From OCR and chart understanding to multi-image reasoning and video comprehension, Seed1.5-VL demonstrates that ByteDance&amp;rsquo;s research team has achieved something genuinely comprehensive in the multimodal space.&lt;/p&gt;</description></item></channel></rss>