<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>TinyZero on SoloSoft</title><link>https://www.solosoft.dev/tags/tinyzero/</link><description>Recent content in TinyZero on SoloSoft</description><generator>Hugo</generator><language>en-us</language><lastBuildDate>Fri, 01 May 2026 00:00:00 +0000</lastBuildDate><atom:link href="https://www.solosoft.dev/tags/tinyzero/index.xml" rel="self" type="application/rss+xml"/><item><title>TinyZero: Reproducing DeepSeek R1-Zero's Reasoning with RL for Under $30</title><link>https://www.solosoft.dev/post/tinyzero-r1-reproduction-2026/</link><pubDate>Fri, 01 May 2026 00:00:00 +0000</pubDate><guid>https://www.solosoft.dev/post/tinyzero-r1-reproduction-2026/</guid><description>&lt;p&gt;DeepSeek R1-Zero was widely regarded as a breakthrough when it was released in January 2025. The model demonstrated that pure reinforcement learning — without any supervised fine-tuning on human reasoning examples — could produce advanced chain-of-thought reasoning, self-correction, and even surprising &amp;ldquo;aha moments&amp;rdquo; where the model independently discovered better reasoning strategies mid-conversation. The catch? The training infrastructure was assumed to require massive compute clusters and budgets in the tens of millions of dollars.&lt;/p&gt;
&lt;p&gt;Jiayi Pan&amp;rsquo;s TinyZero shatters that assumption entirely.&lt;/p&gt;
&lt;p&gt;TinyZero is an open-source, minimal reproduction of the DeepSeek R1-Zero methodology that runs on a single GPU for under $30 in cloud compute costs. Using the &lt;code&gt;veRL&lt;/code&gt; framework — a versatile reinforcement learning library for language models — TinyZero applies PPO (Proximal Policy Optimization) to small base models like Qwen-2.5-1.5B-Instruct and Qwen-2.5-7B. The training task is deceptively simple: given four numbers, the model must combine them using arithmetic operations (+, -, *, /) to reach a target value. Yet from this humble starting point, the same emergent reasoning behaviors that made DeepSeek R1-Zero famous begin to appear.&lt;/p&gt;</description></item></channel></rss>