<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>PPO on SoloSoft</title><link>https://www.solosoft.dev/tags/ppo/</link><description>Recent content in PPO on SoloSoft</description><generator>Hugo</generator><language>en-us</language><lastBuildDate>Fri, 01 May 2026 00:00:00 +0000</lastBuildDate><atom:link href="https://www.solosoft.dev/tags/ppo/index.xml" rel="self" type="application/rss+xml"/><item><title>TinyZero: Reproducing DeepSeek R1-Zero's Reasoning with RL for Under $30</title><link>https://www.solosoft.dev/post/tinyzero-r1-reproduction-2026/</link><pubDate>Fri, 01 May 2026 00:00:00 +0000</pubDate><guid>https://www.solosoft.dev/post/tinyzero-r1-reproduction-2026/</guid><description>&lt;p&gt;DeepSeek R1-Zero was widely regarded as a breakthrough when it was released in January 2025. The model demonstrated that pure reinforcement learning — without any supervised fine-tuning on human reasoning examples — could produce advanced chain-of-thought reasoning, self-correction, and even surprising &amp;ldquo;aha moments&amp;rdquo; where the model independently discovered better reasoning strategies mid-conversation. The catch? The training infrastructure was assumed to require massive compute clusters and budgets in the tens of millions of dollars.&lt;/p&gt;
&lt;p&gt;Jiayi Pan&amp;rsquo;s TinyZero shatters that assumption entirely.&lt;/p&gt;
&lt;p&gt;TinyZero is an open-source, minimal reproduction of the DeepSeek R1-Zero methodology that runs on a single GPU for under $30 in cloud compute costs. Using the &lt;code&gt;veRL&lt;/code&gt; framework — a versatile reinforcement learning library for language models — TinyZero applies PPO (Proximal Policy Optimization) to small base models like Qwen-2.5-1.5B-Instruct and Qwen-2.5-7B. The training task is deceptively simple: given four numbers, the model must combine them using arithmetic operations (+, -, *, /) to reach a target value. Yet from this humble starting point, the same emergent reasoning behaviors that made DeepSeek R1-Zero famous begin to appear.&lt;/p&gt;</description></item><item><title>TRL: Hugging Face's Transformer Reinforcement Learning Library</title><link>https://www.solosoft.dev/post/trl-rlhf-2026/</link><pubDate>Fri, 01 May 2026 00:00:00 +0000</pubDate><guid>https://www.solosoft.dev/post/trl-rlhf-2026/</guid><description>&lt;p&gt;The alignment of large language models with human preferences is one of the most important challenges in AI development. &lt;strong&gt;TRL&lt;/strong&gt; (huggingface/trl on GitHub) &amp;ndash; Hugging Face&amp;rsquo;s Transformer Reinforcement Learning library &amp;ndash; provides a comprehensive toolkit for tackling this challenge, implementing the full spectrum of RLHF (Reinforcement Learning from Human Feedback) algorithms in a production-ready, well-documented package.&lt;/p&gt;
&lt;p&gt;Developed by Hugging Face&amp;rsquo;s research team, TRL has become the standard library for LLM alignment training, with over 10,000 GitHub stars and widespread adoption across both academia and industry. It supports PPO, DPO, KTO, and several other preference optimization algorithms, each offering different trade-offs between training complexity, computational cost, and alignment effectiveness.&lt;/p&gt;</description></item></channel></rss>