<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>ModelCloud on SoloSoft</title><link>https://www.solosoft.dev/tags/modelcloud/</link><description>Recent content in ModelCloud on SoloSoft</description><generator>Hugo</generator><language>en-us</language><lastBuildDate>Fri, 01 May 2026 00:00:00 +0000</lastBuildDate><atom:link href="https://www.solosoft.dev/tags/modelcloud/index.xml" rel="self" type="application/rss+xml"/><item><title>GPTQModel: Production-Ready LLM Quantization Toolkit for GPU and CPU</title><link>https://www.solosoft.dev/post/gptqmodel-quantization-2026/</link><pubDate>Fri, 01 May 2026 00:00:00 +0000</pubDate><guid>https://www.solosoft.dev/post/gptqmodel-quantization-2026/</guid><description>&lt;p&gt;Large language models are powerful, but their size makes them expensive to deploy. A 70-billion-parameter model in 16-bit precision requires 140GB of GPU memory &amp;ndash; well beyond a single consumer GPU. Quantization is the primary solution: reducing numerical precision to shrink memory footprint and accelerate inference. &lt;strong&gt;GPTQModel&lt;/strong&gt;, developed by ModelCloud, is a production-ready quantization toolkit that makes this practical across a wide range of hardware.&lt;/p&gt;
&lt;p&gt;GPTQModel unifies multiple quantization methods &amp;ndash; GPTQ, AWQ, and GGUF &amp;ndash; under a single API, supporting over 30 model architectures on Nvidia, AMD, and Intel GPUs as well as CPU inference. The project at &lt;a href="https://github.com/ModelCloud/GPTQModel"&gt;github.com/ModelCloud/GPTQModel&lt;/a&gt; has rapidly become the go-to quantization library for teams that need to deploy LLMs in production without locking into a single quantization format.&lt;/p&gt;</description></item></channel></rss>