Quantization

Page 2 of 2

  1. ExLlamaV3: High-Performance LLM Inference Engine

    Running large language models on consumer hardware requires efficient inference engines that squeeze every drop of performance from available GPU...

    Open Source
  2. bitsandbytes: Essential k-bit Quantization Library for LLM Training and Inference

    Large language models have grown far beyond the memory capacity of consumer hardware. A 70-billion-parameter model requires 140 gigabytes of GPU...

    AI