Quantization

Page 1 of 2

  1. Auditing Strata: 125B on a 12 GB Card, 62 GB/s of RAM Bandwidth, and a 5× Contradiction in the Docs

    Strata claims a 125B model on a 12 GB graphics card. I measured the repository, the physics and the documentation, and the number holds — because the constraint is not …

    AI
  2. TensorRT-LLM: NVIDIA's Open-Source Library for Optimized LLM Inference

    Deploying large language models in production requires more than just loading weights onto a GPU. To achieve acceptable throughput and latency, you...

    AI
  3. MLX LM: LLM Inference and Fine-Tuning on Apple Silicon

    The promise of running LLMs locally on a MacBook has been seductive but incomplete. Ollama and llama.cpp made it possible, but performance left room...

    AI
  4. llama.cpp: High-Performance LLM Inference on CPU and GPU

    The dream of running powerful language models entirely on your own hardware, without sending data to cloud APIs, was once considered impractical for...

    AI
  5. ik_llama.cpp: Fork of llama.cpp with IQ4_NL and Advanced Quantization

    The ecosystem around llama.cpp has produced numerous forks, each exploring different optimization strategies for running LLMs efficiently on consumer...

    AI
  6. GPTQModel: Production-Ready LLM Quantization Toolkit for GPU and CPU

    Large language models are powerful, but their size makes them expensive to deploy. A 70-billion-parameter model in 16-bit precision requires 140GB of...

    AI