Quantization
Page 1 of 2
- AI
Auditing Strata: 125B on a 12 GB Card, 62 GB/s of RAM Bandwidth, and a 5× Contradiction in the Docs
Strata claims a 125B model on a 12 GB graphics card. I measured the repository, the physics and the documentation, and the number holds — because the constraint is not …
- AI
TensorRT-LLM: NVIDIA's Open-Source Library for Optimized LLM Inference
Deploying large language models in production requires more than just loading weights onto a GPU. To achieve acceptable throughput and latency, you...
- AI
MLX LM: LLM Inference and Fine-Tuning on Apple Silicon
The promise of running LLMs locally on a MacBook has been seductive but incomplete. Ollama and llama.cpp made it possible, but performance left room...
- AI
llama.cpp: High-Performance LLM Inference on CPU and GPU
The dream of running powerful language models entirely on your own hardware, without sending data to cloud APIs, was once considered impractical for...
- AI
ik_llama.cpp: Fork of llama.cpp with IQ4_NL and Advanced Quantization
The ecosystem around llama.cpp has produced numerous forks, each exploring different optimization strategies for running LLMs efficiently on consumer...
- AI
GPTQModel: Production-Ready LLM Quantization Toolkit for GPU and CPU
Large language models are powerful, but their size makes them expensive to deploy. A 70-billion-parameter model in 16-bit precision requires 140GB of...