Consumer GPU
- AI
Auditing Strata: 125B on a 12 GB Card, 62 GB/s of RAM Bandwidth, and a 5× Contradiction in the Docs
Strata claims a 125B model on a 12 GB graphics card. I measured the repository, the physics and the documentation, and the number holds — because the constraint is not …
- AI
PowerInfer: High-Speed LLM Inference on Consumer GPUs via CPU-GPU Hybrid Design
Running large language models locally has always been constrained by a hard wall: GPU memory. A 175-billion parameter model in FP16 requires...