Speculative Decoding
- AI
Auditing Strata: 125B on a 12 GB Card, 62 GB/s of RAM Bandwidth, and a 5× Contradiction in the Docs
Strata claims a 125B model on a 12 GB graphics card. I measured the repository, the physics and the documentation, and the number holds — because the constraint is not …
- AI
KTransformers: Flexible LLM Inference with Advanced Kernel Optimization
The efficiency of LLM inference directly determines the cost, latency, and scalability of AI applications. KTransformers (kvcache-ai/ktransformers on...