Speculative Decoding

  1. Auditing Strata: 125B on a 12 GB Card, 62 GB/s of RAM Bandwidth, and a 5× Contradiction in the Docs

    Strata claims a 125B model on a 12 GB graphics card. I measured the repository, the physics and the documentation, and the number holds — because the constraint is not …

    AI
  2. KTransformers: Flexible LLM Inference with Advanced Kernel Optimization

    The efficiency of LLM inference directly determines the cost, latency, and scalability of AI applications. KTransformers (kvcache-ai/ktransformers on...

    AI