Local LLM

  1. Auditing Strata: 125B on a 12 GB Card, 62 GB/s of RAM Bandwidth, and a 5× Contradiction in the Docs

    Strata claims a 125B model on a 12 GB graphics card. I measured the repository, the physics and the documentation, and the number holds — because the constraint is not …

    AI
  2. Twinny: Local LLM Inference for VS Code

    The tension between cloud-dependent AI tools and developer privacy has become one of the defining debates in AI-assisted software development...

    AI
  3. Ollama: Run Open-Source LLMs Locally with Docker-Like Simplicity

    The world of large language models has evolved at breathtaking speed, but for most users, interacting with these powerful tools still involves...

    Open Source