LLM Inference

Page 2 of 2

  1. llama.cpp: High-Performance LLM Inference on CPU and GPU

    The dream of running powerful language models entirely on your own hardware, without sending data to cloud APIs, was once considered impractical for...

    AI
  2. KTransformers: Flexible LLM Inference with Advanced Kernel Optimization

    The efficiency of LLM inference directly determines the cost, latency, and scalability of AI applications. KTransformers (kvcache-ai/ktransformers on...

    AI
  3. ExLlamaV3: High-Performance LLM Inference Engine

    Running large language models on consumer hardware requires efficient inference engines that squeeze every drop of performance from available GPU...

    Open Source