LLM Inference
Page 2 of 2
- AI
llama.cpp: High-Performance LLM Inference on CPU and GPU
The dream of running powerful language models entirely on your own hardware, without sending data to cloud APIs, was once considered impractical for...
- AI
KTransformers: Flexible LLM Inference with Advanced Kernel Optimization
The efficiency of LLM inference directly determines the cost, latency, and scalability of AI applications. KTransformers (kvcache-ai/ktransformers on...
- Open Source
ExLlamaV3: High-Performance LLM Inference Engine
Running large language models on consumer hardware requires efficient inference engines that squeeze every drop of performance from available GPU...