LLM Inference

Page 1 of 2

  1. vLLM: High-Throughput LLM Inference with PagedAttention

    Serving LLMs in production is fundamentally a memory management problem. The KV cache — the set of attention key-value pairs stored during generation...

    AI
  2. TensorRT-LLM: NVIDIA's Open-Source Library for Optimized LLM Inference

    Deploying large language models in production requires more than just loading weights onto a GPU. To achieve acceptable throughput and latency, you...

    AI
  3. SGLang: Efficient LLM Inference with Structured Generation

    The open-source LLM ecosystem has solved many problems — model quality, fine-tuning, deployment — but one challenge persists: getting models to...

    AI
  4. PowerInfer: High-Speed LLM Inference on Consumer GPUs via CPU-GPU Hybrid Design

    Running large language models locally has always been constrained by a hard wall: GPU memory. A 175-billion parameter model in FP16 requires...

    AI
  5. MLX LM: LLM Inference and Fine-Tuning on Apple Silicon

    The promise of running LLMs locally on a MacBook has been seductive but incomplete. Ollama and llama.cpp made it possible, but performance left room...

    AI
  6. LocalAI: Self-Hosted OpenAI API-Compatible Inference Server

    Running AI models locally offers undeniable advantages: complete data privacy, no API costs, offline operation, and full control over model choice...

    AI