LLM Inference

  1. vLLM: Inferencia de LLMs de Alto Rendimiento con PagedAttention

    Serving LLMs in production is fundamentally a memory management problem. The KV cache — the set of attention key-value pairs stored during generation...

    IA
  2. TensorRT-LLM: La Biblioteca de Codigo Abierto de NVIDIA para Inferencia de LLM Optimizada

    Implementar modelos de lenguaje grandes en produccion requiere mas que solo cargar pesos en una GPU. Para lograr rendimiento y latencia aceptables...

    IA
  3. SGLang: Inferencia Eficiente de LLMs con Generacion Estructurada

    The open-source LLM ecosystem has solved many problems — model quality, fine-tuning, deployment — but one challenge persists: getting models to...

    IA
  4. MLX LM: Inferencia y Ajuste Fino de LLMs en Apple Silicon

    The promise of running LLMs locally on a MacBook has been seductive but incomplete. Ollama and llama.cpp made it possible, but performance left room...

    IA
  5. KTransformers: Flexible LLM Inference with Advanced Kernel Optimization

    The efficiency of LLM inference directly determines the cost, latency, and scalability of AI applications. KTransformers (kvcache-ai/ktransformers on...

    IA
  6. ExLlamaV3: Motor de Inferencia de LLM de Alto Rendimiento

    Ejecutar modelos de lenguaje grandes en hardware de consumo requiere motores de inferencia eficientes que expriman cada gota de rendimiento de la...

    Código Abierto