LLM Inference
- IA
vLLM: Inferencia de LLMs de Alto Rendimiento con PagedAttention
Serving LLMs in production is fundamentally a memory management problem. The KV cache — the set of attention key-value pairs stored during generation...
- IA
TensorRT-LLM: La Biblioteca de Codigo Abierto de NVIDIA para Inferencia de LLM Optimizada
Implementar modelos de lenguaje grandes en produccion requiere mas que solo cargar pesos en una GPU. Para lograr rendimiento y latencia aceptables...
- IA
SGLang: Inferencia Eficiente de LLMs con Generacion Estructurada
The open-source LLM ecosystem has solved many problems — model quality, fine-tuning, deployment — but one challenge persists: getting models to...
- IA
MLX LM: Inferencia y Ajuste Fino de LLMs en Apple Silicon
The promise of running LLMs locally on a MacBook has been seductive but incomplete. Ollama and llama.cpp made it possible, but performance left room...
- IA
KTransformers: Flexible LLM Inference with Advanced Kernel Optimization
The efficiency of LLM inference directly determines the cost, latency, and scalability of AI applications. KTransformers (kvcache-ai/ktransformers on...
- Código Abierto
ExLlamaV3: Motor de Inferencia de LLM de Alto Rendimiento
Ejecutar modelos de lenguaje grandes en hardware de consumo requiere motores de inferencia eficientes que expriman cada gota de rendimiento de la...