VLLM
- IA
vLLM: Inferencia de LLMs de Alto Rendimiento con PagedAttention
Serving LLMs in production is fundamentally a memory management problem. The KV cache — the set of attention key-value pairs stored during generation...
- Código Abierto
IndexTTS-vLLM: Texto a Voz Acelerado de Código Abierto con Inferencia vLLM
IndexTTS-vLLM es una versión acelerada del sistema de texto a voz IndexTTS que porta el pipeline de inferencia del modelo a vLLM. El resultado es una...