README: vLLM is a fast and easy-to-use library for LLM inference and serving.
GitHub 项目简介: A high-throughput and memory-efficient inference and serving engine for LLMs
README: Easy, fast, and cheap LLM serving for everyone
README: vLLM has grown into one of the most active open-source AI projects built and maintained by a diverse community of many dozens of academic institutions and companies from over 2000…
README: Efficient management of attention key and value memory with [**PagedAttention**](https://blog.vllm.ai/2023/06/20/vllm.html)