vLLM
github.com
High-throughput LLM serving
An open-source inference engine for serving large language models at high throughput, built around PagedAttention and continuous batching.
Record
- Category
- AI tools
- Type
- Tool
- Pricing
- Free
- Licence
- Apache-2.0
- Platforms
- web
- Popularity
- ★ 89k
- Link
- ✓ Reached 3 Oct 2026
- Last commit
- 3 Oct 2026
- Found via
- GitHub