vLLM
A high-throughput and memory-efficient inference and serving engine for LLMs
The Linkredibles take
Why it's worth knowing.
A high-throughput and memory-efficient inference and serving engine for LLMs
Topics
#llm#inference#gpu
Open source repository
Explore the source behind vLLM.
Read the source, explore the project, and see how it is being developed on GitHub.
Keep exploring
You might also like.
productivity→
Actual Budget
A local-first personal finance app focused on giving you control over your financial data.
★ 29,139TypeScript
productivity→
AppFlowy
An open-source workspace for notes, tasks, databases, and knowledge management.
★ 76,920Dart
ai→
Atlas
A self-hosted LLM inference platform exposing Anthropic- and OpenAI-compatible APIs over multiple inference engines.
★ 0Go
