vllm
Subscription: Developer
Active
Subscription: Developer
A high-throughput and memory-efficient inference and serving engine for large language models. It utilizes PagedAttention to efficiently manage memory while providing state-of-the-art throughput and continuous batching.
pip install vllm==0.31.0
0.31.0
Linux
0.30.0
Linux
0.29.0
Linux
0.22.0
Linux