vllm python library

vllm

Subscription: Developer
Last updated a day ago
Active
Subscription: Developer

A high-throughput and memory-efficient inference and serving engine for large language models. It utilizes PagedAttention to efficiently manage memory while providing state-of-the-art throughput and continuous batching.

pip install vllm==0.31.0
  • 0.31.0
    Linux
    Last updated a day ago
  • 0.30.0
    Linux
    Last updated a day ago
  • 0.29.0
    Linux
    Last updated a day ago
  • 0.22.0
    Linux
    Last updated a day ago