vllm python library

vllm

Subscription: Developer
No updates yet
Active
Subscription: Developer

A high-throughput and memory-efficient inference and serving engine for large language models. It utilizes PagedAttention to efficiently manage memory while providing state-of-the-art throughput and continuous batching.