vllm
Subscription: Developer
Active
Subscription: Developer
A high-throughput and memory-efficient inference and serving engine for large language models. It utilizes PagedAttention to efficiently manage memory while providing state-of-the-art throughput and continuous batching.