lmcache python library

lmcache

Subscription: Developer
No updates yet
Active
Subscription: Developer

A serving engine extension for large language models that reduces time to first token and increases throughput by reusing KV caches across requests, particularly effective for long-context inference scenarios and multi-turn conversations.