ObjectCache: Layerwise Object-Storage Retrieval for KV Cache Reuse
The paper presents ObjectCache, a new approach for managing KV caching in large language model serving. By utilizing S3-compatible object storage, it aims to alleviate the constraints of GPU memory and local DRAM while minimizing latency. The proposed system demonstrates improved efficiency in data transfer and computation overlap, leading to reduced time to first token in various contexts.
- ▪ObjectCache co-designs the storage protocol and transfer schedule to optimize KV cache retrieval.
- ▪The system adds only 5.6% latency for 64K contexts compared to local DRAM.
- ▪Under shared bandwidth caps, ObjectCache reduces added time to first token by 1.2-1.8x.
arXiv cs.AI files mainly under ai research. We currently carry 1,128 of its stories.
Story provenance
Source · retrieval · rights · ranking — open for full record
inspect →
Story provenance
Attribution is not the same as permission. This drawer separates discovery metadata, excerpts, WeSearch-generated summaries, reuse status, and whether the publisher receives the visit. Nothing here claims a legal grant the publisher has not made.
Record
| Original publisher | arXiv cs.AI |
| Canonical URL | https://arxiv.org/abs/2605.22850 |
| Publication time | Mon, 25 May 2026 00:00:00 -0400 |
| Retrieval time | 2026-05-25T04:07:35.648Z |
| Last seen | 2026-05-25T04:07:35.648Z |
| Headline source | Publisher (no WeSearch rewrite) |
| Excerpt source | publisher body |
| Excerpt method | First ~120 words (~800 chars) of extracted publisher body, fair-use limited. |
| Summary | WeSearch · cerebras-chat (WeSearch summarizer) |
| Summary source text | contentText |
| Citation coverage | Summary is a WeSearch-generated derivative; primary citation is the original publisher URL. |
| Cluster | 46N-Y6Utrhh6 |
| Cluster logic | Grouped by semantic title/content similarity across sources within a rolling window. Same-publisher template collisions are excluded from coverage comparison. |
| Ranking reason | Story pages are not engagement-ranked. Hub feeds use recency, with optional source-diversified chronological ordering (cap consecutive stories per source). No personalized ranking. |
| Publisher visit | Yes — open original |
| Substitutes article? | No — link-out required for full text |
Rights status (four layers)
WeSearch handling by dimension
| Indexing | May the item be indexed (stored, ranked, made findable)? | Allowed |
| Snippet | May a short excerpt of the publisher's text be shown? | Allowed |
| AI summary | May WeSearch generate its own short summary of the article? | Limited |
| Retrieval / RAG | May the content be exposed for third-party retrieval-augmented generation? | Not asserted |
| Model training | May the content be used to train AI models? | Not asserted |
| Commercial reuse | May the content be reused commercially? | Not permitted |
Basis: Derived from the published RSS/Atom feed. Contact: [email protected]. Reviewed: 2026-07-24.
Opening excerpt (first ~120 words) tap to expand
Computer Science > Distributed, Parallel, and Cluster Computing arXiv:2605.22850 (cs) [Submitted on 16 May 2026] Title:ObjectCache: Layerwise Object-Storage Retrieval for KV Cache Reuse Authors:Yu Zhu, Aditya Dhakal, Yunming Xiao, Dejan Milojicic, Gustavo Alonso View a PDF of the paper titled ObjectCache: Layerwise Object-Storage Retrieval for KV Cache Reuse, by Yu Zhu and 4 other authors View PDF HTML (experimental) Abstract:Prefix KV caching has become a key mechanism in LLM serving: it reduces time to first token (TTFT) by avoiding redundant computation across requests that share a prefix (i.e., the system prompt). However, the accumulated KV cache is often larger than what GPU memory and local DRAM can hold.
…
Excerpt limited to ~120 words for fair-use compliance. The full article is at arXiv cs.AI.