Nitsum: Serving Tiered LLM Requests with Adaptive Tensor Parallelism
Nitsum is a new serving system designed to optimize the handling of tiered LLM requests using adaptive tensor parallelism. It allows for dynamic reconfiguration of GPU resources to meet varying latency requirements for different workloads. This approach significantly enhances goodput, achieving up to 5.3 times improvement over existing systems.
- ▪Nitsum treats tensor parallelism as a runtime control surface rather than a fixed deployment choice.
- ▪The system improves service-level objective compliance by dynamically adjusting GPU configurations based on workload changes.
- ▪Nitsum can serve a mix of latency-critical and relaxed background jobs under a fixed GPU budget.
Hacker News (AI / LLM) files mainly under ai. We currently carry 3,270 of its stories.
Story provenance
Source · retrieval · rights · ranking — open for full record
inspect →
Story provenance
Attribution is not the same as permission. This drawer separates discovery metadata, excerpts, WeSearch-generated summaries, reuse status, and whether the publisher receives the visit. Nothing here claims a legal grant the publisher has not made.
Record
| Original publisher | MLSys @ WukLab - Nitsum: Serving Tiered LLM Requests with Adaptive Tensor Parallelism |
| Canonical URL | https://mlsys.wuklab.io/posts/nitsum/ |
| Publication time | Tue, 19 May 2026 01:24:57 +0000 |
| Retrieval time | 2026-05-19T01:29:57.129Z |
| Last seen | 2026-05-19T01:29:57.129Z |
| Headline source | Publisher (no WeSearch rewrite) |
| Excerpt source | publisher body |
| Excerpt method | First ~120 words (~800 chars) of extracted publisher body, fair-use limited. |
| Summary | WeSearch · cerebras-chat (WeSearch summarizer) |
| Summary source text | contentText |
| Citation coverage | Summary is a WeSearch-generated derivative; primary citation is the original publisher URL. |
| Cluster | LbDnrIVhtua1 |
| Cluster logic | Grouped by semantic title/content similarity across sources within a rolling window. Same-publisher template collisions are excluded from coverage comparison. |
| Ranking reason | Story pages are not engagement-ranked. Hub feeds use recency, with optional source-diversified chronological ordering (cap consecutive stories per source). No personalized ranking. |
| Publisher visit | Yes — open original |
| Substitutes article? | No — link-out required for full text |
Rights status (four layers)
WeSearch handling by dimension
| Indexing | May the item be indexed (stored, ranked, made findable)? | Allowed |
| Snippet | May a short excerpt of the publisher's text be shown? | Allowed |
| AI summary | May WeSearch generate its own short summary of the article? | Limited |
| Retrieval / RAG | May the content be exposed for third-party retrieval-augmented generation? | Not asserted |
| Model training | May the content be used to train AI models? | Not asserted |
| Commercial reuse | May the content be reused commercially? | Not permitted |
Basis: Derived from the published RSS/Atom feed. Contact: [email protected]. Reviewed: 2026-07-24.
Opening excerpt (first ~120 words) tap to expand
Nitsum: Serving Tiered LLM Requests with Adaptive Tensor ParallelismMay 16, 2026 - 12 mins readLLMServingTensor ParallelismAuthor: Vikranth Srivatsa, Zijian He, Pu Guo, Dongming Li, and Yiying ZhangTLDR: A single LLM deployment now serves everything from latency-critical chat to relaxed background jobs under a fixed GPU budget, creating a tiered-SLO serving problem. We designed Nitsum [arXiv ‘26], the first serving system that treats tensor parallelism (TP) as a runtime control surface instead of a fixed deployment choice.
…
Excerpt limited to ~120 words for fair-use compliance. The full article is at MLSys @ WukLab - Nitsum: Serving Tiered LLM Requests with Adaptive Tensor Parallelism.