7 stories tagged with #llm-inference, in publish-time order across the WeSearch catalog. Tag pages update as new stories ingest.
⌘ RSS feed for this tag → or search "Llm Inference"
Predictive Speculative KV Replication for Bursty LLM Inference
JW Labs research post.…
Bursty arrivals speed up LLM inference
Bursty arrivals usually cause a drop in performance and are seen as headaches in production systems. In this blog, we investigate a phenomenon where burstiness actually improves pe…
What LLM Inference Costs
blog of what LLM inference actually costs $0.09 to $290.12 per 1M output tokens. almost none of it is the model https://t.co/wPOfOdyKYE…
Ask HN: What are you using for LLM inference in production?
Engy – Verified LLM Inference
OpenAI-compatible inference for frontier open-source LLMs, with a cryptographic proof of correct inference on every response.…
Show HN: NightRun, bare metal LLM inference, no OS, boots from USB
Boot your PC straight into an LLM. Rust, UEFI-resident, no operating system underneath. - hardrave/NIGHTRUN…
How Profitable Is LLM Inference? Doing the Math on Kimi K3
A look at LLM inference economics (batch size, GPU count, and the Pareto frontier that sets token prices) applied to Kimi K3 with back-of-the-envelope math.…