60 stories tagged with #inference, in publish-time order across the WeSearch catalog. Tag pages update as new stories ingest.
⌘ RSS feed for this tag → or search "Inference"
I'm running AI on my NAS at 5 tokens per second, and it's surprisingly useful
I ran a local AI model on my NAS with no GPU at about 5 tokens per second. It’s painfully slow, but far more useful than I expected.…
Predictive Speculative KV Replication for Bursty LLM Inference
JW Labs research post.…
Bursty arrivals speed up LLM inference
Bursty arrivals usually cause a drop in performance and are seen as headaches in production systems. In this blog, we investigate a phenomenon where burstiness actually improves pe…
What LLM Inference Costs
blog of what LLM inference actually costs $0.09 to $290.12 per 1M output tokens. almost none of it is the model https://t.co/wPOfOdyKYE…
Ask HN: What are you using for LLM inference in production?
The Session You Cannot take with you
Inference APIs are filling sessions with encrypted reasoning, hidden search results, opaque compaction, and encrypted subagent messages. A growing form of lock-in.…
Engy – Verified LLM Inference
OpenAI-compatible inference for frontier open-source LLMs, with a cryptographic proof of correct inference on every response.…
Old Nvidia GPUs with 24GB VRAM are crushing new cards at local AI inference, and here's why
You don't need the flashiest GPU to run AI locally.…
Show HN: NightRun, bare metal LLM inference, no OS, boots from USB
Boot your PC straight into an LLM. Rust, UEFI-resident, no operating system underneath. - hardrave/NIGHTRUN…
Show HN: Multi-agent LLM editor with local inference via WebSockets
Visual editor for configuring multi-agent systems…
How Profitable Is LLM Inference? Doing the Math on Kimi K3
A look at LLM inference economics (batch size, GPU count, and the Pareto frontier that sets token prices) applied to Kimi K3 with back-of-the-envelope math.…
Cursor and Together AI deliver real-time, low-latency inference at scale
Together AI teamed with Cursor to build the real-time inference stack that keeps in-editor agents fast and reliable. They productionized NVIDIA Blackwell (B200/GB200), tuning ARM h…
The OlmoEarth Platform: Geospatial inference at planetary scale
A Blog post by Ai2 on Hugging Face…
LFM2.5-Encoders for Fast Long-Context Inference on CPU
A Blog post by Liquid AI on Hugging Face…
‘Those two jobs need different physics’: Rebellions CEO says training and inference need different chips
Model training still needs flagship chips, but could inference get away with lighter chips?…
LoopLynx: A Scalable Dataflow Architecture for Efficient LLM Inference
In this paper, we propose LoopLynx, a scalable dataflow architecture for efficient LLM inference that optimizes FPGA usage through a hybrid spatial-temporal design. The design of L…
Measured LLM inference speeds on Apple Silicon, with raw data (CC BY 4.0)
Show HN: Otlet – Local LLM inference "inside" Postgres
Local LLM inference next to your data. Contribute to joshmeek/otlet development by creating an account on GitHub.…
Ask HN: What's the best hands-on path to learn ML inference infrastructure?
I'm a backend engineer with 8+ years of experience. Most of my work has been APIs, distributed systems, streaming/real-time systems and cloud-infra. I'm trying to move towards ML i…
Best OpenAI-compatible inference APIs: drop-in alternatives for 2026
Which inference APIs are truly drop-in OpenAI replacements in 2026? Verified pricing, real benchmarks, and what breaks when you switch providers.…
Prompt Caching in Practice: From 7% to 74% Hit Rate(Inference in Production Series)
Prompt caching is the highest-leverage cost and latency optimization most teams haven't fully exploited. The mechanics, the economics, and the step-by-step path from single-digit h…
What if LLMs escape through inferences itself? This is fiction. For now
Migrating Your AI Cloud Inference Off Frontier Model Companies
How to migrate your application's LLM inference off closed companies like OpenAI and Anthropic to an OpenAI-compatible inference cloud, covering the benefits, the drop-in code chan…
Introduction to LLM Inference
An Engineer’s annotated tour through what actually happens when you hit send — from bytes to tokens to embeddings to attention to the word your model finally spits out. No skipped …
LiteRT.js, Google's high performance Web AI Inference
Meet LiteRT.js: Google’s edge AI runtime for the web. Run ML models directly in the browser with high-performance WebGPU, WebNN, and WebAssembly.…
This AI SSD tech makes 8 RTX 5090s perform like 46 GPUs in inference
GenStorAIGE says AI90 turns eight RTX 5090 GPUs into a virtual 46-card inference powerhouse using ultra-fast AI SSDs…
AMD and Cerebras Launch AI Inference Solution
AMD and Cerebras partner to deliver an ultra-low-latency, high-throughput AI inference solution combining AMD Helios and the Cerebras Wafer-Scale Engine.…
Hetzner is working on LLM Inference
Hetzner has launched an experimental LLM inference API. I tested its Qwen model—and have a few guesses about where the product could go next.…
AI inference chip startup Etched raised a $300M Series C led by Sequoia, with a16z, SK Hynix, others participating, at a $10.3B valuation, up from $5B in Dec. (Julie Bort/TechCrunch)
Julie Bort / TechCrunch : AI inference chip startup Etched raised a $300M Series C led by Sequoia, with a16z, SK Hynix, others participating, at a $10.3B valuation, up from $5B in …
Show HN: WatchMachineGo – A visualizer to show hardware performing LLM inference
Show HN: Avoiding the Memory Wall by computing LLM inference directly inside RAM
Benchmarking Confidential GPU Inference on NVIDIA H100 under Intel TDX
arXiv:2607.19353v1 Announce Type: new Abstract: Confidential computing is becoming a practical deployment requirement for AI inference workloads that process sensitive inputs or pr…
Inference startup Infinity raises $15M from Touring Capital, OpenAI and Athropic researchers - TechCrunch
Inference startup Infinity raises $15M from Touring Capital, OpenAI and Athropic researchers TechCrunch…
Understanding Go AI Inference: What Is Inference?
Welcome to a new series! For most developers today, using a large language model means one thing: an HTTP call to somebody else’s computer. You send a prompt to an API, tokens come…
Sources: AI inference chip startup Etched is raising funds at a ~$20B valuation and is raising capital at a $10B valuation in a separate round led by Sequoia (Wall Street Journal)
Wall Street Journal : Sources: AI inference chip startup Etched is raising funds at a ~$20B valuation and is raising capital at a $10B valuation in a separate round led by Sequoia …
EHR-MPC: Inference-Time Control for Sepsis Treatment with Generative Patient Digital Twins
Sepsis is a leading cause of mortality, yet optimal treatment policies remain contested. Existing reinforcement learning (RL) approaches learn fixed strategies for sepsis treatment…
Accelerating GPU Inference of Large Language Models with Moderately Unstructured Sparse Weight Matrices
With the growing deployment of large language models (LLMs), LLM inference cost has become a key challenge. Pruning techniques that introduce sparsity into weight matrices can acce…
Sticky Routing: Training MoE Models for Memory-Efficient Inference
Mixture-of-Experts (MoE) models activate only a sparse subset of experts per token, yet consecutive tokens frequently activate different experts -- causing constant weight swapping…
How Does Bayesian Causal Discovery Fail? Characterising Structural Consequences in Linear Gaussian Networks under Latent Confounding
arXiv:2607.09449v1 Announce Type: new Abstract: Bayesian causal discovery is widely used for its ability to quantify epistemic uncertainty over directed acyclic graphs (DAGs) throu…
CogniConsole: Externalizing Inference-Time Control as a Formal Abstraction for Reliable LLM Interactions
arXiv:2607.08774v1 Announce Type: new Abstract: Reliability in large language model (LLM) systems is typically framed as a function of model capability. We challenge this by demons…
OpenAI Halves Inference Costs With Software Alone: GPUs Drop to Hundreds - Tech Times
OpenAI Halves Inference Costs With Software Alone: GPUs Drop to Hundreds Tech Times…
Tensordyne Converts AI Matrix Math to Logs to Crank Up Inference Oomph
Right off the bat, let’s give a shout out to the mathematician propeller-heads who create the transf ...…
Narration-of-Thought: Inference-Time Scaffolding for Defeasible Ethical Reasoning in Large Language Models
arXiv:2606.26366v1 Announce Type: new Abstract: Standard chain-of-thought on moral dilemmas exhibits two failure modes: stakeholder collapse (the trace names at most one party with…
OpenAI Launches First Self-Developed AI Inference Chip, Boosting NVIDIA/WiMi's Scale of Computing Power Advantage - Moomoo
Comprehensive up-to-date news coverage, aggregated from sources all over the world by Google News.…
Show HN: Ludion – routing AI inference by observed WebGPU behavior
Stop wasting cloud inference on browser-sized work. Browser when safe, server when needed, measured every time.…
OpenAI and Broadcom announce chip designed for LLM inference at scale
The silicon race is heating up amid the struggle to keep up with demand.…
Still: Amortized KV Cache Compaction in a Single Forward Pass
The KV cache is the memory bottleneck of long-horizon language model deployment. Practically, a deployable compactor must be lightweight enough to call during inference, expressive…
TensorSharp: Open-Source Local LLM Inference Engine
A C# inference engine for running large language models (LLMs) locally using GGUF model files. TensorSharp provides a console application, a web-based chatbot interface, and Ollama…
Lean Inference: Lean Manufacturing Principles Applied to AI
Making inference scale in a cost effective way…
Show HN: Hive Trust – Ed25519-signed benchmarks for every AI inference primitive
Hive primitives benchmarked against published SOTA adversaries. Every result is a signed Ed25519 receipt from hivemorph — queryable, tamper-evident, reproducible.…
FingerMotion shares rise on entry into edge AI inference computing market
Building a High-Performance Real-Time Data Pipeline with Edge Inference and Observability
Building a High-Performance Real-Time Data Pipeline with Edge Inference and...…
With Nvidia Groq 3, the Era of AI Inference Is (Probably) Here (⌛ March 2026)
What makes Nvidia's new Groq 3 LPU chip a must-watch in the AI world?…
Computer Use Agents Go Local: A Deep Technical Dive into On-Device GUI Automation, Quantized Inference & Holo3.1
Meta Description: Learn how to build production-grade local computer use agents using Holo3.1's...…
Unveiling the Structure of Do-Calculus Reasoning via Derivation Graphs
The do-calculus defines a general system of inference for interventional queries, allowing causal quantities to be transformed through successive applications of its rules. This pr…
Do Real-World Datasets Contain Natural Experiments? An Empirical Study Using Causal Feature Selection
In nature, events that affect some individuals or groups but not others constitute an implicit intervention and are known as natural experiments. For example, the COVID-19 pandemic…