r/LocalLLaMA
social · source
r/LocalLLaMA on WeSearch
Recent social headlines from r/LocalLLaMA.
Lead story
r/LocalLLaMA
Introducing Gemma 4 12B: a unified, encoder-free multimodal model
The latest
r/LocalLLaMA
This day in LLM history….105 years ago today, Qwen 3.6 27b was released open source. /s
r/LocalLLaMA
Gemma 4 Unified is coming
r/LocalLLaMA
Take Three: What’s the rub on memory sessions?
r/LocalLLaMA
ui: Mermaid Diagrams in chat + interactive preview by allozaur · Pull Request #24032 · ggml-org/llama.cpp
r/LocalLLaMA
Gemma 4 is coming - No Vision Tower - No Audio Tower
r/LocalLLaMA
I developed a hard LLM Challenge
r/LocalLLaMA
lipsync possible on mac?
r/LocalLLaMA
Qwen 3.7 Plus just briefly appeared and then disappeared on OpenRouter.
Reddit
Half the top 10 trending GitHub repos right now are "skills" projects, not models
r/LocalLLaMA
Tensor split mode: CUDA error on latest llama.cpp with Qwen-3.6-27b
r/LocalLLaMA
Calling it now Microsoft is buying Unsloth.
r/LocalLLaMA
Helvete-nano
r/LocalLLaMA
Holo3.1 35B/9B/4B/0.8B (Qwen 3.5 finetunes)
r/LocalLLaMA
Mellum & Granite Embedding models are ready on llama.cpp
r/LocalLLaMA
Another shout out to llama.cpp build b9455 2x3090
r/LocalLLaMA
Microsoft Aion 1.0 Instruct and Aion 1.0 Plan models!
r/LocalLLaMA
Nous Research — Hermes Desktop
r/LocalLLaMA
Why do we benchmark quants on perplexity and prose but never on tool call validity?
r/LocalLLaMA
Someone out there likely needs this
r/LocalLLaMA
Everyone here self-hosts inference. Almost nobody self-hosts the tooling around it. That feels backwards to me.
r/LocalLLaMA
Cost Analysis of my $6.4k Local LLM Server
r/LocalLLaMA
Running Qwen 3.6 35b MoE With Zoo Code On M1 Max is Amazing! Fully local, battery-powered coding powerhouse!
r/LocalLLaMA
Would a MacBook M5 16/24/32GB be an upgrade, complement, or waste next to my RTX 4060 laptop?
r/LocalLLaMA
What features dramatically improved your custom memory system?
r/LocalLLaMA
For those creating personal assistants locally - how has short/long term memory impacted your experience?
r/LocalLLaMA
Parallax: Parameterized Local Linear Attention for Language Modeling
r/LocalLLaMA
nvidia/Qwen3.6-35B-A3B-NVFP4 · Hugging Face
r/LocalLLaMA
SupraLabs 50M Parameter Model Just Hit the Trending Page on Hugging Face 🤯
r/LocalLLaMA
Why does Thinking Output More Tokens Than a Response?
r/LocalLLaMA
[LLM analysis challenge] OPERATION: REVERSE ROBOTOMY. We need an LLM Neurosurgeon to extract a password from a fractured artificial mind.
r/LocalLLaMA
Can't get over 250TPS on RTX5090 with Qwen3.5-4B
r/LocalLLaMA
LFM2.5-8B-A1B release
r/LocalLLaMA
anybody got llama-swap working answering concurrent requests for a single model?
r/LocalLLaMA
STT -> LLM -> TTS pipeline
r/LocalLLaMA
Qwen 3.6 coding choice–27B vs 35B quants
r/LocalLLaMA
"What are you good at?"
r/LocalLLaMA
Fulloch V2: 100% Local Voice Assistant for Home Assistant & Obsidian (Runs on 16GB VRAM)
r/LocalLLaMA
MINISFORUM UM790 Pro
r/LocalLLaMA
Gryphe/Pantheon-Reasoning-27B · Hugging Face
r/LocalLLaMA
Open source : Turning vocal imitations into sound effects. (New UX for sound generation)
r/LocalLLaMA
Vidai Community is now available: one Rust binary for cost attribution, guardrails and multi-provider routing on every LLM call
r/LocalLLaMA
The best AI Model for Arabic dialects 🇪🇬🦅🧡
r/LocalLLaMA
made a local voice AI for windows you can talk to in any language. open source, bring your own key
r/LocalLLaMA
I have 2x PC's. One with a 5090 and one with a 4080. Is there an easy way to use both together networked?
r/LocalLLaMA
Keeping multi-GPU rigs cool?
r/LocalLLaMA
Breaking the music supply constraint
r/LocalLLaMA
Uploaded my Qwen3.6 27B based fine tune, after two years of experience fine tuning models
r/LocalLLaMA
Mutating Gemma 4 31B Dense in to a native Gemma 4 additive-MoE model
How WeSearch handles this source
WeSearch's declared handling of r/LocalLLaMA's content. Indexing, snippets, summaries, retrieval and training are separate questions — see the rights registry or read this source's machine-readable record.
Indexing: Allowed
Snippet: Allowed
AI summary: Limited
Retrieval / RAG: Not asserted
Model training: Not asserted
Commercial reuse: Not permitted