“In the next 2 weeks, assuming I can figure out how to mount everything onto some kind of frame, I should have as many as 30 V340Ls (60 GPUs) connected to a single server.”
“Hall of fame material”
Ready or not, personal language models (the smaller, more private, more efficient kind) are coming to the people.
Should the frontier labs be worried? Not immediately. But the band of revolutionaries taking charge are coming from bedrooms all over the world, wiring 220-volt outlets by hand at 2am. Thousands of these makeshift GPU-burning data-centre enthusiasts have clustered in a Discord server named LocalLLM, where the collective wisdom is to not ask permission from OpenAI or Anthropic to access intelligence. This is their story: a year’s worth of chatter and banter, directly from those who turned their favorite hobby into a religion.
July 2025 — “It’s time to get serious”
“It’s time to get serious.” Six thumbs-up and one frowning face. This is how it all started almost a year ago on the 14th of July. What they meant by serious we will never know, but it foreshadows the storm that soon followed: the frantic new open-weight model releases, the benchmarks, and the rise of GPU prices. Serious it became…
The silence was broken a week later, when a curious passerby asked for recommendations on a local coding agent. It did not have to be smart, but most importantly it had “to fit reasonably in 24GB of VRAM.” The answer came later in July: a recommendation for DeepSeek Coder V2 Lite at Q4_K_M with a 30k context window, good enough for autocomplete in VS Code. Space, memory, RAM, compute — whatever you want to call it — would become the community’s most sacred constraint. It did not matter how good the models were if you could not run them on consumer hardware. Enchanted by the promise of infinite intelligence, people upgraded their Macs, repurposed their gaming setups just to get a taste of what it would feel like to truly own your own model, and of course talk about them.
August 2025 — “WE MUST CONSTRUCT ADDITIONAL PYLONS!”
August saw some unglamorous house cleaning. The first mods were appointed; channels for models, hardware, use-cases, and home-labbing were brought into existence. The trickle of hellos slowly became conversations, and Reddit’s r/LocalLLaMA moderators started showing up as regular participants. TrentBot began cross-posting r/LocalLLaMA’s “rising” threads into #general and within two days people were asking for it to be caged.
GPT-5 arrived on August 8th to an unenthused audience, Sonnet 4 still ate their lunch on coding, and corporate benchmarking was dismissed: “if it’s OpenAI I’m going to disregard it like a crazy person screaming in the train station.” Suspicion compounded when several members suspected they were quietly being served 4o. Chat transcripts of the model insisting otherwise were shared: “here you can see it gaslighting me, lmao.” When gpt-oss-20b and 120b arrived mid-month, the community worked on reverse-engineering it: a member discovered that OpenAI’s own Harmony prompt-format repo was broken and rebuilt it themselves, telling the channel, “Nobody has Harmony prompting working right. OpenAI’s own Harmony repo isn’t even right.” That do-it-yourself instinct governed the LocalLLM community. People came from all over to approach technology with their own eyes rather than rely on the wisdom of those above them.
Members really did come from all over. A member with the handle CoralAnchor, already living in a kind of internal exile after his Reddit account got shadowbanned for the crime of editing a comment through a VPN, started offering benchmark data like a man smuggling contraband across a border: “Since all my poor posts got annihilated when my account got shadow banned, I’ll have to send you to github mirrors, but here are some speed numbers for the M2 Ultra and M3 Ultra.”
However, the real excitement that month came from BlueFalcon, newly minted folk hero. He announced on August 17th that he was running something like two hundred simultaneous coding agents off a single RTX 4090 using gpt-oss-20b through vLLM’s live-batching, hitting nine thousand tokens a second, a cartoonish number that went semi-viral on Reddit and dragged the whole Discord into a week of agent-orchestration theology. BlueFalcon delivered this triumph the way a man delivers a toast at his own wedding, unable to resist the bit: “WE MUST CONSTRUCT ADDITIONAL PYLONS,” with an honest admission that “tool calling in oss-20b is broken.” CoralAnchor, watching from the exile bench, replied, “I keep coming back to the idea that I’d probably have a nervous breakdown at having to code review that much stuff.”
The debate over hardware would also become prevalent. People either went the CPU (Apple) or GPU (any random graphics card you had lying around) route. BlueFalcon summed up the Mac-versus-Nvidia debate in one line that month: “Apple computers have the RAM capacity, but not enough compute power; NVIDIA is the opposite.”
By month’s end the model conversation had rotated again. DeepSeek V3.1 landed with a member warning its GGUF quants would occupy his compute for twenty hours, Hermes 4 arrived from Nous Research, and a member’s rapid-fire GLM finetunes brought out humour going back to the “Wizard-Alpaca-Vicuna-Koala-Chronos-Nous-Puffin-13b.ggml” days.
September 2025 — “I like watching the hamsters go brrrrr”
By September, the rhythm of the LocalLLM server had settled and model releases started to flow. On September 3rd, Kimi-K2-Instruct-0905 was released and the same day gpt-oss-120b was crowned top open-source model on Artificial Analysis’s intelligence index.
Then there was MossyBadger, whose late-September buying spree turned #general into a live hardware opera: multiple RTX 6000 Pro Blackwell workstation cards, roughly $45,000 spent by month’s end, and a stated plan to hit four Blackwell 96GB cards by Christmas. “Yo. Where do I buy an H200?” kicked it off and ended with a reality check: “It’s all fun and games until you try to wire the 220 outlet yourself.” In between, CoralAnchor, the server’s resident exiled benchmark priest, talked him down from the datacenter card with an argument that the H200’s real advantage is parallel throughput for many users, not raw speed for one guy alone in a room, so “an individual user would get far more bang for the buck at that price buying 3 of the RTX 6000 96GB.” A spectator chimed in: “You could literally get a used car for that price.” Another member, StaticMagnet, encouraging the whole spectacle to unfold, offered the concluding hardware prayer, noting this was “purely for my entertainment… I like watching the hamsters go brrrrrr.” This wasn’t about productivity anymore; this was about the hamsters.
The month’s model chatter was primarily about Alibaba’s dominance. Qwen releases outpaced everything else with the release of Omni Captioner, Thinking, and Instruct; then Qwen3-Max hit third on the Text Arena leaderboard; then a roadmap teasing million-to-hundred-million-token context and ten-trillion-parameter ambitions was announced. MossyBadger vowed to “shitpost Qwen” once the dust settled. GLM 4.6 was also released at the end of the month and became the coding and agent model people actually wanted to switch to. A member admitted “if it can actually match sonnet 4 i’ll switch, tired of being robbed by Anthropic,” a feeling shared by StaticMagnet, who was burning $250 a week on Anthropic credits. There was disappointment when no Air variant was announced for VRAM-constrained users — “no air no joy” — but within hours a community mlx-6bit quant appeared because the hamsters just keep running.
In September, the consensus was that if you wanted raw capability you went to the cloud providers, but to “tinker and learn” you went local. Issues of privacy were also raised: “anything going to a proprietary AI should be considered semi-public.” A standout highlight was a member impulsively deciding to start a Kimi K2 fan club with the launch of r/kimimania and publishing a Kimi-themed story the same week.
October 2025 — “guys it’s a woke communist server”
In October, the GLM hype train continued. The GGUF dropped on the 1st and by the 6th, it had already hit #1 trending on Hugging Face. The community deduced it was about eight times cheaper than Claude Sonnet 4.5 and, on tool-call accuracy, competitive with Sonnet, GPT-5, and Grok. This was a big win for open-source.
StaticMagnet spent the first days of October assembling a four-card RTX 6000 Pro Blackwell rig atop a Threadripper 7995WX, 512GB of DDR5, and 16TB of NVMe. “God I hate scalpers,” he wrote, “The fact that there are order limits on Blackwells and I can’t even get how many I need for a build because people would buy them all up to resell.” Mid-build he ordered two more cards anyways. By the 7th, he was reporting GLM-4.6 at 8-bit running with EAGLE speculative decoding and FP8: “OMFG I’m getting ~50tok/sec on GLM 4.6 8bit.” This figure was quite good given 44 t/s was achieved on an 8×H200 cluster. GLM 4.6 has about 357B parameters. Even at Q8_0 quantization from Unsloth (a local UI for training and running local models) on 4× Blackwell 6K Pros, it would only leave about 5GB of VRAM for KV Cache, the equivalent of 13,300 input tokens or a 20-page double-spaced history essay. StaticMagnet gave his personal economics: “Someone help me with the math. I spent $64k to save $20/mo on ChatGPT. Am I winning yet?” — “You only need like 266 years to recoup your investment.”
VelvetComet, another member, tested GLM-4.6 for finding pitfalls with his idea and reported it passed brilliantly but “THE SYCOPHANCY!!!! It’s like at the GPT-4o/4.1 era sycophancy levels.” Another member tested sensitive mental-health prompts, and sided with the incumbent: “it’s part of why I like claude in terms of closed models.” Closed models still had an edge with their harnesses.
Not everyone was shelling out thousands on GPUs. CoralAnchor debunked how Mac buyers were deceived by total tokens-per-second without realizing the long prompt processing speeds. Others argued multi-GPU consumer setups (3060s, 3090s, a 4070 Ti paired with a 2070) over tensor-split configs, and quants were also debated: mxfp4 versus fp8 versus int4/NVFP4. SillyTavern roleplay sessions also broke when OpenRouter pulled DeepSeek 3.1’s free tier and dragged one person into a long debugging spiral. StaticFalcon translated an entire web novel and wrote “the quality is so much better than Google Translate.” October also saw the release of DeepSeek-OCR, MiniMax-M2, and Granite 4.0. To end, a member tossed in “let’s be honest guys it’s a woke communist server,” which detonated days of argument over communism and capitalism, VelvetComet adding “I actually lived under commies. I’m originally from Russia and born in 1977.” MoltenCactus summarized the state of LocalLLM aptly: “this server’s got range, llm to commies to umamusume.”
November 2025 — “they gotta negotiate supply like OPEC”
November’s wound was NVIDIA’s DGX Spark, the datacenter that could now sit on everyone’s desk with 128GB of memory for the low, low price of $5k. For three days, from November 3rd through the 5th, the channel performed an autopsy on the machine. Pros were the CUDA ecosystem, which made the box still feel safer than AMD’s Strix Halo, but fears of a cheaper and faster Spark 2.0 on the horizon led many to ponder. The benchmark on performance was as always “how many (insert your favourite hardware) does it take to = a claude plan?”
The problem was that November was the month the RAM markets turned upside down. A 192GB DDR4 kit which cost $900 in late October pushed past $3000 by the end of the month. “They gotta negotiate supply like OPEC,” a member wrote. Another member also noticed something odd: every RTX 5090 for rent on vast.ai had seemingly disappeared within a few hours. The “just build a rig” optimism was being crushed by strange outside forces.
It wasn’t just the markets. Signs of real fear over job security also emerged. VelvetGlacier wrote regarding all the fine-tuning and agentic workflows that “my career is being quickly eaten by agents. I’d rather be the person doing the agents, and still employed.” More cynicism ensued: “get on this ai bubble money... learn enough to get a job in the industry, get overpaid while you really learn how to do it, and then hopefully you can keep it up and outlast the other people trying to do the same thing.” The community joked “I did not know deepseek was a skill.”
Kimi K2 Thinking landing on the 7th of November kept the general channel buzzing for two weeks — the first trillion-parameter open-weight model ever released. The community fired back to the model’s widely shared $4.6 million training-cost claim: “Obviously that doesn’t include the cost of chips or building a datacenter. That’s their electricity bill for the training run.” Nevertheless, GLM 4.6 still held its ground for VRAM-constrained builds. The launch of Google Gemini 3 Pro turned attention to cloud models before a Cloudflare outage that same week vindicated self-hosting.
ArliAI’s GLM-4.5-Air-Derestricted and gpt-oss-20b-Derestricted, using a “norm-preserving biprojected abliteration” method, also drew reactions. MellowCactus humorously commented “what a relief, I need to update my meth recipe.”
December 2025 — “shut up and give me all your money, also all your electricity”
December started with another large model release. Mistral Large 3 with 675 billion parameters released on the 2nd. At FP16, that meant 1.35TB of VRAM. Again, the scarce commodity was yearned for, but a real shortage started to form after November’s debacle. The conspiracy theories started: “it’s nvidia manipulating the market. They are throttling the inventories so they can keep an artificial high price.” Others blamed OpenAI’s wafer-buying instead. G.SKILL (a Taiwanese hardware manufacturer) put out price-hike notices which were forwarded around like earthquake warnings. The mood was summarized by a member: “shut up and give me all your money, also all your electricity.”
Next was the community’s favorite pastime, agonizing in public over what to buy. The DGX Spark vs. RTX 5090 vs. Strix Halo debates continued. LucidCactus announced he still had “$5832.14 left in my quarterly budget to burn through” and others spoke of their craigslist finds: “theres a guy local selling 4 of the 395 systems for $1600 each with 128gb vram each.”
The money flowed but it didn’t mean once you had bought the hardware you actually had the chance to run the models. StaticMagnet, sitting on four RTX 6000 Blackwell cards, wrote “this is depressing, but I can’t get sglang or vllm to work with any models.” As models got bigger, hardware requirements increased and the software stack had a hard time cooperating whether it was vLLM’s config hell or litellm exploding the second it touched openai-agents-sdk.
The server had started expanding, and so did its cast of new characters. A steady flow of curious people started using small models like Devstral-Small-2-24B and Qwen3 Coder as Claude substitutes. “Vibe coding” also started becoming ordinary labor. It was not without its critics. The chatter led a member to jab, “you don’t actually have a purpose for using the LLM other than telling people.” DeepSeek-V3.2 and Nemotron 3 Nano were also benchmarked to death on a made-up RuneScape puzzle and Devstral 2 was bullied for refusing to say bad words.
Then everyone held their breath as GLM 4.7 was released on the 22nd. CoralAnchor, still exiled, ran the q8 on a Mac Studio: “Since I can’t say it on reddit, I’ll say it here: GLM 4.7 q8 is insanely good to me... there is something so inherently different about the quality.” A member credited it as the first local model to beat Claude Opus 4.1 and Gemini 2.5 Pro. It was the hot release before Nvidia’s $20b acquisition of Groq, “a clever architecture that scaled up like a toddler with Lincoln logs.” Bright days were ahead!
January 2026 — “this feels like a war crime, ngl”
Yann LeCun on his way out the door at Meta told a reporter on the 2nd that the Llama 4 benchmarks “were fudged a little bit.” This revelation served only to confirm the community’s existing suspicions, which naturally meant distrusting the big labs. CES 2026 also arrived a few days later on the 5th but Nvidia announced no new consumer GPU, no 50 Super refresh, nothing to replace the 5090. AMD used its CES slot to reveal a Strix Halo mini PC to rival Nvidia’s DGX Spark, and that was about it. VRAM and RAM would just have to keep getting rationed.
The hardware crunch pushed people to be more clever. On the 9th, one member wrote roughly 1,500 lines of custom NCCL networking code to cluster three DGX Sparks together, going beyond Nvidia’s official two-node limit, just to route around a triangle-mesh subnet problem the stock software couldn’t handle. This was the essence of homelabbing: you didn’t have to buy more, you just had to be scrappier. A member wrote “are you really homelabbing if it isn’t ultra janky?” The jokes extended to students crushing a 5090 into decade-old Xeon v4 servers. “This feels like a war crime, ngl,” one wrote, capturing the spirit of the extreme tinkerers who always found ways to make things work.
Then, GLM 4.7 Flash’s release on the 15th triggered a mountain of llama.cpp and vLLM breakages: repetition loops, VRAM blowups, a prefill regression that dumped prompt processing on the CPU of the Blackwell cards. One member wrote bluntly: “Even with all of the llama.cpp updates... and using Unsloth’s direct recommendations, GLM 4.7 Flash is just bad. Like really bad.” Ten days later, this same model was hailed as a breakthrough. Fixes and patches came in waves for the next few days, and a member even took it into their own hands to independently get quantized KV cache plus flash attention working for 131k context on a single 4090.
Local models also found ways to stay competitive on coding tasks. StaticHarbor ran a fifty-sample real-world harness and reported MiniMax 2.1 scoring nearly as well as Sonnet 4.5. Then later in the month Kimi K2.5 arrived and scored 76.8% on the SWE-bench Verified benchmark and was priced at a tenth of Opus for similar output. However, that still meant 240GB of RAM even when quantized to 1.8 bits… The community also accidentally created its own benchmark. A member one-shotted Flappy Bird with Qwen3 Coder and found the game apparently “hard-baked into its parameters,” which led to the “Flappy Bird Olympics.” GLM 4.7 Flash and MiniMax 2.1 suspiciously gave near-identical results which signaled overfitting rather than genuine coding skills.
The month closed with Yann LeCun, now fully departed from Meta, telling Davos that the best open-weights models were coming from China rather than the West. And it was only getting started.
to be continued . . .