locallist

“The hamsters go brrrrr”
qwen3.6-35b-a3b
quant6bit
rank#15 / 154 open-weight
open-weight#15 / 154
AA Index31.6 pts #118/442
GPQA84.1% #81/437
HLE20.2% #83/425
AA-LCR63.7 #82/371
Terminal44.9% #62/143
MMMU-Pro75 #43/203
SciCode35.8% #157/415
benchmarklist
hardware2× RTX 3060 12GB
price$700
memory31.8 +kv 2.5 / 24 GB
speed260t/s pp
infollama.cpp, Qwen3.6-35B-A3B-UD-Q6_K_XL, MoE offloaded to CPU, -c 131072. Prefill ~260t/s on CUDA0 (PCIe x16 to CPU) vs ~170t/s on CUDA1 (PCIe x8); tg/sec similar on both. Asking whether the x16 vs x8 link explains the pp gap.
qwen3.6-35b-a3b
quant4bit
rank#15 / 154 open-weight
open-weight#15 / 154
AA Index31.6 pts #118/442
GPQA84.1% #81/437
HLE20.2% #83/425
AA-LCR63.7 #82/371
Terminal44.9% #62/143
MMMU-Pro75 #43/203
SciCode35.8% #157/415
benchmarklist
hardwareRTX 3060 12GB · Xeon E5-2650 v4 · 32GB RAM
price$850
memory17.3 +kv 1.71661 / 44 GB
speed35t/s
infollama-server, APEX-MTP-I-Compact.gguf, -ngl all, --n-cpu-moe 15, -c 90000, -b 2048 -ub 1024, -fa on, -ctk/-ctv turbo3_tcq. Reported 35.20 t/s (2.301 tokens in ~65s sample). Asking if that's good for this hybrid GPU+CPU MoE setup.
qwen3.6-27b
quant8bit
rank#18 / 154 open-weight
open-weight#18 / 154
AA Index37.1 pts #92/442
GPQA84.2% #79/437
HLE21.6% #79/425
SWE-bench70% #43/68
AA-LCR68.7 #40/371
Terminal60.7% #44/143
MMMU-Pro74.6 #47/203
SciCode39.8% #103/415
benchmarklist
hardware4× RTX 5060 Ti 16GB
price$2400
memory36.8 +kv 12.207 / 64 GB
speed128t/s · 177t/s pp
infoSGLang with Minachist/Qwen3.6-27B-INT8-AutoRound — 8 concurrency, TP=4, MTP, 16K batch tokens, 200K context, bfloat16 KV cache. 4× RTX 5060 Ti (4 lanes of gen4 per card via bifurcation, ~8GB/s each). Serving bench: input 177.35 tok/s, output 127.63 tok/s, mean TTFT 7.5s. Lower TTFT than vLLM at the same concurrency on this rig.
qwen3.6-27b
quant4bit
rank#18 / 154 open-weight
open-weight#18 / 154
AA Index37.1 pts #92/442
GPQA84.2% #79/437
HLE21.6% #79/425
SWE-bench70% #43/68
AA-LCR68.7 #40/371
Terminal60.7% #44/143
MMMU-Pro74.6 #47/203
SciCode39.8% #103/415
benchmarklist
hardwareRTX Quadro 4000 8GB · AMD BC-250 16GB
price$400
memory13.5 +kv 2 / 24 GB
speed20t/s · 100t/s pp
infoRunning with llama.cpp RPC (cuda backend on the RTX, vulkan on the BC-250). llama-server --backend-sampling --n-gpu-layers -1 --rpc 192.168.100.10:50052 --jinja --cache-ram 32768 -fa on --model /opt/models/bottlecapai/ThinkingCap-Qwen3.6-27B-GGUF/ThinkingCap-Qwen3.6-27B-Q4_K_M.gguf --cache-type-k q8_0 --cache-type-v q8_0 --ctx-size 65536 --temp 1.0 --top-p 0.95 --top-k 64 -b 4096 -ub 1024 --spec-type draft-mtp --spec-draft-n-max 2
qwen3.6-27b
quant4bit
rank#18 / 154 open-weight
open-weight#18 / 154
AA Index37.1 pts #92/442
GPQA84.2% #79/437
HLE21.6% #79/425
SWE-bench70% #43/68
AA-LCR68.7 #40/371
Terminal60.7% #44/143
MMMU-Pro74.6 #47/203
SciCode39.8% #103/415
benchmarklist
hardwareRTX 3090 24GB
price$1400
memory13.5 +kv 2.7 / 24 GB
speed30-35t/s
infollama.cpp, Qwen3.6-27B Q4_M on a 3090 — 30–35 t/s out of the box. No ctx/KV details given.
qwen3.5-397b
quant4bit
rank#10 / 154 open-weight
open-weight#10 / 154
AA Index33.7 pts #103/442
GPQA89.3% #38/437
HLE27.3% #56/425
SWE-bench69.9 #28/29
AA-LCR65.7 #64/371
Terminal51.3% #54/143
MMMU-Pro77.3 #30/203
SciCode42% #71/415
LiveCode79.3% #13/44
benchmarklist
hardwareRTX PRO 6000 Max-Q 96GB · RTX A6000 48GB · 128GB DDR4
price$17000
memory244 +kv 1.875 / 272 GB
speed13-14t/s
infoQwen3.5-397B-A17B Q4_K_M (~244GB) with 128K Q8 KV (~1.9GB) squeezed across PRO 6000 Max-Q 96GB + A6000 48GB + 128GB DDR4 (PCIe Gen3) — ~13–14 t/s.
qwen3.5-122b
quant4bit
rank#9 / 154 open-weight
open-weight#9 / 154
AA Index32.3 pts #114/442
GPQA85.7% #62/437
HLE23.4% #72/425
AA-LCR66.7 #53/371
Terminal47.6% #58/143
MMMU-Pro75 #45/203
SciCode42% #72/415
benchmarklist
hardwareApple M3 Ultra 512GB (80 GPU cores)
price$10000
memory69.6 +kv 0.0117188 / 512 GB
speed54t/s · 605t/s pp
infoMLX 4bit on M3 Ultra 512GB — ~54 t/s tg at short ctx (~70GB resident), falling to ~32 t/s / 91.7GB at 128k. Also tables an M5 Max 128GB for comparison (faster pp/tg).
qwen3.5-122b
quant3bit
rank#9 / 154 open-weight
open-weight#9 / 154
AA Index32.3 pts #114/442
GPQA85.7% #62/437
HLE23.4% #72/425
AA-LCR66.7 #53/371
Terminal47.6% #58/143
MMMU-Pro75 #45/203
SciCode42% #72/415
benchmarklist
hardwareRTX 5060 Ti 16GB
price$600
memory58.2 +kv 11.64 / 16 GB
speed14.4t/s · 140t/s pp
infollama.cpp + tools. 122B Q3_K_M on single 5060 Ti: 140t/s pp @2.7k / 14.4t/s tg @2k.
qwen3.5-35b-a3b
quant8bit
rank#16 / 154 open-weight
open-weight#16 / 154
AA Index29.3 pts #131/442
GPQA84.5% #75/437
HLE19.7% #86/425
AA-LCR62.7 #89/371
Terminal40.8% #69/143
MMMU-Pro72.7 #62/203
SciCode37.7% #128/415
benchmarklist
hardwareMacBook M1 Max 64GB
price$2000
memory37.7 +kv 0.693722 / 64 GB
speed41t/s · 483t/s pp
infoMLX mlx-community/Qwen3.5-35B-A3B-8bit on M1 Max 64GB. Baseline (no KV quant): prefill 483.3 / decode 41.0 t/s on ~36k prompt. TurboQuant 3.5bit KV tanks decode to 6.6 t/s.
qwen3.5-35b-a3b
quant8bit
rank#16 / 154 open-weight
open-weight#16 / 154
AA Index29.3 pts #131/442
GPQA84.5% #75/437
HLE19.7% #86/425
AA-LCR62.7 #89/371
Terminal40.8% #69/143
MMMU-Pro72.7 #62/203
SciCode37.7% #128/415
benchmarklist
hardware2× RTX 5060 Ti 16GB
price$1200
memory36.9 +kv 7.38 / 32 GB
speed58t/s · 826t/s pp
infoQwen3.5-35B-A3B Q8 on dual 5060 Ti — prompt eval 826 t/s, gen 58 t/s. PP limited by x1 PCIe 4.0 link.
qwen3.5-35b-a3b
quant8bit
rank#16 / 154 open-weight
open-weight#16 / 154
AA Index29.3 pts #131/442
GPQA84.5% #75/437
HLE19.7% #86/425
AA-LCR62.7 #89/371
Terminal40.8% #69/143
MMMU-Pro72.7 #62/203
SciCode37.7% #128/415
benchmarklist
hardware2× RTX 4090 24GB
price$4800
memory36.9 +kv 7.38 / 48 GB
speed94t/s · 4505t/s pp
infoDual 4090, Q8 KV. 35B-A3B Q8_0: 4505 t/s pp / 94 t/s tg. Also quotes 27B Q8_0 at 1462 pp / 23.45 tg on the same rig.
qwen3.5-35b-a3b
quant8bit
rank#16 / 154 open-weight
open-weight#16 / 154
AA Index29.3 pts #131/442
GPQA84.5% #75/437
HLE19.7% #86/425
AA-LCR62.7 #89/371
Terminal40.8% #69/143
MMMU-Pro72.7 #62/203
SciCode37.7% #128/415
benchmarklist
hardwareMacBook M1 Max 64GB
price$2000
memory37.7 +kv 0.82016 / 64 GB
speed39t/s
infoM1 Max 64GB multi-model note. Headline in post: 35B-A3B@8bit 38.54 t/s at ~43k ctx. Also 9B MLX bf16 17.3 and 9B@9bit 26.5 at ~40–44k (row model tagged 9b).
qwen3.5-35b-a3b
quant5.5bit
rank#16 / 154 open-weight
open-weight#16 / 154
AA Index29.3 pts #131/442
GPQA84.5% #75/437
HLE19.7% #86/425
AA-LCR62.7 #89/371
Terminal40.8% #69/143
MMMU-Pro72.7 #62/203
SciCode37.7% #128/415
benchmarklist
hardwareMacBook Pro M1 Max 64GB
price$2000
memory22.3 +kv 0.400543 / 64 GB
speed45t/s
infoM1 Max 64GB, ~40–44k context (70-page PDF summarize). Best: inferencerlabs 5.5bit MLX 45 t/s (8bit KV in LM Studio). Also 9bit 39, mlx 6bit 23, GGUF MXFP4 10, UD-Q4_K_XL 20 t/s.
qwen3.5-35b-a3b
quant4bit
rank#16 / 154 open-weight
open-weight#16 / 154
AA Index29.3 pts #131/442
GPQA84.5% #75/437
HLE19.7% #86/425
AA-LCR62.7 #89/371
Terminal40.8% #69/143
MMMU-Pro72.7 #62/203
SciCode37.7% #128/415
benchmarklist
hardwareNVIDIA DGX Spark 128GB
price$4800
memory22 +kv 4.4 / 128 GB
speed70t/s
infoDGX Spark running Qwen3.5-35B Q4_K_M at 70+ t/s full context, alongside Kokoro TTS, Parakeet ASR, z-image-turbo, and qwen-image-edit.
qwen3.5-35b-a3b
quant2bit
rank#16 / 154 open-weight
open-weight#16 / 154
AA Index29.3 pts #131/442
GPQA84.5% #75/437
HLE19.7% #86/425
AA-LCR62.7 #89/371
Terminal40.8% #69/143
MMMU-Pro72.7 #62/203
SciCode37.7% #128/415
benchmarklist
hardwareRX 9070 XT 16GB
price$700
memory12.2 +kv 1.25 / 16 GB
speed111t/s
infoUnsloth UD-Q2_K_XL, 64k context, fully on 16GB RX 9070 XT with no offloading — 111 t/s.
qwen3.5-27b
quant8bit
rank#11 / 154 open-weight
open-weight#11 / 154
AA Index33.8 pts #103/442
GPQA85.8% #60/437
HLE22.2% #78/425
AA-LCR67.3 #47/371
MMMU-Pro75 #43/203
SciCode39.5% #108/415
benchmarklist
hardware2× AMD Instinct MI50/MI60 16GB
price$400
memory28.6 +kv 1 / 32 GB
speed28t/s · 256t/s pp
infollama.cpp on 2× AMD Instinct MI50/MI60, Qwen 27B Q8_0. Tensor (Full) bench: pp2048 256.29 t/s, tg256 28.45 t/s; at 16k context pp 233.83 / tg 27.43. Tensor over PCIe 1.0 drops prefill ~30% (pp 180 → still ok decode). MTP wasn't a big win on MI50 — compute-bound more than interconnect.
qwen3.5-27b
quant5bit
rank#11 / 154 open-weight
open-weight#11 / 154
AA Index33.8 pts #103/442
GPQA85.8% #60/437
HLE22.2% #78/425
AA-LCR67.3 #47/371
MMMU-Pro75 #43/203
SciCode39.5% #108/415
benchmarklist
hardwareRTX 5090 32GB
price$4200
memory21.7 +kv 4.34 / 32 GB
speed52t/s
infoRTX 5090 — Qwen3.5-27B Q5_K_L at 52 t/s generation (comparison datapoint).
qwen3.5-0.8b
quant8bit
rank#153 / 154 open-weight
open-weight#153 / 154
AA Index5.5 pts #451/464
GPQA23.6% #444/454
HLE4.9% #306/443
AA-LCR6.7 #280/370
Terminal0.4% #137/142
MMMU-Pro25.8 #181/197
SciCode2.9% #419/433
benchmarklist
hardwarei5-1145G7 (CPU only)
price$300
memory0.8 +kv 0.16 / 16 GB
speed12t/s
infoqwen3.5-0.8b Q8_0 in LM Studio — 12 t/s CPU-only on i5-1145G7.
qwen3-coder-next
quant4bit
rank#63 / 154 open-weight
open-weight#63 / 154
AA Index21.1 pts #193/464
GPQA73.7% #170/454
HLE9.3% #173/443
AA-LCR40 #168/370
Terminal38.2% #73/142
SciCode32.3% #200/433
benchmarklist
hardwareRTX 5060 Ti 16GB · 96GB DDR5
price$2200
memory48.5 +kv 9.7 / 112 GB
speed14.75t/s
infoqwen3-coder-next Q4_K_M on 5060 Ti 16GB + 96GB DDR5. ik_llama 14.75t/s vs llama.cpp 8t/s vs krasis 26.5t/s. Side note: krasis OpenAI API issues with openclaw.
qwen3-coder-next
quant4bit
rank#63 / 154 open-weight
open-weight#63 / 154
AA Index21.1 pts #193/464
GPQA73.7% #170/454
HLE9.3% #173/443
AA-LCR40 #168/370
Terminal38.2% #73/142
SciCode32.3% #200/433
benchmarklist
hardwareNVIDIA DGX Spark 128GB
price$4800
memory45 +kv 9 / 128 GB
speed95t/s · 45000t/s pp
infoQwen3-Coder-Next NVFP4 on DGX Spark — 45000 t/s PP2048, 95 t/s tg512.
qwen3-coder-next
quant4bit
rank#63 / 154 open-weight
open-weight#63 / 154
AA Index21.1 pts #193/464
GPQA73.7% #170/454
HLE9.3% #173/443
AA-LCR40 #168/370
Terminal38.2% #73/142
SciCode32.3% #200/433
benchmarklist
hardwareRTX 4070 12GB · 64GB DDR5 · 7950X
price$1500
memory48.5 +kv 0.9375 / 76 GB
speed26t/s · 307t/s pp
infollama-server unsloth Coder-Next Q4_K_M, ctx 40k, flash-attn, -ctk/-ctv q4_0, --fit on. 30k-token prompt: 307 t/s pp / 25.54 t/s tg.
glm-5.2
quant1bit
rank#2 / 154 open-weight
open-weight#2 / 154
AA Index51.1 pts #28/442
GPQA91.2% #31/112
HLE40.1% #24/425
SWE-bench82.8% #6/68
AA-LCR71.3 #17/371
Terminal82.7% #11/143
SciCode50.5% #26/415
LiveCode69.5% #75/119
benchmarklist
hardwareRTX 5070 Laptop 8GB · Ryzen AI 7 350 · 32GB RAM
price$1500
memory228 +kv 45.6 / 40 GB
speed0.2t/s · 0.3t/s pp
infoGLM-5.2 UD-IQ1_M (~228GB) streaming experts from NVMe on a 5070 Laptop 8GB + 32GB RAM — ~0.3 t/s pp / ~0.2 t/s tg (~3.7GB VRAM + ~31GB RAM).
glm-4.7
quant4bit
rank#43 / 154 open-weight
open-weight#43 / 154
AA Index26.6 pts #149/442
GPQA66.4% #53/112
HLE6.1% #221/425
SWE-bench69.4% #48/68
AA-LCR36.3 #172/371
SciCode35.4% #160/415
LiveCode82.2% #40/119
benchmarklist
hardwareRTX PRO 6000 96GB · 9950X · 256GB DDR5-6000
price$20000
memory216 +kv 11.5 / 352 GB
speed5.8t/s
infoGLM-4.7 Q4_K_M (~216GB) on 9950X + 256GB DDR5-6000 + RTX PRO 6000 — ~5.8 t/s.
glm-4.7
quant3bit
rank#43 / 154 open-weight
open-weight#43 / 154
AA Index26.6 pts #149/442
GPQA66.4% #53/112
HLE6.1% #221/425
SWE-bench69.4% #48/68
AA-LCR36.3 #172/371
SciCode35.4% #160/415
LiveCode82.2% #40/119
benchmarklist
hardware2× Radeon AI PRO R9700 32GB · EPYC 7532 · 256GB DDR4-3200
price$6000
memory171 +kv 11.5 / 320 GB
speed12t/s
infoGLM-4.7 Q3_K_L (~171GB, ~16GB active) on EPYC 7532 + 256GB DDR4 + 2× R9700 — 12 t/s with -ncmoe 0 (attention on GPU, experts in RAM).
glm-4.5
quant3bit
rank#62 / 154 open-weight
open-weight#62 / 154
AA Index19.5 pts #204/442
GPQA78.2% #69/112
HLE12.2% #125/425
AA-LCR48.3 #148/371
SciCode34.8% #166/415
LiveCode67.4% #76/119
benchmarklist
hardware2× RX 7900 XTX 24GB · 192GB DDR5
price$11000
memory145 +kv 11.5 / 240 GB
speed5t/s
infoGLM-4.5 IQ3_XXS (~145GB) on 2× 7900 XTX + 192GB DDR5 — only ~5 t/s; quality not acceptable — first prompt was "Hello" and it responded in Chinese.
gemma4-31b
quant16bit
rank#20 / 154 open-weight
open-weight#20 / 154
AA Index29.4 pts #131/442
GPQA85.7% #62/437
HLE26.5% #61/425
AA-LCR62 #92/371
Terminal43.4% #66/143
MMMU-Pro73.4 #55/203
SciCode43.4% #58/415
LiveCode80% #12/44
benchmarklist
hardwareRTX PRO 6000 96GB
price$12000
memory63 +kv 20 / 96 GB
speed23t/s
infoGemma 4 31B BF16 (~63GB) on RTX PRO 6000 — 23 t/s at max context, ~10GB VRAM spare, GPU power capped ~440W / 600W.
gemma4-31b
quant4bit
rank#20 / 154 open-weight
open-weight#20 / 154
AA Index29.4 pts #131/442
GPQA85.7% #62/437
HLE26.5% #61/425
AA-LCR62 #92/371
Terminal43.4% #66/143
MMMU-Pro73.4 #55/203
SciCode43.4% #58/415
LiveCode80% #12/44
benchmarklist
hardwareRTX 5060 Ti 16GB · RTX 3060 12GB
price$950
memory16.7 +kv 5 / 28 GB
speed14.6t/s
infoGemma 4 31B MeroMero IQ4_XS (~16.7GB) on KoboldCPP — 5060 Ti + 3060 layer split, 14.6 t/s.
gemma4-31b
quant4bit
rank#20 / 154 open-weight
open-weight#20 / 154
AA Index29.4 pts #131/442
GPQA85.7% #62/437
HLE26.5% #61/425
AA-LCR62 #92/371
Terminal43.4% #66/143
MMMU-Pro73.4 #55/203
SciCode43.4% #58/415
LiveCode80% #12/44
benchmarklist
hardware2× RTX 3090 24GB
price$2800
memory14.7 +kv 2.94 / 48 GB
speed50t/s · 1200t/s pp
infoGemma 4 31B (~14.7GB) with llama.cpp speculative decoding on 2× 3090 — up to ~50 t/s / ~1200 t/s pp on some tasks (~2× vs Qwen3.5-27B for this user).
gemma4-12b
quant4bit
rank#46 / 154 open-weight
open-weight#46 / 154
AA Index22 pts #188/464
GPQA75.3% #159/454
HLE14.8% #112/443
AA-LCR55.3 #121/370
MMMU-Pro69.7 #65/197
SciCode38.2% #125/433
LiveCode72% #20/41
benchmarklist
hardwareASUS ROG G700 · RTX 5080 16GB · Ultra 7 265KF · 32GB RAM
price$3000
memory6.72 +kv 2 / 16 GB
speed96t/s
infogoogle/gemma-4-12b-qat (~6.72GB) on ASUS ROG G700 (RTX 5080 16GB, Ultra 7 265KF, 32GB) — 96 t/s.
minimax-m2.5
quant3bit
rank#34 / 154 open-weight
open-weight#34 / 154
AA Index33.7 pts #104/442
GPQA84.8% #44/112
HLE19.1% #89/425
SWE-bench74.2% #30/68
AA-LCR66 #60/371
SciCode42.6% #67/415
LiveCode79.2% #58/119
benchmarklist
hardwareRTX PRO 6000 96GB
price$12000
memory109 +kv 7.75 / 96 GB
speed20t/s · 200t/s pp
infoMiniMax M2.5 Q3 (~109GB) on a single RTX PRO 6000 — 20 t/s tg / 200 t/s pp. Planning to add a 5090 for layer offload.
mistral-small-4-119b
quant4bit
rank#88 / 154 open-weight
open-weight#88 / 154
AA Index19.6 pts #208/464
GPQA76.9% #143/454
HLE9.5% #172/443
AA-LCR44.7 #157/370
Terminal21% #92/142
MMMU-Pro56.8 #117/197
SciCode38% #130/433
benchmarklist
hardwareRTX PRO 6000 96GB
price$12000
memory58.1 +kv 18 / 96 GB
speed189.7t/s · 3494t/s pp
infoMistral Small 4 119B UD-IQ4_XS (~58.1GB) on RTX PRO 6000 — llama-bench pp512 3494 t/s / tg128 190 t/s, no RAM offload.