mindpool.io
$ ls ./models

The open models worth running

Quality on four axes — reasoning, coding, agentic, long-context — every number cited to its source. Filter by your rig to see what fits, how fast it decodes, and where each model stands on the frontier.

~/mindpool/models --explore
rig
modelReasoningCodingAgenticLong-contextfits my rig
Gemma 4 E2B
2B · 3 GB · Apache-2.0 · source ↗
604424.519.1~120 tok/s
Gemma 4 E4B
4B · 5 GB · Apache-2.0 · source ↗
69.45242.225.4~60 tok/s
Gemma 4 12B
12B · 7 GB · Apache-2.0 · source ↗
77.2726943.4~20 tok/s
Gemma 4 26B-A4B
26B (4B active) · 15 GB · Apache-2.0 · source ↗
82.677.168.244.1~60 tok/s
Gemma 4 31B
31B · 18 GB · Apache-2.0 · source ↗
85.28076.966.4— too big
Qwen3 8B
8B · 6 GB · Apache-2.0 · source ↗
87.557.568.1~30 tok/s
Qwen3 14B
14B · 9 GB · Apache-2.0 · source ↗
88.6~17 tok/s
Qwen3 30B-A3B
30B (3B active) · 18 GB · Apache-2.0 · source ↗
89.5— too big
Llama 8B
8B · 6 GB · Llama Community · source ↗
48.372.676.1~30 tok/s
Llama 70B
70B · 42 GB · Llama Community · source ↗
68.988.4— too big
Llama 405B
405B · 230 GB · Llama Community · source ↗
73.3— too big
DeepSeek-V2-Lite
16B (2.4B active) · 10 GB · DeepSeek · source ↗
57.3~100 tok/s
OLMo 2 13B
13B · 8 GB · Apache-2.0 · source ↗
68.5~18 tok/s
tok/s ≈ bandwidth ÷ (active-params × bytes/param) — analytic decode ceiling; real-world is lower (attention, KV, overhead) and ignores prefill.
Scores are curated from primary sources (each number links to its source ↗); axes a model doesn't publish show "—".
Reasoning mixes MMLU-Redux %, MMLU-Pro %, MMLU-Pro CoT %, MMLU-Pro CoT 5-shot %, MMLU 5-shot % — different benchmarks, not directly rankable. Hover a score for its exact metric.
models as of 2026-06 · benchmarks as of 2026-06 — curated, not exhaustive; re-verify at source.