mindpool.io
$ df -h ./hardware

Pick the rig that runs the work

Bandwidth, not capacity, sets your tokens/sec. Match the machine to the model size you must load, then accept the speed its bandwidth allows.

T2 is the enrollment floor — it runs the full curriculum end to end.

Local hardware runs ~$1,000–$5,000 one-time — a single purchase, not a subscription.

Local vs cloud — See the local-vs-cloud comparison ↓

Check your rig before you enroll

Install mpl and run one command — it prints a curriculum verdict for your machine (full / partial / not completable) so you know before you pay.

$ mpl doctor --profile

Two ways in. Not ready to buy hardware? Rent first — mpl cloud (beta) spins up a GPU box you control, billed by the second; stop it when you're not learning, own the metal when you're ready. Or own it from day one — pick a rig below. Either path runs the full curriculum on the same sovereign mpl.

~/mindpool/cloud-fork
rentmpl cloud · beta
Rent first, own later.

Stop/start works on GCP/Azure; RunPod is terminate-and-restore.

mpl cloud docs →

mpl cloud is beta. It rents and controls GPU compute on your behalf, and like any beta that depends on third-party cloud providers and their APIs, mpl cloud can fail to launch, stop, start, or resume a cluster if something upstream changes or breaks — and work on a stopped cluster may be at risk until you recover it. We recommend mpl cloud for learners who accept this, back up their checkpoints to a network volume or object store, and don't mind a rough edge. Want full control with no beta? Bring any SSH-able GPU box, run mpl setup after enrolling, and connect it with mpl fleet add — any provider, fully sovereign.

own
Own the metal from day one.

Pick a rig below — a single purchase, not a subscription. Full control, no beta.

pick a tier ↓
~/mindpool/hardware
T1Entry
Apple
Mac mini M4 16–24GB
120 GB/s
Mini-PC
DIY NVIDIA
RTX 5060 Ti 16GB
448 GB/s
T2Floor
Apple
MBP 14" M5 Pro 64GB
307 GB/s
Mini-PC
AMD Ryzen AI Max+ 395 · Strix Halo 128GB (Beelink GTR9 Pro / GMKtec EVO-X2)
~180 GB/s eff
DIY NVIDIA
used RTX 4090 24GB
1,008 GB/s
T3Pro
Apple
Mac Studio M3 Ultra 96GB
819 GB/s
Mini-PC
NVIDIA DGX Spark (GB10) 128GB · or ASUS Ascent GX10
273 GB/s
DIY NVIDIA
RTX 5090 32GB
1,792 GB/s
T4Research
Apple
MBP M5 Max 128GB
614 GB/s
Mini-PC
2× GB10 clustered → 256GB (DGX Spark / Ascent GX10)
273 GB/s
DIY NVIDIA
dual RTX 3090 = 48GB
936 GB/s/card
policy · CPU-only machines (no Metal, CUDA, or ROCm): you can read every lesson and dry-run the labs, but you cannot complete the curriculum locally — below the tier floor mpl blocks local execution, and the supported path is the cloud (your machine becomes the cockpit).
$ ls ./models --by-tier

What your box can actually run

Each model at its featured quant, against the four tiers: whether it fits and how fast it decodes. Speed is the bandwidth truth made literal — active-param reads explain why MoE models far exceed their total size.

~/mindpool/models
# ── Gemma 4 ──
Gemma 4 E2B
2B · QAT UD-Q4_K_XL · 3 GB · multimodal · Apache-2.0 · source ↗
T1
~120–448 tok/s
T2
~180–1008 tok/s
T3
~273–1792 tok/s
T4
~273–936 tok/s
Gemma 4 E4B
4B · QAT UD-Q4_K_XL · 5 GB · multimodal · Apache-2.0 · source ↗
T1
~60–224 tok/s
T2
~90–504 tok/s
T3
~137–896 tok/s
T4
~137–468 tok/s
Gemma 4 12B
12B · QAT UD-Q4_K_XL · 7 GB · multimodal · Apache-2.0 · source ↗
T1
~20–75 tok/s
T2
~30–168 tok/s
T3
~46–299 tok/s
T4
~46–156 tok/s
Gemma 4 26B-A4B
26B (4B active) · QAT UD-Q4_K_XL · 15 GB · multimodal · Apache-2.0 · source ↗
T1
~60–224 tok/s
T2
~90–504 tok/s
T3
~137–896 tok/s
T4
~137–468 tok/s
Gemma 4 31B
31B · QAT UD-Q4_K_XL · 18 GB · multimodal · Apache-2.0 · source ↗
T1
— too big
T2
~12–65 tok/s
T3
~18–116 tok/s
T4
~18–60 tok/s
# ── Qwen3 ──
Qwen3 8B
8B · Q4_K_M · 6 GB · text · Apache-2.0 · source ↗
T1
~30–112 tok/s
T2
~45–252 tok/s
T3
~68–448 tok/s
T4
~68–234 tok/s
Qwen3 14B
14B · Q4_K_M · 9 GB · text · Apache-2.0 · source ↗
T1
~17–64 tok/s
T2
~26–144 tok/s
T3
~39–256 tok/s
T4
~39–134 tok/s
Qwen3 30B-A3B
30B (3B active) · Q4_K_M · 18 GB · text · Apache-2.0 · source ↗
T1
— too big
T2
~120–672 tok/s
T3
~182–1195 tok/s
T4
~182–624 tok/s
# ── Llama ──
Llama 8B
8B · Q4_K_M · 6 GB · text · Llama Community · source ↗
T1
~30–112 tok/s
T2
~45–252 tok/s
T3
~68–448 tok/s
T4
~68–234 tok/s
Llama 70B
70B · Q4_K_M · 42 GB · text · Llama Community · source ↗
T1
— too big
T2
~5–9 tok/s · 2/3 rigs
T3
~8–23 tok/s · 2/3 rigs
T4
~8–27 tok/s
Llama 405B
405B · Q4_K_M · 230 GB · text · Llama Community · source ↗
T1
— too big
T2
— too big
T3
— too big
T4
~1 tok/s · 1/3 rigs
# ── DeepSeek ──
DeepSeek-V2-Lite
16B (2.4B active) · Q4_K_M · 10 GB · text · DeepSeek · source ↗
T1
~100–373 tok/s
T2
~150–840 tok/s
T3
~228–1493 tok/s
T4
~228–780 tok/s
# ── OLMo 2 ──
OLMo 2 13B
13B · Q4_K_M · 8 GB · text · Apache-2.0 · source ↗
T1
~18–69 tok/s
T2
~28–155 tok/s
T3
~42–276 tok/s
T4
~42–144 tok/s
formula tok/s ≈ bandwidth ÷ (active-params × bytes/param)
Estimated decode ceiling — real-world is lower (attention, KV, overhead) and this ignores prefill. Uses active params, so MoE reads fast-and-small.
as of 2026-06 — curated, not exhaustive; models churn, re-verify each footprint at its source.

Why activeparams, not total, set speed — unpacked in the quantization & MoE deep dive (coming with the curriculum).

$ cat ./if-you-are-X-buy-Y

The recommendation matrix

Tightest budget, want 70B in memory, OK with Linux + ROCm
Beelink GTR9 Pro — AMD Ryzen AI Max+ 395 (Strix Halo, 128 GB)

Cheapest 128 GB box in volume; 70B Q4 ~5–8 tok/s.

Want a turnkey Windows 128 GB box, no building
GMKtec EVO-X2 — AMD Ryzen AI Max+ 395 (Strix Halo)

Windows-native, OCuLink eGPU escape hatch.

Want AMD's own supported Halo box, not a budget mini-PC
AMD Ryzen AI Halo Developer Platform (Ryzen AI Max+ 395, 128 GB)

AMD's first-party dev box — 128 GB LPDDR5x-8000, Radeon 8060S, Linux or Win11 Pro; same Strix Halo silicon as the budget mini-PCs, vendor-supported.

Smoothest experience, value macOS + portability, models ≤34B
MacBook Pro 14" M5 Pro, 64 GB

MLX just works for LoRA/QLoRA; 307 GB/s; ships now.

Fine-tuning-heavy, want fast tokens + GPU upgrade path
DIY used RTX 4090 24 GB

1,008 GB/s; cleanest CUDA path; best value for the capability.

Want NVIDIA's stack + 70B QLoRA in one mini box
NVIDIA DGX Spark (GB10, 128 GB)

NVIDIA's own GB10 box — 1 PFLOP FP4, full CUDA, fine-tune ≤70B. ASUS Ascent GX10 is the cheaper GB10 sibling.

Max memory + fast tokens, single vendor desk box
Mac Studio M3 Ultra, 96 GB

819 GB/s — highest bandwidth shipping; 70B ~20–30 tok/s.

$ ls ./certified-silicon

Certified silicon — bring your own

Commodity hardware, silicon-neutral. Systems are certified with real hardware profiling and benchmark runs on the actual box — AMD first.

AMD
First certified
Apple Silicon (MLX)
Certified
NVIDIA
Runtime validated (T4) · certification in progress
OEM boxes
On request

Representative — certification in progress. Real benchmarked throughput figures are published as each certification run lands.

$ cat ./local-vs-cloud

Own it, or rent it?

Running the same open models on hardware you own versus cloud GPUs you rent, across four axes. This lays out the trade-offs so you can match a choice to your own usage, privacy needs, and budget — it is not a recommendation either way. The right answer depends on how many hours you'll run, your data-sovereignty requirements, and the largest model you need to reach.

~/mindpool/local-vs-cloud
Local owned rigown
Cost shape
One-time ~$1k–$5k capex; ~$10–70/mo power. Near-zero marginal cost after purchase; strong mid-2026 resale.
Control & privacy
Full sovereignty — weights, data, traffic never leave the box. Offline / air-gap capable.
Setup & ops
One-time setup (hours–days); low upkeep (Apple ≈ none, NVIDIA driver/CUDA). No per-session friction; doesn't burst-scale.
Capability ceiling
70B Q4 inference; 128GB+ rigs reach 405B-capable Q4 / 70B QLoRA. Hard memory ceiling — no 70B FP16 or 405B training.
GPU specialist / marketplacerent
RunPod, Vast.ai, Lambda, CoreWeave, Crusoe, TensorDock
Cost shape
Pure opex. Consumer ~$0.1–0.7/hr · A100 ~$0.7–2.5/hr · H100 ~$1–4.25/hr · B200 ~$2–6/hr. Spot −40–80%; egress often free.
Control & privacy
Your stack, but data/weights sit on rented infra during the job. Mostly US jurisdiction; marketplace hosts add unknown location. No offline.
Setup & ops
First token ~1–3 min managed, ~10+ min marketplace. No hardware upkeep; watch spot interruption, idle billing, ephemeral storage.
Capability ceiling
Budget-bound, not memory-bound: single A100/H100 for 70B; 4–8× clusters for 405B / full 70B fine-tune. Far past any single rig.
Hyperscaler GPUrent
AWS, GCP, Azure
Cost shape
Pure opex. ~$0.7–1.7/hr small GPU to ~$55–98/hr per 8×H100 node. H100 on-demand ~$7–11/GPU-hr; spot ~$2.25–3.7. Egress + surcharges add ~20–40%.
Control & privacy
Region-pinned with VPC isolation, CMEK/BYOK, residency controls — but US CLOUD Act always applies. No air-gap.
Setup & ops
Highest setup friction: GPU quota requests (1–7 days), IAM/VPC/driver setup. Elastic once running; auto-stop discipline needed.
Capability ceiling
Effectively unlimited within quota/spend: 8×H100 node (640GB) runs 405B FP8; multi-node trains beyond 405B.
Managed / serverless GPUrent
Modal, RunPod Serverless, Replicate, HF Endpoints, Baseten
Cost shape
Pure opex. Per-second serverless scales to zero (cheapest for bursty); dedicated managed bills continuously, +30–75% over bare metal. ~$0.5–9.25/hr.
Control & privacy
Processed on third-party managed infra; no offline. Dedicated endpoints cut multi-tenant bleed; compliance varies. Must upload weights.
Setup & ops
Lowest friction — first token <~5 min, deploy by function/container. Cold starts <200ms–~60s. Minimal ops; limited low-level tuning.
Capability ceiling
Single-GPU ceiling matches raw rental (H100/H200 → 70B; B200 → larger). 405B / full fine-tune needs cluster configs near raw-cloud pricing.
~/mindpool/cloud --by-gpu-class
Consumer GPU (RTX 4090/5090-class, 24–32GB) ~$0.10–0.70/hr
Spot/marketplace as low as ~$0.10–0.34/hr; on-demand toward the top. 8B–13B comfortably, 70B only at aggressive quant. Specialist/marketplace clouds only.
Datacenter single GPU (A100 80GB / L40S) ~$0.67–2.50/hr
Marketplace ~$0.67–1.50/hr; managed on-demand ~$1.40–2.50/hr; hyperscaler ~$3.50–4.10/hr. Handles 70B Q4 inference + 70B QLoRA on one card.
Datacenter H100/H200 (80–141GB) ~$1.00–7.00/hr per GPU
Marketplace spot ~$1.00–1.30/hr, on-demand ~$2.50–4.25/hr; hyperscaler ~$6.9–11/GPU-hr (GCP spot ~$2.25–3.7). Volatile.
Blackwell-tier (B200/B300, 192–288GB) ~$2.10–12/hr per GPU
Spot from ~$2.10/hr; specialist on-demand ~$5.9–10.5/hr; hyperscaler/managed often sales-quoted. 405B Q4 in a few cards.

Break-even

Break-even is utilization-driven, not binary. Light use (under ~30 hrs/month): cloud stays cheaper for the whole 3-year ownership cycle. Moderate use (~6 hrs/day): crossover against an A100-equivalent rental lands around ~10–11 months; against the cheapest consumer-GPU spot it can stretch to multiple years. Always-on serving (~720 hrs/month): local typically wins by month ~2–4. Rough anchor: a ~$3,000 rig ÷ ~$1.50/hr A100 ≈ 2,000 hrs (~11 months at 6 hrs/day); ÷ ~$0.34/hr consumer spot ≈ 8,800 hrs (~4 years).

  • Mid-2026 GPU/RAM pricing is volatile (H100 swung from ~$8+/hr peak to ~$1.70 trough to ~$2.35 by early 2026) — treat every figure as a coarse range and verify live before budgeting.
  • Provider choice matters as much as own-vs-rent: hyperscaler H100 (~$7–11/GPU-hr) runs 3–6× community/specialist rates (~$1–4/hr).
  • Hidden cloud costs add ~20–40% on hyperscalers: egress (~$0.09–0.12/GB), checkpoint storage, managed-platform surcharges, spot-interruption waste. Community clouds mostly waive egress.
  • Spot pricing is the biggest cost lever (−40–80%) but requires checkpointing; interruption warnings range ~2 min (AWS) to ~15 sec (Vast.ai).
  • US CLOUD Act applies to essentially all major US providers regardless of region; for strict sovereignty / air-gap / regulated data, local is the only fully offline option.
  • Local capability is memory-bound: the $1k–$5k tier tops out near 405B-capable Q4 inference / 70B QLoRA; full 70B FP16 fine-tunes and 405B training need rented multi-GPU clusters — mpl burst is the relief valve, on your own cluster or managed cloud.
$ man bandwidth

The bandwidth truth

Generating each token reads the active weights out of memory, so your tok/s ceiling ≈ bandwidth ÷ active-weight-bytes. Capacity decides whether a model fits; bandwidth decides how fast it runs once it fits — and mpl burst is the fallback for the rare 70B full fine-tune — run it on hardware you own or rent managed cloud, your call; see the local-vs-cloud comparison below.

Mid-2026 figures; the 2026 RAM/GPU shortage adds volatility — re-verify before buying or renting.

Hardware ready.

See which open models run on your hardware — and how each performs.