Pick the rig that runs the work

Bandwidth, not capacity, sets your tokens/sec. Match the machine to the model size you must load, then accept the speed its bandwidth allows.

▸ T2 is the enrollment floor — it runs the full curriculum end to end.

▸ Local hardware runs ~$1,000–$5,000 one-time — a single purchase, not a subscription.

▸ Local vs cloud — See the local-vs-cloud comparison ↓

Check your rig before you enroll

Install mpl and run one command — it prints a curriculum verdict for your machine (full / partial / not completable) so you know before you pay.

$ mpl doctor --profile

Two ways in. Not ready to buy hardware? Rent first — mpl cloud (beta) spins up a GPU box you control, billed by the second; stop it when you're not learning, own the metal when you're ready. Or own it from day one — pick a rig below. Either path runs the full curriculum on the same sovereign mpl.

~/mindpool/cloud-fork
rentmpl cloud · beta
Rent first, own later.

Stop/start works on GCP/Azure; RunPod is terminate-and-restore.

mpl cloud docs →

mpl cloud is beta. It rents and controls GPU compute on your behalf, and like any beta that depends on third-party cloud providers and their APIs, mpl cloud can fail to launch, stop, start, or resume a cluster if something upstream changes or breaks — and work on a stopped cluster may be at risk until you recover it. We recommend mpl cloud for learners who accept this, back up their checkpoints to a network volume or object store, and don't mind a rough edge. Want full control with no beta? Bring any SSH-able GPU box, run mpl setup after enrolling, and connect it with mpl fleet add — any provider, fully sovereign.

own
Own the metal from day one.

Pick a rig below — a single purchase, not a subscription. Full control, no beta.

pick a tier ↓
~/mindpool/hardware
T1Entry
Apple
Mac mini M4 16–24GB
120 GB/s
Mini-PC
—
DIY NVIDIA
RTX 5060 Ti 16GB
448 GB/s
T2Floor
Apple
MBP 14" M5 Pro 64GB
307 GB/s
Mini-PC
AMD Ryzen AI Max+ 395 · Strix Halo 128GB (Beelink GTR9 Pro / GMKtec EVO-X2)
~180 GB/s eff
DIY NVIDIA
used RTX 4090 24GB
1,008 GB/s
T3Pro
Apple
Mac Studio M3 Ultra 96GB
819 GB/s
Mini-PC
NVIDIA DGX Spark (GB10) 128GB · or ASUS Ascent GX10
273 GB/s
DIY NVIDIA
RTX 5090 32GB
1,792 GB/s
T4Research
Apple
MBP M5 Max 128GB
614 GB/s
Mini-PC
2× GB10 clustered → 256GB (DGX Spark / Ascent GX10)
273 GB/s
DIY NVIDIA
dual RTX 3090 = 48GB
936 GB/s/card
policy · CPU-only machines (no Metal, CUDA, or ROCm): you can read every lesson and dry-run the labs, but you cannot complete the curriculum locally — below the tier floor mpl blocks local execution, and the supported path is the cloud (your machine becomes the cockpit).

What your box can actually run

Each model at its featured quant, against the four tiers: whether it fits and how fast it decodes. Speed is the bandwidth truth made literal — active-param reads explain why MoE models far exceed their total size.

~/mindpool/models
# ── Gemma 4 ──
Gemma 4 E2B
2B · QAT UD-Q4_K_XL · 3 GB · multimodal · Apache-2.0 · source ↗
T1
~120–448 tok/s
T2
~180–1008 tok/s
T3
~273–1792 tok/s
T4
~273–936 tok/s
Gemma 4 E4B
4B · QAT UD-Q4_K_XL · 5 GB · multimodal · Apache-2.0 · source ↗
T1
~60–224 tok/s
T2
~90–504 tok/s
T3
~137–896 tok/s
T4
~137–468 tok/s
Gemma 4 12B
12B · QAT UD-Q4_K_XL · 7 GB · multimodal · Apache-2.0 · source ↗
T1
~20–75 tok/s
T2
~30–168 tok/s
T3
~46–299 tok/s
T4
~46–156 tok/s
Gemma 4 26B-A4B
26B (4B active) · QAT UD-Q4_K_XL · 15 GB · multimodal · Apache-2.0 · source ↗
T1
~60–224 tok/s
T2
~90–504 tok/s
T3
~137–896 tok/s
T4
~137–468 tok/s
Gemma 4 31B
31B · QAT UD-Q4_K_XL · 18 GB · multimodal · Apache-2.0 · source ↗
T1
— too big
T2
~12–65 tok/s
T3
~18–116 tok/s
T4
~18–60 tok/s
# ── Qwen3 ──
Qwen3 8B
8B · Q4_K_M · 6 GB · text · Apache-2.0 · source ↗
T1
~30–112 tok/s
T2
~45–252 tok/s
T3
~68–448 tok/s
T4
~68–234 tok/s
Qwen3 14B
14B · Q4_K_M · 9 GB · text · Apache-2.0 · source ↗
T1
~17–64 tok/s
T2
~26–144 tok/s
T3
~39–256 tok/s
T4
~39–134 tok/s
Qwen3 30B-A3B
30B (3B active) · Q4_K_M · 18 GB · text · Apache-2.0 · source ↗
T1
— too big
T2
~120–672 tok/s
T3
~182–1195 tok/s
T4
~182–624 tok/s
# ── Llama ──
Llama 8B
8B · Q4_K_M · 6 GB · text · Llama Community · source ↗
T1
~30–112 tok/s
T2
~45–252 tok/s
T3
~68–448 tok/s
T4
~68–234 tok/s
Llama 70B
70B · Q4_K_M · 42 GB · text · Llama Community · source ↗
T1
— too big
T2
~5–9 tok/s · 2/3 rigs
T3
~8–23 tok/s · 2/3 rigs
T4
~8–27 tok/s
Llama 405B
405B · Q4_K_M · 230 GB · text · Llama Community · source ↗
T1
— too big
T2
— too big
T3
— too big
T4
~1 tok/s · 1/3 rigs
# ── DeepSeek ──
DeepSeek-V2-Lite
16B (2.4B active) · Q4_K_M · 10 GB · text · DeepSeek · source ↗
T1
~100–373 tok/s
T2
~150–840 tok/s
T3
~228–1493 tok/s
T4
~228–780 tok/s
# ── OLMo 2 ──
OLMo 2 13B
13B · Q4_K_M · 8 GB · text · Apache-2.0 · source ↗
T1
~18–69 tok/s
T2
~28–155 tok/s
T3
~42–276 tok/s
T4
~42–144 tok/s
formula tok/s ≈ bandwidth ÷ (active-params × bytes/param)
Estimated decode ceiling — real-world is lower (attention, KV, overhead) and this ignores prefill. Uses active params, so MoE reads fast-and-small.
as of 2026-06 — curated, not exhaustive; models churn, re-verify each footprint at its source.

▸ Why activeparams, not total, set speed — unpacked in the quantization & MoE deep dive (coming with the curriculum).

The recommendation matrix

Tightest budget, want 70B in memory, OK with Linux + ROCm
Beelink GTR9 Pro — AMD Ryzen AI Max+ 395 (Strix Halo, 128 GB)

Cheapest 128 GB box in volume; 70B Q4 ~5–8 tok/s.

Want a turnkey Windows 128 GB box, no building
GMKtec EVO-X2 — AMD Ryzen AI Max+ 395 (Strix Halo)

Windows-native, OCuLink eGPU escape hatch.

Want AMD's own supported Halo box, not a budget mini-PC
AMD Ryzen AI Halo Developer Platform (Ryzen AI Max+ 395, 128 GB)

AMD's first-party dev box — 128 GB LPDDR5x-8000, Radeon 8060S, Linux or Win11 Pro; same Strix Halo silicon as the budget mini-PCs, vendor-supported.

Smoothest experience, value macOS + portability, models ≤34B
MacBook Pro 14" M5 Pro, 64 GB

MLX just works for LoRA/QLoRA; 307 GB/s; ships now.

Fine-tuning-heavy, want fast tokens + GPU upgrade path
DIY used RTX 4090 24 GB

1,008 GB/s; cleanest CUDA path; best value for the capability.

Want NVIDIA's stack + 70B QLoRA in one mini box
NVIDIA DGX Spark (GB10, 128 GB)

NVIDIA's own GB10 box — 1 PFLOP FP4, full CUDA, fine-tune ≤70B. ASUS Ascent GX10 is the cheaper GB10 sibling.

Max memory + fast tokens, single vendor desk box
Mac Studio M3 Ultra, 96 GB

819 GB/s — highest bandwidth shipping; 70B ~20–30 tok/s.

Certified silicon — bring your own

Commodity hardware, silicon-neutral. Systems are certified with real hardware profiling and benchmark runs on the actual box — AMD first.

AMD
First certified
Apple Silicon (MLX)
Certified
NVIDIA
Runtime validated (T4) · certification in progress
OEM boxes
On request

▸ Representative — certification in progress. Real benchmarked throughput figures are published as each certification run lands.

Own it, or rent it?

Running the same open models on hardware you own versus cloud GPUs you rent, across four axes. This lays out the trade-offs so you can match a choice to your own usage, privacy needs, and budget — it is not a recommendation either way. The right answer depends on how many hours you'll run, your data-sovereignty requirements, and the largest model you need to reach.

~/mindpool/local-vs-cloud
Local owned rigown
Cost shape
One-time ~$1k–$5k capex; ~$10–70/mo power. Near-zero marginal cost after purchase; strong mid-2026 resale.
Control & privacy
Full sovereignty — weights, data, traffic never leave the box. Offline / air-gap capable.
Setup & ops
One-time setup (hours–days); low upkeep (Apple ≈ none, NVIDIA driver/CUDA). No per-session friction; doesn't burst-scale.
Capability ceiling
70B Q4 inference; 128GB+ rigs reach 405B-capable Q4 / 70B QLoRA. Hard memory ceiling — no 70B FP16 or 405B training.
GPU specialist / marketplacerent
RunPod, Vast.ai, Lambda, CoreWeave, Crusoe, TensorDock
Cost shape
Pure opex. Consumer ~$0.1–0.7/hr · A100 ~$0.7–2.5/hr · H100 ~$1–4.25/hr · B200 ~$2–6/hr. Spot −40–80%; egress often free.
Control & privacy
Your stack, but data/weights sit on rented infra during the job. Mostly US jurisdiction; marketplace hosts add unknown location. No offline.
Setup & ops
First token ~1–3 min managed, ~10+ min marketplace. No hardware upkeep; watch spot interruption, idle billing, ephemeral storage.
Capability ceiling
Budget-bound, not memory-bound: single A100/H100 for 70B; 4–8× clusters for 405B / full 70B fine-tune. Far past any single rig.
Hyperscaler GPUrent
AWS, GCP, Azure
Cost shape
Pure opex. ~$0.7–1.7/hr small GPU to ~$55–98/hr per 8×H100 node. H100 on-demand ~$7–11/GPU-hr; spot ~$2.25–3.7. Egress + surcharges add ~20–40%.
Control & privacy
Region-pinned with VPC isolation, CMEK/BYOK, residency controls — but US CLOUD Act always applies. No air-gap.
Setup & ops
Highest setup friction: GPU quota requests (1–7 days), IAM/VPC/driver setup. Elastic once running; auto-stop discipline needed.
Capability ceiling
Effectively unlimited within quota/spend: 8×H100 node (640GB) runs 405B FP8; multi-node trains beyond 405B.
Managed / serverless GPUrent
Modal, RunPod Serverless, Replicate, HF Endpoints, Baseten
Cost shape
Pure opex. Per-second serverless scales to zero (cheapest for bursty); dedicated managed bills continuously, +30–75% over bare metal. ~$0.5–9.25/hr.
Control & privacy
Processed on third-party managed infra; no offline. Dedicated endpoints cut multi-tenant bleed; compliance varies. Must upload weights.
Setup & ops
Lowest friction — first token <~5 min, deploy by function/container. Cold starts <200ms–~60s. Minimal ops; limited low-level tuning.
Capability ceiling
Single-GPU ceiling matches raw rental (H100/H200 → 70B; B200 → larger). 405B / full fine-tune needs cluster configs near raw-cloud pricing.
~/mindpool/cloud --by-gpu-class
Consumer GPU (RTX 4090/5090-class, 24–32GB) ~$0.10–0.70/hr
Spot/marketplace as low as ~$0.10–0.34/hr; on-demand toward the top. 8B–13B comfortably, 70B only at aggressive quant. Specialist/marketplace clouds only.
Datacenter single GPU (A100 80GB / L40S) ~$0.67–2.50/hr
Marketplace ~$0.67–1.50/hr; managed on-demand ~$1.40–2.50/hr; hyperscaler ~$3.50–4.10/hr. Handles 70B Q4 inference + 70B QLoRA on one card.
Datacenter H100/H200 (80–141GB) ~$1.00–7.00/hr per GPU
Marketplace spot ~$1.00–1.30/hr, on-demand ~$2.50–4.25/hr; hyperscaler ~$6.9–11/GPU-hr (GCP spot ~$2.25–3.7). Volatile.
Blackwell-tier (B200/B300, 192–288GB) ~$2.10–12/hr per GPU
Spot from ~$2.10/hr; specialist on-demand ~$5.9–10.5/hr; hyperscaler/managed often sales-quoted. 405B Q4 in a few cards.

Break-even

Break-even is utilization-driven, not binary. Light use (under ~30 hrs/month): cloud stays cheaper for the whole 3-year ownership cycle. Moderate use (~6 hrs/day): crossover against an A100-equivalent rental lands around ~10–11 months; against the cheapest consumer-GPU spot it can stretch to multiple years. Always-on serving (~720 hrs/month): local typically wins by month ~2–4. Rough anchor: a ~$3,000 rig ÷ ~$1.50/hr A100 ≈ 2,000 hrs (~11 months at 6 hrs/day); ÷ ~$0.34/hr consumer spot ≈ 8,800 hrs (~4 years).

  • ▸Mid-2026 GPU/RAM pricing is volatile (H100 swung from ~$8+/hr peak to ~$1.70 trough to ~$2.35 by early 2026) — treat every figure as a coarse range and verify live before budgeting.
  • ▸Provider choice matters as much as own-vs-rent: hyperscaler H100 (~$7–11/GPU-hr) runs 3–6× community/specialist rates (~$1–4/hr).
  • ▸Hidden cloud costs add ~20–40% on hyperscalers: egress (~$0.09–0.12/GB), checkpoint storage, managed-platform surcharges, spot-interruption waste. Community clouds mostly waive egress.
  • ▸Spot pricing is the biggest cost lever (−40–80%) but requires checkpointing; interruption warnings range ~2 min (AWS) to ~15 sec (Vast.ai).
  • ▸US CLOUD Act applies to essentially all major US providers regardless of region; for strict sovereignty / air-gap / regulated data, local is the only fully offline option.
  • ▸Local capability is memory-bound: the $1k–$5k tier tops out near 405B-capable Q4 inference / 70B QLoRA; full 70B FP16 fine-tunes and 405B training need rented multi-GPU clusters — mpl burst is the relief valve, on your own cluster or managed cloud.

The bandwidth truth

Generating each token reads the active weights out of memory, so your tok/s ceiling ≈ bandwidth ÷ active-weight-bytes. Capacity decides whether a model fits; bandwidth decides how fast it runs once it fits — and mpl burst is the fallback for the rare 70B full fine-tune — run it on hardware you own or rent managed cloud, your call; see the local-vs-cloud comparison below.

Mid-2026 figures; the 2026 RAM/GPU shortage adds volatility — re-verify before buying or renting.

Hardware ready.

See which open models run on your hardware — and how each performs.