Pick the rig that runs the work
Bandwidth, not capacity, sets your tokens/sec. Match the machine to the model size you must load, then accept the speed its bandwidth allows.
▸ T2 is the enrollment floor — it runs the full curriculum end to end.
▸ Local hardware runs ~$1,000–$5,000 one-time — a single purchase, not a subscription.
▸ Local vs cloud — See the local-vs-cloud comparison ↓
Install mpl and run one command — it prints a curriculum verdict for your machine (full / partial / not completable) so you know before you pay.
$ mpl doctor --profileTwo ways in. Not ready to buy hardware? Rent first — mpl cloud (beta) spins up a GPU box you control, billed by the second; stop it when you're not learning, own the metal when you're ready. Or own it from day one — pick a rig below. Either path runs the full curriculum on the same sovereign mpl.
Stop/start works on GCP/Azure; RunPod is terminate-and-restore.
mpl cloud docs →mpl cloud is beta. It rents and controls GPU compute on your behalf, and like any beta that depends on third-party cloud providers and their APIs, mpl cloud can fail to launch, stop, start, or resume a cluster if something upstream changes or breaks — and work on a stopped cluster may be at risk until you recover it. We recommend mpl cloud for learners who accept this, back up their checkpoints to a network volume or object store, and don't mind a rough edge. Want full control with no beta? Bring any SSH-able GPU box, run mpl setup after enrolling, and connect it with mpl fleet add — any provider, fully sovereign.
Pick a rig below — a single purchase, not a subscription. Full control, no beta.
pick a tier ↓| tier | Apple | Mini-PC | DIY NVIDIA |
|---|---|---|---|
| T1 Entry | Mac mini M4 16–24GB 120 GB/s | — | RTX 5060 Ti 16GB 448 GB/s |
| T2 Floor | MBP 14" M5 Pro 64GB 307 GB/s | AMD Ryzen AI Max+ 395 · Strix Halo 128GB (Beelink GTR9 Pro / GMKtec EVO-X2) ~180 GB/s eff | used RTX 4090 24GB 1,008 GB/s |
| T3 Pro | Mac Studio M3 Ultra 96GB 819 GB/s | NVIDIA DGX Spark (GB10) 128GB · or ASUS Ascent GX10 273 GB/s | RTX 5090 32GB 1,792 GB/s |
| T4 Research | MBP M5 Max 128GB 614 GB/s | 2× GB10 clustered → 256GB (DGX Spark / Ascent GX10) 273 GB/s | dual RTX 3090 = 48GB 936 GB/s/card |
- Apple
- Mac mini M4 16–24GB120 GB/s
- Mini-PC
- —
- DIY NVIDIA
- RTX 5060 Ti 16GB448 GB/s
- Apple
- MBP 14" M5 Pro 64GB307 GB/s
- Mini-PC
- AMD Ryzen AI Max+ 395 · Strix Halo 128GB (Beelink GTR9 Pro / GMKtec EVO-X2)~180 GB/s eff
- DIY NVIDIA
- used RTX 4090 24GB1,008 GB/s
- Apple
- Mac Studio M3 Ultra 96GB819 GB/s
- Mini-PC
- NVIDIA DGX Spark (GB10) 128GB · or ASUS Ascent GX10273 GB/s
- DIY NVIDIA
- RTX 5090 32GB1,792 GB/s
- Apple
- MBP M5 Max 128GB614 GB/s
- Mini-PC
- 2× GB10 clustered → 256GB (DGX Spark / Ascent GX10)273 GB/s
- DIY NVIDIA
- dual RTX 3090 = 48GB936 GB/s/card
What your box can actually run
Each model at its featured quant, against the four tiers: whether it fits and how fast it decodes. Speed is the bandwidth truth made literal — active-param reads explain why MoE models far exceed their total size.
| model | T1 | T2 | T3 | T4 |
|---|---|---|---|---|
| # ── Gemma 4 ── | ||||
Gemma 4 E2B 2B · QAT UD-Q4_K_XL · 3 GB · multimodal · Apache-2.0 · source ↗ | ~120–448 tok/s | ~180–1008 tok/s | ~273–1792 tok/s | ~273–936 tok/s |
Gemma 4 E4B 4B · QAT UD-Q4_K_XL · 5 GB · multimodal · Apache-2.0 · source ↗ | ~60–224 tok/s | ~90–504 tok/s | ~137–896 tok/s | ~137–468 tok/s |
Gemma 4 12B 12B · QAT UD-Q4_K_XL · 7 GB · multimodal · Apache-2.0 · source ↗ | ~20–75 tok/s | ~30–168 tok/s | ~46–299 tok/s | ~46–156 tok/s |
Gemma 4 26B-A4B 26B (4B active) · QAT UD-Q4_K_XL · 15 GB · multimodal · Apache-2.0 · source ↗ | ~60–224 tok/s | ~90–504 tok/s | ~137–896 tok/s | ~137–468 tok/s |
Gemma 4 31B 31B · QAT UD-Q4_K_XL · 18 GB · multimodal · Apache-2.0 · source ↗ | — too big | ~12–65 tok/s | ~18–116 tok/s | ~18–60 tok/s |
| # ── Qwen3 ── | ||||
Qwen3 8B 8B · Q4_K_M · 6 GB · text · Apache-2.0 · source ↗ | ~30–112 tok/s | ~45–252 tok/s | ~68–448 tok/s | ~68–234 tok/s |
Qwen3 14B 14B · Q4_K_M · 9 GB · text · Apache-2.0 · source ↗ | ~17–64 tok/s | ~26–144 tok/s | ~39–256 tok/s | ~39–134 tok/s |
Qwen3 30B-A3B 30B (3B active) · Q4_K_M · 18 GB · text · Apache-2.0 · source ↗ | — too big | ~120–672 tok/s | ~182–1195 tok/s | ~182–624 tok/s |
| # ── Llama ── | ||||
Llama 8B 8B · Q4_K_M · 6 GB · text · Llama Community · source ↗ | ~30–112 tok/s | ~45–252 tok/s | ~68–448 tok/s | ~68–234 tok/s |
Llama 70B 70B · Q4_K_M · 42 GB · text · Llama Community · source ↗ | — too big | ~5–9 tok/s · 2/3 rigs | ~8–23 tok/s · 2/3 rigs | ~8–27 tok/s |
Llama 405B 405B · Q4_K_M · 230 GB · text · Llama Community · source ↗ | — too big | — too big | — too big | ~1 tok/s · 1/3 rigs |
| # ── DeepSeek ── | ||||
DeepSeek-V2-Lite 16B (2.4B active) · Q4_K_M · 10 GB · text · DeepSeek · source ↗ | ~100–373 tok/s | ~150–840 tok/s | ~228–1493 tok/s | ~228–780 tok/s |
| # ── OLMo 2 ── | ||||
OLMo 2 13B 13B · Q4_K_M · 8 GB · text · Apache-2.0 · source ↗ | ~18–69 tok/s | ~28–155 tok/s | ~42–276 tok/s | ~42–144 tok/s |
- T1
- ~120–448 tok/s
- T2
- ~180–1008 tok/s
- T3
- ~273–1792 tok/s
- T4
- ~273–936 tok/s
- T1
- ~60–224 tok/s
- T2
- ~90–504 tok/s
- T3
- ~137–896 tok/s
- T4
- ~137–468 tok/s
- T1
- ~20–75 tok/s
- T2
- ~30–168 tok/s
- T3
- ~46–299 tok/s
- T4
- ~46–156 tok/s
- T1
- ~60–224 tok/s
- T2
- ~90–504 tok/s
- T3
- ~137–896 tok/s
- T4
- ~137–468 tok/s
- T1
- — too big
- T2
- ~12–65 tok/s
- T3
- ~18–116 tok/s
- T4
- ~18–60 tok/s
- T1
- ~30–112 tok/s
- T2
- ~45–252 tok/s
- T3
- ~68–448 tok/s
- T4
- ~68–234 tok/s
- T1
- ~17–64 tok/s
- T2
- ~26–144 tok/s
- T3
- ~39–256 tok/s
- T4
- ~39–134 tok/s
- T1
- — too big
- T2
- ~120–672 tok/s
- T3
- ~182–1195 tok/s
- T4
- ~182–624 tok/s
- T1
- ~30–112 tok/s
- T2
- ~45–252 tok/s
- T3
- ~68–448 tok/s
- T4
- ~68–234 tok/s
- T1
- — too big
- T2
- ~5–9 tok/s · 2/3 rigs
- T3
- ~8–23 tok/s · 2/3 rigs
- T4
- ~8–27 tok/s
- T1
- — too big
- T2
- — too big
- T3
- — too big
- T4
- ~1 tok/s · 1/3 rigs
- T1
- ~100–373 tok/s
- T2
- ~150–840 tok/s
- T3
- ~228–1493 tok/s
- T4
- ~228–780 tok/s
- T1
- ~18–69 tok/s
- T2
- ~28–155 tok/s
- T3
- ~42–276 tok/s
- T4
- ~42–144 tok/s
▸ Why activeparams, not total, set speed — unpacked in the quantization & MoE deep dive (coming with the curriculum).
The recommendation matrix
Cheapest 128 GB box in volume; 70B Q4 ~5–8 tok/s.
Windows-native, OCuLink eGPU escape hatch.
AMD's first-party dev box — 128 GB LPDDR5x-8000, Radeon 8060S, Linux or Win11 Pro; same Strix Halo silicon as the budget mini-PCs, vendor-supported.
MLX just works for LoRA/QLoRA; 307 GB/s; ships now.
1,008 GB/s; cleanest CUDA path; best value for the capability.
NVIDIA's own GB10 box — 1 PFLOP FP4, full CUDA, fine-tune ≤70B. ASUS Ascent GX10 is the cheaper GB10 sibling.
819 GB/s — highest bandwidth shipping; 70B ~20–30 tok/s.
Certified silicon — bring your own
Commodity hardware, silicon-neutral. Systems are certified with real hardware profiling and benchmark runs on the actual box — AMD first.
▸ Representative — certification in progress. Real benchmarked throughput figures are published as each certification run lands.
Own it, or rent it?
Running the same open models on hardware you own versus cloud GPUs you rent, across four axes. This lays out the trade-offs so you can match a choice to your own usage, privacy needs, and budget — it is not a recommendation either way. The right answer depends on how many hours you'll run, your data-sovereignty requirements, and the largest model you need to reach.
| Cost shape | Control & privacy | Setup & ops | Capability ceiling | |
|---|---|---|---|---|
Local owned rig own | One-time ~$1k–$5k capex; ~$10–70/mo power. Near-zero marginal cost after purchase; strong mid-2026 resale. | Full sovereignty — weights, data, traffic never leave the box. Offline / air-gap capable. | One-time setup (hours–days); low upkeep (Apple ≈ none, NVIDIA driver/CUDA). No per-session friction; doesn't burst-scale. | 70B Q4 inference; 128GB+ rigs reach 405B-capable Q4 / 70B QLoRA. Hard memory ceiling — no 70B FP16 or 405B training. |
GPU specialist / marketplace RunPod, Vast.ai, Lambda, CoreWeave, Crusoe, TensorDock rent | Pure opex. Consumer ~$0.1–0.7/hr · A100 ~$0.7–2.5/hr · H100 ~$1–4.25/hr · B200 ~$2–6/hr. Spot −40–80%; egress often free. | Your stack, but data/weights sit on rented infra during the job. Mostly US jurisdiction; marketplace hosts add unknown location. No offline. | First token ~1–3 min managed, ~10+ min marketplace. No hardware upkeep; watch spot interruption, idle billing, ephemeral storage. | Budget-bound, not memory-bound: single A100/H100 for 70B; 4–8× clusters for 405B / full 70B fine-tune. Far past any single rig. |
Hyperscaler GPU AWS, GCP, Azure rent | Pure opex. ~$0.7–1.7/hr small GPU to ~$55–98/hr per 8×H100 node. H100 on-demand ~$7–11/GPU-hr; spot ~$2.25–3.7. Egress + surcharges add ~20–40%. | Region-pinned with VPC isolation, CMEK/BYOK, residency controls — but US CLOUD Act always applies. No air-gap. | Highest setup friction: GPU quota requests (1–7 days), IAM/VPC/driver setup. Elastic once running; auto-stop discipline needed. | Effectively unlimited within quota/spend: 8×H100 node (640GB) runs 405B FP8; multi-node trains beyond 405B. |
Managed / serverless GPU Modal, RunPod Serverless, Replicate, HF Endpoints, Baseten rent | Pure opex. Per-second serverless scales to zero (cheapest for bursty); dedicated managed bills continuously, +30–75% over bare metal. ~$0.5–9.25/hr. | Processed on third-party managed infra; no offline. Dedicated endpoints cut multi-tenant bleed; compliance varies. Must upload weights. | Lowest friction — first token <~5 min, deploy by function/container. Cold starts <200ms–~60s. Minimal ops; limited low-level tuning. | Single-GPU ceiling matches raw rental (H100/H200 → 70B; B200 → larger). 405B / full fine-tune needs cluster configs near raw-cloud pricing. |
- Cost shape
- One-time ~$1k–$5k capex; ~$10–70/mo power. Near-zero marginal cost after purchase; strong mid-2026 resale.
- Control & privacy
- Full sovereignty — weights, data, traffic never leave the box. Offline / air-gap capable.
- Setup & ops
- One-time setup (hours–days); low upkeep (Apple ≈ none, NVIDIA driver/CUDA). No per-session friction; doesn't burst-scale.
- Capability ceiling
- 70B Q4 inference; 128GB+ rigs reach 405B-capable Q4 / 70B QLoRA. Hard memory ceiling — no 70B FP16 or 405B training.
- Cost shape
- Pure opex. Consumer ~$0.1–0.7/hr · A100 ~$0.7–2.5/hr · H100 ~$1–4.25/hr · B200 ~$2–6/hr. Spot −40–80%; egress often free.
- Control & privacy
- Your stack, but data/weights sit on rented infra during the job. Mostly US jurisdiction; marketplace hosts add unknown location. No offline.
- Setup & ops
- First token ~1–3 min managed, ~10+ min marketplace. No hardware upkeep; watch spot interruption, idle billing, ephemeral storage.
- Capability ceiling
- Budget-bound, not memory-bound: single A100/H100 for 70B; 4–8× clusters for 405B / full 70B fine-tune. Far past any single rig.
- Cost shape
- Pure opex. ~$0.7–1.7/hr small GPU to ~$55–98/hr per 8×H100 node. H100 on-demand ~$7–11/GPU-hr; spot ~$2.25–3.7. Egress + surcharges add ~20–40%.
- Control & privacy
- Region-pinned with VPC isolation, CMEK/BYOK, residency controls — but US CLOUD Act always applies. No air-gap.
- Setup & ops
- Highest setup friction: GPU quota requests (1–7 days), IAM/VPC/driver setup. Elastic once running; auto-stop discipline needed.
- Capability ceiling
- Effectively unlimited within quota/spend: 8×H100 node (640GB) runs 405B FP8; multi-node trains beyond 405B.
- Cost shape
- Pure opex. Per-second serverless scales to zero (cheapest for bursty); dedicated managed bills continuously, +30–75% over bare metal. ~$0.5–9.25/hr.
- Control & privacy
- Processed on third-party managed infra; no offline. Dedicated endpoints cut multi-tenant bleed; compliance varies. Must upload weights.
- Setup & ops
- Lowest friction — first token <~5 min, deploy by function/container. Cold starts <200ms–~60s. Minimal ops; limited low-level tuning.
- Capability ceiling
- Single-GPU ceiling matches raw rental (H100/H200 → 70B; B200 → larger). 405B / full fine-tune needs cluster configs near raw-cloud pricing.
Break-even
Break-even is utilization-driven, not binary. Light use (under ~30 hrs/month): cloud stays cheaper for the whole 3-year ownership cycle. Moderate use (~6 hrs/day): crossover against an A100-equivalent rental lands around ~10–11 months; against the cheapest consumer-GPU spot it can stretch to multiple years. Always-on serving (~720 hrs/month): local typically wins by month ~2–4. Rough anchor: a ~$3,000 rig ÷ ~$1.50/hr A100 ≈ 2,000 hrs (~11 months at 6 hrs/day); ÷ ~$0.34/hr consumer spot ≈ 8,800 hrs (~4 years).
- ▸Mid-2026 GPU/RAM pricing is volatile (H100 swung from ~$8+/hr peak to ~$1.70 trough to ~$2.35 by early 2026) — treat every figure as a coarse range and verify live before budgeting.
- ▸Provider choice matters as much as own-vs-rent: hyperscaler H100 (~$7–11/GPU-hr) runs 3–6× community/specialist rates (~$1–4/hr).
- ▸Hidden cloud costs add ~20–40% on hyperscalers: egress (~$0.09–0.12/GB), checkpoint storage, managed-platform surcharges, spot-interruption waste. Community clouds mostly waive egress.
- ▸Spot pricing is the biggest cost lever (−40–80%) but requires checkpointing; interruption warnings range ~2 min (AWS) to ~15 sec (Vast.ai).
- ▸US CLOUD Act applies to essentially all major US providers regardless of region; for strict sovereignty / air-gap / regulated data, local is the only fully offline option.
- ▸Local capability is memory-bound: the $1k–$5k tier tops out near 405B-capable Q4 inference / 70B QLoRA; full 70B FP16 fine-tunes and 405B training need rented multi-GPU clusters — mpl burst is the relief valve, on your own cluster or managed cloud.
The bandwidth truth
Generating each token reads the active weights out of memory, so your tok/s ceiling ≈ bandwidth ÷ active-weight-bytes. Capacity decides whether a model fits; bandwidth decides how fast it runs once it fits — and mpl burst is the fallback for the rare 70B full fine-tune — run it on hardware you own or rent managed cloud, your call; see the local-vs-cloud comparison below.
Mid-2026 figures; the 2026 RAM/GPU shortage adds volatility — re-verify before buying or renting.
Hardware ready.
See which open models run on your hardware — and how each performs.