Four builds, priced in the open.
Nobody in this market publishes what a private AI machine costs. So here it is — the tiers, what they run, what they draw, and the case where buying one is the wrong decision.
| Unit | Users | Comfortably runs | Form factor | Draw | From (CAD) |
|---|---|---|---|---|---|
Desk A quiet box for a small firm | 2–4 | 14B dense at Q8, 20–30B MoE, 32K context | Sound-dampened mid-tower | ~300 W | $6,300 |
Workgroup A team's daily driver | 10–30 | 70B dense at Q4/Q5, 64–128K context | Full tower, sound-dampened | ~1,070 W | $17,900 |
Department Serious throughput and long context | 25–40 | 70B at FP8, or 120B-class MoE, 128K+ | Tower or 4U rack | ~1,400 W · dedicated 20 A circuit | $44,900 |
Enterprise Survives a node failure | 60–100 | 235B MoE, or 70B FP8 with a large KV cache | Two rack nodes, load balanced | ~2,100 W · 208/240 V | $123,800 |
Indicative build cost in CAD, hardware only, before installation and GST, priced against Canadian retail on 29 July 2026. GPU and memory pricing moved sharply through 2026 — quotes are re-priced at time of order and held firm for seven days.
What fits in how much memory.
Three numbers pick the tier: the largest model you need, how many people hit it at once, and whether they're chatting or running agents.
| Memory | Largest sensible model | Precision | Context | Concurrent chat |
|---|---|---|---|---|
| 24 GB | 14B dense, 20–30B MoE | Q4 / MXFP4 | 32K | 2–4 |
| 32 GB | 32B dense | Q4_K_M | 32–64K | 5–10 |
| 48 GB | 32B at Q8, or 70B at Q4 | Q8 / Q4 | 64K | 8–15 |
| 64 GB | 70B dense | Q4 / Q5 | 64–128K | 15–25 |
| 96 GB | 70B at FP8, 120B-class MoE | FP8 / MXFP4 | 128K+ | 25–40 |
| 128 GB | 235B MoE, or 70B FP8 + large KV | FP8 / MXFP4 | 256K | 60–100 |
Concurrency assumes bursty chat, not agents in a loop — agentic workloads consume 10–50× the tokens. We size against what the assessment measured, not headcount. Throughput is benchmarked on your build before any number goes into a contract.
When buying one is the wrong call.
A 30-person firm doing ordinary document Q&A generates about 63 million tokens a month. On a commercial API that is roughly $300. Against the Department tier, the payback period runs to decades.
We would rather lose the hardware sale than sell you a machine that sits idle. If your usage stays at chat volumes and your data can lawfully leave, an API subscription is the better buy — and the assessment will say so in writing.
The part that sinks self-built deployments.
The machine is the easy bit. Where these projects fail is the room it goes in.
Get a specification with real prices.
The assessment produces an itemised bill of materials at supplier cost, with current quoted prices and honest lead times — plus the cloud comparison, so you can see both.