Local Kimi K3 cluster — UK cost brief
Bottom line. Kimi K3 is a real open-weight 2.8T MoE model (~1.56 TB MXFP4). The clean minimum to load it is 8× NVIDIA B300. For 20–30 people with real concurrency headroom, plan on 16× B300 (two 8-GPU nodes).
UK list CapEx for that path is about £0.9–1.2m ex VAT, plus Manchester colo OpEx of roughly £100–200k / year. Payback vs API only works if utilisation stays high.
1. What you are buying into
- Moonshot AI Kimi K3 — open weights 27 Jul 2026.
- Mixture-of-Experts: 2.8T total / 104B active per token.
- Native quant: MXFP4 weights / MXFP8 activations (QAT).
- HF download ≈ 1.56 TB (96 shards).
- Context window up to 1M tokens (hybrid KDA + Gated MLA).
- Serve with vLLM / SGLang / TokenSpeed.
- Licence is custom (not MIT). Internal use is generally fine; MaaS / large commercial use can trigger Moonshot commercial terms. Read the HF licence before productising.
Memory reality: weights alone need ~1.5–1.7 TB on GPU. Concurrent users then eat remaining HBM for KV / batch. That is why 8× B300 (2.3 TB) is “can load”, and 16× B300 (4.6 TB) is the safer 20–30-user plan.
2. Hardware options
| Config | HBM | Fit for K3 | Verdict for 20–30 users |
|---|---|---|---|
| 8× B300 | 2.3 TB | Yes — minimum single-node | Viable lab / small team. Concurrency risk if long context. |
| 16× B300 ★ | 4.6 TB | Comfortable | Recommended. Two nodes + fabric. Matches the “16× B300” instinct. |
| 16× B200 | ~3 TB | Yes (documented path) | Works; less memory / FP4 headroom than B300. |
| 16× H200 | 2.26 TB | Bare fit | Cheaper CapEx; slower / tighter. UK list exists. |
| 8× MI350X / MI355X | 2.3 TB | Supported in ROCm recipes | Lower list price; higher software-day-0 risk. |
| GB200 / GB300 NVL72 | Rack-scale | Excellent | Overkill CapEx/power for this user count. |
B300 itself: 288 GB HBM3e, ~1,400 W TGP, liquid-friendly density. One DGX/HGX node = 8 GPUs ≈ 14 kW IT. Two nodes ≈ 28–32 kW before PUE.
3. CapEx — UK GBP
QUOTE = Broadberry public “configure from” as of 22 Aug 2026. EST = market estimate. Prices ex VAT unless noted. UK VAT 20% (usually reclaimable if VAT-registered).
UK list anchors (Broadberry)
| SKU | Configure from | |
|---|---|---|
| CyberServe 8× HGX H200 | £269,060 | QUOTE |
| DGX H200 8× | £346,157 | QUOTE |
| CyberServe 8× HGX B300 | £420,895 | QUOTE |
| DGX B200 8× | £469,574 | QUOTE |
| DGX B300 8× | £504,381 | QUOTE |
Scan UK (Bolton / Greater Manchester) lists DGX B300 as enquire-only — good channel for Manchester delivery and MSP wrap. Sales: 01204 474210.
Recommended build: 16× B300
| Line | Low–High | |
|---|---|---|
| 2× HGX B300 CyberServe | £841,790 | QUOTE×2 |
| or 2× DGX B300 | £1,008,762 | QUOTE×2 |
| InfiniBand / Ethernet fabric + optics | £40,000–£120,000 | EST / partial quote |
| Rack, PDUs, rails | £8,000–£25,000 | EST |
| Shared NVMe / NAS (≥4–8 TB) | £8,000–£30,000 | EST |
| Spares (no spare B300 GPU) | £5,000–£40,000 | EST |
| HGX path subtotal | ~£0.90–1.06m | ex VAT |
| DGX path subtotal | ~£1.07–1.22m | ex VAT |
| + VAT if not reclaimable | ×1.20 |
Cheaper / alternate CapEx
- 8× B300 only: £421–504k ex VAT — min viable, concurrency risk.
- 16× H200: £538–692k ex VAT (2× Broadberry H200 quotes + fabric).
- 16× B200: roughly £0.75–1.0m ex VAT EST.
- NVL72: multi-million + liquid facility — wrong shape for 20–30 users.
Lead times. H200 often 8–12 weeks; B200 12–20 weeks; B300 shipping but UK ETA is quote-gated. Allocation is the critical path, not software. Prefer UK reseller (Broadberry / Scan) over US street + import.
4. OpEx — Manchester
- IT load for 16× B300: ~28–32 kW; with PUE 1.3–1.5 → ~36–48 kW site.
- UK DC power typically 18–28 p/kWh; Manchester colo often 15–30% cheaper than London/M4.
- High-density / liquid colo: about £300–420 / kW / month (bundled).
- At 32 kW × £220–320/kW/mo → ~£7–10k / month (~£85–123k / year) for space+power EST.
- Add internet, support renewal (~5–12% of CapEx/yr), remote hands → total OpEx ballpark £100–200k / year.
Manchester operators named in market guides: Equinix MA1/MA2/MA3, Pulsant, TeleData / Digital Realty corridor. Density ceiling (~30–50 kW retrofit) is the constraint vs Slough.
5. Soft costs to “live”
- UK install / burn-in: £2–15k.
- Engineering to production: 40–120 person-days (vLLM/SGLang K3, fabric, auth, monitoring, load test).
- At £600–1,200/day → £24–144k labour EST.
- Software: OSS stack free; optional NVIDIA AI Enterprise with DGX.
6. Itinerary — order to 20–30 users
| Phase | Weeks | Gate |
|---|---|---|
| 0. Scope lock (SLO, max ctx, internal vs product) | 0–1 | Write the concurrency target |
| 1. Facility survey (colo ≥35 kW, cooling, floor load) | 1–3 | Manchester cage available |
| 2. RFQ Broadberry + Scan (+ optional EU OEM) | 2–5 | Written quotes + lead times |
| 3. Decision gate A | ~5 | 8× vs 16×; HGX vs DGX; colo signed |
| 4. Order / deposit / NVIDIA allocation | 5–6 | PO confirmed |
| 5. Parallel soft prep (licence, images, IdP, observability) | 5–12 | Stack ready offline |
| 6. Delivery, rack, fabric burn-in | 12–24 | Nodes green |
| 7. Model load (TP8 smoke → multi-node) | +1–2 | K3 serving |
| 8. Decision gate B — load test 30 concurrent | +1–2 | Pass SLO or cut scope / add node |
| 9. Pilot 10 users | +1–2 | Security + runbooks |
| 10. Production 20–30 | +1 | Quotas + on-call |
7. Investor pitch — why buy vs rent
- Kimi K3 API: about $3 / $15 per 1M tokens in/out (cache ~$0.30).
- Frontier API peers sit in a similar band; cloud GPU rental for B300 is roughly $7–18 / GPU-hr on-demand.
- Illustrative: 50–150M tokens/day for 20–30 knowledge workers → roughly £7–40k / month API spend EST.
- Self-host CapEx ~£1.0m + OpEx ~£10k/mo → cash payback often 2–8 years, highly sensitive to utilisation.
- Buy if: UK data residency, predictable high utilisation, control, and API rate-limit risk matter.
- Rent if: utilisation is bursty, you want zero CapEx, or you are still proving the product.
Risks to say out loud
- Blackwell allocation / lead time.
- Obsolescence — 18–36 month GPU cycle; residual value uncertain.
- Power / liquid cooling in Manchester cages.
- Custom model licence if you productise or hit MaaS revenue gates.
- Under-utilisation — 20–30 users may not saturate 16× B300; starting at 8× with a colo expansion option is the conservative capital move.
- K3’s hybrid attention is new — ops complexity on day-0 stacks.
8. Recommended ask of the co-investor
- Fund a dual RFQ now: Broadberry (020 8997 6000) and Scan (01204 474210) for 2× B300 + fabric, with written lead times and VAT treatment.
- Parallel Manchester colo survey for ≥35 kW liquid-ready space.
- Decision gate in ~5 weeks: 8× pilot vs 16× production, based on quotes and a written concurrency SLO.
- Do not buy NVL72 for this use case. Do not skip the load-test gate before inviting all 30 users.