Private inference for one industry: what £50m buys, what to charge, and the path to 50,000 seats
Open-weight LLMs served on our own GB300 NVL72 racks, single-tenant, no training on client data. The founder has not chosen a model: Kimi K3, GLM-5.3 and a Hybrid of both are costed identically as Routes 1 / 2 / 3, and the route choice turns out to be the capacity choice. This brief ties the financial model to the diligence work and answers the questions the founder and his investors will ask first.
Revised 27 Aug after three commissioned adversarial reviews (technical · financial · buyer-side); every accepted fix is live in the engine and workbook — decisions in work/14-fix-log.md. How to read this: every number traces to 07-model-spec.json / 07-base-results.json; where the Excel is referenced, the sheet is named. Nothing on this page is invented outside those files and the sources in the appendix.
The answer
Viable on one route, not three. On the base case, one GB300 NVL72 rack serves 2,194 seats on GLM-5.3, 1,095 on the 35% Hybrid, 638 on Kimi K3 — the model choice is a capacity choice. £50m buys 12 racks of iron and nothing else; the fundable build is 6 racks plus contract-triggered leases. Every number below is the post-adversarial engine (three hostile reviews, all accepted fixes live in the model — §09 lists the verdicts).
Build 6 racks now (£24.6m / $33.4m all-in) on GLM-5.3, Route 2 pure — 5 live racks serve ~11,000 seats — and gate every further tranche of 3 leased racks on the anchor’s signed take-or-pay.
On that plan the model reaches ARR $7.0m / $18.3m / $36.0m / $48.8m (£5.2m / £13.4m / £26.5m / £35.9m) at M12 / 18 / 24 / 36, EBITDA break-even at M17, and cash never below £7.6m — the peak funding gap is £0: the plan never needs money beyond the £50m. If all four lease tranches fire, the aggregate obligation is $39.1m principal (~$1.3m / £1.0m per month at full service, 36-month terms) secured on the anchor contract, plus $3.9m of lessor deposits — all inside the cash path shown in §07. The recommendation is a view, not a verdict on the models: §04 shows exactly what flips the route — a measured K3+DSpark curve at ≥12 streams/replica at 100 tok/s, or an anchor paying ≥$300/seat for K3 with Moonshot terms signed.
The question, and the honest constraint
The founder’s question is “can £50m build a private LLM service for a 50,000-seat anchor?” The honest constraint: 12 racks is not 50,000 seats on any route — seats convert to racks through peak concurrency, and the per-replica stream count each model holds at the SLA.
Every seat number on this page comes from one chain. Licensed seats are not concurrent users: on a Copilot-class service ~30% of seats touch it daily, under half of those overlap the peak window, and a fifth of those are actually streaming tokens at any instant — roughly 3% in-flight today. But the demand file’s own agent-share ramp reaches ~6.5% by M36, so the model sizes at 6% (the M36 agentic point) and protects it with contractual per-seat stream caps: 2 concurrent streams per volume seat, 5 per premium. Est source brief §1 · 01-demand Little’s law (1.6% hour-avg M12 → ~6.5% M36) · 12-adversarial A1
Tax factor 0.39 = (1 − 25% p99 headroom, incl. rolling-upgrade drain) × 0.65 long-context factor ÷ (1 + 25% prefill tax at ≥85% prefix-cache hit). Seats per live rack = usable streams ÷ 6% = 2,194 (GLM-5.3) / 1,095 (Hybrid 35%) / 638 (K3). These cells were re-based as one bundle after adversarial review — the old +100% prefill tax was ~4× too high and was quietly compensating for optimism elsewhere; never relax one cell without the others. 07-model-spec.md §0, §2.1 · 12-adversarial A7 · Excel: Capacity sheet
Racks, fabric, storage, colo, power, staff, software and runway all come out of the same £50m. Spending £46.7m on 12 racks (Option A) leaves the company insolvent at delivery — cash goes negative at M4 on every route and the plan needs ~£57m. Applying the reserve rule (cash ≥ £4m after 24 months of zero revenue), £50m safely carries 6 racks; with base-case revenue it carries 9. 07-base-results.json runs["base|route2|A"] · racks_buyable
The 26 Aug diligence brief sized 12 racks at ~8,000 seats — that was K3 at 4 streams/replica with no prefill tax. The full chain (02-capacity), anchored to GLM’s published conc=16 point and with every tax applied, gives 7,020 seats on K3 and 24,131 on GLM-5.3 for the same 12 racks. Same iron, ~3.4× the service — that is why §04 exists.
Capacity: seats vs racks, route by route
The stream count per 8-GPU replica is published at one point for GLM-5.3 (16 streams at 184 tok/s) and only at batch-1 for K3. Everything else is extrapolation, tagged as such — and measuring it is a condition precedent of the raise, not a nice-to-have.
Naming note: GLM-5.3 open weights were gated as of 26 Aug 2026; GLM-5.2’s weights are the same 744B / 40B-active architecture, so the published decode curve and the capacity maths carry over. Pub source brief §8 · huggingface.co/zai-org/GLM-5.3
| Metric | Route 1 · Kimi K3 only | Route 2 · GLM-5.3 only | Route 3 · Hybrid (35% K3) |
|---|---|---|---|
| Streams per replica at ≥60 tok/sthe number the volume SLA hangs on | 4 / 12 / 24 Est | 30 / 40 / 56 Est | blended 20.2 (base) |
| Streams per replica at ≥100 tok/sthe premium SLA — K3 sells as “target 100 / floor 60” until measured | 1 Pub / 6 / 12 Est | 16 Pub / 24 / 32 Est | mix of both |
| Published anchor point | 111–118 tok/s at bs=1; 331–370 with DSpark; no conc>1 curve Pub | 184 tok/s/user at conc=16, 8×B200 NVFP4 Pub | inherits both |
| Seats per live rack (base, 6%) | 638 | 2,194 | 1,095 |
| Seats, 12 racks (11 live), 10% premium mix | 7,020 | 24,131 | 12,045 |
| … all seats at ≥60 tok/s | 7,722 | 25,740 | 13,106 |
| … all seats at ≥100 tok/s | 3,861 | 15,444 | 6,969 |
| Racks for 25,000 seats (incl. spares) | 43 | 13 | 25 |
| Racks for 50,000 seats (incl. spares) | 86 | 25 | 50 |
| Iron for 25k / 50k, all-in | $219m / $434m£161m / £319m | $68.5m / $129m£50m / £95m | $129m / $254m£95m / £187m |
| Scenario range, seats per live rackConservative / Base / Upside | 81 / 638 / 3,415 | 727 / 2,194 / 8,154 | 131 / 1,095 / 6,063 |
On the published-only floor (GLM conc=16, nothing extrapolated) 11 live racks are a 10,296-seat service — and that floor now carries a full named P&L in §07: it dies at M30 and needs ~$349 list. Dotted underline = row-wise best number, not a recommendation. 07-model-spec.json results.capacity_table, twelve_racks_serve_base, named_runs · 02-capacity §6 · Excel: Capacity + Scenarios sheets
Racks required by seat target
Same SLA (≥60 tok/s at 6% peak), same taxes, same iron — only the model changes. Includes spares.
K3’s base cell (12 streams/replica at ≥60 tok/s) is a bandwidth model, not a benchmark; its ≥100 tok/s floor is 1. GLM’s base of 40 is a no-spec-decode fit plus a 1.2–1.3× MTP retention haircut — 02-capacity’s own caveat that the linear fit turns optimistic above c≈32 is load-bearing. The four one-day benchmarks — GLM-5.3 at conc 32 / 48 / 64 × 8K / 32K ISL, K3+DSpark at conc 4 / 8 / 12, prefix-cache hit rate on a real anchor trace, and a written rack quote — are closing conditions of the raise, and the anchor term sheet carries a capacity acceptance test: the measured curve re-baselines seats-per-rack. 02-capacity §2b, §9 · 12-adversarial A2 · 13b buyer review · sensitivity grid 2, §08
Which model: Kimi K3, GLM-5.3, or both?
Three routes, costed identically end-to-end. The founder chooses; this section shows where each route wins, what each costs, and precisely what would flip the decision. Nothing here says a tier “has to be” a given model.
| Criterion | Route 1 · Kimi K3 only | Route 2 · GLM-5.3 only | Route 3 · Hybrid (35% K3) |
|---|---|---|---|
| Answer quality / brandAA Intelligence Index; benchmark split | AA 60 — #1 open on AA; GPQA 93.5, native vision Pub | AA 60 — tie; leads Terminal-Bench 3.0, GDPval Elo; flagship is text-only Pub | best-of-both per task |
| Seats per live rack (≥60, 6%) | 638 | 2,194 | 1,095 |
| Blended list price | $190£140 (incl. $25 K3 uplift) | $165£121 | $190£140 (incl. $25 K3 uplift) |
| Price required at fleet-full, cash / full costlist $/seat-mo at which EBITDA / EBIT = 0 | $400 / $512 | $117 / $144 | $233 / $298 |
| Full-cost floor at 25k / 50k seats | $330 / $313 | $123 / $103 | $206 / $189 |
| Option C: seats supportable M36 | 7,020 | 30,712 | 12,045 |
| Option C: ARR M12 / 18 / 24 / 36, $m | 6.5 / 10.3 / 14.2 / 12.5 | 7.0 / 18.3 / 36.0 / 48.8 | 8.1 / 17.7 / 24.4 / 21.5 |
| Option C: seats turned away at M36the route’s growth ceiling | 26,980 | 3,288 | 21,955 |
| Option C: EBITDA break-even | never | M17 | M16 |
| Option C: cash low (month) | −£9.1m — dead M26 | +£7.6m (M23) | +£3.8m (M36) — brushing the £4m floor |
| Gross margin M36, cash / GAAP | 13% / −103% | 72% / 35% | 49% / −18% |
| Licence | Kimi K3 License — not open source; separate Moonshot agreement once aggregate MaaS revenue > $20m T12M; the C fleet’s capacity ceiling keeps it un-triggered in-horizon — any scale-up fires it Pub | MIT (5.2); 5.3 unstated until repo opens Pub | clause reads aggregate revenue — any K3 share carries the same single-counterparty negotiation Pub |
| Measurement riskwhat the seat counts rest on | no published conc>1 curve at all; DSpark needed above ~118 tok/s | published to conc=16; extrapolated 16→40 (floor P&L named in §07) | both of the above + routing layer |
| Ops burden | one model | one model | two eval suites, two templates, ~+2 FTE, 7.5% split-pool penalty |
Dotted underline = row-wise best number, never a recommendation. Quality, safety and the anchor’s own eval set are outside this table. 07-model-spec.json results.route_comparison["base|C"] · 06b-model-comparison · Excel: Scenarios sheet (routes side by side)
The decision criteria
- Quality is a published tie, split underneath. Artificial Analysis Intelligence Index 60 vs 60. K3 leads long-horizon SWE, tool marathons, GPQA (93.5) and is the only flagship with native vision; GLM-5.3 leads Terminal-Bench 3.0 (28.3 vs 17.4), HLE-with-tools and GDPval Elo. Only the anchor’s own eval set can separate them. Pub
- Capacity is ~3.4× and it compounds. A K3 stream costs the fleet ~3.4× a GLM stream at the same SLA (4× at ≥100 tok/s). That single fact drives the seat counts, floors and margins above.
- Price required vs price the niche pays. At the $150/$300 book, GLM clears its full-cost floor with ~15% headroom ($165 blended vs $144 required) — and at the 25k ladder tier the anchor pays ~$124 net, essentially at the 25k full-cost floor ($123): the discount ladder is the whole negotiation. The 35% Hybrid needs $298 against $190 charged; K3-only needs $512 — a boutique price for ≤7k seats, not a 25k-seat anchor product.
- Licence exposure is asymmetric. The K3 trigger is on aggregate company revenue, so Hybrid triggers exactly as K3-only does; the cost is not the fee (5% base on K3 share) but negotiating with a single counterparty mid-ramp. GLM is MIT — which is why the recommendation launches Route 2 pure and adds K3 only after Moonshot terms are in writing. Pub
- Optionality is one-way. GLM-today can add K3 replicas later in a weekend (weights public, day-0 engines). K3-today cannot cheaply become GLM if the licence talk goes badly — the whole price book would be K3-shaped.
The flip chart: what a K3 traffic share costs
| K3 share of streams | 0% (= GLM + uplift) | 10% | 20% | 35% (base) | 50% | 100% (= K3) |
|---|---|---|---|---|---|---|
| Seats per live rack | 2,194 | 1,632 | 1,364 | 1,095 | 915 | 638 |
| Seats at M36 (fleet the trigger builds) | 37,294 | 22,841 | 15,006 | 12,045 | 10,060 | 7,020 |
| ARR M36, $m | 62.8* | 40.8 | 26.8 | 21.5 | 18.0 | 12.5 |
| Cash low, £m | 8.9 | 7.5 | 5.3 | 3.8 | −1.1 (dead M34) | −9.1 (dead M26) |
| Full-cost list price required, $ | 136 | 193 | 248 | 298 | 357 | 512 |
| GAAP GM M36 | 41% | 22% | 4% | −18% | −42% | −103% |
*The 0% row carries the $25 K3-access uplift on GLM-only seats; the extra revenue also unlocks the 4th tranche in-horizon (17 live racks vs 14), hence above Route 2’s $48.8m. A free model picker drifts K3-ward — every seat paying a flat uplift pulls toward the 50% attractor, which is cash-negative on this sweep — so the share must be enforced (router policy + per-seat allowance, or metered K3 pricing), and split pools pay a 7.5% quantisation penalty while both exist. Even the 10% point is marginal, not comfortable: $193 required vs $190 charged, still carrying the Moonshot aggregate-revenue surface. 07-model-spec.json results.hybrid_k3_share_sweep_base (C rows) · 12-adversarial A6 · 06b §8 · Excel: Scenarios sheet, K3-share input
(a) A measured K3+DSpark curve at ≥12 streams/replica at 100 tok/s — in the Upside capacity column the K3 routes out-earn GLM ($102m vs $81m ARR M24) because the uplift and premium price carry it once capacity stops binding; until that bench exists, no deck slide may show K3 out-earning GLM. (b) The anchor pays ≥$300/seat for K3 across the estate and Moonshot terms are signed in writing. (c) The anchor’s eval set is vision-heavy — GLM’s flagship cannot see images.
(a) Counsel confirms the $20m trigger is on aggregate revenue and Moonshot won’t pre-agree terms. (b) A measured 8×B300 GLM curve holding ≥40 streams at ≥60 tok/s. (c) Procurement requiring OSI-approved licences (MIT passes; the Kimi K3 License does not). If GLM-5.3’s weights stay gated past the procurement window, the fallback is GLM-5.2’s identical-architecture weights, not K3.
The wider field: Qwen, DeepSeek and the rest — verified 27 Aug
The founder asked about Qwen3.8-Max and DeepSeek V4 Pro by name; a buyer will too. Verdict up front: nothing in the field flips the route — but two entries strengthen the plan and one is a trap. Because the fleet is model-agnostic iron, every candidate maps onto one of the two capacity classes the engine already prices (GLM-class ~40 streams/replica, K3-class ~12, at ≥60 tok/s) — so “which models will his clients want?” is a menu question, not a capex question: adding one is a weights swap plus an eval pass, not new iron.
| Model | Licence | Total / active | Fits 8×B300 replica? | Class · streams/rack | AA | Verdict |
|---|---|---|---|---|---|---|
| Qwen3.8-Max (2.4T-A95B, CN) · open ckpt is text-only | custom “qwen3.8-max”, unreviewed | 2.4T / 95B | No at BF16/FP8 (4.8 / 2.4 TB vs 2,304 GB); community FP4 only | K3-class · ~32 | 58 | Near-peer to K3 on paper; blocked by fit, licence and no vision in the weights. On-request only — never a committed tier. |
| DeepSeek V4 Pro (0813, CN) | MIT | 1.6T / 49B | Yes (~893 GB) | GLM-class · ~110 | 53 | Best licence + the only published GB300 rack-scale curve of any open model; 7 AA points under GLM-5.3. Menu slot 3. |
| GLM-5.3-Flash (CN) · vision | MIT | 320B / 18B | Yes | GLM-class · ~130 | 57 | Vision at MIT in the volume class; engine support immature — add when vLLM/SGLang mainline it. |
| Mistral Large 3 (EU) | Apache 2.0 | 675B / 41B | Yes | GLM-class · ~130 | 16 | The non-PRC fallback on paper; quality nowhere near the class bar. |
| MiniMax M3 (CN) | community, $20m revenue trigger | 428B / 23B | Yes | GLM-class · ~130 | 45 | Dominated: K3’s licence risk at GLM-minus-15 quality. |
| gpt-oss-120b (US) | Apache 2.0 | 117B / 5.1B | Yes | GLM-class · very high | 24 | Only US open model with real serving footprint; a utility/draft model, not a seat model. |
Also checked: MiniMax M2.7 (licence is non-commercial — excludes MaaS outright), Llama 4 Maverick (Apr 2025, superseded), Meta’s Muse Glimmer 30B (AA 35, a distillation), and Muse Spark 1.2 (API-only — do not plan around it). 06c-extended-model-landscape.md · landscape verified via Artificial Analysis, Hugging Face model cards, vendor blogs, 27 Aug 2026
Contract tok/s SLAs and capacity classes, never model names, and publish a menu: 1) GLM-5.3 (volume default) · 2) Kimi K3 (premium/vision, metered, post-Moonshot-terms) · 3) DeepSeek V4 Pro (MIT understudy, kept hot — also the honest answer when procurement asks “what if Z.ai is a problem”, though it shares the CN origin) · 4) Qwen3.8-Max FP4 on request, after licence review and an 8×B300 benchmark · 5) GLM-5.3-Flash when engines mature. Marginal cost per added model is eval and ops surface, not iron: ~0.4–1.6 TB storage per checkpoint (noise against the storage line), a one-day rented benchmark, an anchor-eval + red-team pass, and gateway integration — ~2–4 weeks, ~1 FTE-month. Menu share is a weekly replica rebalance behind one gateway, not a build. Est 06c §3b
Where the money goes
A rack is $5.01m all-in, not $4.0m — fabric, spares, install and contingency add 25%. Build options differ only in how many racks are bought with equity on day one, and that is the difference between insolvency and a 33-month runway.
| Line | USD | GBP |
|---|---|---|
| GB300 NVL72 rack Est | $4.00m | £2.94m |
| IB fabric — leaf + optics (8%) | $320k | £235k |
| Spares (1.5%) | $60k | £44k |
| NVMe staging | $30k | £22k |
| Shipping & insurance | $45k | £33k |
| Install & commissioning | $75k | £55k |
| Software, year 1 (OSS stack) | $25k | £18k |
| Per-rack all-in, + 10% contingency | $5.01m | £3.69m |
| Fixed site, one-off (spine, parallel FS, gateway, spares kit), + 10% Est | $3.36m | £2.47m |
| Day one — C: 6 racks | $33.4m | £24.6m |
| Day one — B: 8 · A: 12 | $43.4m · $63.5m | £32.0m · £46.7m |
| Milestone | USD / mo | GBP / mo |
|---|---|---|
| M12 · 6 racks, 21 heads | $1.14m | £0.84m |
| M24 · 12 racks, 31 heads | $2.32m | £1.70m |
| M36 · 15 racks, 41 heads | $3.04m | £2.24m |
| Colo + power per rack-month (140 kW committed) Est | $60.4k | £44.4k |
| Depreciation per rack-month (4-yr, 10% residual) | $96.1k | £70.7k |
Monthly cost stack at M36
Base · GLM-5.3 · Option C, 15 racks. Cash costs + lease; depreciation excluded.
| Role | Loaded £/yr | Q4 | Q8 | Q12 |
|---|---|---|---|---|
| CEO / founder · CTO | £188k · £225k | 2 | 2 | 2 |
| Inference / platform engineering | £163k | 5 | 7 | 8 |
| ML engineering | £138k | 2 | 4 | 5 |
| SRE / on-call (5-primary rota through the pilot SLA window) | £131k | 5 | 6 | 6 |
| Security / compliance (1 head from Q1 — buyer questionnaires precede the pilot) | £125k | 2 | 2 | 3 |
| Sales / CS / solutions | £150k | 3 | 6 | 8 |
| Finance / ops · DC ops | £106k · £81k | 2 | 4 | 5 |
| Total heads | — | 21 | 31 | 37 |
All roles Est from Glassdoor / Robert Half bands (04-opex §2). SRE and DC-ops people sit in opex, not COGS — gross margins on this page exclude them from COGS by definition. New opex line after the buyer review: model assurance £150k/yr (per-checkpoint red-team, model bill of materials, egress audit). Excel: Opex sheet, Staff sub-table · 13b-adversarial
A $6.0m supply-chain quote instead of $4.0m kills Option A at M4 and Option B at M7 on every route. Option C on GLM-5.3 survives it with an £8.4m cash low — the forward-looking gate defers tranches until cash clears (ARR M24 $19.3m instead of $36.0m); C on the Hybrid survives at £2.6m; C on K3 dies at M29. Do not commit an anchor seat count before a written rack quote exists. 07-model-spec.json results.stress_6m_rack_base · §08 grid 1
Pricing: what the market charges, what he must charge
The privacy premium is real but bounded: $150 volume / $300 premium list is 2.5× ChatGPT Enterprise. On GLM-5.3 that price clears the full-cost floor with ~15% headroom; on the 35% Hybrid it does not; on K3-only it is not close.
| Line | Conservative | Base | Upside | Basis |
|---|---|---|---|---|
| Market anchor: ChatGPT Enterprise Pub | — | ~$60~£44 | — | the seat the buyer already knows |
| Volume tier, ≥60 tok/s Asm | $120£88 | $150£110 | $175£129 | 2.0 / 2.5 / 2.9 × ChatGPT Ent. |
| Premium tier, ≥100 tok/s Asm | $250£184 | $300£221 | $400£294 | 2× volume; 10% of seats |
| K3-access uplift, every seat (Routes 1 & 3) Est | $0 | $25£18 | $50£37 | public-API token delta K3 vs GLM; the metered-overage design replaces it if K3 ships later |
| Blended list (90/10 mix + uplift) | $133 | R2 $165 · R1/R3 $190 | R2 $198 · R1/R3 $248 | net base: $140.25 blended (R2) |
| Anchor net at 25k contracted seats (ladder −25%) | — | ~$124~£91 | — | essentially at the 25k full-cost floor ($123) — the ladder is the negotiation |
| Full-cost list required at the M36 fleet — GLM-5.3 | — | $144£106 (cash $117) | — | list price at which EBIT after interest = 0 |
| … Hybrid 35% · Kimi K3 only | — | $298 · $512 | — | same definition, same build |
| Overage (info) | — | vol $2/$6 · prem $4/$18 per M tok in/out | — | ~6% of seat revenue at base |
Revenue also carries SLA credits at 1.5% as contra-revenue (availability ≥99.9% + tok/s p95 over 5-minute peak windows, buyer-run probes as evidence, credit ladder 5/15/30%), and a DSO input: base = anchor billed quarterly in advance as a contractual condition; plain net-60 is the Conservative and would strand ~$8m in receivables at the trough. 07-model-spec.md §1.6 · results.runs_summary price_required per route · 13/13b-adversarial · Excel: Revenue sheet
The price ladder: market, our book, and each route’s floor
$ per seat-month, list. A route whose floor sits above the book cannot serve the anchor at this price.
Sell the anchor a 3-year take-or-pay for ≥10,000 seats at $140 net — $50m / £37m TCV — with 25% of year-1 value prepaid at signature, billed quarterly in advance, with the capacity acceptance test attached. In the model the M12 prepay is $4.62m / £3.40m (it caps at ordered capacity); that plus reserve funds the equity slice of the first 3-rack tranche outright (£3.9m), and the contract — assigned to the lessor, with parent guarantee — is the collateral for the 65% advance. Unit check: one GLM-5.3 rack at fleet-full bills ~$308k/mo against $60k colo + $108k lease + $6k support — ~57% fill covers cash cost. 07-model-spec.json results.phase2_arithmetic_base · Excel: Revenue + Monthly sheets
A prepared buyer team will find the floor from the sensitivity grid regardless, so present it: open at $165–175 volume list with the published seat ladder; expect the buyer to target $120–130 net blended. The model survives the corridor — at a $120 volume list the cash low is +£0.8m, which breaches the plan’s own £4m floor rule and is therefore the walk-away; $100 needs new money (−£5.3m). Signature at M15 instead of M12 (the buyer’s realistic calendar) moves break-even to M20 with a £6.1m low — survivable, and shown in §07 before diligence finds it. §08 grid 1 · results.named_runs buyer_realistic · 13b-adversarial
36-month P&L and cash
Base case, GLM-5.3, Option C: EBITDA break-even at M17, cash bottoms at £7.6m in M23 (the only trough — tranche 3 waits behind the forward-looking gate until it clears), and ends at £10.8m with the fleet 100% full. Peak funding gap: £0.
Revenue vs cash cost run-rate, monthly
Base · GLM-5.3 · Option C. Cost = cash COGS + opex + lease payments. £ per month. The M36 revenue step-down is the anchor crossing 25k contracted seats onto the −25% ladder tier.
Closing cash by route, Option C
Base scenario. Same build, same prices — only the model changes. £m.
| Metric | M12 | M18 | M24 | M36 |
|---|---|---|---|---|
| Racks total / live | 6 / 5 | 9 / 8 | 12 / 11 | 15 / 14 |
| Seats served / demand | 4,000 / 4,000 | 10,400 / 10,400 | 20,500 / 20,500 | 30,712 / 34,000 |
| Seat-fill of capacity | 36% | 59% | 85% | 100% |
| ARR | $7.0m£5.2m | $18.3m£13.4m | $36.0m£26.5m | $48.8m£35.9m |
| EBITDA, monthly | −$553k | +$193k | +$1.34m | +$2.00m |
| Gross margin, cash / GAAP | 20% / −88% | 58% / −2% | 70% / 30% | 72% / 35% |
| Closing cash | £19.8m$27.0m | £11.8m$16.0m | £7.8m$10.6m | £10.8m$14.7m |
| Heads · anchor share of seats | 21 · 100% | 27 · 77% | 31 · 78% | 41 · 81% |
Year totals, $m: Y1 revenue 1.8, EBITDA −7.6, capex 36.8 · Y2 revenue 21.0, EBITDA +4.3, capex 9.4, lease 4.5 · Y3 revenue 45.0, EBITDA +23.4, capex 7.7, lease 9.1. Expansion tranches: PO M12 / 16 / 29 / 35, live M16 / 20 / 33 / 39 (the 4th serves beyond the horizon), each 3 racks = $15.0m installed, $9.77m leased (65% at 12%, 36 months), $5.3m / £3.9m equity + $0.98m / £0.72m lessor deposit at PO. The M20→M29 gap is the forward-looking cash gate holding tranche 3 — expansion waits for cash, not the other way round. All four tranches together: $39.1m principal, ~$1.3m/mo aggregate service, $3.9m deposits. M36 demand runs 3,288 seats ahead of capacity — the growth case for tranche 5, not a Phase-1 problem. runs["base|route2|C"] year_totals, racks.expansions
| Metric | A · 12 racks all-in | B · 8 racks + runway | C · 6 racks + lease expansion |
|---|---|---|---|
| Day-one capex | £46.7m$63.5m | £32.0m$43.4m | £24.6m$33.4m |
| Final fleet (racks) / seats supportable | 12 / 24,131 | 14 / 28,519 | 15 (18 ordered) / 30,712 |
| EBITDA break-even | M19 | M17 | M17 |
| Cash low (month) | −£6.8m (M21) | +£9.0m (M34) | +£7.6m (M23) |
| Runway with base revenue | insolvent at delivery (M4) — needs ~£57m | >36 months | >36 months |
| Zero-revenue runway | dead M4 | dead M23 | dead M33 |
| ARR M24 / M36 | $36.0m / $37.4m (capped) | $27.0m / $44.9m (tranche fires M31) | $36.0m / $48.8m |
| Metric | Conservative | Base | Upside |
|---|---|---|---|
| Seats per live rack | 727 | 2,194 | 8,154 |
| EBITDA break-even | never (in 36) | M17 | M8 |
| Cash low | −£25.5m — dead M14, gap £25.5m | +£7.6m (M23) | +£27.8m (M7) |
| ARR M24 / M36 | $3.9m / $3.9m | $36.0m / $48.8m | $81.2m / $151.2m |
| K3-route comparison | dead M14, gap £29.9m | dead M26, needs $512 seats | K3/Hybrid out-earn GLM ($102m ARR M24) if the curve measures at 24/12 |
Conservative £ figures are at its own stacked FX of 0.809. 07-model-spec.json results.runs_summary (27 combinations) · Excel: Scenarios sheet shows Base × three routes side by side, then Conservative and Upside
| Run | What it holds | Break-even | Cash low | ARR M36 | Verdict |
|---|---|---|---|---|---|
| Defended base | everything in §A1 | M17 | +£7.6m (M23) | $48.8m | gap £0 |
| Published-only floor | GLM held at the one PUBLISHED point (conc=16 both tiers) → 936 seats/live rack | never | −£4.2m — dead M30 | $16.0m | needs ~$349 list — why the benchmark is a condition precedent |
| Buyer-realistic | signature M15, schedules +3 months, prepay released in thirds (signature / acceptance / SOC 2 Type II) | M20 | +£6.1m (M26) | $42.4m | gap £0 — the buyer’s calendar does not kill the plan |
| FX stress | GBP at 0.78 / 0.809 all 36 months | M15 | +£6.2m / +£5.2m | $48.8m | gap £0; mitigation: forward-buy the initial PO’s ~$33m at close |
07-base-results.json named_runs · Excel: Scenarios sheet, named-runs block
Sensitivity
Two grids decide the build: rack price × seat price sets whether Option C survives, and concurrency × measured streams sets how many seats a rack really is. Both are pasted from the engine; the Excel Sensitivity sheet carries all three grids per route.
| Cash low £m · rack ↓ / volume list $ → | $100 | $120 | $150 | $175 | $200 |
|---|---|---|---|---|---|
| $3.7m | −1.8 | 4.3 | 10.1 | 12.1 | 13.7 |
| $4.0m | −5.3 | 0.8 | 7.6 | 9.7 | 11.4 |
| $4.5m | −11.2 | −5.0 | 3.4 | 5.6 | 7.4 |
| $5.0m | −17.2 | −10.8 | −1.8 | 1.5 | 3.3 |
| $6.0m | −29.0 | −22.6 | −13.4 | −6.8 | −4.8 |
GLM-5.3 · Option C · tranche months held fixed at the base plan so the grid is monotonic. Any negative cell needs new money. Premium = 2× volume throughout. The live model does better at the $6m print (£8.4m low, §05) because its forward-looking gate defers tranches — this grid deliberately does not. 07-model-spec.json results.sensitivity_base.route2.rack_price_x_seat_price (C rows)
| Seats / live rack · conc ↓ / measured streams → | 0.5× | 0.75× | 1.0× | 1.25× | 1.5× |
|---|---|---|---|---|---|
| 3% | 2,194 | 3,291 | 4,388 | 5,484 | 6,581 |
| 4% | 1,645 | 2,468 | 3,291 | 4,113 | 4,936 |
| 5% | 1,316 | 1,974 | 2,632 | 3,291 | 3,949 |
| 6% | 1,097 | 1,645 | 2,194 | 2,742 | 3,291 |
| 8% | 823 | 1,234 | 1,645 | 2,057 | 2,468 |
The “measure before you buy” grid: multiplier scales all stream cells (the rented-node benchmark lands somewhere on this axis); concurrency is the agent-workload axis — base now sits at 6% with contractual per-seat caps. Even at 0.5× and 8% concurrency a GLM rack (823 seats) outserves the K3 base (638). results.sensitivity_base.route2.concurrency_x_streams_mult
The path, the investor narrative, and the risks
Three phases, each gated on something signed or measured — no phase starts on a forecast. The story for the next investor is contracted ARR against leased iron, not GPUs on a balance sheet.
-
Phase 1 · M1–M9Stand it up, prove itGates out: the four one-day benchmarks (condition precedent of the raise); anchor production take-or-pay signed — base M12 conditional on signed pilot paper at close; the buyer-realistic M15 case survives (§07)
-
Phase 2 · M12–M24Scale with the contractGate per tranche: anchor signed · demand > 85% of ordered capacity · forward-looking cash test: cash after equity slice, deposit and 6 months of the new lease service ≥ £4m floor
-
Phase 3 · M25–M36Fill, then fund the next legGate: 3+ logos; Series B at 4–6× contracted ARR while the anchor is >60% of revenue — not 8× Asm
Top risks and the mitigations already in the plan
| Risk | Impact if it lands | Mitigation in the plan |
|---|---|---|
| Rack quote prints $6.0m, not $4.0m Pub print exists | A dead M4, B dead M7 on every route | Option C on GLM-5.3 survives it (cash low £8.4m — the forward gate defers tranches); written quote is decision #1; HGX-B300 8-GPU boxes are the procurement fallback and the residual-value hedge |
| Streams measure below the base cells (GLM 40, K3 12) | seat counts fall toward the published floor (dead M30 at conc=16) | The benchmark is a condition precedent of the raise; the anchor term sheet carries a capacity acceptance test that re-baselines seats/rack; §08 grid 2 prices every landing point |
| GLM-5.3 weights stay gated or arrive non-MIT | route 2’s model unavailable at go-live | GLM-5.2 weights are the same 744B/40B architecture under MIT — same curve, same maths; K3 is not the fallback |
| K3 licence: $20m aggregate-revenue trigger Pub | an open-ended single-counterparty negotiation on any route serving K3 | Launch is Route 2 pure — zero licence surface; a K3 tier ships only after Moonshot commercial terms are in writing, as metered overage that keeps the share self-limiting |
| Anchor signs late (buyer-realistic M15; Conservative M21) or never | stacked-downside case is dead M14 | Named buyer-realistic run: BE M20, low £6.1m, ARR M36 $42.4m, gap £0. 6-rack build holds a 33-month zero-revenue runway; no contract = no tranche = no new cash out |
| Anchor concentration: 77–81% of seats vs ≤60% target | one procurement decision owns the P&L; cross-defaults the leases | Take-or-pay assigned to the lessor with parent guarantee; Series B presented at 4–6× contracted ARR; the non-anchor schedule must roughly double — 8 sales/CS heads by Y3 are costed in; say the miss out loud |
| Ordinary net-60 billing terms | ~$8m stuck in receivables at the cash trough — the base case dies on payment terms alone | Anchor billed quarterly in advance as a contractual condition (the DSO input prices the failure); prepay released in tranches against milestones |
| Single site fails the buyer’s BC/DR review | pilot blocked at the questionnaire stage | Tiered DR: named bare-metal provider under identical controls in the DPA, RTO 72h degraded; the standing cost (×3 from anchor signature) is in opex; a second site is priced as a Phase-2 option |
| Lessors refuse tranche 2+ | fleet stops at 9 racks | 9 racks ≈ 17,550 GLM seats — still ahead of base demand to M18; two lessor term sheets are decision #6, with non-disturbance + step-in negotiated at tranche 1 |
What a sceptic will say — and our answer
Every row is a verdict from the three adversarial reviews commissioned on this plan (12-technical, 13-financial, 13b buyer-side). The answers are fixes now live in the engine, not debate points; the fix log is work/14-fix-log.md.
Recommendation, and this month’s decisions
Commit to Option C on Route 2 pure — 6 racks (£24.6m / $33.4m) on GLM-5.3, MIT weights, zero licence surface — gate every leased tranche on the anchor’s take-or-pay, and add a K3 opt-in tier only after commercial terms with Moonshot are agreed in writing, as metered overage (~3–4× GLM token rates with a premium-seat allowance), not a per-seat uplift.
This is a view with named flip conditions, not a foregone conclusion: a measured K3+DSpark curve at ≥12 streams/replica at 100 tok/s, or an anchor paying ≥$300/seat for K3 with Moonshot signed, reopens the route choice — the model reproduces both by moving two inputs (Excel: Inputs sheet, Route selector and K3-share cell; the metered future state is Route 3 at the 10% point). What flips C back to B: a rack quote ≤$3.7m and the anchor contracting ≥15k seats at signature. Option A is rejected on every route: £46.7m of iron on £50m is insolvent at delivery.
Decisions the founder must make this month
- Get the rack quote in writing. $4.0m is a planning price; the $6.0–6.5m print is real. Every seat promise waits on this number.
- Run the four benchmarks — they are closing conditions of the raise. GLM-5.3 at conc 32 / 48 / 64 × 8K / 32K ISL and K3+DSpark at conc 4 / 8 / 12 on a rented 8×B300 node, prefix-cache hit rate on a real anchor trace, plus the written rack quote. One–two weeks; it moves more money than any negotiation.
- Put the two binary gates to the anchor in meeting 1. (a) Provenance: does their (or their US affiliates’) policy exclude PRC-origin open weights — get pre-clearance in writing; (b) residency: UK / EU / UAE position fixed in the DPA. Either answer reshapes the plan more than any price.
- Sign the colo LOI now. Power-on (M6) is the go-live gate, not rack delivery (M4). £44.4k/rack-month base at 140 kW committed per rack; 3-month deposit.
- Adopt the route posture. Launch GLM-5.3 volume + premium tiers (contracted on tok/s SLAs, never model names); K3 later as metered overage, only after Moonshot terms in writing.
- Table the anchor offer. 3-year take-or-pay, ≥10,000 seats at $140 net ($50m / £37m TCV), 25% year-1 prepay released in thirds, billed quarterly in advance, capacity acceptance test attached; open at $165–175 list with the published ladder; walk away below $120 list.
- Get two lessor term sheets. 65% advance at 12% over 36 months against the anchor contract, 10% deposits — with non-disturbance + step-in for the anchor negotiated at tranche 1. Verify before it is load-bearing.
- Forward-buy the initial PO’s ~$33m at close. Every input is dollar-denominated except the revenue’s discount; GBP −10% costs £2.4m of trough on its own.
Appendix
Assumptions register, the original 26 Aug diligence brief re-presented in full (its sources retained verbatim), and the model-comparison sources.
A1 · Assumptions register — the inputs that move the answer
| Input | Cons | Base | Upside | Tag | Source |
|---|---|---|---|---|---|
| FX, GBP per USD | 0.809 | 0.7357 | 0.7357 | Pub base / Asm cons | Wise, 26 Aug 2026; Conservative stacks GBP −10% — mitigation: forward-buy the initial PO ~$33m |
| Design concurrency (in-flight ÷ seats) · per-seat stream caps | 8% · 2/5 | 6% · 2/5 | 3% · 2/5 | Est | 01-demand Little’s law M36 agentic point; caps contractual (12-adv A1) |
| K3 streams / replica @ ≥60 tok/s | 4 | 12 | 24 | Est | 02-capacity §2b; no published curve |
| K3 streams / replica @ ≥100 tok/s | 1 | 6 | 12 | Pub floor, else Est | vllm.ai/blog/2026-07-27-k3; tier sells “target 100 / floor 60” until measured |
| GLM streams / replica @ ≥60 tok/s | 30 | 40 | 56 | Est | no-spec fit + 1.2–1.3× MTP retention (12-adv A2) |
| GLM streams / replica @ ≥100 tok/s | 16 | 24 | 32 | Pub at 16, else Est | sglang PR 29674 (184 tok/s @ conc 16) |
| Hybrid K3 traffic share · split-pool penalty | 50% · 10% | 35% · 7.5% | 20% · 5% | Asm | BRIEF; enforced by router + per-seat allowance (12-adv A6) |
| Prefill tax (extra GPUs, fn of prefix-cache hit ≥70/85/90%) | +60% | +25% | +15% | Est | 12-adv A7 re-derivation; cache-hit bench is a closing condition |
| p99 headroom (incl. upgrade drain) · long-context factor | 25% · 0.50 | 25% · 0.65 | 25% · 0.80 | Asm | 02-capacity §4; 12-adv A8/A9; volume seats capped 256K context |
| Rack price (stress $6.0m held separately) | $4.5m | $4.0m | $3.7m | Est | Loop Capital $3.7–4.0m; analyst $3.5–4.5m |
| Depreciation (racks) · residual | 3 yr · 20% | 4 yr · 10% | 5 yr · 5% | Asm | 03-capex §7; 3-yr column shown to lenders unprompted |
| Colo per rack-month, UK (140 kW committed) | £68.2k | £44.4k | £29.6k | Est | 04-opex §1.2; 12-adv A11; UAE saves ~£2.8m/yr at 12 racks (anchor-consent-gated) |
| Volume / premium seat price, $ list | 120 / 250 | 150 / 300 | 175 / 400 | Asm | 2.0–2.9× ChatGPT Enterprise ~$60 |
| Achieved discount · anchor seat ladder (≥5k/10k/25k) | 30% · 15/20/30% | 15% · 10/15/25% | 5% · 5/10/15% | Asm | 05-pricing §5.1; anchor pays max(achieved, ladder) — 13b blocking 1 |
| Price escalator from M25 · SLA credits | 0% · 3% | 0% · 1.5% | 8% · 0.5% | Asm | token prices fell ~10× in 2 yrs (13-adv); credits contra-revenue (13b) |
| Billing terms (DSO days) | 60 | 15 | 0 | Asm | base = anchor quarterly-in-advance, contractual (13-adv blocking 2) |
| Anchor production signature month | M21 | M12 | M9 | Asm | 05 §6 / 06 §1; M12 conditional on signed pilot paper at close; buyer-realistic M15 is a named run |
| Lease: advance · rate · deposit | 55% · 15% · 15% | 65% · 12% · 10% | 75% · 10% · 5% | Asm | 06 §2; CoreWeave SOFR+2.25–5.50% Pub; deposits at PO (13-adv) |
| K3 licence rate on K3 revenue (trigger $20m T12M Pub) | 10% | 5% | 2% | Asm | no published Moonshot terms; clause reads aggregate revenue |
| Model assurance (red-team, MBOM, egress audit), £/yr | 200k | 150k | 100k | Est | 13b-adversarial; nothing in 04-opex covered it |
| Go-live month (colo power-on gate) | M10 | M7 | M5 | Est | 04 §1.3; racks delivered M4 |
| Cash floor (reserve rule) · expansion gate | £4m | £4m | £4m | Asm | 06 §1; gate is forward-looking: post-tranche cash incl. 6 months of new lease service ≥ floor (13-adv) |
Files: work/07-model-spec.md (full register and formulas), 07-model-spec.json (machine-readable), 07-model-calc.py (engine), 07-base-results.json (all 27 scenario × route × build runs, monthly). Excel: Inputs sheet, blue cells; Scenario / Route / Build selectors.
A2 · The original diligence brief (26 Aug 2026), tidied
Internal diligence brief · hardware-envelope note · dated 26 Aug 2026 · web edition 27 Aug 2026 · not for index, search, or public share. Working numbers from the 26 Aug conversation are tagged ESTIMATE. Lab or vendor-published figures are tagged PUBLISHED. There is no NVIDIA list price in either set. Context: an inference-serving company; 50,000 R&D employees would get access, not all generating at once; model Kimi K3 on GB300 NVL72. Figures in §§1–4 are the 26 Aug conversation transcribed; they are not a published bake-off.
How the new model departs from this brief
- Model name. The new model is named GLM-5.3 throughout. GLM-5.3 shares the same 744B / 40B-active base as GLM-5.2; its open weights were still gated as of 26 Aug 2026, so the published GLM-5.2 decode curve (§8) carries over unchanged. Where this brief writes GLM-5.2 in §§8, 9, 11 and 13, read GLM-5.3 in the model.
- Routes, not tiers. Capacity is costed per route — Route 1 Kimi K3 only, Route 2 GLM-5.3 only, Route 3 Hybrid (both models offered to every seat) — rather than assuming the volume / premium split that §13 leans to. That lean is research, not a decision; the choice is the founder's.
- Full BOM. The “$2m for storage and IB fabric” placeholder in §§4 and 12 is replaced by a full bill of materials in the new model.
- Beyond iron. Colo, power, staff and runway — which this brief only notes “sit on top of the $50m” — are modelled explicitly on top of the iron line.
Honest offer on a $50m iron envelope. Twelve GB300 NVL72 racks at a ESTIMATE $4m planning price is $48m of racks and $2m left for storage and IB fabric. N+1 (11 live, 1 spare) supports an ~8,000-seat research service at 50–100 tok/s — not 50,000 seats.
Fifty thousand seats at 5% peak and 50 tok/s is ~70 racks, ~$280m of iron, before the building. Colo, staff, and power sit on top of the $50m. Do not spend the whole envelope on GPUs and then discover networking and colo are still unpaid.
- Planning price / rack
- $4m — until a quote. Loop Capital $3.7–4.0m. EST
- $50m envelope
- 12 racks — $48m iron + $2m storage / IB. EST
- Honest seats
- ~8,000 — @ 50 tok/s, 5% peak, N+1. EST
- IT power, 12 racks
- ~1.6 MW — ~135 kW / rack. Colo on top. EST
- Resilience
- N+1 — 11 live, 1 spare. One rack down, no user-visible loss.
1. Concurrency design
Design at 3% of seats generating; size comfortable at 5%.
- Copilot-class: ~30% of seats touch on a given day.
- Global R&D overlap in one peak window.
- Most of a research session is reading, not generating.
- 50k × 30% daily × ~45% in peak window × ~20% streaming ≈ 1,500 in-flight = 3%.
- Size 5% (2,500 generating) so a 200-person org after a meeting does not queue.
- Agent loops can push 5–8%.
All of the above is ESTIMATE — a utilisation model, not a measured fleet.
2. One GB300 NVL72 rack
- 72 B300s, one NVLink domain. PUBLISHED NVIDIA product page: 72 Blackwell Ultra + 36 Grace, one NVLink-5 domain.
- No NVIDIA list price. Credible $3.7–4.0m (Loop Capital). Supply-chain print $6.0–6.5m. Plan $4m until a quote. Power ~135 kW. ESTIMATE
- K3 packs as 9 replicas of 8 GPUs (72 ÷ 8). Packing math matches the research replica definition; interactivity numbers below are conversation working points.
| Working point (one rack) | In-flight | Seats @ 5% peak | Seats @ 3% peak | Basis |
|---|---|---|---|---|
| People generating at 100 tok/s | 9 | 180 | 300 | EST |
| People generating at 50 tok/s | 36 | 720 | 1,200 | EST |
The 36 figure assumes the published 8-GPU curve still holds ~4 streams at 50 tok/s (9 replicas × 4). That concurrency point is not in the published K3 table — see §6. 50 tok/s is the promise when the rack is busy. 100 tok/s is what a user sees when that replica is not sharing.
Multiply racks to multiply capability if each rack runs its own K3 replicas behind a load balancer. This does not hold if one giant model copy is smashed across racks. The first rack also pays for gateway, weight store, and monitoring.
3. Fifty thousand seats at 5% peak (2,500 generating)
| Target | Racks | Basis |
|---|---|---|
| 50 tok/s | 70 | EST |
| +20% spare | 85 | EST |
| + long-context prefill | ~100 | EST |
| 100 tok/s guaranteed | 280 | EST |
Reasonable buy for full 50k: ~85–100 NVL72s (~6,000–7,200 GPUs, ~11–14 MW, ~$300m of racks). ESTIMATE Do not go below 60 at 3% unless willing to shed load.
4. What $50m of racks actually buys
- At $4m: 12 racks ($48m) and $2m left for storage and IB fabric.
- At a $6m quote: only 8 racks.
- Do not spend the whole $50m on GPUs then discover colo / networking still needed.
12 racks, N+1 (11 live, 1 spare)
| Promise | In-flight | Seats @ 5% | Seats @ 3% | User speed |
|---|---|---|---|---|
| 50 tok/s | 396 | ~8,000 | ~13,000 | 50 guaranteed at peak; 80–100 typical. One rack down, no user-visible loss. |
| 100 tok/s | 99 | ~2,000 | ~3,300 | 100 always. |
All rows ESTIMATE from the 26 Aug conversation.
- Honest offer: 8,000-seat research service at 50–100 tok/s, not 50,000 seats.
- 50,000 seats at 5% and 50 tok/s is ~70 racks, ~$280m of iron, before the building.
- Power for 12 racks ~1.6 MW. Colo, staff, power sit on top of the $50m.
5. Research appendix (26 Aug 2026)
Facts in §§6–14 are only those in research.md / scenarios.csv. No invented tok/s or prices. Two price books are shown on purpose — do not mash them into one fake list price.
6. Kimi K3
| Item | Value | Basis |
|---|---|---|
| Total / active | 2.8T / 104B (tech report also 104.2B active) | PUBLISHED |
| Serve precision | Native MXFP4 weights / MXFP8 activations (QAT from SFT) | PUBLISHED |
| Checkpoint | 1,560.9 GB (1,560,936,091,448 bytes; 1,453.7 GiB) across 96 shards | PUBLISHED |
| Minimum node to hold | 8× B300 / 8× GB300 / 8× MI355X (288 GB class, 2,304 GB). 8× H200 does not fit (~433 GB short). | PUBLISHED |
| Context | 1,048,576 tokens; 896 routed experts (top-16) + 2 shared | PUBLISHED |
vLLM wording: tensor-parallel is the interactivity recipe; wide expert-parallel is the throughput recipe and hurts per-user tok/s. An NVL72 can pack 9 TP8 replicas (72/8) if the fabric is partitioned that way — that changes server/rack count, not GPU count.
7. Published K3 decode — no concurrency>1 at 100 / 200
| Engine | Hardware | Batch | Spec | tok/s per user |
|---|---|---|---|---|
| vLLM | GB300 NVL72, TP8 | bs=1 | none | 111 |
| vLLM | GB300 NVL72, TP16 | bs=1 | none | 118 |
| vLLM | same, TP8 | bs=1 | DSpark (7 draft) | 331 |
| vLLM | same, TP16 | bs=1 | DSpark | 370 |
| SGLang | day-0 kernel ladder | bs=1 | none | ~113 |
| SGLang | same | bs=1 | DSpark | ~423 |
All rows PUBLISHED (vLLM 27 Jul 2026; SGLang day-0). 200 tok/s requires DSpark (or equivalent). Without speculation, 111–118 does not reach 200 at bs=1.
Not published as of 26 Aug 2026: a K3 decode-speed curve vs concurrency that still holds ≥100 or ≥200 tok/s/user. vLLM’s Pareto figure is prose only (“100+ TPS/user” at the low-latency end; “2K+ TPGS” at high-throughput). Therefore conservative max concurrent sequences per replica at those SLAs = 1. No higher-concurrency estimate is made — the research rule is estimate only from a published curve. K3 has none at this SLA.
The conversation’s “~4 streams at 50 tok/s” is a working assumption for a busy-rack 50 tok/s promise, not a published cell. Treat §2’s 36-in-flight / 720-seat row as ESTIMATE until that curve is measured.
Hosted Artificial Analysis (26 Aug 2026 window): K3 comparison median 35.5 t/s; fastest listed provider Makora 162.6 t/s; Parasail 151.5; Modal 148.2. Hosted multi-tenant ≠ dedicated batch-1 replica. PUBLISHED
8. GLM-5.x — largest open-weight is GLM-5.2
The flagship name as of 26 Aug 2026 is GLM-5.3 (Z.ai, 14 Aug 2026). Same 744B / 40B-active MoE base as GLM-5.2; gains are post-training only. Open weights were not public as of 26 Aug (Z.ai: “in two weeks” after 14 Aug; HF zai-org/GLM-5.3 still gated mid-late August). Largest open-weight GLM is GLM-5.2 (MIT, 753B incl. MTP / ~744B backbone, 40B active).
SGLang NVFP4, 8×B200, low-latency + MTP — the only two-point numeric TPOT table:
| Concurrency | TPOT | tok/s per user |
|---|---|---|
| 1 | 1.85 ms | 541 |
| 16 | 5.43 ms | 184 |
So ≥100 is published at 16 concurrent sequences per 8-GPU replica; ≥200 is published only at concurrency=1 (conc=16 is 184, below 200). PUBLISHED SGLang also states 500+ tok/s/user on 8×B300 at bs=1 and 450 on 4×GB300 at bs=1 — no conc>1 table on those exact SKUs.
Hosted AA: GLM-5.3 (Z.ai) 90.0 t/s. PUBLISHED Z.ai’s 14 Aug footnote used 115 TPS (GLM-5.3) and 40 TPS (Kimi K3) from AA at write time.
9. Conservative 25k / 50k / 100k concurrent is not this $50m story
Because the SLA is per-user, replicas = ceil(clients / max_conc_at_SLA), not aggregate tok/s ÷ peak GPU tok/s. For K3, conservative conc = 1 at both 100 and 200 (200 requires DSpark).
| K3 clients | SLA | Replicas | GPUs | NVL72 (÷9) |
|---|---|---|---|---|
| 25,000 | 100 or 200 | 25,000 | 200,000 | 2,778 |
| 50,000 | 100 or 200 | 50,000 | 400,000 | 5,556 |
| 100,000 | 100 or 200 | 100,000 | 800,000 | 11,112 |
PUBLISHED hardware math on a conservative conc=1 reading. That is a national AI-factory, not the $50m / 12-rack envelope. GLM-5.2 at 100 tok/s with published conc=16 is 3–16× smaller in replica count (see §11); the 200 tok/s conservative GLM case returns to the same order of magnitude as K3.
10. Two price books — do not mash
NVIDIA publishes no MSRP for HGX / DGX / NVL72. Two separate books are in the sources. They are not one list price.
| Book | 8-GPU B300 | GB300 NVL72 | Used where |
|---|---|---|---|
| Research bookends (scenarios.csv) | $381k–$650k | $3.0–4.3M | Low: Rillor/Supermicro $381k; Thunder Compute high $650k. NVL72: market low $3.0M; Wolfe ~$4.3M (GPUSmith). |
| Proposal planning price | — | $4m | Loop Capital $3.7–4.0m cited in the 26 Aug conversation. Supply-chain print $6.0–6.5m. Plan $4m until a quote. Used for the 12-rack / $50m arithmetic. |
Other cited 8-GPU points (not used as the CSV bookends): GPUSmith/Turbo Max ~$420k; Exeton “starting at” $795k; Petronella bare baseboard from $485k; DGX B300 ~$400k (GPUSmith citing FT via PCMag). GB200 NVL72 Wolfe ~$3M, market $2–3M. FX used in the CSV GBP columns: 1 USD = 0.7357 GBP (Wise, 26 Aug 2026).
11. Scenario table — conservative K3 and GLM @ 100 only
Rows copied from scenarios.csv. Interpolated GLM-@-200 and 4×GB300 alternate rows are omitted. USD columns are the research bookends in §10, hardware-only, no network / colo / spare.
| Model | Clients | SLA | Conc / replica | Replicas | GPUs | NVL72 | $ 8-GPU $381k–$650k | $ NVL72 $3.0–4.3M |
|---|---|---|---|---|---|---|---|---|
| K3 PUB | 25k | 100 | 1 | 25,000 | 200,000 | 2,778 | $9.53B–$16.25B | $8.33B–$11.95B |
| K3 PUB | 50k | 100 | 1 | 50,000 | 400,000 | 5,556 | $19.05B–$32.50B | $16.67B–$23.89B |
| K3 PUB | 100k | 100 | 1 | 100,000 | 800,000 | 11,112 | $38.10B–$65.00B | $33.34B–$47.78B |
| K3 PUB | 25k | 200 | 1 | 25,000 | 200,000 | 2,778 | $9.53B–$16.25B | $8.33B–$11.95B |
| K3 PUB | 50k | 200 | 1 | 50,000 | 400,000 | 5,556 | $19.05B–$32.50B | $16.67B–$23.89B |
| K3 PUB | 100k | 200 | 1 | 100,000 | 800,000 | 11,112 | $38.10B–$65.00B | $33.34B–$47.78B |
| GLM-5.2 PUB | 25k | 100 | 16 | 1,563 | 12,504 | 174 | $0.60B–$1.02B | $0.52B–$0.75B |
| GLM-5.2 PUB | 50k | 100 | 16 | 3,125 | 25,000 | 348 | $1.19B–$2.03B | $1.04B–$1.50B |
| GLM-5.2 PUB | 100k | 100 | 16 | 6,250 | 50,000 | 695 | $2.38B–$4.06B | $2.09B–$2.99B |
K3 200 tok/s rows require DSpark (331–370 published at bs=1; 200 not reached without spec). GLM @ 100 applies the SGLang B200 NVFP4 conc=16 point (184 tok/s/user) to an 8×B300 replica as the same recipe class. GBP at 0.7357: K3 100k conservative is £28.03–£47.82B at the 8-GPU bookends and £24.53–£35.15B at the NVL72 bookends.
12. What this sizing ignores
Do not treat replica math as a BOM. All of the following are real and not in the GPU-count tables:
- Prefill. Production recipes are P/D disaggregated. A +50% to +200% GPU tax for prefill is typical; not included in the conservative tables. The 26 Aug conversation adds ~15 racks (~100 vs 85) for long-context prefill on the 50k / 50 tok/s path — still an estimate.
- Speculative decoding is not free. DSpark / MTP acceptance collapses on high-entropy tokens (vLLM: 4.73 accept on coding vs 2.61 on creative writing). 331–423 and 541 tok/s are workload-specific.
- Networking / IB fabric. Multi-rack decode needs Quantum-X800 or Spectrum-X 800 Gb/s. Often tens of percent of GPU capex. Not priced. The $50m plan parks $2m for storage and IB — a placeholder, not a quote.
- Spare / N+1 / rolling deploy. The 12-rack offer includes one spare. The conservative factory tables do not.
- Storage. K3 checkpoint 1.56 TB × replica count if local, or a shared parallel FS plus staging NVMe.
- Colo and power. Conversation: ~135 kW / rack, ~1.6 MW for 12, ~11–14 MW for 85–100. Official GPU TDP is up to 1,400 W (NVL72) / 1,100 W (HGX B300). Facility 1.5–2× is the usual next tax and is not in the $50m.
- Software headroom. Published tok/s is a lab point, not p99 under churn.
- Licence. K3 License is not MIT (third-party notes a ~USD 20M MaaS revenue trigger). GLM-5.2 is MIT; GLM-5.3 license unconfirmed until the repo opens.
13. Recommendation (from the 26 Aug research)
Research lean as written on 26 Aug. Item 3 is not a decision — the new model costs all three routes (K3 only / GLM-5.3 only / Hybrid) identically and leaves the choice to the founder.
- Default metal: GB300 NVL72 (or HGX B300 8-GPU if racks slip). Do not buy B200 as the long-term default for K3.
- Productize two SKUs: (a) interactive 100–200 tok/s, concurrency 1–16, TP8/TP16 + spec decode; (b) batch / throughput wide-EP. Do not promise (a) economics on (b) hardware efficiency.
- Treat GLM-5.2 / 5.3 as the volume model (3–16× fewer replicas at 100 tok/s) and Kimi K3 as the premium flagship.
- Re-measure the missing K3 concurrency curve and the GLM B300 curve before ordering. The entire 200 tok/s GLM estimate in the full memo hangs on two SGLang B200 cells; it is omitted from §11 on purpose.
- Budget prefill + fabric + power + spare as a separate line at least as large as the decode-GPU line.
- On a $50m envelope: sell the 8,000-seat service, or raise the iron budget toward ~70–100 racks for 50k seats. Do not claim 50k seats from 12 racks.
14. Sources
URLs as cited in the 26 Aug research memo. No fabricated citations.
- Kimi K3 model card — huggingface.co/moonshotai/Kimi-K3
- Kimi K3 tech report — arxiv.org/html/2607.24653v1
- vLLM K3 day-0 (111 / 118 / 331 / 370) — vllm.ai/blog/2026-07-27-k3
- SGLang K3 day-0 (~113 / ~423) — lmsys.org/blog/2026-07-27-kimi-k3-day0-support
- Checkpoint byte sum — particula.tech K3 GPU sizing
- GLM-5.2 blog — z.ai/blog/glm-5.2
- GLM-5.3 blog — z.ai/blog/glm-5.3
- SGLang NVFP4 TPOT table (541 / 184) — sglang PR 29674
- AA GLM-5.3 providers — artificialanalysis.ai/models/glm-5-3/providers
- AA Kimi K3 providers — artificialanalysis.ai/models/kimi-k3/providers
- AA GLM-5.3 vs K3 — comparisons/glm-5-3-vs-kimi-k3
- NVIDIA GB300 NVL72 — nvidia.com/.../gb300-nvl72
- 8-GPU B300 tape — rillor.com/skus ($381k–$528k)
- B300 street / cloud — thundercompute.com/blog/nvidia-b300-pricing ($550k–$650k)
- NVL72 Wolfe / market — gpusmith.com DGX vs HGX vs NVL72 (GB300 ~$4.3M Wolfe; market $3–4M)
- USD/GBP — wise.com USD-GBP 26 Aug 2026 (0.7357)
Loop Capital $3.7–4.0m and the supply-chain $6.0–6.5m print are from the 26 Aug conversation, not from a URL collected in the research memo. They are not given a fake link here.
Diligence brief dated 26 August 2026 (Asia/Dubai). Web edition 27 August 2026. Prepared for internal review of an inference-serving hardware envelope. Numbers that are not in the 26 Aug conversation, research.md, or scenarios.csv are not in this brief.
This is a planning note, not a quote and not a BOM. Allocation, colo power, and a measured K3 concurrency curve must exist before a capital commit.
A3 · Model-comparison sources (27 Aug session)
- AA head-to-head (Intelligence Index 60 = 60; hosted speed/price) — artificialanalysis.ai/models/comparisons/glm-5-3-vs-kimi-k3
- Kimi K3 License clause ($20m aggregate MaaS trigger) — simonwillison.net/2026/Jul/27/kimi-k3 · implicator.ai
- GLM-5.3 gated-weights placeholder ("expected 28 Aug 2026") — huggingface.co/zai-org/GLM-5.3; GLM-5.2 MIT card — huggingface.co/zai-org/GLM-5.2; NVIDIA NVFP4 checkpoint — huggingface.co/nvidia/GLM-5.2-NVFP4
- GLM-5.3 benchmark chart (as reported) — marktechpost.com · explainx.ai
- Serving recipes — vllm.ai/blog/2026-07-27-k3 · lmsys.org GLM-5.2 optimisation · sglang PR 29674
- Enterprise adoption of Chinese open weights — computing.co.uk (Vercel gateway data) · cnbc.com
- Gulf sovereignty precedent — iiss.org strategic comment
- Wider-field landscape (§04 table, verified 27 Aug) — Qwen3.8-Max: huggingface.co/Qwen/Qwen3.8-2.4T-A95B · community NVFP4; DeepSeek V4 Pro: huggingface.co/deepseek-ai/DeepSeek-V4-Pro · AA model page · SemiAnalysis GB300 rack-scale day-0; field: AA open-source leaderboard · AA Mistral Large 3 · AA MiniMax M3 · openai.com gpt-oss. Full list: work/06c-extended-model-landscape.md