Decision brief · 27 Aug 2026 · Private & confidential

Private inference for one industry: what £50m buys, what to charge, and the path to 50,000 seats

Open-weight LLMs served on our own GB300 NVL72 racks, single-tenant, no training on client data. The founder has not chosen a model: Kimi K3, GLM-5.3 and a Hybrid of both are costed identically as Routes 1 / 2 / 3, and the route choice turns out to be the capacity choice. This brief ties the financial model to the diligence work and answers the questions the founder and his investors will ask first.

Prepared for the founder Model Inference-DC-Financial-Model.xlsx Engine 07-model-calc.py → 07-base-results.json Horizon 36 months, monthly Scenarios Conservative / Base / Upside
FX 1 USD = 0.7357 GBP · Wise, 26 Aug 2026 · £50m = $67.96m · Conservative stacks GBP −10% (0.809)
Published lab or vendor figure, URL in appendix Estimate derived from published points, method stated Assumption our input; changes the answer if wrong

Revised 27 Aug after three commissioned adversarial reviews (technical · financial · buyer-side); every accepted fix is live in the engine and workbook — decisions in work/14-fix-log.md. How to read this: every number traces to 07-model-spec.json / 07-base-results.json; where the Excel is referenced, the sheet is named. Nothing on this page is invented outside those files and the sources in the appendix.

§ 01

The answer

Viable on one route, not three. On the base case, one GB300 NVL72 rack serves 2,194 seats on GLM-5.3, 1,095 on the 35% Hybrid, 638 on Kimi K3 — the model choice is a capacity choice. £50m buys 12 racks of iron and nothing else; the fundable build is 6 racks plus contract-triggered leases. Every number below is the post-adversarial engine (three hostile reviews, all accepted fixes live in the model — §09 lists the verdicts).

Recommendation
Option C6 racks
£24.6m / $33.4m day one
Launch Route 2 pure (GLM-5.3, MIT); lease-financed 3-rack tranches to 15 racks by M33; K3 later as a metered opt-in, only after Moonshot terms in writing
£50m buys
12racks
£46.7m / $63.5m all-in
and nothing else — Option A is cash-negative at delivery (M4) on every route; needs ~£57m Est
Seats · 12 racks · ≥60 tok/s
24,131GLM-5.3
Hybrid 12,045 · K3 7,020
11 live racks, 6% peak, 10% premium mix; per live rack 2,194 / 1,095 / 638 Est
Charge (list)
$150/ $300
£110 / £221 per seat-month
net blended $140 = 2.5× ChatGPT Enterprise; full-cost list required at the M36 fleet: GLM $144 · Hybrid $298 · K3 $512 Asm
Break-even & cash
M17EBITDA
cash low £7.6m / $10.4m (M23)
GLM-5.3 · Option C · Base; peak funding gap £0; zero-revenue runway 33 months; survives a $6m rack quote (low £8.4m)
Racks for 25k / 50k seats
13 / 25GLM-5.3
£50m / £95m iron
Hybrid 25 / 50 · K3 43 / 86, incl. spares — 25k firm inside this raise; 50k is the Phase-3 / Series-B option Est
Recommended path

Build 6 racks now (£24.6m / $33.4m all-in) on GLM-5.3, Route 2 pure — 5 live racks serve ~11,000 seats — and gate every further tranche of 3 leased racks on the anchor’s signed take-or-pay.

On that plan the model reaches ARR $7.0m / $18.3m / $36.0m / $48.8m (£5.2m / £13.4m / £26.5m / £35.9m) at M12 / 18 / 24 / 36, EBITDA break-even at M17, and cash never below £7.6m — the peak funding gap is £0: the plan never needs money beyond the £50m. If all four lease tranches fire, the aggregate obligation is $39.1m principal (~$1.3m / £1.0m per month at full service, 36-month terms) secured on the anchor contract, plus $3.9m of lessor deposits — all inside the cash path shown in §07. The recommendation is a view, not a verdict on the models: §04 shows exactly what flips the route — a measured K3+DSpark curve at ≥12 streams/replica at 100 tok/s, or an anchor paying ≥$300/seat for K3 with Moonshot terms signed.

§ 02

The question, and the honest constraint

The founder’s question is “can £50m build a private LLM service for a 50,000-seat anchor?” The honest constraint: 12 racks is not 50,000 seats on any route — seats convert to racks through peak concurrency, and the per-replica stream count each model holds at the SLA.

Every seat number on this page comes from one chain. Licensed seats are not concurrent users: on a Copilot-class service ~30% of seats touch it daily, under half of those overlap the peak window, and a fifth of those are actually streaming tokens at any instant — roughly 3% in-flight today. But the demand file’s own agent-share ramp reaches ~6.5% by M36, so the model sizes at 6% (the M36 agentic point) and protects it with contractual per-seat stream caps: 2 concurrent streams per volume seat, 5 per premium. Est source brief §1 · 01-demand Little’s law (1.6% hour-avg M12 → ~6.5% M36) · 12-adversarial A1

Tax factor 0.39 = (1 − 25% p99 headroom, incl. rolling-upgrade drain) × 0.65 long-context factor ÷ (1 + 25% prefill tax at ≥85% prefix-cache hit). Seats per live rack = usable streams ÷ 6% = 2,194 (GLM-5.3) / 1,095 (Hybrid 35%) / 638 (K3). These cells were re-based as one bundle after adversarial review — the old +100% prefill tax was ~4× too high and was quietly compensating for optimism elsewhere; never relax one cell without the others. 07-model-spec.md §0, §2.1 · 12-adversarial A7 · Excel: Capacity sheet

£50m is everything, not an iron budget

Racks, fabric, storage, colo, power, staff, software and runway all come out of the same £50m. Spending £46.7m on 12 racks (Option A) leaves the company insolvent at delivery — cash goes negative at M4 on every route and the plan needs ~£57m. Applying the reserve rule (cash ≥ £4m after 24 months of zero revenue), £50m safely carries 6 racks; with base-case revenue it carries 9. 07-base-results.json runs["base|route2|A"] · racks_buyable

Where the source brief’s “~8,000 seats” went

The 26 Aug diligence brief sized 12 racks at ~8,000 seats — that was K3 at 4 streams/replica with no prefill tax. The full chain (02-capacity), anchored to GLM’s published conc=16 point and with every tax applied, gives 7,020 seats on K3 and 24,131 on GLM-5.3 for the same 12 racks. Same iron, ~3.4× the service — that is why §04 exists.

§ 03

Capacity: seats vs racks, route by route

The stream count per 8-GPU replica is published at one point for GLM-5.3 (16 streams at 184 tok/s) and only at batch-1 for K3. Everything else is extrapolation, tagged as such — and measuring it is a condition precedent of the raise, not a nice-to-have.

Naming note: GLM-5.3 open weights were gated as of 26 Aug 2026; GLM-5.2’s weights are the same 744B / 40B-active architecture, so the published decode curve and the capacity maths carry over. Pub source brief §8 · huggingface.co/zai-org/GLM-5.3

Seats and racks by route Base scenario: 6% peak, 9 TP8 replicas/rack, tax factor 0.39, 10% premium mix. Streams cells show Conservative / Base / Upside.
Metric Route 1 · Kimi K3 only Route 2 · GLM-5.3 only Route 3 · Hybrid (35% K3)
Streams per replica at ≥60 tok/sthe number the volume SLA hangs on4 / 12 / 24 Est30 / 40 / 56 Estblended 20.2 (base)
Streams per replica at ≥100 tok/sthe premium SLA — K3 sells as “target 100 / floor 60” until measured1 Pub / 6 / 12 Est16 Pub / 24 / 32 Estmix of both
Published anchor point111–118 tok/s at bs=1; 331–370 with DSpark; no conc>1 curve Pub184 tok/s/user at conc=16, 8×B200 NVFP4 Pubinherits both
Seats per live rack (base, 6%)6382,1941,095
Seats, 12 racks (11 live), 10% premium mix7,02024,13112,045
… all seats at ≥60 tok/s7,72225,74013,106
… all seats at ≥100 tok/s3,86115,4446,969
Racks for 25,000 seats (incl. spares)431325
Racks for 50,000 seats (incl. spares)862550
Iron for 25k / 50k, all-in$219m / $434m£161m / £319m$68.5m / $129m£50m / £95m$129m / $254m£95m / £187m
Scenario range, seats per live rackConservative / Base / Upside81 / 638 / 3,415727 / 2,194 / 8,154131 / 1,095 / 6,063

On the published-only floor (GLM conc=16, nothing extrapolated) 11 live racks are a 10,296-seat service — and that floor now carries a full named P&L in §07: it dies at M30 and needs ~$349 list. Dotted underline = row-wise best number, not a recommendation. 07-model-spec.json results.capacity_table, twelve_racks_serve_base, named_runs · 02-capacity §6 · Excel: Capacity + Scenarios sheets

Racks required by seat target

Same SLA (≥60 tok/s at 6% peak), same taxes, same iron — only the model changes. Includes spares.

07-model-spec.json results.capacity_table (racks_for_25k / racks_for_50k) Estimate
Measure before ordering — now a condition precedent of the round

K3’s base cell (12 streams/replica at ≥60 tok/s) is a bandwidth model, not a benchmark; its ≥100 tok/s floor is 1. GLM’s base of 40 is a no-spec-decode fit plus a 1.2–1.3× MTP retention haircut — 02-capacity’s own caveat that the linear fit turns optimistic above c≈32 is load-bearing. The four one-day benchmarks — GLM-5.3 at conc 32 / 48 / 64 × 8K / 32K ISL, K3+DSpark at conc 4 / 8 / 12, prefix-cache hit rate on a real anchor trace, and a written rack quote — are closing conditions of the raise, and the anchor term sheet carries a capacity acceptance test: the measured curve re-baselines seats-per-rack. 02-capacity §2b, §9 · 12-adversarial A2 · 13b buyer review · sensitivity grid 2, §08

§ 04

Which model: Kimi K3, GLM-5.3, or both?

Three routes, costed identically end-to-end. The founder chooses; this section shows where each route wins, what each costs, and precisely what would flip the decision. Nothing here says a tier “has to be” a given model.

Route comparison — Base scenario, Option C build Identical racks, taxes, staffing and pricing discipline on all three. Estimate unless tagged.
Criterion Route 1 · Kimi K3 only Route 2 · GLM-5.3 only Route 3 · Hybrid (35% K3)
Answer quality / brandAA Intelligence Index; benchmark splitAA 60 — #1 open on AA; GPQA 93.5, native vision PubAA 60 — tie; leads Terminal-Bench 3.0, GDPval Elo; flagship is text-only Pubbest-of-both per task
Seats per live rack (≥60, 6%)6382,1941,095
Blended list price$190£140 (incl. $25 K3 uplift)$165£121$190£140 (incl. $25 K3 uplift)
Price required at fleet-full, cash / full costlist $/seat-mo at which EBITDA / EBIT = 0$400 / $512$117 / $144$233 / $298
Full-cost floor at 25k / 50k seats$330 / $313$123 / $103$206 / $189
Option C: seats supportable M367,02030,71212,045
Option C: ARR M12 / 18 / 24 / 36, $m6.5 / 10.3 / 14.2 / 12.57.0 / 18.3 / 36.0 / 48.88.1 / 17.7 / 24.4 / 21.5
Option C: seats turned away at M36the route’s growth ceiling26,9803,28821,955
Option C: EBITDA break-evenneverM17M16
Option C: cash low (month)−£9.1m — dead M26+£7.6m (M23)+£3.8m (M36) — brushing the £4m floor
Gross margin M36, cash / GAAP13% / −103%72% / 35%49% / −18%
LicenceKimi K3 License — not open source; separate Moonshot agreement once aggregate MaaS revenue > $20m T12M; the C fleet’s capacity ceiling keeps it un-triggered in-horizon — any scale-up fires it PubMIT (5.2); 5.3 unstated until repo opens Pubclause reads aggregate revenue — any K3 share carries the same single-counterparty negotiation Pub
Measurement riskwhat the seat counts rest onno published conc>1 curve at all; DSpark needed above ~118 tok/spublished to conc=16; extrapolated 16→40 (floor P&L named in §07)both of the above + routing layer
Ops burdenone modelone modeltwo eval suites, two templates, ~+2 FTE, 7.5% split-pool penalty

Dotted underline = row-wise best number, never a recommendation. Quality, safety and the anchor’s own eval set are outside this table. 07-model-spec.json results.route_comparison["base|C"] · 06b-model-comparison · Excel: Scenarios sheet (routes side by side)

The decision criteria

The flip chart: what a K3 traffic share costs

Hybrid K3-share sweep — Base scenario, Option C Share = % of peak in-flight streams routed to K3, enforced by router policy + per-seat K3 allowance — not a free model picker. 0% and 100% reproduce Routes 2 and 1.
K3 share of streams0% (= GLM + uplift)10%20%35% (base)50%100% (= K3)
Seats per live rack2,1941,6321,3641,095915638
Seats at M36 (fleet the trigger builds)37,29422,84115,00612,04510,0607,020
ARR M36, $m62.8*40.826.821.518.012.5
Cash low, £m8.97.55.33.8−1.1 (dead M34)−9.1 (dead M26)
Full-cost list price required, $136193248298357512
GAAP GM M3641%22%4%−18%−42%−103%

*The 0% row carries the $25 K3-access uplift on GLM-only seats; the extra revenue also unlocks the 4th tranche in-horizon (17 live racks vs 14), hence above Route 2’s $48.8m. A free model picker drifts K3-ward — every seat paying a flat uplift pulls toward the 50% attractor, which is cash-negative on this sweep — so the share must be enforced (router policy + per-seat allowance, or metered K3 pricing), and split pools pay a 7.5% quantisation penalty while both exist. Even the 10% point is marginal, not comfortable: $193 required vs $190 charged, still carrying the Moonshot aggregate-revenue surface. 07-model-spec.json results.hybrid_k3_share_sweep_base (C rows) · 12-adversarial A6 · 06b §8 · Excel: Scenarios sheet, K3-share input

What flips the route toward K3

(a) A measured K3+DSpark curve at ≥12 streams/replica at 100 tok/s — in the Upside capacity column the K3 routes out-earn GLM ($102m vs $81m ARR M24) because the uplift and premium price carry it once capacity stops binding; until that bench exists, no deck slide may show K3 out-earning GLM. (b) The anchor pays ≥$300/seat for K3 across the estate and Moonshot terms are signed in writing. (c) The anchor’s eval set is vision-heavy — GLM’s flagship cannot see images.

What locks in GLM-5.3

(a) Counsel confirms the $20m trigger is on aggregate revenue and Moonshot won’t pre-agree terms. (b) A measured 8×B300 GLM curve holding ≥40 streams at ≥60 tok/s. (c) Procurement requiring OSI-approved licences (MIT passes; the Kimi K3 License does not). If GLM-5.3’s weights stay gated past the procurement window, the fallback is GLM-5.2’s identical-architecture weights, not K3.

The wider field: Qwen, DeepSeek and the rest — verified 27 Aug

The founder asked about Qwen3.8-Max and DeepSeek V4 Pro by name; a buyer will too. Verdict up front: nothing in the field flips the route — but two entries strengthen the plan and one is a trap. Because the fleet is model-agnostic iron, every candidate maps onto one of the two capacity classes the engine already prices (GLM-class ~40 streams/replica, K3-class ~12, at ≥60 tok/s) — so “which models will his clients want?” is a menu question, not a capex question: adding one is a weights swap plus an eval pass, not new iron.

Open-weight field beyond K3 and GLM-5.3 Verified 27 Aug 2026 (URLs in appendix A3). Streams/rack = est. usable streams per live rack at ≥60 tok/s by class mapping Est; specs Pub. Reference points: GLM-5.3 AA 60 · ~130 streams/rack; Kimi K3 AA 60 · ~32.
ModelLicenceTotal / activeFits 8×B300 replica?Class · streams/rackAAVerdict
Qwen3.8-Max (2.4T-A95B, CN) · open ckpt is text-onlycustom “qwen3.8-max”, unreviewed2.4T / 95BNo at BF16/FP8 (4.8 / 2.4 TB vs 2,304 GB); community FP4 onlyK3-class · ~3258Near-peer to K3 on paper; blocked by fit, licence and no vision in the weights. On-request only — never a committed tier.
DeepSeek V4 Pro (0813, CN)MIT1.6T / 49BYes (~893 GB)GLM-class · ~11053Best licence + the only published GB300 rack-scale curve of any open model; 7 AA points under GLM-5.3. Menu slot 3.
GLM-5.3-Flash (CN) · visionMIT320B / 18BYesGLM-class · ~13057Vision at MIT in the volume class; engine support immature — add when vLLM/SGLang mainline it.
Mistral Large 3 (EU)Apache 2.0675B / 41BYesGLM-class · ~13016The non-PRC fallback on paper; quality nowhere near the class bar.
MiniMax M3 (CN)community, $20m revenue trigger428B / 23BYesGLM-class · ~13045Dominated: K3’s licence risk at GLM-minus-15 quality.
gpt-oss-120b (US)Apache 2.0117B / 5.1BYesGLM-class · very high24Only US open model with real serving footprint; a utility/draft model, not a seat model.

Also checked: MiniMax M2.7 (licence is non-commercial — excludes MaaS outright), Llama 4 Maverick (Apr 2025, superseded), Meta’s Muse Glimmer 30B (AA 35, a distillation), and Muse Spark 1.2 (API-only — do not plan around it). 06c-extended-model-landscape.md · landscape verified via Artificial Analysis, Hugging Face model cards, vendor blogs, 27 Aug 2026

The model menu — the answer to “we don’t know what clients want”

Contract tok/s SLAs and capacity classes, never model names, and publish a menu: 1) GLM-5.3 (volume default) · 2) Kimi K3 (premium/vision, metered, post-Moonshot-terms) · 3) DeepSeek V4 Pro (MIT understudy, kept hot — also the honest answer when procurement asks “what if Z.ai is a problem”, though it shares the CN origin) · 4) Qwen3.8-Max FP4 on request, after licence review and an 8×B300 benchmark · 5) GLM-5.3-Flash when engines mature. Marginal cost per added model is eval and ops surface, not iron: ~0.4–1.6 TB storage per checkpoint (noise against the storage line), a one-day rented benchmark, an anchor-eval + red-team pass, and gateway integration — ~2–4 weeks, ~1 FTE-month. Menu share is a weekly replica rebalance behind one gateway, not a build. Est 06c §3b

§ 05

Where the money goes

A rack is $5.01m all-in, not $4.0m — fabric, spares, install and contingency add 25%. Build options differ only in how many racks are bought with equity on day one, and that is the difference between insolvency and a 33-month runway.

Capex bill of materials, Base USD; £ = × 0.7357. Per GB300 NVL72 rack unless marked one-off.
LineUSDGBP
GB300 NVL72 rack Est$4.00m£2.94m
IB fabric — leaf + optics (8%)$320k£235k
Spares (1.5%)$60k£44k
NVMe staging$30k£22k
Shipping & insurance$45k£33k
Install & commissioning$75k£55k
Software, year 1 (OSS stack)$25k£18k
Per-rack all-in, + 10% contingency$5.01m£3.69m
Fixed site, one-off (spine, parallel FS, gateway, spares kit), + 10% Est$3.36m£2.47m
Day one — C: 6 racks$33.4m£24.6m
Day one — B: 8 · A: 12$43.4m · $63.5m£32.0m · £46.7m
Opex run-rate, Base · GLM-5.3 · Option C Cash COGS + opex + lease, per month. UK site; import duty 0% (HS 8471).
MilestoneUSD / moGBP / mo
M12 · 6 racks, 21 heads$1.14m£0.84m
M24 · 12 racks, 31 heads$2.32m£1.70m
M36 · 15 racks, 41 heads$3.04m£2.24m
Colo + power per rack-month (140 kW committed) Est$60.4k£44.4k
Depreciation per rack-month (4-yr, 10% residual)$96.1k£70.7k

Monthly cost stack at M36

Base · GLM-5.3 · Option C, 15 racks. Cash costs + lease; depreciation excluded.

07-base-results.json runs["base|route2|C"].monthly (m=36 row) · Excel: Opex + Monthly sheets
Staff plan, compact UK loaded cost = base salary × 1.25. Heads at end of Q4 / Q8 / Q12; +4 ops heads when the fleet passes 12 racks. SRE and security pulled forward ~2 quarters after the buyer-side review; on-call paid from Q3.
RoleLoaded £/yrQ4Q8Q12
CEO / founder · CTO£188k · £225k222
Inference / platform engineering£163k578
ML engineering£138k245
SRE / on-call (5-primary rota through the pilot SLA window)£131k566
Security / compliance (1 head from Q1 — buyer questionnaires precede the pilot)£125k223
Sales / CS / solutions£150k368
Finance / ops · DC ops£106k · £81k245
Total heads213137

All roles Est from Glassdoor / Robert Half bands (04-opex §2). SRE and DC-ops people sit in opex, not COGS — gross margins on this page exclude them from COGS by definition. New opex line after the buyer review: model assurance £150k/yr (per-checkpoint red-team, model bill of materials, egress audit). Excel: Opex sheet, Staff sub-table · 13b-adversarial

The $6.0m rack print

A $6.0m supply-chain quote instead of $4.0m kills Option A at M4 and Option B at M7 on every route. Option C on GLM-5.3 survives it with an £8.4m cash low — the forward-looking gate defers tranches until cash clears (ARR M24 $19.3m instead of $36.0m); C on the Hybrid survives at £2.6m; C on K3 dies at M29. Do not commit an anchor seat count before a written rack quote exists. 07-model-spec.json results.stress_6m_rack_base · §08 grid 1

§ 06

Pricing: what the market charges, what he must charge

The privacy premium is real but bounded: $150 volume / $300 premium list is 2.5× ChatGPT Enterprise. On GLM-5.3 that price clears the full-cost floor with ~15% headroom; on the 35% Hybrid it does not; on K3-only it is not close.

Per-seat price book, $ per seat-month list £ pair beneath at 0.7357. Non-anchor net = list × (1 − 15%); the anchor pays max(15%, seat ladder −10/15/25% at ≥5k/10k/25k seats). Escalator 0% base — token prices fell ~10× in two years.
LineConservativeBaseUpsideBasis
Market anchor: ChatGPT Enterprise Pub~$60~£44the seat the buyer already knows
Volume tier, ≥60 tok/s Asm$120£88$150£110$175£1292.0 / 2.5 / 2.9 × ChatGPT Ent.
Premium tier, ≥100 tok/s Asm$250£184$300£221$400£2942× volume; 10% of seats
K3-access uplift, every seat (Routes 1 & 3) Est$0$25£18$50£37public-API token delta K3 vs GLM; the metered-overage design replaces it if K3 ships later
Blended list (90/10 mix + uplift)$133R2 $165 · R1/R3 $190R2 $198 · R1/R3 $248net base: $140.25 blended (R2)
Anchor net at 25k contracted seats (ladder −25%)~$124~£91essentially at the 25k full-cost floor ($123) — the ladder is the negotiation
Full-cost list required at the M36 fleet — GLM-5.3$144£106 (cash $117)list price at which EBIT after interest = 0
… Hybrid 35% · Kimi K3 only$298 · $512same definition, same build
Overage (info)vol $2/$6 · prem $4/$18 per M tok in/out~6% of seat revenue at base

Revenue also carries SLA credits at 1.5% as contra-revenue (availability ≥99.9% + tok/s p95 over 5-minute peak windows, buyer-run probes as evidence, credit ladder 5/15/30%), and a DSO input: base = anchor billed quarterly in advance as a contractual condition; plain net-60 is the Conservative and would strand ~$8m in receivables at the trough. 07-model-spec.md §1.6 · results.runs_summary price_required per route · 13/13b-adversarial · Excel: Revenue sheet

The price ladder: market, our book, and each route’s floor

$ per seat-month, list. A route whose floor sits above the book cannot serve the anchor at this price.

Floors = full-cost price required at fleet-full (Option C, Base). 07-model-spec.json results.runs_summary["base|route{1,2,3}|C"].price_required.full_cost_list_usd
The capacity-block offer for the anchor

Sell the anchor a 3-year take-or-pay for ≥10,000 seats at $140 net — $50m / £37m TCV — with 25% of year-1 value prepaid at signature, billed quarterly in advance, with the capacity acceptance test attached. In the model the M12 prepay is $4.62m / £3.40m (it caps at ordered capacity); that plus reserve funds the equity slice of the first 3-rack tranche outright (£3.9m), and the contract — assigned to the lessor, with parent guarantee — is the collateral for the 65% advance. Unit check: one GLM-5.3 rack at fleet-full bills ~$308k/mo against $60k colo + $108k lease + $6k support — ~57% fill covers cash cost. 07-model-spec.json results.phase2_arithmetic_base · Excel: Revenue + Monthly sheets

The negotiation corridor, stated openly

A prepared buyer team will find the floor from the sensitivity grid regardless, so present it: open at $165–175 volume list with the published seat ladder; expect the buyer to target $120–130 net blended. The model survives the corridor — at a $120 volume list the cash low is +£0.8m, which breaches the plan’s own £4m floor rule and is therefore the walk-away; $100 needs new money (−£5.3m). Signature at M15 instead of M12 (the buyer’s realistic calendar) moves break-even to M20 with a £6.1m low — survivable, and shown in §07 before diligence finds it. §08 grid 1 · results.named_runs buyer_realistic · 13b-adversarial

§ 07

36-month P&L and cash

Base case, GLM-5.3, Option C: EBITDA break-even at M17, cash bottoms at £7.6m in M23 (the only trough — tranche 3 waits behind the forward-looking gate until it clears), and ends at £10.8m with the fleet 100% full. Peak funding gap: £0.

Revenue vs cash cost run-rate, monthly

Base · GLM-5.3 · Option C. Cost = cash COGS + opex + lease payments. £ per month. The M36 revenue step-down is the anchor crossing 25k contracted seats onto the −25% ladder tier.

07-base-results.json runs["base|route2|C"].monthly · Excel: Monthly sheet

Closing cash by route, Option C

Base scenario. Same build, same prices — only the model changes. £m.

Kimi K3 goes through zero at M26; the Hybrid grinds down to £3.8m; GLM-5.3 troughs at M23 and recovers. runs["base|route{1,2,3}|C"].monthly.cash
The defended case: GLM-5.3 × Option C, Base What we would put in front of a sceptical CFO. Excel: Summary sheet.
MetricM12M18M24M36
Racks total / live6 / 59 / 812 / 1115 / 14
Seats served / demand4,000 / 4,00010,400 / 10,40020,500 / 20,50030,712 / 34,000
Seat-fill of capacity36%59%85%100%
ARR$7.0m£5.2m$18.3m£13.4m$36.0m£26.5m$48.8m£35.9m
EBITDA, monthly−$553k+$193k+$1.34m+$2.00m
Gross margin, cash / GAAP20% / −88%58% / −2%70% / 30%72% / 35%
Closing cash£19.8m$27.0m£11.8m$16.0m£7.8m$10.6m£10.8m$14.7m
Heads · anchor share of seats21 · 100%27 · 77%31 · 78%41 · 81%

Year totals, $m: Y1 revenue 1.8, EBITDA −7.6, capex 36.8 · Y2 revenue 21.0, EBITDA +4.3, capex 9.4, lease 4.5 · Y3 revenue 45.0, EBITDA +23.4, capex 7.7, lease 9.1. Expansion tranches: PO M12 / 16 / 29 / 35, live M16 / 20 / 33 / 39 (the 4th serves beyond the horizon), each 3 racks = $15.0m installed, $9.77m leased (65% at 12%, 36 months), $5.3m / £3.9m equity + $0.98m / £0.72m lessor deposit at PO. The M20→M29 gap is the forward-looking cash gate holding tranche 3 — expansion waits for cash, not the other way round. All four tranches together: $39.1m principal, ~$1.3m/mo aggregate service, $3.9m deposits. M36 demand runs 3,288 seats ahead of capacity — the growth case for tranche 5, not a Phase-1 problem. runs["base|route2|C"] year_totals, racks.expansions

Build options A / B / C — Base scenario, GLM-5.3 route Same demand, same prices. The only variable is how much iron is bought with equity on day one.
MetricA · 12 racks all-inB · 8 racks + runwayC · 6 racks + lease expansion
Day-one capex£46.7m$63.5m£32.0m$43.4m£24.6m$33.4m
Final fleet (racks) / seats supportable12 / 24,13114 / 28,51915 (18 ordered) / 30,712
EBITDA break-evenM19M17M17
Cash low (month)−£6.8m (M21)+£9.0m (M34)+£7.6m (M23)
Runway with base revenueinsolvent at delivery (M4) — needs ~£57m>36 months>36 months
Zero-revenue runwaydead M4dead M23dead M33
ARR M24 / M36$36.0m / $37.4m (capped)$27.0m / $44.9m (tranche fires M31)$36.0m / $48.8m
Scenarios — GLM-5.3 route, Option C Conservative is every adverse input at once — including GBP −10% (FX 0.809) and plain net-60 billing — not a forecast; the grids in §08 isolate single inputs.
MetricConservativeBaseUpside
Seats per live rack7272,1948,154
EBITDA break-evennever (in 36)M17M8
Cash low−£25.5m — dead M14, gap £25.5m+£7.6m (M23)+£27.8m (M7)
ARR M24 / M36$3.9m / $3.9m$36.0m / $48.8m$81.2m / $151.2m
K3-route comparisondead M14, gap £29.9mdead M26, needs $512 seatsK3/Hybrid out-earn GLM ($102m ARR M24) if the curve measures at 24/12

Conservative £ figures are at its own stacked FX of 0.809. 07-model-spec.json results.runs_summary (27 combinations) · Excel: Scenarios sheet shows Base × three routes side by side, then Conservative and Upside

Named stress runs beside the defended case All Base · GLM-5.3 · Option C. These are the runs diligence would otherwise build itself — shown before it has to.
RunWhat it holdsBreak-evenCash lowARR M36Verdict
Defended baseeverything in §A1M17+£7.6m (M23)$48.8mgap £0
Published-only floorGLM held at the one PUBLISHED point (conc=16 both tiers) → 936 seats/live racknever−£4.2m — dead M30$16.0mneeds ~$349 list — why the benchmark is a condition precedent
Buyer-realisticsignature M15, schedules +3 months, prepay released in thirds (signature / acceptance / SOC 2 Type II)M20+£6.1m (M26)$42.4mgap £0 — the buyer’s calendar does not kill the plan
FX stressGBP at 0.78 / 0.809 all 36 monthsM15+£6.2m / +£5.2m$48.8mgap £0; mitigation: forward-buy the initial PO’s ~$33m at close

07-base-results.json named_runs · Excel: Scenarios sheet, named-runs block

§ 08

Sensitivity

Two grids decide the build: rack price × seat price sets whether Option C survives, and concurrency × measured streams sets how many seats a rack really is. Both are pasted from the engine; the Excel Sensitivity sheet carries all three grids per route.

Cash low point in £m by rack price and volume seat price, GLM-5.3 route, Option C
Cash low £m · rack ↓ / volume list $ →$100$120$150$175$200
$3.7m−1.84.310.112.113.7
$4.0m−5.30.87.69.711.4
$4.5m−11.2−5.03.45.67.4
$5.0m−17.2−10.8−1.81.53.3
$6.0m−29.0−22.6−13.4−6.8−4.8
insolventfunded · outlined = base

GLM-5.3 · Option C · tranche months held fixed at the base plan so the grid is monotonic. Any negative cell needs new money. Premium = 2× volume throughout. The live model does better at the $6m print (£8.4m low, §05) because its forward-looking gate defers tranches — this grid deliberately does not. 07-model-spec.json results.sensitivity_base.route2.rack_price_x_seat_price (C rows)

Seats per live rack by peak concurrency and measured stream multiplier, GLM-5.3
Seats / live rack · conc ↓ / measured streams →0.5×0.75×1.0×1.25×1.5×
3%2,1943,2914,3885,4846,581
4%1,6452,4683,2914,1134,936
5%1,3161,9742,6323,2913,949
6%1,0971,6452,1942,7423,291
8%8231,2341,6452,0572,468
fewermore seats · outlined = base

The “measure before you buy” grid: multiplier scales all stream cells (the rented-node benchmark lands somewhere on this axis); concurrency is the agent-workload axis — base now sits at 6% with contractual per-seat caps. Even at 0.5× and 8% concurrency a GLM rack (823 seats) outserves the K3 base (638). results.sensitivity_base.route2.concurrency_x_streams_mult

§ 09

The path, the investor narrative, and the risks

Three phases, each gated on something signed or measured — no phase starts on a forecast. The story for the next investor is contracted ARR against leased iron, not GPUs on a balance sheet.

  1. Phase 1 · M1–M9
    Stand it up, prove it
    6 racks ordered M1, delivered M4, live M7
    £24.6m / $33.4m all-in · colo signed M1, power-on M6
    pilot on rented nodes M1–M6; ISO 27001 issued M9 is the pilot gate; anchor pilot 2,000 seats from M9 at list
    Gates out: the four one-day benchmarks (condition precedent of the raise); anchor production take-or-pay signed — base M12 conditional on signed pilot paper at close; the buyer-realistic M15 case survives (§07)
  2. Phase 2 · M12–M24
    Scale with the contract
    +6 leased racks · tranches of 3, PO M12 / M16
    each $15.0m installed, 65% leased at 12%, £3.9m equity + £0.72m deposit
    12 racks, 20,500 seats served, ARR $36.0m by M24
    Gate per tranche: anchor signed · demand > 85% of ordered capacity · forward-looking cash test: cash after equity slice, deposit and 6 months of the new lease service ≥ £4m floor
  3. Phase 3 · M25–M36
    Fill, then fund the next leg
    15 racks by M33 (tranche 3 PO M29 — the gate holds it until cash clears; tranche 4 on order M35)
    30,712 seats served, fleet 100% full · ARR $48.8m / £35.9m · cash £10.8m and rising
    25k seats firm inside this raise (13 racks / £50m of iron); 50k = Phase-3 / Series-B option (25 racks / £95m — +10 racks / ~£37m of iron from this fleet), financeable against contracts
    Gate: 3+ logos; Series B at 4–6× contracted ARR while the anchor is >60% of revenue — not 8× Asm

Top risks and the mitigations already in the plan

RiskImpact if it landsMitigation in the plan
Rack quote prints $6.0m, not $4.0m Pub print existsA dead M4, B dead M7 on every routeOption C on GLM-5.3 survives it (cash low £8.4m — the forward gate defers tranches); written quote is decision #1; HGX-B300 8-GPU boxes are the procurement fallback and the residual-value hedge
Streams measure below the base cells (GLM 40, K3 12)seat counts fall toward the published floor (dead M30 at conc=16)The benchmark is a condition precedent of the raise; the anchor term sheet carries a capacity acceptance test that re-baselines seats/rack; §08 grid 2 prices every landing point
GLM-5.3 weights stay gated or arrive non-MITroute 2’s model unavailable at go-liveGLM-5.2 weights are the same 744B/40B architecture under MIT — same curve, same maths; K3 is not the fallback
K3 licence: $20m aggregate-revenue trigger Puban open-ended single-counterparty negotiation on any route serving K3Launch is Route 2 pure — zero licence surface; a K3 tier ships only after Moonshot commercial terms are in writing, as metered overage that keeps the share self-limiting
Anchor signs late (buyer-realistic M15; Conservative M21) or neverstacked-downside case is dead M14Named buyer-realistic run: BE M20, low £6.1m, ARR M36 $42.4m, gap £0. 6-rack build holds a 33-month zero-revenue runway; no contract = no tranche = no new cash out
Anchor concentration: 77–81% of seats vs ≤60% targetone procurement decision owns the P&L; cross-defaults the leasesTake-or-pay assigned to the lessor with parent guarantee; Series B presented at 4–6× contracted ARR; the non-anchor schedule must roughly double — 8 sales/CS heads by Y3 are costed in; say the miss out loud
Ordinary net-60 billing terms~$8m stuck in receivables at the cash trough — the base case dies on payment terms aloneAnchor billed quarterly in advance as a contractual condition (the DSO input prices the failure); prepay released in tranches against milestones
Single site fails the buyer’s BC/DR reviewpilot blocked at the questionnaire stageTiered DR: named bare-metal provider under identical controls in the DPA, RTO 72h degraded; the standing cost (×3 from anchor signature) is in opex; a second site is priced as a Phase-2 option
Lessors refuse tranche 2+fleet stops at 9 racks9 racks ≈ 17,550 GLM seats — still ahead of base demand to M18; two lessor term sheets are decision #6, with non-disturbance + step-in negotiated at tranche 1

What a sceptic will say — and our answer

Every row is a verdict from the three adversarial reviews commissioned on this plan (12-technical, 13-financial, 13b buyer-side). The answers are fixes now live in the engine, not debate points; the fix log is work/14-fix-log.md.

What a sceptic will say
Our answer
You promised 50,000 seats and day one serves ~11,000.
Correct, and by design. £50m was never 50k seats on any model — that is 25 racks (£95m of iron) even on GLM-5.3. 25k seats are firm inside this raise (13 racks / £50m of iron); 50k is the Phase-3 / Series-B option, financeable against contracts.
Your seat counts are contracted on an extrapolated curve — GLM is published at conc=16 only.
The floor is named, priced and visible. The published-only run (936 seats/rack) dies M30 and needs $349 list — it is in the model, the route table and the deck. That is exactly why the one-day 8×B300 benchmark is a condition precedent of the round and the anchor term sheet carries a capacity acceptance test with the vendor holding shortfall risk.
5% flat concurrency contradicts your own demand file — agents push it past 6.5% by M36.
Fixed: base is now 6% (the M36 agentic point) with contractual per-seat stream caps (2 volume / 5 premium) protecting the sizing. The headline M36 ARR fell from $63.7m to $48.8m — inside the technical review’s honest $47–58m band.
Your +100% prefill tax was a fat compensating error hiding thin optimism elsewhere.
Correct — it was ~4× too high. It is now +25%, re-derived as a function of prefix-cache hit (≥85%), adopted simultaneously with tighter streams (48→40), long-context (0.75→0.65) and headroom (20→25%). The spec flags the bundle never-unbundle: dropping the tax alone after a benchmark would re-inflate seats/rack to ~3,800.
You book revenue as same-month cash. Enterprise pays net-60.
A DSO input now exists. Base = anchor billed quarterly in advance as a contractual condition of the deal; plain net-60 is priced in Conservative — it strands ~$8m in receivables at the trough, which is why advance billing is a term, not a hope.
The Hybrid’s 35% K3 share is a hope, not a dial — and any K3 re-imports the Moonshot licence trigger.
Agreed on both. A free model picker drifts K3-ward, and the 50% attractor is cash-negative on our own sweep; the clause reads on aggregate revenue. So launch is Route 2 pure (MIT, zero licence surface); a K3 tier ships only after Moonshot terms are in writing, as metered overage with a per-seat allowance — self-limiting by construction.
M12 signature with $4–5m prepaid unsecured into a startup? Regulated buyers close M15–18.
The buyer-realistic case is a named run: signature M15, schedules +3 months, prepay released in thirds against milestones — break-even M20, cash low £6.1m, ARR M36 $42.4m, funding gap £0. The realistic calendar does not kill the plan, and we show it rather than let diligence discover it.
Your own ladder gives −25% at 25k seats — that prices the anchor at your cost floor.
Yes: ~$124 net vs the $123 full-cost floor at 25k. The ladder is now in the engine (anchor pays max(15%, tier)), the M24–36 ARR carries the haircut, and the corridor is presented openly in §06 — open $165–175, walk away below $120 list. A buyer team would have found the floor from the grid anyway.
78% anchor concentration and a premium ARR multiple don’t co-exist.
Series B is modelled at 4–6× contracted ARR, not 8×, until the ≤60% gate is met. The take-or-pay is assigned to the lessor with a parent guarantee, and the plan states its own milestone misses (seat-fill 85% vs 90%; anchor share 81% vs ≤60%) before you find them.
30% GAAP gross margin at M24 is not a software business.
That is 4-year depreciation on 85%-filled racks. Cash GM is 70% at M24; GAAP reaches 35% at M36 at full fleet. The 3-year depreciation column goes to lenders unprompted — policy alone swings M24 GAAP GM by ±10 points, so we show the range rather than defend a point.
What happens to my data and service if you fail?
Open weights make failure survivable for the buyer — configs, evals and runbooks in escrow; lessor non-disturbance + step-in at tranche 1; minimum-cash covenant at the £4m floor; and the estate can run the same MIT-licensed model on its own iron as the board-acceptable worst case. No closed vendor can offer that.
§ 10

Recommendation, and this month’s decisions

Decision requested

Commit to Option C on Route 2 pure — 6 racks (£24.6m / $33.4m) on GLM-5.3, MIT weights, zero licence surface — gate every leased tranche on the anchor’s take-or-pay, and add a K3 opt-in tier only after commercial terms with Moonshot are agreed in writing, as metered overage (~3–4× GLM token rates with a premium-seat allowance), not a per-seat uplift.

This is a view with named flip conditions, not a foregone conclusion: a measured K3+DSpark curve at ≥12 streams/replica at 100 tok/s, or an anchor paying ≥$300/seat for K3 with Moonshot signed, reopens the route choice — the model reproduces both by moving two inputs (Excel: Inputs sheet, Route selector and K3-share cell; the metered future state is Route 3 at the 10% point). What flips C back to B: a rack quote ≤$3.7m and the anchor contracting ≥15k seats at signature. Option A is rejected on every route: £46.7m of iron on £50m is insolvent at delivery.

Decisions the founder must make this month

  1. Get the rack quote in writing. $4.0m is a planning price; the $6.0–6.5m print is real. Every seat promise waits on this number.
  2. Run the four benchmarks — they are closing conditions of the raise. GLM-5.3 at conc 32 / 48 / 64 × 8K / 32K ISL and K3+DSpark at conc 4 / 8 / 12 on a rented 8×B300 node, prefix-cache hit rate on a real anchor trace, plus the written rack quote. One–two weeks; it moves more money than any negotiation.
  3. Put the two binary gates to the anchor in meeting 1. (a) Provenance: does their (or their US affiliates’) policy exclude PRC-origin open weights — get pre-clearance in writing; (b) residency: UK / EU / UAE position fixed in the DPA. Either answer reshapes the plan more than any price.
  4. Sign the colo LOI now. Power-on (M6) is the go-live gate, not rack delivery (M4). £44.4k/rack-month base at 140 kW committed per rack; 3-month deposit.
  5. Adopt the route posture. Launch GLM-5.3 volume + premium tiers (contracted on tok/s SLAs, never model names); K3 later as metered overage, only after Moonshot terms in writing.
  6. Table the anchor offer. 3-year take-or-pay, ≥10,000 seats at $140 net ($50m / £37m TCV), 25% year-1 prepay released in thirds, billed quarterly in advance, capacity acceptance test attached; open at $165–175 list with the published ladder; walk away below $120 list.
  7. Get two lessor term sheets. 65% advance at 12% over 36 months against the anchor contract, 10% deposits — with non-disturbance + step-in for the anchor negotiated at tranche 1. Verify before it is load-bearing.
  8. Forward-buy the initial PO’s ~$33m at close. Every input is dollar-denominated except the revenue’s discount; GBP −10% costs £2.4m of trough on its own.
§ A

Appendix

Assumptions register, the original 26 Aug diligence brief re-presented in full (its sources retained verbatim), and the model-comparison sources.

A1 · Assumptions register — the inputs that move the answer Conservative / Base / Upside; full register in 07-model-spec.json (103 rows) and the Excel Inputs sheet
InputConsBaseUpsideTagSource
FX, GBP per USD0.8090.73570.7357Pub base / Asm consWise, 26 Aug 2026; Conservative stacks GBP −10% — mitigation: forward-buy the initial PO ~$33m
Design concurrency (in-flight ÷ seats) · per-seat stream caps8% · 2/56% · 2/53% · 2/5Est01-demand Little’s law M36 agentic point; caps contractual (12-adv A1)
K3 streams / replica @ ≥60 tok/s41224Est02-capacity §2b; no published curve
K3 streams / replica @ ≥100 tok/s1612Pub floor, else Estvllm.ai/blog/2026-07-27-k3; tier sells “target 100 / floor 60” until measured
GLM streams / replica @ ≥60 tok/s304056Estno-spec fit + 1.2–1.3× MTP retention (12-adv A2)
GLM streams / replica @ ≥100 tok/s162432Pub at 16, else Estsglang PR 29674 (184 tok/s @ conc 16)
Hybrid K3 traffic share · split-pool penalty50% · 10%35% · 7.5%20% · 5%AsmBRIEF; enforced by router + per-seat allowance (12-adv A6)
Prefill tax (extra GPUs, fn of prefix-cache hit ≥70/85/90%)+60%+25%+15%Est12-adv A7 re-derivation; cache-hit bench is a closing condition
p99 headroom (incl. upgrade drain) · long-context factor25% · 0.5025% · 0.6525% · 0.80Asm02-capacity §4; 12-adv A8/A9; volume seats capped 256K context
Rack price (stress $6.0m held separately)$4.5m$4.0m$3.7mEstLoop Capital $3.7–4.0m; analyst $3.5–4.5m
Depreciation (racks) · residual3 yr · 20%4 yr · 10%5 yr · 5%Asm03-capex §7; 3-yr column shown to lenders unprompted
Colo per rack-month, UK (140 kW committed)£68.2k£44.4k£29.6kEst04-opex §1.2; 12-adv A11; UAE saves ~£2.8m/yr at 12 racks (anchor-consent-gated)
Volume / premium seat price, $ list120 / 250150 / 300175 / 400Asm2.0–2.9× ChatGPT Enterprise ~$60
Achieved discount · anchor seat ladder (≥5k/10k/25k)30% · 15/20/30%15% · 10/15/25%5% · 5/10/15%Asm05-pricing §5.1; anchor pays max(achieved, ladder) — 13b blocking 1
Price escalator from M25 · SLA credits0% · 3%0% · 1.5%8% · 0.5%Asmtoken prices fell ~10× in 2 yrs (13-adv); credits contra-revenue (13b)
Billing terms (DSO days)60150Asmbase = anchor quarterly-in-advance, contractual (13-adv blocking 2)
Anchor production signature monthM21M12M9Asm05 §6 / 06 §1; M12 conditional on signed pilot paper at close; buyer-realistic M15 is a named run
Lease: advance · rate · deposit55% · 15% · 15%65% · 12% · 10%75% · 10% · 5%Asm06 §2; CoreWeave SOFR+2.25–5.50% Pub; deposits at PO (13-adv)
K3 licence rate on K3 revenue (trigger $20m T12M Pub)10%5%2%Asmno published Moonshot terms; clause reads aggregate revenue
Model assurance (red-team, MBOM, egress audit), £/yr200k150k100kEst13b-adversarial; nothing in 04-opex covered it
Go-live month (colo power-on gate)M10M7M5Est04 §1.3; racks delivered M4
Cash floor (reserve rule) · expansion gate£4m£4m£4mAsm06 §1; gate is forward-looking: post-tranche cash incl. 6 months of new lease service ≥ floor (13-adv)

Files: work/07-model-spec.md (full register and formulas), 07-model-spec.json (machine-readable), 07-model-calc.py (engine), 07-base-results.json (all 27 scenario × route × build runs, monthly). Excel: Inputs sheet, blue cells; Scenario / Route / Build selectors.

A2 · The original diligence brief (26 Aug 2026), tidied sections 1–14 condensed; sources list retained verbatim; superseded numbers noted at the top

Internal diligence brief · hardware-envelope note · dated 26 Aug 2026 · web edition 27 Aug 2026 · not for index, search, or public share. Working numbers from the 26 Aug conversation are tagged ESTIMATE. Lab or vendor-published figures are tagged PUBLISHED. There is no NVIDIA list price in either set. Context: an inference-serving company; 50,000 R&D employees would get access, not all generating at once; model Kimi K3 on GB300 NVL72. Figures in §§1–4 are the 26 Aug conversation transcribed; they are not a published bake-off.

How the new model departs from this brief

  1. Model name. The new model is named GLM-5.3 throughout. GLM-5.3 shares the same 744B / 40B-active base as GLM-5.2; its open weights were still gated as of 26 Aug 2026, so the published GLM-5.2 decode curve (§8) carries over unchanged. Where this brief writes GLM-5.2 in §§8, 9, 11 and 13, read GLM-5.3 in the model.
  2. Routes, not tiers. Capacity is costed per route — Route 1 Kimi K3 only, Route 2 GLM-5.3 only, Route 3 Hybrid (both models offered to every seat) — rather than assuming the volume / premium split that §13 leans to. That lean is research, not a decision; the choice is the founder's.
  3. Full BOM. The “$2m for storage and IB fabric” placeholder in §§4 and 12 is replaced by a full bill of materials in the new model.
  4. Beyond iron. Colo, power, staff and runway — which this brief only notes “sit on top of the $50m” — are modelled explicitly on top of the iron line.

Honest offer on a $50m iron envelope. Twelve GB300 NVL72 racks at a ESTIMATE $4m planning price is $48m of racks and $2m left for storage and IB fabric. N+1 (11 live, 1 spare) supports an ~8,000-seat research service at 50–100 tok/s — not 50,000 seats.

Fifty thousand seats at 5% peak and 50 tok/s is ~70 racks, ~$280m of iron, before the building. Colo, staff, and power sit on top of the $50m. Do not spend the whole envelope on GPUs and then discover networking and colo are still unpaid.

Planning price / rack
$4m — until a quote. Loop Capital $3.7–4.0m. EST
$50m envelope
12 racks — $48m iron + $2m storage / IB. EST
Honest seats
~8,000 — @ 50 tok/s, 5% peak, N+1. EST
IT power, 12 racks
~1.6 MW — ~135 kW / rack. Colo on top. EST
Resilience
N+1 — 11 live, 1 spare. One rack down, no user-visible loss.

1. Concurrency design

Design at 3% of seats generating; size comfortable at 5%.

  • Copilot-class: ~30% of seats touch on a given day.
  • Global R&D overlap in one peak window.
  • Most of a research session is reading, not generating.
  • 50k × 30% daily × ~45% in peak window × ~20% streaming ≈ 1,500 in-flight = 3%.
  • Size 5% (2,500 generating) so a 200-person org after a meeting does not queue.
  • Agent loops can push 5–8%.

All of the above is ESTIMATE — a utilisation model, not a measured fleet.

2. One GB300 NVL72 rack

  • 72 B300s, one NVLink domain. PUBLISHED NVIDIA product page: 72 Blackwell Ultra + 36 Grace, one NVLink-5 domain.
  • No NVIDIA list price. Credible $3.7–4.0m (Loop Capital). Supply-chain print $6.0–6.5m. Plan $4m until a quote. Power ~135 kW. ESTIMATE
  • K3 packs as 9 replicas of 8 GPUs (72 ÷ 8). Packing math matches the research replica definition; interactivity numbers below are conversation working points.
Working point (one rack)In-flightSeats @ 5% peakSeats @ 3% peakBasis
People generating at 100 tok/s9180300EST
People generating at 50 tok/s367201,200EST

The 36 figure assumes the published 8-GPU curve still holds ~4 streams at 50 tok/s (9 replicas × 4). That concurrency point is not in the published K3 table — see §6. 50 tok/s is the promise when the rack is busy. 100 tok/s is what a user sees when that replica is not sharing.

Multiply racks to multiply capability if each rack runs its own K3 replicas behind a load balancer. This does not hold if one giant model copy is smashed across racks. The first rack also pays for gateway, weight store, and monitoring.

3. Fifty thousand seats at 5% peak (2,500 generating)

TargetRacksBasis
50 tok/s70EST
+20% spare85EST
+ long-context prefill~100EST
100 tok/s guaranteed280EST

Reasonable buy for full 50k: ~85–100 NVL72s (~6,000–7,200 GPUs, ~11–14 MW, ~$300m of racks). ESTIMATE Do not go below 60 at 3% unless willing to shed load.

4. What $50m of racks actually buys

  • At $4m: 12 racks ($48m) and $2m left for storage and IB fabric.
  • At a $6m quote: only 8 racks.
  • Do not spend the whole $50m on GPUs then discover colo / networking still needed.

12 racks, N+1 (11 live, 1 spare)

PromiseIn-flightSeats @ 5%Seats @ 3%User speed
50 tok/s396~8,000~13,00050 guaranteed at peak; 80–100 typical. One rack down, no user-visible loss.
100 tok/s99~2,000~3,300100 always.

All rows ESTIMATE from the 26 Aug conversation.

  • Honest offer: 8,000-seat research service at 50–100 tok/s, not 50,000 seats.
  • 50,000 seats at 5% and 50 tok/s is ~70 racks, ~$280m of iron, before the building.
  • Power for 12 racks ~1.6 MW. Colo, staff, power sit on top of the $50m.

5. Research appendix (26 Aug 2026)

Facts in §§6–14 are only those in research.md / scenarios.csv. No invented tok/s or prices. Two price books are shown on purpose — do not mash them into one fake list price.

  1. Kimi K3 facts
  2. Published K3 decode — and the missing curve
  3. GLM-5.x (largest open is 5.2)
  4. Conservative 25k / 50k / 100k — a national AI-factory
  5. Two price books
  6. Scenario table (conservative K3 + GLM @ 100 only)
  7. What the sizing ignores
  8. Recommendation
  9. Sources

6. Kimi K3

ItemValueBasis
Total / active2.8T / 104B (tech report also 104.2B active)PUBLISHED
Serve precisionNative MXFP4 weights / MXFP8 activations (QAT from SFT)PUBLISHED
Checkpoint1,560.9 GB (1,560,936,091,448 bytes; 1,453.7 GiB) across 96 shardsPUBLISHED
Minimum node to hold8× B300 / 8× GB300 / 8× MI355X (288 GB class, 2,304 GB). 8× H200 does not fit (~433 GB short).PUBLISHED
Context1,048,576 tokens; 896 routed experts (top-16) + 2 sharedPUBLISHED

vLLM wording: tensor-parallel is the interactivity recipe; wide expert-parallel is the throughput recipe and hurts per-user tok/s. An NVL72 can pack 9 TP8 replicas (72/8) if the fabric is partitioned that way — that changes server/rack count, not GPU count.

7. Published K3 decode — no concurrency>1 at 100 / 200

EngineHardwareBatchSpectok/s per user
vLLMGB300 NVL72, TP8bs=1none111
vLLMGB300 NVL72, TP16bs=1none118
vLLMsame, TP8bs=1DSpark (7 draft)331
vLLMsame, TP16bs=1DSpark370
SGLangday-0 kernel ladderbs=1none~113
SGLangsamebs=1DSpark~423

All rows PUBLISHED (vLLM 27 Jul 2026; SGLang day-0). 200 tok/s requires DSpark (or equivalent). Without speculation, 111–118 does not reach 200 at bs=1.

Not published as of 26 Aug 2026: a K3 decode-speed curve vs concurrency that still holds ≥100 or ≥200 tok/s/user. vLLM’s Pareto figure is prose only (“100+ TPS/user” at the low-latency end; “2K+ TPGS” at high-throughput). Therefore conservative max concurrent sequences per replica at those SLAs = 1. No higher-concurrency estimate is made — the research rule is estimate only from a published curve. K3 has none at this SLA.

The conversation’s “~4 streams at 50 tok/s” is a working assumption for a busy-rack 50 tok/s promise, not a published cell. Treat §2’s 36-in-flight / 720-seat row as ESTIMATE until that curve is measured.

Hosted Artificial Analysis (26 Aug 2026 window): K3 comparison median 35.5 t/s; fastest listed provider Makora 162.6 t/s; Parasail 151.5; Modal 148.2. Hosted multi-tenant ≠ dedicated batch-1 replica. PUBLISHED

8. GLM-5.x — largest open-weight is GLM-5.2

The flagship name as of 26 Aug 2026 is GLM-5.3 (Z.ai, 14 Aug 2026). Same 744B / 40B-active MoE base as GLM-5.2; gains are post-training only. Open weights were not public as of 26 Aug (Z.ai: “in two weeks” after 14 Aug; HF zai-org/GLM-5.3 still gated mid-late August). Largest open-weight GLM is GLM-5.2 (MIT, 753B incl. MTP / ~744B backbone, 40B active).

SGLang NVFP4, 8×B200, low-latency + MTP — the only two-point numeric TPOT table:

ConcurrencyTPOTtok/s per user
11.85 ms541
165.43 ms184

So ≥100 is published at 16 concurrent sequences per 8-GPU replica; ≥200 is published only at concurrency=1 (conc=16 is 184, below 200). PUBLISHED SGLang also states 500+ tok/s/user on 8×B300 at bs=1 and 450 on 4×GB300 at bs=1 — no conc>1 table on those exact SKUs.

Hosted AA: GLM-5.3 (Z.ai) 90.0 t/s. PUBLISHED Z.ai’s 14 Aug footnote used 115 TPS (GLM-5.3) and 40 TPS (Kimi K3) from AA at write time.

9. Conservative 25k / 50k / 100k concurrent is not this $50m story

Because the SLA is per-user, replicas = ceil(clients / max_conc_at_SLA), not aggregate tok/s ÷ peak GPU tok/s. For K3, conservative conc = 1 at both 100 and 200 (200 requires DSpark).

K3 clientsSLAReplicasGPUsNVL72 (÷9)
25,000100 or 20025,000200,0002,778
50,000100 or 20050,000400,0005,556
100,000100 or 200100,000800,00011,112

PUBLISHED hardware math on a conservative conc=1 reading. That is a national AI-factory, not the $50m / 12-rack envelope. GLM-5.2 at 100 tok/s with published conc=16 is 3–16× smaller in replica count (see §11); the 200 tok/s conservative GLM case returns to the same order of magnitude as K3.

10. Two price books — do not mash

NVIDIA publishes no MSRP for HGX / DGX / NVL72. Two separate books are in the sources. They are not one list price.

Book8-GPU B300GB300 NVL72Used where
Research bookends (scenarios.csv)$381k–$650k$3.0–4.3MLow: Rillor/Supermicro $381k; Thunder Compute high $650k. NVL72: market low $3.0M; Wolfe ~$4.3M (GPUSmith).
Proposal planning price$4mLoop Capital $3.7–4.0m cited in the 26 Aug conversation. Supply-chain print $6.0–6.5m. Plan $4m until a quote. Used for the 12-rack / $50m arithmetic.

Other cited 8-GPU points (not used as the CSV bookends): GPUSmith/Turbo Max ~$420k; Exeton “starting at” $795k; Petronella bare baseboard from $485k; DGX B300 ~$400k (GPUSmith citing FT via PCMag). GB200 NVL72 Wolfe ~$3M, market $2–3M. FX used in the CSV GBP columns: 1 USD = 0.7357 GBP (Wise, 26 Aug 2026).

11. Scenario table — conservative K3 and GLM @ 100 only

Rows copied from scenarios.csv. Interpolated GLM-@-200 and 4×GB300 alternate rows are omitted. USD columns are the research bookends in §10, hardware-only, no network / colo / spare.

ModelClientsSLAConc / replicaReplicasGPUsNVL72$ 8-GPU $381k–$650k$ NVL72 $3.0–4.3M
K3 PUB25k100125,000200,0002,778$9.53B–$16.25B$8.33B–$11.95B
K3 PUB50k100150,000400,0005,556$19.05B–$32.50B$16.67B–$23.89B
K3 PUB100k1001100,000800,00011,112$38.10B–$65.00B$33.34B–$47.78B
K3 PUB25k200125,000200,0002,778$9.53B–$16.25B$8.33B–$11.95B
K3 PUB50k200150,000400,0005,556$19.05B–$32.50B$16.67B–$23.89B
K3 PUB100k2001100,000800,00011,112$38.10B–$65.00B$33.34B–$47.78B
GLM-5.2 PUB25k100161,56312,504174$0.60B–$1.02B$0.52B–$0.75B
GLM-5.2 PUB50k100163,12525,000348$1.19B–$2.03B$1.04B–$1.50B
GLM-5.2 PUB100k100166,25050,000695$2.38B–$4.06B$2.09B–$2.99B

K3 200 tok/s rows require DSpark (331–370 published at bs=1; 200 not reached without spec). GLM @ 100 applies the SGLang B200 NVFP4 conc=16 point (184 tok/s/user) to an 8×B300 replica as the same recipe class. GBP at 0.7357: K3 100k conservative is £28.03–£47.82B at the 8-GPU bookends and £24.53–£35.15B at the NVL72 bookends.

12. What this sizing ignores

Do not treat replica math as a BOM. All of the following are real and not in the GPU-count tables:

  1. Prefill. Production recipes are P/D disaggregated. A +50% to +200% GPU tax for prefill is typical; not included in the conservative tables. The 26 Aug conversation adds ~15 racks (~100 vs 85) for long-context prefill on the 50k / 50 tok/s path — still an estimate.
  2. Speculative decoding is not free. DSpark / MTP acceptance collapses on high-entropy tokens (vLLM: 4.73 accept on coding vs 2.61 on creative writing). 331–423 and 541 tok/s are workload-specific.
  3. Networking / IB fabric. Multi-rack decode needs Quantum-X800 or Spectrum-X 800 Gb/s. Often tens of percent of GPU capex. Not priced. The $50m plan parks $2m for storage and IB — a placeholder, not a quote.
  4. Spare / N+1 / rolling deploy. The 12-rack offer includes one spare. The conservative factory tables do not.
  5. Storage. K3 checkpoint 1.56 TB × replica count if local, or a shared parallel FS plus staging NVMe.
  6. Colo and power. Conversation: ~135 kW / rack, ~1.6 MW for 12, ~11–14 MW for 85–100. Official GPU TDP is up to 1,400 W (NVL72) / 1,100 W (HGX B300). Facility 1.5–2× is the usual next tax and is not in the $50m.
  7. Software headroom. Published tok/s is a lab point, not p99 under churn.
  8. Licence. K3 License is not MIT (third-party notes a ~USD 20M MaaS revenue trigger). GLM-5.2 is MIT; GLM-5.3 license unconfirmed until the repo opens.

13. Recommendation (from the 26 Aug research)

Research lean as written on 26 Aug. Item 3 is not a decision — the new model costs all three routes (K3 only / GLM-5.3 only / Hybrid) identically and leaves the choice to the founder.

  1. Default metal: GB300 NVL72 (or HGX B300 8-GPU if racks slip). Do not buy B200 as the long-term default for K3.
  2. Productize two SKUs: (a) interactive 100–200 tok/s, concurrency 1–16, TP8/TP16 + spec decode; (b) batch / throughput wide-EP. Do not promise (a) economics on (b) hardware efficiency.
  3. Treat GLM-5.2 / 5.3 as the volume model (3–16× fewer replicas at 100 tok/s) and Kimi K3 as the premium flagship.
  4. Re-measure the missing K3 concurrency curve and the GLM B300 curve before ordering. The entire 200 tok/s GLM estimate in the full memo hangs on two SGLang B200 cells; it is omitted from §11 on purpose.
  5. Budget prefill + fabric + power + spare as a separate line at least as large as the decode-GPU line.
  6. On a $50m envelope: sell the 8,000-seat service, or raise the iron budget toward ~70–100 racks for 50k seats. Do not claim 50k seats from 12 racks.

14. Sources

URLs as cited in the 26 Aug research memo. No fabricated citations.

Loop Capital $3.7–4.0m and the supply-chain $6.0–6.5m print are from the 26 Aug conversation, not from a URL collected in the research memo. They are not given a fake link here.

Diligence brief dated 26 August 2026 (Asia/Dubai). Web edition 27 August 2026. Prepared for internal review of an inference-serving hardware envelope. Numbers that are not in the 26 Aug conversation, research.md, or scenarios.csv are not in this brief.

This is a planning note, not a quote and not a BOM. Allocation, colo power, and a measured K3 concurrency curve must exist before a capital commit.

A3 · Model-comparison sources (27 Aug session) what §04 rests on; full list in 06b-model-comparison.md