The numbers, with methodology

We publish our own measurements because vendor tables don't tell you what a $0.99/h endpoint actually does at 260K context. All runs: llama.cpp fork, Qwen3.8-27B, single server, no other tenants. Reproduce anything — configs below.

SKUQuantContexttok/s (decode)TTFT$/hr (us)
24GBMXFP4 + Q4_0 KV260,000 ×1~15–40$0.99
32GBMXFP4 + Q4_0 KV500,000 ×1 · 260,000 ×2$1.49
48GBMXFP4 + Q4_0 KV1,000,000 ×1 · 500,000 ×2 · 260,000 ×4$1.99
96GBMXFP4 + Q4_0 KV1,000,000 ×3 · 750,000 ×4 · 500,000 ×6 · 260,000 ×12$3.49

Reference points (community, same model): 218 tok/s single-stream (short ctx) · 140–260 tok/s p50, 0.156s TTFT. TODO at launch: fill TTFT column + link to raw benchmark logs.

Measured capacity per card (Qwen3.8-27B, llama.cpp)

CardCapacity
Blackwell 24 GB1 user @ 260K
Blackwell 32 GB1 user @ 500K · or 2 @ 260K
Blackwell 48 GB1 user @ 1M · or 2 @ 500K · or 4 @ 260K
Blackwell 96 GB3 @ 1M · 4 @ 750K · 6 @ 500K · 12 @ 260K

These are the numbers behind the SKU capacity column on the pricing page. Concurrent users = concurrent in-flight generations (max-num-seqs).

Quality (vendor-reported, Qwen model card)

BenchmarkQwen3.6-27BQwen3.8-27B
Terminal-Bench 2.163.473.0
DeepSWE 1.113.342.2
OSWorld-Verified63.984.3
SWE-bench Pro61.7
LiveCodeBench v690.3

Artificial Analysis Intelligence Index: 52 (GLM-5.2: 53, Kimi K3: higher but cluster-scale). #9 on Code Arena WebDev — the only small model in the top 10. Scores are Alibaba's; we flag that instead of hiding it.

Cost vs per-token APIs (agentic workload)

10 turns/hr × 200K context + 2K output:

Option$/hr
noTOKEN.cloud 32GB (dedicated)$1.49
GLM-5.2 ($1.40/$4.40 per M)$3.02
Kimi K3 ($3/$15 per M)$6.75

And this is a light workload — 10 turns an hour. At 100 turns/hr × 100K context (10M input tokens/hr), GLM-5.2 reads $14.88/hr and Kimi K3 reads $33/hr, while the dedicated endpoint stays at $1.49. Per-token pricing bills every context re-read; dedicated pricing doesn't care how hard your agent works.