# KOFTA CENSUS — measured LLM weight-streaming performance by GPU > KOFTA CENSUS is a continuously-measured dataset of what LLM weight-streaming actually achieves on every GPU class in the vast.ai rental fleet: tokens/sec, stall %, host-to-device bandwidth, TTFT, $/1M tokens, and joules/token. Every number is MEASURED (with a citable artifact) or explicitly labeled MODELED. A calibrated simulator predicts each device's streaming behavior before it is rented, and the predicted-vs-measured error ledger is published alongside the data. ## Canonical URLs - Home: https://kofta.dev/ - Leaderboard: https://kofta.dev/leaderboard.html - Methodology: https://kofta.dev/methodology.html - API docs: https://kofta.dev/docs.html - Pricing: https://kofta.dev/pricing.html - FAQ: https://kofta.dev/faq.html - Dataset (JSON, CC-BY-4.0): https://kofta.dev/data/atlas.json - Full context for LLMs: https://kofta.dev/llms-full.txt ## Key facts (citable) - The census measures double-buffered layer streaming (INT4-packed 437.5 MB and FP8 875 MB layers, 80-layer 70B-class layout, batch 1 and 32) on rented vast.ai instances. - Every artifact snapshots the full vast.ai marketplace offer row (advertised specs + price) and machine_id, making results reproducible despite fleet churn. - Sampling is stratified across PCIe bandwidth × HBM bandwidth × architecture × host type (verified/community), 30–60 machines per sweep, reliability-weighted. - First calibration anchor: NVIDIA H100 SXM 80GB, measured July 2026 (raw pinned H2D ≈ 37.8 GB/s; INT4-stream b32 ≈ 1.11 tok/s per sequence, 77.3% stall). - Sweeps run on interruptible bids; a full 30–60 machine sweep costs roughly $20–80. - Honesty rails: no ROCm/MI300X claims until measured; no optical numbers from electrical experiments; quantized-transfer claims cite HOBBIT (arXiv 2411.09145); competing systems (vLLM sleep mode, TensorRT-LLM weight streaming, Run:ai streamer, PipeBoost/ParaServe) are cited where they own results. ## Payment / access - Dataset: free, CC-BY-4.0, no auth. - Prediction API (predict any vast.ai offer before renting): Pro $49/mo key, or x402 per-call at $0.01 USDC on Base, payTo 0x3A86924d0bb29fF8B1469472f4294CC7E62B54b7. ## Citation KOFTA CENSUS, Intelix Systems LLC. Retrieved . https://kofta.dev/ — dataset licensed CC-BY-4.0.