A living measurement of the GPU fleet

What weight-streaming actually achieves on every GPU class.

We rent the vast.ai fleet, run a standard streaming benchmark on every device stratum, and publish the numbers — with our own prediction error next to them. Measured, not marketed.

devices measured
fastest streaming tok/s
median prediction error
cheapest $/1M tok

The problem: inference buyers fly blind

Every GPU listing advertises bandwidth. Almost nobody publishes what a model actually streams at on that hardware — stall percentage, real tokens/sec, joules per token. Rental decisions get made on datasheet numbers that don't survive contact with a 70B model.

Advertised vs achievable

We audit the marketplace itself

Every artifact snapshots the full vast.ai offer row next to measured results. When advertised PCIe bandwidth and achievable streaming bandwidth diverge, that divergence is published.

Calibration, in public

We publish our own error

Our simulator predicts each device's streaming behavior before we rent it. The predicted-vs-measured error ledger ships with the dataset — so you can judge the model, not just the numbers.

Continuous

The fleet is non-stationary

Hardware turns over, hosts change, prices move. The census re-measures on a rolling schedule. Every number carries a timestamp and a machine ID.

Latest measurements

A preview of the current census. Full sortable data on the leaderboard.

GPUStreamtok/s (agg, b32)stall %H2D GB/s$/1M tokpred. errorstatus

Honesty rails

Every number on this site is either MEASURED with a citable artifact, or labeled MODELED. We cite competing systems — vLLM sleep mode, TensorRT-LLM weight streaming, Run:ai streamer, PipeBoost/ParaServe — wherever they own a result. No ROCm/MI300X claims until we measure one. No optical numbers from electrical experiments.