How every number on this site is produced. If you can't reproduce it, we don't publish it.
Every device runs the same artifact (gpu_bench_streamos.py):
nvidia-smi sampling) average watts and joules/token.The census targets 30–60 machines per sweep, stratified across the live vast.ai marketplace by PCIe bandwidth × HBM bandwidth × architecture × verified/community. Sampling is reliability-weighted and deliberately includes the ugly tail — community hosts, PCIe risers — because that's where calibration degrades and buyers get burned. Stated bias: the rentable fleet is not an unbiased census of the world's GPUs; results describe the marketplace, not the population.
Every run snapshots the full marketplace offer row (advertised specs + price + geolocation) and the machine_id into the artifact. The marketplace is non-stationary; an artifact without its offer row is unreproducible and doesn't ship. Schema: kofta.atlas.v1 (see API docs).
Before renting any device, our simulator predicts its streaming behavior from its public listing: effective H2D bandwidth, stall %, tokens/sec, TTFT. Predictions are anchored to measured constants (first anchor: NVIDIA H100 SXM, July 2026) and scaled by device specs. The predicted-vs-measured error ledger is published per device — the median error on this site's homepage is ours, not a competitor's.
Sweeps run on interruptible bids with precomputed minimum bids and retry-elsewhere on preemption. Weight caches live on local volumes pinned to machine IDs, making 50-run sweeps practical in an evening. Every dollar is ledgered per experiment and published in aggregate.