Arena

Pick two machines. The top half compares what they are — memory, bandwidth, power, price. The bottom half compares how they run one model you choose.

A row only takes a winner when both sides are comparable claims. Otherwise it shows both numbers and declares none.

When a winner is and is not declared

A speed row is only marked with a winner when both sides were actually measured, and both were measured close to the context length you selected — within a factor of two either way. Where one side is an estimate, has nothing at all, or was measured too far from that context, the row still shows what is known but declares no winner: results that answer different questions don’t get ranked against each other just because they sit in the same row.

The per-unit rows — bandwidth and memory per price — take a winner only when both sides’ prices share the same condition (new or used) and the same basis (sourced or converted). A used price against a new one, or a sourced listing against a currency conversion, are different kinds of claim too.

Running one model

Decode is token generation — the speed you feel while an answer streams. Prefill is prompt processing — the wait before it starts. Why bandwidth predicts decode speed.

Choose a machine