inference: your browser · votes: 0
blind evaluation · model A vs model B · you judge

Does size still matter?

Two open sub-1B models answer your prompt side-by-side, identities hidden. The inference itself — llama.cpp compiled to WebAssembly — runs entirely in your browser. No server, no GPU, no API keys. Your CPU is the benchmark rig.

engine llama.cpp → WASM
sampling temp 0.7 · top-p 0.9
budget $0.00
your votes stored locally
LOADING FIGHTERSdownloading…
MODEL A · ~108 MB0%
MODEL B · ~484 MB0%
Models stream from the Hugging Face CDN and are cached by your browser (OPFS) — next visits start instantly. One Web Worker per fighter: both generate in parallel, one thread lane each, symmetric load.
0/500
SIDE-BY-SIDE · SAME PROMPT · SAME SAMPLING
MODEL Ahidden
waiting…
MODEL Bhidden
waiting…
Which response is better?

Your arena

#ModelEloVotesWin rateTTFBtok/s
No votes yet — the first duel decides the first Elo point.