A capability report and a set of micro-benchmarks for the kernels that dominate on-device LLM inference: streaming reads and the vector×matrix product. Everything runs locally in your tab. Nothing is sent anywhere until you read the payload and press send.
Logos identify the hardware, software and models being measured. They are not endorsements, partnerships or sponsorships, and no vendor listed here is affiliated with this site.
| Model | Format | Download | Verdict |
|---|
Judged on what the browser actually tells us: the largest buffer WebGPU will allocate, the storage quota this origin is granted, and whether 16-bit shaders are available. The browser does not report how much memory your GPU has, so nothing here is guessed from your card's specs.
| Model | Q4_K_M | Q8_0 | f16 | Verdict at Q4_K_M |
|---|
Sizes are weights plus a 15% runtime overhead; the KV cache grows on top of that with your context length and depends on the model's architecture, so treat these as a floor. Run them with llama.cpp, Ollama or LM Studio. For the detailed per-layer breakdown, including KV cache and finetuning, gpu_poor does it properly and this page does not try to replace it.
| status | not probed yet |
| status | not probed yet |
| Kernel | Workgroup | Median | Throughput | Spread | Validity |
|---|---|---|---|---|---|
| Press “Run the benchmark”. | |||||
The workgroup sweep is the point: the size that wins here is a property of your architecture, not of the kernel. A run whose spread exceeds 15%, or whose rounds climb monotonically, is flagged and excluded from the public aggregate — that is thermal drift or another process competing for the GPU, not your hardware.
Nothing yet.
UtopiaIA builds Elffuss, a WebGPU inference engine. Its kernels were tuned on one architecture, and we have no way to tune for NVIDIA, AMD, Intel, Qualcomm or ARM without measurements from machines we do not own. That is what this page is for, and we would rather say so than pretend to be a neutral referee.