Run the benchmark

It downloads the models you pick (up to 9.7 MB), finds your fastest thread count, then measures steady-state speed and time to first token, and writes a sample email so you can watch it. Nothing leaves your browser until you press submit.

Models

Takes about 2–6 minutes and uses all your CPU cores: close heavy tabs first, and plug in a laptop.

What your browser shows

How the score is measured

Tuning. Each model is loaded with 1, 2, 4, 8… threads (up to your core count) and run for a short steady-state test; the fastest thread count is kept. Browsers can't pin threads to cores, so more threads isn't always faster.

Tokens / second. The model is reloaded at the best thread count (the load includes a warm-up so the WebAssembly is fully optimised), then decodes to its full 256-token context, 10 times, ignoring end-of-email. The score is the median of the 10 runs; the range shows the slowest and fastest.

Time to first token. 10 real emails for “write email to john saying that is all good”: the time from pressing Generate to the first output token (request parsing, prompt processing and sampling included), median of 10.

The sample email is one extra, unscored generation. It is written in a few milliseconds, so the page replays it at reading speed and shows the real time it took.

Your CPU. Browsers hide the CPU model (only Apple silicon reveals it, through the graphics name), so it's filled in from what your browser shows, your last submission, or what machines with the same core count and graphics picked; otherwise it says "Unknown CPU" with your thread count, and you can correct it.

Engine. The same hand-written WebAssembly SIMD engine everywhere (relaxed SIMD where your browser has it, plain SIMD otherwise): . Scores are self-reported by the browser, so they're for fun.

A 1.58-bit (ternary weights) email-writing transformer, running on your CPU in this tab via WebAssembly SIMD and threads. Scores are self-reported by browsers: for fun, not for procurement.