Glossary

The terms behind the benchmark: technical, but at a high level. Dotted words throughout the site link here; hover one for a one-line definition.

Four weights share one byte; shifts and masks unpack them just before the multiply.

Weights are stored as 0, 1, 2 (meaning −1, 0, +1). Four fit in a byte, laid out so a vector load plus a few shifts and masks yields vectors of weights ready for the dot product. The +1 offset is corrected afterwards by subtracting the sum of the activations: Σ(w+1)·x − Σx = Σw·x.

Being 0 to 2 also means every weight fits the 7-bit operand of the relaxed int8 dot instruction.

Attention

Model

How each token looks back at earlier tokens: it compares itself with them, then blends what they contain.

For the token being written, the model forms a query, compares it with the stored keys of all earlier tokens, turns the comparison into weights (softmax) and takes a weighted sum of their values. Several heads do this in parallel with different focuses.

Its cost grows with how much text came before, and it is the reason the KV cache exists. In our profiles attention is 15 to 30% of a small model’s time.

x86 vector instruction sets: AVX2 is 256 bits wide, AVX-512 is 512, VNNI adds a fused int8 dot product.

The native C engine has an AVX2 path. Newer CPUs (Zen 4 and Zen 5, recent Intel server chips) also have AVX-512 and VNNI, which could roughly double int8 throughput, but no AVX-512 kernel has been written yet, and browsers cannot use it anyway (WebAssembly is capped at 128-bit vectors).

On Arm the equivalents are NEON and the sdot instruction.

BitNet b1.58

Numbers

The research recipe for language models with ternary weights and int8 activations.

Introduced by Microsoft Research in 2024 (“The Era of 1-bit LLMs”), it showed that networks trained with weights in {−1, 0, +1} can be competitive with full-precision ones at scale. These models follow the recipe (ternary weights with a scale, 8-bit activations, RMSNorm, no biases) but at tiny size.

The inference engine is our own, not Microsoft’s bitnet.cpp: it had to run in a browser as WebAssembly.

CPUs change frequency with load, temperature and power source, so the same code gets faster or slower.

A laptop on battery may run far slower than plugged in, and a cold CPU takes a moment to ramp up. Benchmarks here run warm-up work first and report the median of ten runs, plus the slowest and fastest, so these effects are visible rather than hidden.

The free hosting behind this site: static files and functions on Pages, scores in D1 (serverless SQLite).

Pages serves the page, the engine and the models (about 22 MB of static files) and runs the small API. D1 is an SQLite database next to it; the free tier allows 5 million row reads and 100,000 writes a day, far more than a hobby leaderboard needs.

See also Turnstile

Browsers first compile WebAssembly quickly, then recompile the hot parts with an optimiser: the first run is slower.

A baseline compiler makes code in milliseconds; an optimising compiler (V8 calls them Liftoff and TurboFan) recompiles frequently used functions in the background. Until that finishes, code runs at perhaps half speed. The first email after loading used to be 2 to 3 times slower than the tenth.

The engine now runs a short warm-up (up to 200 ms) when a model loads, which also touches the KV cache memory, wakes the threads and lets the CPU speed up.

Compute-bound

Hardware

Limited by arithmetic and overhead rather than memory: what small models become once they fit in cache.

When the weights sit in cache the memory limit vanishes and other costs show up: thread synchronisation between layers, the scoring of the vocabulary, attention, unpacking the 2-bit weights. On the cache-rich server we tested, the 4M model ran at a small fraction of its memory ceiling, so it was compute- and overhead-bound.

This is where the engine optimisations (fused kernels, fast logits, a leaner thread pool) pay off.

How many tokens the model can hold at once: 256 here, request plus email.

The v5 models were trained on sequences up to 256 tokens. The benchmark decodes all the way to the full window, which is the slowest and most cache-hungry case, because attention and the KV cache are at their largest.

Copy slot

Model

A placeholder (<WORD_1> … <WORD_4>) for a rare word you typed, such as “gazebo”, so the model can repeat it without knowing it.

If a word in your request is outside the vocabulary, the parser replaces it with <WORD_1> and remembers it. The model learned to write the tag in the right place, and the page puts the original word back.

It is a cheap way to give a tiny model an open vocabulary: it learns the pattern “copy this thing here”, not the thing.

Logical “threads” are not full cores: two hardware threads on one core share its execution units.

A 4-core CPU with SMT (hyper-threading) reports 8 logical cores, which is what the browser shows. Two busy threads on one core each run slower than one alone, and a web page cannot tell the OS which cores to use. The thread tuning in the benchmark exists because “more threads” is sometimes slower.

Native programs can pin one thread per physical core; browsers cannot, which is part of the gap between the native and browser numbers.

Small, fast memory on the CPU: each level is bigger and slower; RAM is bigger and slower again.

L1 is per core and tiny (about 32 to 64 KB), L2 per core is 0.5 to 2 MB, L3 is shared by all cores (4 MB on a laptop, 64 MB or more on servers), then RAM. Reading from L1/L2 is many times faster than RAM: on the Ryzen laptop used to build this, four threads read about 80 GB/s from cache but only about 6 GB/s from RAM.

Writing one word makes the model read every weight once, so where the weights live sets the speed. That is why the models are small: the 4M model (1.9 MB) fits in L2-plus-L3 on almost any modern CPU, while the 40M (12.6 MB) wants a big L3.

Two response headers (COOP and COEP) that unlock shared memory, at the price of restricting third-party content.

The page must be served with Cross-Origin-Opener-Policy: same-origin and Cross-Origin-Embedder-Policy: require-corp. Then everything it embeds must opt in to being embedded. That is why this site serves its own models and scripts, and why the bot check (Turnstile) had to be tested under these headers.

Without isolation the page still works, but single-threaded.

Embedding

Model

The table that turns each token into a list of numbers the network can work with.

Each of the 3,596 tokens has a row of d numbers (d = the model width, 128 to 640). Reading a token is a table lookup. The same table is reused in reverse at the end to score the next word (“tied” embeddings), which saves memory, and is stored as int8 with one scale per row.

Scores every word cheaply in int8 with a provable error bound, then recomputes exactly only the few words that could win.

Scoring the whole vocabulary is a big fraction of each token. The engine first scores all words using the int8 activation and gets, for each, an upper and lower bound on the exact score from the known rounding error. Only words whose upper bound reaches the k-th best lower bound (about 1.5% of the vocabulary) are rescored in full precision; sampling considers only those.

Because the bounds are guaranteed, the chosen word, tie-breaking and random numbers are identical to the slow path (checked on 480 emails across models, threads, temperatures and seeds).

Doing several steps in one pass over the data, so it is read once and threads synchronise less.

Examples: RoPE and the KV-cache write happen inside the attention job, SwiGLU inside the gate/up matrix multiply, and “add the residual, normalise, quantise to int8” is one loop. Each fusion removes a trip through memory and a wait for other threads; together they were worth tens of percent for the small models.

Before each matrix multiply the input vector is scaled and rounded to whole numbers from −128 to 127.

Per token the vector is scaled so its largest value becomes 127 and rounded. Then the multiply against ternary weights is pure integer arithmetic, and one float multiply at the end undoes the scale.

Because one huge value would squash everything else after scaling, RMSNorm runs first to keep vectors well-behaved.

A CPU instruction that multiplies many pairs of small integers and adds them up in one go.

x86 has pmaddubsw (AVX2) and VNNI’s vpdpbusd; Arm has sdot/udot. They are what make int8 networks fast. WebAssembly relaxed SIMD exposes them as i16x8.relaxed_dot_i8x16_i7x16, which the browser compiles to the best instruction the CPU has.

The “i7” means one operand must fit in 7 bits (0..127): fine for our 0/1/2 weights, which keeps results exact on every CPU.

KV cache

Model

The stored keys and values of every earlier token, so each new token only computes its own.

Without it the model would redo the whole history for every new word. Its size is layers × 2 × context × width × bytes: for the 4M model at the full 256-token context, 6 × 2 × 256 × 256 × 4 bytes ≈ 3.1 MB in float32.

That is larger than the 4M model’s own 1.9 MB of weights, so the KV cache, not the weights, decides whether the model still fits in cache late in an email. Storing it as int8 would shrink it 4×.

Logits

Model

One score per vocabulary word at the last step: the highest (or a sampled one) becomes the next token.

After the last layer the model compares its final vector with every word in the vocabulary: 3,596 dot products of width d. For small models that is 10 to 30% of a token’s time, because the matrix is large relative to the rest.

The engine’s fast logits trick avoids most of this work without changing the result.

Memory-bound

Hardware

Limited by how fast data can be fetched, not by how fast the CPU can calculate.

Tokens per second ≈ memory bandwidth ÷ bytes read per token. The 21M model reads roughly 10 MB per token; at 6 GB/s from RAM that caps it near 600 tokens per second, whatever the code does. From cache the same model can go many times faster.

This is the main reason a cheap laptop and a big-cache server differ so much on this benchmark.

These models write plausible short emails but make mistakes. That is the price of fitting in a CPU cache, and it is not the point of the project.

Everything here is sized to live in a few megabytes of cache, so the models are tiny: 1 to 40 million parameters, where useful assistants have billions. Expect the right overall shape (greeting, the facts you gave, sign-off) with errors: filler sentences nobody asked for, a dropped or swapped detail, trouble when a request contains two values of the same kind (“3 and 4”), and odd output for unusual or very short requests.

The project asks how fast a language model can run when its whole working set stays in fast on-chip memory, using ordinary consumer hardware (see the Cerebras idea). Quality is being improved over time with better data and training, but the goal is the speed experiment, and a larger, smarter model would simply not fit in the cache that makes the speed possible. Bigger is also not automatically better here: the 4M scores about the same as the 21M and 40M on held-out data (validation loss).

A stand-in like <NAME_1> or <TIME> that replaces a real value in your request; the page swaps the real value back into the email.

Before the model sees “tell Mike the meeting is at 3pm” the parser (try it on the How it works page) turns it into “tell <NAME_1> the meeting is at <TIME>”. The model writes “Hi <NAME_1>, … at <TIME> …” and the page fills “Mike” and “3pm” back in.

That way a 1M-parameter model never has to memorise names, dates, amounts or numbers, only where they belong in a sentence. Training data was built the same way, which is why the model is reliable at it.

Prefill reads your request in one go; decode then writes the email one token at a time.

Prefill processes the prompt tokens, which can share work. Decode is the slow part: every new token needs all the weights and the whole KV cache, and depends on the previous token, so it cannot be parallelised across tokens.

That sequential nature makes decoding memory-bound on most hardware: tokens per second ≈ how fast the weights can be streamed from wherever they live.

Quantisation

Numbers

Storing numbers with fewer bits (32-bit float → 8-bit integer → 2 bits) to save memory and speed up arithmetic.

Fewer bits means less data to move and simpler arithmetic, at the cost of precision. Here weights are 2-bit (ternary), activations 8-bit (int8), embeddings 8-bit, and only a few per-row scales and norms stay in float.

Quantisation-aware training teaches the model to work with the rounding, which is why it keeps quality at such low precision.

Relaxed SIMD

Browser

A WebAssembly extension whose instructions may differ slightly between CPUs, which lets them map to the fastest native instruction.

It adds fused multiply-add and integer dot products. “Relaxed” means the standard allows hardware-specific results in edge cases (overflow, rounding) so the browser can emit the single best instruction. We only use it where the maths is exact, so results are identical everywhere.

The benchmark uses it when the browser supports it (most desktop browsers) and falls back to plain SIMD otherwise; the leaderboard records which one ran. It was worth 10 to 20% in our measurements.

RMSNorm

Model

Rescales a vector so its typical size is 1 before each block: simpler than LayerNorm (no mean, no bias).

Divides each number by the root-mean-square of the vector and multiplies by a learned weight. It keeps values in a predictable range, which matters a lot when the next step rounds them to 8 bits.

RoPE

Model

Rotary position embedding: position is encoded by rotating pairs of numbers in queries and keys by an angle that depends on position.

Instead of adding a learned “position 17” vector, each query/key pair of numbers is rotated by an angle proportional to the position. When two tokens are compared, the result depends on how far apart they are, which is what attention needs, with no position table to store.

In the engine it is a few multiplies per head, done inside the attention step rather than as a separate pass.

How the next token is chosen from the scores: temperature 0 always takes the best; higher values add variety.

Temperature divides the scores before they become probabilities: 0 means “always the top word” (greedy, repeatable), 1 is the model’s own uncertainty. Top-k keeps only the k best candidates (30 here) and picks among them at random by probability. The seed fixes the randomness so a run can be repeated.

The benchmark uses temperature 0 for the scored emails so every machine does identical work.

The browser computes and sends its own numbers, so a determined person can fake them.

The server rejects impossible values, limits submissions per network, strips markup and lets an admin delete rows, but cannot prove a timing it did not measure. The board is for fun. Stronger options (outlier flagging by CPU, one-time run tokens) exist and are cheap to add if it ever matters.

See also Turnstile

Memory that several web workers can read and write at once: the basis of WebAssembly threads.

The engine’s threads are web workers that share one block of memory and coordinate with atomic operations (spin-waiting, with a short sleep when idle). Browsers only enable SharedArrayBuffer on cross-origin isolated pages, because shared memory plus a precise timer was the ingredient of the Spectre attacks.

SIMD

Hardware

Single Instruction, Multiple Data: one instruction works on a whole vector of numbers at once.

A 128-bit register holds 16 int8 values or 4 float32s; AVX2 doubles that to 256 bits and AVX-512 to 512. The engine’s matrix multiply is SIMD end to end.

Browsers currently give WebAssembly 128-bit vectors (WASM SIMD), so the browser version handles half the width of the native AVX2 engine per instruction, which is the main reason native can be faster on the same CPU.

A cheap draft guesses several tokens ahead and the real model checks them all in one pass. Not used here.

It helps when decoding is memory-bound: checking four guessed tokens costs about the same memory traffic as generating one, so when the guesses are right you get several tokens for the price of one. The output is unchanged because the real model has the final say.

We do not use it: the small models already run from cache, where the bottleneck is overhead per token rather than memory, and a checker for several tokens at once would need new kernels. It may make sense for the 21M and 40M models on small-cache CPUs.

Speed once everything is warm: decoding to the full context, repeated, ignoring the first cold run.

The score is not a single lucky run: after a warm-up the model decodes all 256 positions ten times in a row. That is slower per token than a short email (attention and the cache are larger) but far more repeatable, and it is the same protocol the native engine’s benchmark uses, so the two can be compared.

The training trick that lets a network learn through rounding: pretend the rounding was not there when computing gradients.

Rounding to −1/0/+1 has zero gradient almost everywhere, so nothing would learn. The forward pass uses the rounded ternary weights, but the backward pass treats the rounding as the identity and updates hidden full-precision “shadow” weights, which are re-rounded every step. Only the ternary weights are kept for inference.

SwiGLU

Model

The feed-forward block: two projections, one passed through a smooth gate (SiLU) and multiplied into the other, then projected back.

A token’s vector is projected to a wider one twice (“gate” and “up”); the gate goes through SiLU (x · sigmoid(x)) and multiplies the up-projection; a third matrix (“down”) projects back. In the profiler these are the w13 and w2 matrices, and together with the attention projections they are most of the arithmetic.

Training examples written by bigger models, then filtered, instead of collected from real people.

About 175,000 request → email pairs: one LLM pass invents realistic requests (casual, typo-ridden, with negations, conditions, family words, titles), our own parser converts them to placeholder form, and a second pass writes the email from that request alone. Filters discard emails that invent numbers, drop a placeholder, use stock phrases, or break the layout.

Building the request side with the same parser used at run time is what makes the small model dependable on real input.

Every weight in the big matrices is −1, 0 or +1, so “multiplying” is add, subtract or skip.

Three values carry log₂(3) ≈ 1.58 bits of information, hence “1.58-bit”. The model is trained with this restriction (see straight-through estimator), not compressed afterwards. Each matrix has one floating-point scale to restore the magnitude.

Stored as 2 bits each (four to a byte) a 21M model is 7 MB instead of 84 MB in float32, and the arithmetic becomes small-integer additions that SIMD does very fast.

Keep the whole model in fast on-chip memory so generating text is not limited by memory bandwidth: this project tests a tiny version of that on ordinary CPUs.

Generating a word means reading every weight once, so on most hardware the speed limit is how fast the weights can be fetched from external memory (memory-bound). Cerebras’s wafer-scale chips attack that directly by holding the model’s weights in on-chip SRAM instead, which is why they can generate text far faster than GPUs reading from external memory. This project is inspired by that idea and is not affiliated with Cerebras.

The question here: what happens if you do the same thing at a tiny scale on hardware everyone owns? A CPU’s L2 and L3 caches are on-chip SRAM too, just megabytes instead of gigabytes. So the models are made small enough (0.6 to 12.6 MB) to live there, and the leaderboard measures how fast different CPUs run them. The trade-off is that models this small are far less capable (see quality); the experiment is about speed.

Helper threads that wait for work, run their slice of each matrix, and report back: tens of times per token.

Each layer has several parallel steps, so a token involves about 30 hand-offs between the main thread and helpers. At 0.2 ms per token the cost of each hand-off matters: helpers spin briefly instead of sleeping, each keeps its “done” flag on its own cache line (so they do not fight over memory), and idle helpers sleep so an idle tab uses almost no CPU.

More threads give diminishing returns: each thread gets a smaller slice, but the hand-off cost stays.

Thread tuning

Benchmark

The benchmark tries 1, 2, 4, 8… threads and keeps the fastest for each model.

Because browsers cannot choose cores and hyper-threads share hardware, the best thread count is not obvious (8 was slower than 4 on the 4-core laptop). Each model is tuned separately, and the thread count used is shown on the board.

How long from pressing Generate until the first word exists: parsing, reading the prompt and one sampling step.

It is the latency you feel. The benchmark measures it on ten real short emails and reports the median. On a fast CPU it is under a millisecond for the smallest model; on a phone a few milliseconds.

Browsers deliberately blur the clock; isolated pages get about 5 µs, others much coarser.

performance.now() is rounded to limit timing attacks. A single token takes about 200 µs, so the benchmark times whole runs and divides, and reports the median of ten so a single coarse reading cannot distort the result.

Token

Model

The unit the model reads and writes: a common word, a punctuation mark, a placeholder, or a single letter.

Most tokens are whole words from a 3,500-word vocabulary. Others are punctuation, a capital-letter marker, a placeholder like <NAME_1>, or (for a word the vocabulary lacks) individual letters between start/end markers.

An email is 25 to 60 tokens, so “5,000 tokens per second” is roughly 100 short emails a second. Time to first token (TTFT) and tokens per second are the two numbers on the leaderboard.

The network design behind modern language models: layers that let words look at each other, then transform each word on its own.

These are small decoder-only transformers: 4 to 10 layers, each 128 to 640 numbers wide. A layer is: normalise → attention (mix information from earlier tokens) → add back → normalise → a feed-forward block (transform each token alone) → add back.

The model reads your request and the email written so far, and predicts one token at a time. Everything else on this site is about doing that fast.

Turnstile

Benchmark

Cloudflare’s free, privacy-friendly captcha alternative, used to keep bots out of the submit form.

A small widget proves the visitor is a real browser (usually invisibly) and gives a one-time token; the server verifies it with Cloudflare before accepting a score. Free with no challenge limit.

It loads a third-party iframe, which cross-origin isolation can block, so it was tested under this site’s headers.

How surprised the model is, on average, by examples it was not trained on: lower is better.

It is the average cross-entropy per token on held-out emails: 0.59 for the 1M model down to about 0.48 for the 4M and 21M (the 40M is 0.49: bigger did not help on this data). When training loss keeps falling but validation loss rises the model is memorising, called overfitting, and the best checkpoint is kept instead of the last one.

The 4M matching the 21M is why it is the interesting model here: nearly the same quality at a fraction of the size and cache footprint.

The list of 3,596 tokens the model knows: 3,500 words plus punctuation, letters and placeholders.

A big model uses 30,000+ tokens. Here the table is kept tiny on purpose: the final step scores every vocabulary entry for every word written (logits), and the table is also the embedding table, so its size directly costs speed and cache space. At width 512 and 1 byte per number it is about 1.8 MB, a quarter of the 21M model.

Rare words that do not fit are handled by copy slots instead of growing the vocabulary.

Web worker

Browser

A background thread for JavaScript and WebAssembly, so heavy work does not freeze the page or disturb timings.

The engine runs inside a worker, and each helper thread is another worker. The page only receives results, so scrolling, rendering and the progress bar never compete with the model for the same thread.

A compact binary format that browsers compile to near-native machine code.

The engine is about a thousand lines of C, compiled to a ~95 KB .wasm file. The browser compiles it to machine code for your CPU when the page loads. It runs in a sandbox with no access to anything outside its own memory.

Good fit: predictable performance and SIMD. Limits: 128-bit vectors only, and threads need shared memory.

The 128-bit vector instructions WebAssembly offers, supported by all major browsers.

One v128 register holds 16 int8, 8 int16, 4 int32 or 4 float32 values. The engine’s kernels (weight unpacking, dot products, the softmax exponential) are written with these. If a browser lacks them the page cannot run the model at all, so it checks at start.

The name of your graphics chip, which browsers expose and which sometimes reveals the CPU.

Browsers hide the CPU model but show the GPU name. On Apple silicon that name is the chip (“Apple M2”), so the leaderboard can fill the CPU in for you. Elsewhere it only hints (for example “AMD Radeon Vega 8” implies a Ryzen APU) and is stored with the score as a clue.

See also WebGPU

WebGPU

Browser

The browser API for running compute on the GPU. This project does not use it; here is why and when it would make sense.

WebGPU lets a page run shader programs on the graphics card. For big models it is the right tool. For a model of a few megabytes, the work per token is small: a GPU would need dozens of separate dispatches per token (one per layer operation), and getting the chosen word back to JavaScript costs about a millisecond on its own, which is larger than the 0.2 ms the CPU engine takes per token.

On integrated graphics (most laptops) the GPU also reads the same RAM, so it gets no bandwidth advantage. Our estimate (not measured: we have not built a WebGPU path) is that a CPU running from cache beats it for models this size; a discrete GPU with the whole generation loop kept on the GPU could win. The graphics string is only read to hint at your CPU.

A 1.58-bit (ternary weights) email-writing transformer, running on your CPU in this tab via WebAssembly SIMD and threads. Scores are self-reported by browsers: for fun, not for procurement.