Dev log

Can a language model live in a CPU cache?

The dev log of a tiny, from-scratch LLM that writes emails in your browser, and a leaderboard to see how fast your CPU can run it.

11 minute read · a speed experiment, not a product

The question

Cerebras has a beautiful idea. A language model writes one word at a time, and to write each word it has to read every one of its weights. On most hardware those weights sit in external memory, so the speed limit is not the arithmetic, it is how fast the weights can be fetched. Cerebras builds a chip the size of a dinner plate with tens of gigabytes of on-chip memory, keeps the whole model on it, and the bottleneck goes away (the Cerebras idea; this project is inspired by it and not affiliated with them).

I kept wondering what a tiny version of that looks like on hardware I already own. Because your CPU has on-chip memory too. It is called cache: a megabyte or two of L2 per core, and anywhere from 4 to 96 MB of shared L3 depending on the chip. No useful language model fits in there. Unless you build one that does.

So that became the experiment: make language models small enough to live in a CPU cache, write an engine that runs them flat out, and measure how fast they go, in a browser tab, on whatever CPU someone happens to have. The leaderboard at the front of this site is the measuring stick.

A warning up front: the quality is low, and that is expected. These models write plausible short emails and then make mistakes: filler sentences, a dropped or swapped detail, nonsense on unusual requests. That is the price of fitting in a cache, and it is not what the project is about. I am improving it over time, but a bigger, smarter model would not fit in the cache that makes the speed possible. The question was always how fast, not how smart.

The rules I set myself

  • Cache-sized. The whole model between 0.6 and 12.6 MB, so the smaller ones fit in L2 and all of them in a decent L3.
  • Built from scratch. My own tokenizer, model, training loop and inference engine. No llama.cpp, no PyTorch at runtime, not even Microsoft’s bitnet.cpp.
  • No GPU, no server. Inference happens in a static web page, on the visitor’s CPU.
  • A narrow job. Something small enough for a tiny model to do decently: turn a one-line request into a short, friendly email.

The last rule is the one that makes the rest possible. A general chatbot needs billions of parameters. “Write email to john saying that is all good” needs a few million, if you are smart about what you ask the model to remember.

your request→request parser→the model→fill the values back→your email

A model built from nothing

The model is a small decoder-only transformer: between 4 and 10 layers, 128 to 640 numbers wide, with rotary positions, RMSNorm, a SwiGLU feed-forward block and a 256-token context. The interesting part is the weights. Following the BitNet b1.58 recipe, nearly every weight is just −1, 0 or +1 (ternary). That is about 1.58 bits of information per weight, so four of them pack into one byte and a 21-million-parameter model weighs 7 MB instead of 84. “Multiplying” by a ternary weight is adding, subtracting or skipping, which CPUs are very good at.

The second squeeze is the vocabulary. A normal model knows 30,000+ tokens. Mine knows 3,596: the 3,500 most common words, punctuation, single letters as a fallback, and some special markers. That matters more than it sounds, because the last step of every word scores the whole vocabulary (logits), and the table doing it is a big slice of a tiny model.

The trick that makes a tiny model usable

A model with a few million parameters cannot memorise names, dates and amounts. So it never sees them. Before the model reads your request, a small parser pulls the specifics out and swaps in placeholders. The model learns where a name goes in a sentence, not what the name is, and afterwards the page puts your real values back. Words it does not know, like “gazebo”, travel the same way as copy slots. Try the real parser below; it is the same JavaScript file the benchmark uses.

loading the parser…

This parser is a contract. The training data was built with the exact same code that runs in your browser, which is why a model this small is dependable on real input. Most of the bugs I found later were the contract being broken in some small way; more on that below.

Teaching it with other models’ homework

My first plan was the Enron email corpus: real emails, redact the names with spaCy, invert each one into an instruction. It gave me stiff corporate prose, instructions that did not match the emails, and a dependency (spaCy) that cannot run in a browser anyway. The first model trained on it was cheerfully off-topic.

So I switched to synthetic data from bigger models through OpenRouter, which lets you pick the cheapest model that is good enough. I settled on Ling 3.0 Flash. The recipe has two stages on purpose:

  • an LLM invents realistic requests (casual, typo-ridden, with negations, conditions, titles, family words);
  • my own parser converts each request to placeholder form, exactly as the browser will;
  • a second pass writes the email from that placeholder request only, so it can never lean on details the model will not see;
  • filters throw away emails that invent numbers, drop a placeholder, use stock phrases (“I hope this email finds you well”), repeat themselves or break the layout.

Two lessons cost me real money before I learned them. Ling quietly spends thousands of hidden “reasoning” tokens per answer unless you switch that off, which made early runs both slow and expensive. And reusing HTTP connections per thread, instead of opening a fresh one per request, was worth about a 9× speed-up. With both fixed, roughly five dollars of credit bought about 190,000 raw pairs, of which about 175,000 survived the filters.

Training on rented GPUs

Training happens on RunPod: a rented GPU for 20 to 35 cents an hour, used for well under an hour per run. I wrote a launcher that does the boring parts every time: create a pod, check that the GPU genuinely works, upload the data once, start training, poll, download the results, and always delete the pod, even when something crashes. The weights are trained in full precision but rounded to ternary on every forward pass, learning through the rounding with a straight-through estimator.

Renting from a marketplace is a lottery, and the launcher grew a scar for each loss. One host’s GPU drivers were broken (“CUDA unknown error”). One never answered SSH. One accepted the connection and then could not finish uploading 12 MB, three times running, so the launcher learned to skip hosts by address. Sometimes there simply are no machines. Each of those is now a retry instead of a wasted hour.

Overlap everything

The workflow turned into a pipeline with three things in flight at once. While one set of models trained on a RunPod GPU, the next dataset was already being generated on OpenRouter, and I was testing the previous model locally. The pod itself ran two models side by side on one GPU (the 21M and the 40M), which cost barely more than one. A run is about 50 minutes and about 25 cents, so the whole project’s GPU bill is a dollar or two.

A chart taught me something counter-intuitive. On the first 66,000-pair dataset, validation loss bottomed out after about six passes over the data and then got worse (0.496 → 0.596) as the model started memorising. “Train it for longer” was the wrong instinct. The fix was more varied data, not more steps, and the best checkpoint is kept instead of the last.

Every failure became a dataset

The loop that actually improved the model was dull and effective: try to break it, find a category of failure, generate data aimed at that category, retrain, repeat.

VersionWhat went wrongWhat I changed
v1Enron-era data: off-topic, invented details, long waffle.Switched to synthetic, two-stage, filtered “faithful” pairs: short, friendly, only the facts given.
v2Recipients that were not first names (“the team”, “my landlord”) produced greetings with an unfilled <NAME_1> tag. Unusual words were swapped for common ones (“the gazebo” became “the doctor”).A second dataset of edge cases (titles, groups, roles, negations, conditions) and copy slots for rare words.
v3“Hi john that’s great thanks” was not even recognised as addressed to John.Ten request styles (greeting-led, thanks, apologies, “turn into email:”, keywords, replies, rambling) and family words, plus parser fixes.
v4“tell mike …” lost Mike. The name had leaked into the vocabulary as an ordinary word.Discovered names from the data, added a hand-written list of common first names, patched the tokenizer.
v5The 4M and 1M sizes were needed to test the cache hypothesis properly.Retrained four sizes (1M, 4M, 21M, 40M) on the merged data with a 256-token context.
Some bugs were embarrassingly basic. The parser only recognised “tell Mike” when the verb was lowercase, so a capitalised “Tell Mike” never matched until I finally tried it.

It is still far from perfect. Two values of the same kind in one request (“3 and 4”) confuse it, “ping the boss” produces nonsense, and every model pads with filler now and then. Those are the next round of data. They are also why the quality note above is not modesty.

An engine in a thousand lines of C

Running a ternary model properly needs a custom engine, so I wrote one: about a thousand lines of C that compile to a 95 KB WebAssembly file. It unpacks the 2-bit weights with shifts and masks and multiplies them against 8-bit activations with SIMD integer dot-product instructions. Helper threads share memory through a SharedArrayBuffer, which is why the site has to be cross-origin isolated. The same C also builds natively, which gave me a reference to test the browser against.

The first browser numbers were sobering: a few hundred tokens a second for the 21M model. So I added a profiler to the page that times each part of a token, and the optimisation became a game of finding whatever was biggest:

  • A faster kernel. Using the browser’s relaxed SIMD dot product, summing in 16 bits like the native code, was worth about 20%.
  • Dodging what the browser emulates. Some WebAssembly instructions have no x86 equivalent and expand into many. The saturating float-to-int conversion and 8-bit shifts were two: rewriting them (a rounding trick, 16-bit shifts plus a mask) gave identical results much faster.
  • Fewer, fatter steps. RoPE and the cache write moved inside the attention step, SwiGLU inside the matrix multiply, and normalise-and-quantise became one pass (fused kernels).
  • A leaner [[thread-pool|thread pool]]. An idle open tab used to burn about 20% of a CPU core spinning. It now uses 0.1%.
  • [[fast-logits|Fast logits]]. Score the whole vocabulary cheaply in int8 with a provable error bound, and recompute exactly only the ~1.5% of words that could win. Same output, less work.
  • [[jit-tiers|Warm-up]]. The browser compiles WebAssembly twice, quickly and then well, so the first email after loading was 2 to 3 times slower than the tenth. The page now runs a short warm-up before it says “ready”.
Steady-state tok/s, 4 threads, on the dev laptopBeforeAfter
1M model10,40013,200
4M model3,1004,200
21M model1,0701,110
40M model660675
The gains shrink as the model grows: the 21M and 40M live in RAM on this laptop (4 MB of L3), so they are limited by memory, not by my code. That is the hypothesis showing up in the data.

Every change had to keep the output identical, so the test suite became part of the craft: the request parser is compared against the native one on 44 requests, the browser output against the native engine, one thread against four, and fast logits against the full computation on 480 emails across models, seeds and temperatures. When two words tie within 0.001 in the score (out of a range of about 16), floating-point order alone can pick either, so the test now reports that as a near-tie with the gap, instead of crying wolf.

What a server CPU does with it

The dev laptop has 4 MB of L3, so most of the models spill out of cache. To see the hypothesis properly I rented a CPU-only RunPod machine: an AMD EPYC 4564P, 16 cores, 1 MB of L2 per core and 64 MB of L3, about 56 cents an hour. A script ran the same benchmark the site uses (three repeats of ten runs each in this first pass) for the native engine and for the browser build in headless Chromium, and the pod was gone about two minutes later.

4 threads, tok/sDev laptop (browser)EPYC (browser)EPYC (native)
1M13,20024,10024,800
4M4,20012,50013,200
21M1,1103,9006,100
40M6802,8003,800
The 4M model writes an email (25 to 60 tokens) in roughly 2 to 5 milliseconds on the server.

That is the answer, at least a first one: yes. Put a model in a big enough cache and even a browser writes thousands, on a good CPU tens of thousands, of words a second. The native engine is faster still because it can use 256-bit vectors and pin threads to cores, which browsers cannot (cores and threads). This EPYC also has AVX-512 with VNNI, which my native engine does not use yet, so there is headroom left.

Shipping it for free

The leaderboard is a SvelteKit app with DaisyUI, hosted on Cloudflare Pages, with scores in D1 (serverless SQLite) and Turnstile bot protection built in, ready to switch on. All of it fits in the free tiers, and the models are just static files. The benchmark page tries 1, 2, 4, 8 threads, keeps the fastest, runs ten steady-state passes for the score, and measures time to first token on ten real emails.

The gotchas were instructive. Cross-origin isolation restricts what a page may embed, so I tested the captcha under the same headers before trusting it. Browsers hide the CPU model, so the site fills it in where it can (Apple silicon names itself through the graphics string) and otherwise asks you once. And scores are self-reported: the server rejects impossible numbers and rate-limits, but it is a fun board, not a benchmark authority.

What is next

  • An AVX-512/VNNI kernel for the native engine, since server CPUs are sitting on instructions I am not using.
  • An int8 [[kv-cache|KV cache]]. It would shrink the cache 4× so the 4M model fits in L3 again even on a small laptop.
  • A [[webgpu|WebGPU]] experiment to replace my guess (“the CPU wins at this size”) with a measurement.
  • [[speculative-decoding|Speculative decoding]] for the bigger models on small-cache CPUs, where they are memory-bound.
  • Better data, so the failures above stop happening. The quality will keep creeping up; the speed question is already answered.

The real data, though, has to come from other people’s machines. If you have an unusual CPU, an old laptop, a new Mac or a phone, run the benchmark and put it on the board. I would love to see where the cache cliffs are.

A 1.58-bit (ternary weights) email-writing transformer, running on your CPU in this tab via WebAssembly SIMD and threads. Scores are self-reported by browsers: for fun, not for procurement.