I believe knowledge is a right. Every person on earth deserves access to the best information they can get, on their own terms, and local AI is the closest thing to that freedom anyone has ever had. Even a 4B model, small enough to run on an ordinary laptop, holds an insane amount of knowledge, and once it’s on your own machine nobody can meter it, change the terms on you, or take it away. There are two catches. You need enough RAM to run it, and the tools around it have to be good enough to make it worth running. On everyday hardware, a big part of that is something pretty basic: managing context well. commonllama is my answer to that challenge, so the knowledge in these models actually reaches the people who could use it.
Here’s what I mean by context. Picture a small shop or a classroom running a local model. There’s a body of material it works from that barely changes, like the handbook, the course, the product list, and other sources of truth. Then there’s the task in front of it, which changes all the time, such as a question, a document, or yesterday’s conversation. Before a model can answer anything, it has to read its instructions and material (a step called prefill), and if it throws that reading away every time the job changes, it spends most of its life re-reading things it already knew. On a fast machine that’s an annoyance. On the machines most people actually have, it’s the difference between a tool you can use and one you give up on.
Retain what the model already read
When a model reads something, it builds an internal memory of the material, usually called the KV cache. commonllama maintains that memory. You can hold it in RAM or save it to disk, bring it back when you need it, and carry on from where the model left off instead of paying for the reading again.
We organize memory in three layers: Core Memory for the part that rarely changes, like that handbook, a codebase, or a set of standing instructions; Working Memory for each task built on top of it; and Tail Memory for the live conversation at the end. One model stays loaded while different Working Memories swap in and out over the same Core so that a small side job, like a safety check on what an agent is about to do, can use the same model instead of loading a second copy.
Keeping memory also means you decide when to use it. When a reviewer sends back notes, you want the same coding agent (with everything it already knows about the code) to make the fixes, and you want the reviewer looking with fresh eyes every time. And because a saved memory marks an exact point in a conversation, you can go back to it and branch off in a new direction without starting over.
Agent work multiplies re-reading
Last month I was running a multi-agent code review on an RTX 5090 (a high-end GPU), which is about as fast as consumer hardware gets. I’d left a probe in that timed prefill, and when the review finished, about 20% of the total time had gone to reading, much of it spent on agents going back over documents the model had seen only minutes before. On slower hardware the reading takes an even bigger share, so I prioritized connecting commonllama to an agent coding harness (OpenCode) and adding support for the Responses API, allowing agent tools to talk to it directly. Re-reading piles up quickly in agent work, which is where commonllama can make a big difference.
A saved memory, tested on three machines
What finally convinced me wasn’t the speed. It was getting the same answers back.
We gave Qwen3.8 27B a set of standing instructions (about 12,000 tokens), had a GPU read them once, and saved that Core Memory as a single encrypted file. Then we copied the file and woke it up on other machines: a desktop running on its CPU alone, a mini PC running on its CPU alone, and the same mini PC on its integrated GPU. All three picked up from the same point in the conversation, despite running on different hardware and backends, and reached the same decision on each of the 15 test questions. Getting there took months of work, and it proved the idea behind commonllama: a memory made once is a memory you can hand to anyone’s machine and trust.
The time difference is what makes running a local model practical. On the CPUs, loading the saved memory took two to two and a half seconds, and reading the instructions from scratch took more than eight minutes. A 27B model on a CPU only writes a few tokens a second (nobody is calling that fast), but waiting around 13 seconds for the first word is workable while waiting eight minutes isn’t. People usually talk about speed as getting answers sooner. I care more about what speed does for less powerful machines, because every minute you take out of the wait makes everyday machines good enough to use.
Do the heavy work once
The time saved is only half the story. Those eight minutes keep the processor running at full power, burning energy just to rebuild something that already exists, while loading a saved memory skips that step every single time. Loading a file off a drive is more economical on every level than redoing the reading, even on a fast GPU. That’s the shift from wasteful to efficient, and it doesn’t even have to happen while anyone is waiting since a memory can be prepared ahead of time, like overnight or whenever a machine is sitting idle.
Let whichever machine is best at the heavy lifting do it once, and let everything else pick it up.
RAM as the constraint
RAM is getting more expensive. Big AI companies have bought and locked up an enormous share of the world’s memory for years to come, and everyone else is paying for it. If local AI only works for people who can buy a lot of RAM, knowledge stops being a right and becomes a privilege. So I’d rather make small jobs work well on modest machines, keep one model loaded and swap Working Memories over the same Core, and let machines share work with each other.
Your memory is your data
A saved memory is a record of what you were working on: your documents, your conversations, your decisions. I treat it as personal data. If a laptop goes missing, the memories on it shouldn’t be readable by whoever ends up with it. That’s why everything commonllama writes to disk is encrypted with AES-256, checked when it’s loaded, and wiped between swaps so that one job never picks up what another left behind.
Where commonllama stands
commonllama is built on llama.cpp, which is the engine that actually runs the model. What we’re building is everything around it that keeps, moves, and protects the work the model has already done. I think of commonllama as the vehicle rather than the engine, and as llama.cpp and the models keep improving, it carries those improvements along.
The project is in closed alpha, open source under the Apache 2.0 license, and it runs on CPUs and GPUs from NVIDIA, AMD, and Apple, on Windows, Mac, and Linux. In my next post, I’ll lay out our plans for version 1.0: a fast cache that reuses prepared memory automatically, and cheap branching built on a shared Core that never has to be read twice.
When modest machines can handle more, the everyday person’s access to knowledge expands, on their own terms. That’s my goal for commonllama.
