Reusable memory for local AI
Come back to the work without starting over.
You found the right material. You worked out what matters. Then the next session asks you to feed it all back in and wait. That gets old when you have a class to prepare, a decision to revisit, or a client waiting on an answer.
commonllama lets you save a model’s prepared starting point and bring it back when you need it. Your knowledge stays useful; the wait gets smaller on hardware you own.
Keep what you prepared
A starting point you can check and keep.
A course reader, a set of project notes, the rules behind a safety check: you know which material belongs in the room. Preparing it for a model takes time. commonllama can save that prepared context as a memory.
Ask the questions you would ask before trusting it. If the answers show a gap, change the material and prepare it again. Once it works for you, save it.
Next week, or next term, load that memory and pick up from the same prepared starting point. Keep different versions for different work; bring back the one you need.
This is useful for continuity. It also changes the wait: when loading a saved memory is quicker than preparing the same material again, you get back to the question sooner.
Prepare. Check. Save. Return.
How it works
Keep the foundation. Change the task.
locked
locked
Animations licensed CC BY 4.0. Attribution: commonllama project, clrbx.org.
The time tradeoff
When is it quicker to return?
There are two ways back to the work: prepare the material again, or load the memory you saved. The crossing point depends on your model, the amount of context, your machine, and where the memory is stored.
The chart below illustrates that comparison with sample data while we prepare measured records.
Charts licensed CC BY 4.0. Attribution: commonllama project, clrbx.org.
Built on llama.cpp.
We build on llama.cpp so you can use the models and hardware that fit your work. Then we add the part that lets you keep prepared context and come back to it.
llama.cpp provides
- Model loading and inference
- Quantization
- KV cache primitives
commonllama adds
- Save and load prepared model memory
- Keep a foundation in place while tasks change
- Encrypt saved sessions and clear working memory
- Manage multiple memories from one runtime
Strata keeps the right memory close.
Some work calls for a memory you chose. Some work just needs to pick up where you left off. Strata gives both a place in your day.
Ready when you return.
You open the weekly report, the familiar set of forms, or the unit you teach again. Strata recognizes prepared work and brings its saved memory close, so you can spend your time on what changed.
You decide how much space it can use. The work you reach for most stays close at hand.
The one you chose.
Some work needs a deliberate starting point. Name the memory you checked and load that one for an exam review, a safety check, or a decision you may need to explain later.
When the material changes, make another version. You can return to the earlier one and see which starting point you used.
Many seats
Let several people work from one prepared starting point, including a class sharing a local machine.
Across the building
Prepare on the strongest machine you have and make that work available to other machines on your own network.
Hardware profiles
Fit the memory to the machine in front of you, from a small local computer to a workstation.
Honest ETAs
See an estimate based on your machine before you start preparing a large body of material.
Your material is still your material.
A saved memory can reflect your documents, instructions, and conversation. We treat it as work you would want to keep under your control.
Sealed on disk.
Saved memories, checkpoints, and adapters are encrypted under your key with AES-256 and the platform’s own crypto.
Cleared from RAM.
When a session closes, the runtime clears the working memory it held. The details of your work should not linger just because you closed a window.
Reports to you alone.
The runtime watches for known trackers in the app around it and warns you when it finds one. You should know when another piece of software may be listening.
Whole on one machine.
The runtime can work without a network connection. A classroom, a workshop, or your desk should not have to depend on someone else’s server.
The next time you need it
Pick up where you left off.
You finish a lesson and hear the question you wish you had asked while preparing it. Next time, you load the material you already checked and ask that question. You can keep the useful work and change what needs changing.
That is the promise behind commonllama: knowledge you worked to gather can stay in use, and local AI can be practical to return to on hardware you own.