commonllama benchmark · model
Qwen3.5 4B
Every number on this page is an observed measurement from a benchmark run, shown with the run that produced it.
Pre-release: these numbers come from commonllama harness 0.0.0, a build not yet released. They are real measurements, not estimates; released-harness numbers will replace them.
Observed facts
What the harness read from the file.
| Model file digest | c6f6972469c72b8081f3e5b07c1f24379330fccfc2f7f4fd3b70e330d98465ec |
|---|---|
| Architecture | not observed |
| Quantization | not observed |
| Max context | not observed |
| Model file size | 2.52 GiB (2,707,514,208 bytes) |
| Fact provenance | unknown |
Context capacity
How much context this model has carried.
Qualified to at least 131,072 tokens of context, an observed lower bound, not a maximum.
Observed on goldspier-rtx5090 vulkan. A frontier is a property of one run on one machine class, never a promise about another.
Saved context
What a saved context costs on disk.
Each curve keeps to one named key: this model, one backend. Unlike measurements are never folded together.
The KV cache is the AI’s memory of the conversation so far. Saving it to disk lets the AI return to the conversation and load that memory back in, rather than read everything from the start again.
| Backend | Context | Bytes per token | Saved bytes |
|---|---|---|---|
| cuda | 4,096 | unavailable | unavailable |
| cuda | 8,192 | unavailable | unavailable |
| cuda | 16,384 | unavailable | unavailable |
| cuda | 32,768 | unavailable | unavailable |
| cuda | 65,536 | unavailable | unavailable |
| cuda | 131,072 | unavailable | unavailable |
| cuda | 4,096 | unavailable | unavailable |
| cuda | 8,192 | unavailable | unavailable |
| cuda | 16,384 | unavailable | unavailable |
| cuda | 32,768 | unavailable | unavailable |
| cuda | 65,536 | unavailable | unavailable |
| cuda | 131,072 | unavailable | unavailable |
| vulkan | 4,096 | unavailable | unavailable |
| vulkan | 8,192 | unavailable | unavailable |
| vulkan | 16,384 | unavailable | unavailable |
| vulkan | 32,768 | unavailable | unavailable |
| vulkan | 65,536 | unavailable | unavailable |
| vulkan | 131,072 | unavailable | unavailable |
| vulkan | 4,096 | unavailable | unavailable |
| vulkan | 8,192 | unavailable | unavailable |
| vulkan | 16,384 | unavailable | unavailable |
| vulkan | 32,768 | unavailable | unavailable |
| vulkan | 65,536 | unavailable | unavailable |
| vulkan | 131,072 | unavailable | unavailable |
Provenance
Every run behind the numbers.
| Run | Provenance | Status | Hardware class | RAM bucket | Backend | Storage | Frontier |
|---|---|---|---|---|---|---|---|
0032b87abd93 | migration | complete | goldspier-rtx5090 | unavailable | vulkan | ssd | 131,072+ |
6cb47bc304bf | migration | complete | goldspier-rtx5090 | unavailable | cuda | ssd | 131,072+ |
A partial run is content, not a failure: its rungs carry typed reasons on the run page. Partial and complete runs are never mixed inside one aggregate on this page.
commonllama benchmark records are dedicated to the public domain under CC0 1.0. Use them for anything.