commonllama benchmark

Every run, one chart.

Sort any column, search any name, and set the context depth. A measured number says measured. A computed number names its method. The chart never blends the two silently.

loading
Published runs
Model Architecture Quantization Hardware class Site class RAM Backend Status Storage Provenance Fit Qualified at depth Cook Reload Memory estimate Saved context
Loading chart data…

Cook is the whole way in, measured end to end: process start, model load, context create, and the cold prefill that reads the document. Reload is the whole way back to that same context: process start, model load, context create, and the import of the saved state. Fit is an observed lower bound from one run on one machine class, not a maximum. Qualified at depth answers the slider position itself: qualified when at or under an observed passing rung, failed when at or past an observed failure, unresolved between the two, or not measured when no rung has been observed for this row. Speeds at a detent are measured; between detents they are interpolated log linear between observed rungs and labelled so. The saved-context and memory-estimate rate varies with context (it is a property of the model, not a fixed number), so both interpolate the per-token rate from the row's density curve: exact at a measured point, interpolated strictly between two, and not observed outside the measured range, never extrapolated.

Models

Models, most active first.

Each model keeps its own page of runs. The order follows recent publishing activity.

Models on the commonllama benchmark
ModelRunsMost recent run
Granite 4.0 H Tiny122026-09-09
Qwen3 4B112026-08-12
Nemotron 3 Nano 4B22026-08-12
Qwen3.5 4B22026-08-12

commonllama benchmark records are dedicated to the public domain under CC0 1.0. Use them for anything.