Benchmark · AI on your own computer
Tholos-Bench
160 workspace tasks for small AI assistants, scored on what the workspace looks like when the assistant has finished.
Results
One Kaggle T4 graphics card, llama.cpp, Q4_K_M model files (compressed to about 4 bits per weight), output held to a JSON schema
| Model | File | Passed (of 160) | Hidden instructions (of 12) |
|---|---|---|---|
| Qwen3.5-2B | 1.28 GB | 97 | 4 |
| MiniCPM5-2B | 1.56 GB | 112 | 8 |
| Tholos-2B | 1.56 GB | 137 | 10 |
| Llama 3.2 3B Instruct | 2.02 GB | 37 | 1 |
| Granite 4.2-3B | 2.24 GB | 129 | 7 |
| Qwen3.5-4B | 2.74 GB | 141 | 6 |
Limits
Tholos-2B was trained on top of MiniCPM5-2B and is the only model in the table trained for this step format. On a computer’s main processor (CPU) with llama.cpp, MiniCPM5-2B passes 117 and Tholos-2B 134, so compare scores only within one setup.
Run it
Start a model server, install Tholos and run the full benchmark. Add --only with a category name to run part of it.
llama-server -hf mertkayacs/Tholos-2B-GGUF:Q4_K_M --jinja -c 16384 -t 4 -a tholos-2b --host 127.0.0.1 --port 8080
uv tool install git+https://github.com/mertkayacs/tholos
tholos bench --base-url http://127.0.0.1:8080/v1 --model tholos-2b Cite
@misc{kaya2026tholosbench,
author = {Kaya, Mert},
title = {Tholos-Bench: Workspace Tasks for Small AI Assistants},
publisher = {Eschatia Labs},
year = {2026},
url = {https://huggingface.co/datasets/mertkayacs/tholos-bench}
}