Benchmark · Decision models
jevalt-bench
3,116 written situations in English, Turkish and German for decision models that pick an option and give a probability.
Dataset (Hugging Face)
1,300 downloads on Hugging Face, as of October 10, 2026.
Results
Accuracy and Brier score on the test questions
| Language | Kev-4B | JevAlt | Brier, Kev-4B | Brier, JevAlt |
|---|---|---|---|---|
| English | 84.7% | Deem-4B 94.7% | 0.256 | 0.091 |
| Turkish | 87.1% | Karar-4B 96.8% | 0.214 | 0.058 |
| German | 81.1% | Wähler-4B 92.0% | 0.298 | 0.119 |
Limits
Test and training rows come from the same data pipeline, so these scores measure JevAlt’s own task. On the German news set 10kGNAD, Kev-4B scores 65.3% and Wähler-4B 62.5%.
Run it
Load one language with the Hugging Face datasets library. To score your own model on these rows, point JevOss at any server that speaks the Jev API.
from datasets import load_dataset
test = load_dataset("mertkayacs/jevalt-bench", "en", split="test")
print(test[0]) Cite
@techreport{kaya2026jevaltreport,
author = {Kaya, Mert},
title = {JevAlt: Technical report},
institution = {Eschatia Labs},
year = {2026},
url = {https://eschatialabs.com/research/jevalt-report/},
note = {Technical report, not peer-reviewed}
}