All benchmarks

Benchmark · Decision models

jevalt-bench

3,116 written situations in English, Turkish and German for decision models that pick an option and give a probability.

Dataset (Hugging Face)

1,300 downloads on Hugging Face, as of October 10, 2026.

Each row describes a situation, such as an email, a policy or a chat message, and asks one or more typed questions about it: choose an option, give a rating or estimate the probability of yes.

A model is scored on how often it answers right and on how close its probabilities come to how often it is right. The second measure is the Brier score, where lower is better.

Rows
3,116: English 1,277, Turkish 1,019, German 820
Questions
8,434
Format
JSON Lines, one subset per language
Licence
Rows written for the project Apache-2.0; converted rows keep their source licence (MASSIVE CC BY 4.0, Open-Jev CC0, PAWS-X free to use with credit to Google, typed-decisions Apache-2.0)
Released
1 October 2026

Results

Accuracy and Brier score on the test questions

LanguageKev-4BJevAltBrier, Kev-4BBrier, JevAlt
English84.7%Deem-4B 94.7%0.2560.091
Turkish87.1%Karar-4B 96.8%0.2140.058
German81.1%Wähler-4B 92.0%0.2980.119

Limits

Test and training rows come from the same data pipeline, so these scores measure JevAlt’s own task. On the German news set 10kGNAD, Kev-4B scores 65.3% and Wähler-4B 62.5%.

Run it

Load one language with the Hugging Face datasets library. To score your own model on these rows, point JevOss at any server that speaks the Jev API.

from datasets import load_dataset

test = load_dataset("mertkayacs/jevalt-bench", "en", split="test")
print(test[0])

Cite

@techreport{kaya2026jevaltreport,
  author = {Kaya, Mert},
  title = {JevAlt: Technical report},
  institution = {Eschatia Labs},
  year = {2026},
  url = {https://eschatialabs.com/research/jevalt-report/},
  note = {Technical report, not peer-reviewed}
}