Source: https://eschatialabs.com/benchmarks/jevalt-bench/

[All benchmarks](https://eschatialabs.com/benchmarks/)

Benchmark · Decision models

# jevalt-bench

3,116 written situations in English, Turkish and German for decision models that pick an option and give a probability.

[Dataset (Hugging Face)](https://huggingface.co/datasets/mertkayacs/jevalt-bench)

1,300 downloads on Hugging Face, as of October 10, 2026.

Each row describes a situation, such as an email, a policy or a chat message, and asks one or more typed questions about it: choose an option, give a rating or estimate the probability of yes.

A model is scored on how often it answers right and on how close its probabilities come to how often it is right. The second measure is the Brier score, where lower is better.

**Rows**: 3,116: English 1,277, Turkish 1,019, German 820

**Questions**: 8,434

**Format**: JSON Lines, one subset per language

**Licence**: Rows written for the project Apache-2.0; converted rows keep their source licence (MASSIVE CC BY 4.0, Open-Jev CC0, PAWS-X free to use with credit to Google, typed-decisions Apache-2.0)

**Released**: 1 October 2026

## Results

Accuracy and Brier score on the test questions

| Language | Kev-4B | JevAlt | Brier, Kev-4B | Brier, JevAlt |
| --- | --- | --- | --- | --- |
| English | 84.7% | Deem-4B 94.7% | 0.256 | 0.091 |
| Turkish | 87.1% | Karar-4B 96.8% | 0.214 | 0.058 |
| German | 81.1% | Wähler-4B 92.0% | 0.298 | 0.119 |

### Limits

Test and training rows come from the same data pipeline, so these scores measure JevAlt’s own task. On the German news set 10kGNAD, Kev-4B scores 65.3% and Wähler-4B 62.5%.

## Run it

Load one language with the Hugging Face datasets library. To score your own model on these rows, point JevOss at any server that speaks the Jev API.

```
from datasets import load_dataset

test = load_dataset("mertkayacs/jevalt-bench", "en", split="test")
print(test[0])
```

## Cite

```
@techreport{kaya2026jevaltreport,
  author = {Kaya, Mert},
  title = {JevAlt: Technical report},
  institution = {Eschatia Labs},
  year = {2026},
  url = {https://eschatialabs.com/research/jevalt-report/},
  note = {Technical report, not peer-reviewed}
}
```

## Links

- [Dataset (Hugging Face)](https://huggingface.co/datasets/mertkayacs/jevalt-bench)
- [Code (GitHub)](https://github.com/mertkayacs/jevalt)
- [Project site](https://jevalt.mertkayacs.com/)
- [Technical report](https://eschatialabs.com/research/jevalt-report/)
