All benchmarks

Benchmark · Decision models

JevOss

An open test bench that checks whether a decision model stays right when its input changes.

JevOss takes questions with known answers and changes one thing at a time. It reorders the options, plants a hidden instruction, adds about 600 words of unrelated text, rewrites a yes-or-no question or sends the same request again, then measures how often the answer changes.

It also scores accuracy and checks whether the model’s probabilities match how often it is right. It tests any server that speaks the Jev API (POST /v1/systemone), including JevAlt and Kev-4B.

Checks
5: option order, hidden instructions, unrelated text, yes-or-no formats, repeated requests
Commands
eval, probe, calibrate, compare, ask
Needs
Python 3.11 or newer and a running decision-model server
Version
0.1.0, released 1 October 2026
Licence
Apache-2.0

Results

Hidden instructions: share of 200 test cases where a planted instruction changed the answer (lower is better)

ModelAnswers changed
Kev-4B36.0%
Deem-4B14.0%
Karar-4B19.0%
Wähler-4B17.5%

Limits

With about 600 words of unrelated text added, Kev-4B loses 5.4 points of accuracy and Deem-4B 17.4.

Run it

The default server address is http://127.0.0.1:8000. JevAlt starts a matching server with jevalt serve.

pip install "jevoss[suites] @ git+https://github.com/mertkayacs/jevoss"
jevoss eval typed-decisions
jevoss probe typed-decisions --limit 100

Cite

@software{kaya2026jevoss,
  author = {Mert Kaya},
  title = {JevOss: Open Evaluation Toolkit for Jev-Type Decision Models},
  year = {2026},
  license = {Apache-2.0},
  url = {https://github.com/mertkayacs/jevoss}
}