JevBench by Benchmark Heaven: benchmark and leaderboard for AI decision models, measuring typed decision accuracy, calibration, latency and cost.
-
Updated
Oct 10, 2026 - Python
JevBench by Benchmark Heaven: benchmark and leaderboard for AI decision models, measuring typed decision accuracy, calibration, latency and cost.
JevK5: open-weight alternative to TypeSafe Jev. Typed decisions with probabilities in one forward pass; Apache-2.0 weights and code.
Benchmark Heaven: AI model benchmarks with sources and modeled costs, plus the JevBench and ImageJevBench decision model leaderboards.
Constrained Decoding for Diffusion Language Models via Efficient Inference over Finite Automata (NeurIPS 2026)
Wald-Q4B: open-weight 4B decision model. Calibrated probability for every option, Jev-compatible /v1/systemone API, self-hosted. Weights on Hugging Face.
Code and data for evaluating Jev, a System One model, on scientific decisions and how its choices affect downstream results.
Find out whether a smaller model could handle some of your AI agent’s routine choices.
Convert any causal LM into a Jev-style typed decision model — no training, no new weights. Measured honestly against JevBench, negative results included.
Local frozen-backbone decision readout: evidence + criterion + options in, a probability per option out, in one forward pass. JevBench public set 0.805 / hard 0.604 with Qwen3.5-4B, no training.
XAYA-2B — probabilities, not prose. A 2B multimodal decision model for text and images, built on Qwen3.5. Python SDK for candidate probabilities, ordinal scores and yes/no decisions; historical JevBench public-dev evaluations.
Serving package for Torchcast Decision 12B (JevBench typed decisions)
One question, every public Strands decider at once — probabilities, confidence, cost, and the full JevBench board. Hugging Face Space + code.
A decision model: give it a document and a typed question, get calibrated probabilities for every allowed answer. No text generation. Qwen3.5-4B + LoRA, /v1/systemone wire format.
Open-source Jev harness for TypeSafe, OpenJev, LocalJev, Ollama, vLLM, LM Studio, llama.cpp. Typed decisions, calibration, verification, abstention, CI gates.
To associate your repository with the jevbench topic, visit your repo's landing page and select "manage topics."