Decision Models (Jev)¶
Since v1.3.0, LatteReview can review with System One decision models such as TypeSafe's Jev, in addition to LLMs. This page explains what they are, how they differ from LLM reviewers, and how to use them well.
What is a decision model?¶
An LLM reviewer reads a prompt and generates text: LatteReview asks it for JSON with a score and a written reasoning. A decision model generates nothing. It reads a text state (here, an article's title and abstract) and a set of typed questions, and returns a probability for every allowed answer. Three question types exist:
| Question | You define | The answer |
|---|---|---|
Noul (yes/no) |
an instruction, optionally what "true" and "false" mean | the probability of "yes" |
Choice |
named options (up to 255 with Jev), optionally with descriptions | the chosen option, a probability per option, and a confidence |
Score |
2-10 ordered levels | the expected level, the most likely level, a probability per level, and a confidence |
Jev was released by TypeSafe AI in September 2026. It bills $0.042 per million input tokens (output is free) and
answers in about 0.2 seconds. Screening 1,000 abstracts costs about $0.05-0.07: roughly half the cost of a small LLM
such as gpt-6-luna (about $0.13), and one to two orders of magnitude less than larger models.
Jev vs LLM reviewers¶
LLM reviewers (TitleAbstractReviewer, ...) |
Decision reviewers (DecisionTitleAbstractReviewer, ...) |
|
|---|---|---|
| Output | JSON generated by the model | probabilities; answers can never be off-schema |
| Written reasoning | yes (reasoning="brief" or "cot") |
no. The reasoning column is generated by LatteReview from the probabilities |
| Certainty | the model's self-reported number | a probability, which can be thresholded for a target recall |
| Inputs | text and images | text only (English-optimized, up to 32k tokens for the state plus the longest question) |
| Multi-step reasoning | yes | no multi-hop reasoning; keep hard, multi-step criteria for an LLM |
| Prompting | prompt templates, examples, model_args |
typed questions only; examples, reasoning, model_args and prompt templates raise an error |
| Cost per 1,000 abstracts | about $0.13 (gpt-6-luna) to several dollars (gpt-6-sol, gpt-6-astra) |
about $0.05-0.07 |
| Latency per article | one to several seconds | about 0.2 seconds (about 16 articles per second at the default pace) |
| Repeatability | answers can change between runs | confident answers repeat; uncertain ones vary between identical requests (we saw 0.62-0.84 for one article) |
Decision reviewers are BasicReviewer subclasses, so they work in any ReviewWorkflow, next to LLM reviewers, with
the usual round-{R}_{name}_{key} columns, filters, cost tracking and memory. backstory is accepted but ignored.
Backends¶
SystemOneProvider speaks the /v1/systemone protocol, so every backend works the same way:
from lattereview.providers import SystemOneProvider
provider = SystemOneProvider() # TypeSafe, jev-latest, reads TYPESAFE_API_KEY
provider = SystemOneProvider(backend="openrouter") # OpenRouter, ~typesafe/jev-latest, reads OPENROUTER_API_KEY
provider = SystemOneProvider(model="jev-1.13.0") # pin a version instead of the -latest alias
provider = SystemOneProvider(base_url="http://localhost:3000") # a self-hosted server (see below)
provider = SystemOneProvider(
base_url="https://your-gateway.example/v1/systemone", # any other /v1/systemone gateway
api_key="...",
model="...",
input_price_per_million=0.042, # for cost tracking (custom URLs default to 0)
)
| Backend | How to get access | Default model | Key variable |
|---|---|---|---|
typesafe (default) |
an API key from the TypeSafe console | jev-latest |
TYPESAFE_API_KEY |
openrouter |
an OpenRouter API key; Jev is billed through your OpenRouter credits | ~typesafe/jev-latest |
OPENROUTER_API_KEY |
custom base_url |
your own server or gateway; the URL may end with or without /v1/systemone |
the server's default | none (pass api_key= if needed) |
Other options: timeout (default 60 s), max_retries (default 3; retries 429, 529, 5xx and timeouts with
exponential backoff, honoring Retry-After), and requests_per_minute. Jev allows 1,200 requests per minute, so the
two named backends pace requests to 1,000 per minute by default; custom URLs are not paced. Pass
requests_per_minute=0 to turn pacing off.
On OpenRouter, the cost comes from the response; elsewhere it is computed from the reported input tokens.
Screening titles and abstracts¶
from lattereview.agents import DecisionTitleAbstractReviewer
reviewer = DecisionTitleAbstractReviewer(
provider=SystemOneProvider(),
name="Jev",
inclusion_criteria={1: "The study must involve CT scans.", 2: "The study must use deep learning."},
exclusion_criteria={1: "The study must not include PET scans."},
)
Criteria can be a string (one criterion), a list, or a numbered dict. For each article, the reviewer sends one
request with a 5-level evaluation score (mirroring the 1-5 rubric of TitleAbstractReviewer), an overall
include question, and one yes/no question per criterion. It returns these columns:
| Column | Meaning |
|---|---|
evaluation |
int 1-5, the most likely rubric level; existing filters such as evaluation >= 4 keep working |
include_probability |
the probability that the article meets all inclusion and no exclusion criteria |
confidence |
Jev's confidence in the evaluation score (None if the backend does not report one) |
criteria |
{"inclusion": {criterion: p(met)}, "exclusion": {criterion: p(applies)}} |
reasoning |
a sentence generated by LatteReview from the probabilities, e.g., P(include)=0.03; rated 'absolutely exclude'. Likely fails inclusion 1 (p=0.02: 'The study must involve CT scans.'). |
The raw normalized answers (per-level probabilities, confidence) are kept under _answers in the output column.
Scoring and extraction¶
DecisionScoringReviewer is the counterpart of ScoringReviewer. score_descriptions become the answer levels,
and certainty is the confidence x 100, on the same 0-100 scale:
from lattereview.agents import DecisionScoringReviewer
reviewer = DecisionScoringReviewer(
provider=SystemOneProvider(),
name="Validation",
scoring_task="How strong is the validation of the model in this study?",
scoring_set=[1, 2, 3],
score_descriptions={1: "no validation", 2: "internal validation only", 3: "external validation"},
)
# Columns: score (a value from scoring_set), certainty (0-100), probabilities ({score: p})
For your own questions, use DecisionReviewer. Each question becomes a column: the chosen option for Choice, the
probability of "yes" for Noul, and the expected 0-based level for Score:
from lattereview.agents import DecisionReviewer
from lattereview.providers import Noul, Choice, Score
reviewer = DecisionReviewer(
provider=SystemOneProvider(),
name="Extractor",
questions={
"modality": Choice("What is the main imaging modality?", ["CT", "MRI", "X-ray", "ultrasound"]),
"deep_learning": Noul("Does the study use deep learning?"),
"quality": Score("How rigorous is the evaluation?", ["weak", "moderate", "strong"]),
},
)
This covers categorical extraction. Free-text extraction (e.g., "list the reported complications") still needs an LLM
AbstractionReviewer.
You can also call the provider directly:
result = await provider.decide(state="Title: ...\nAbstract: ...", questions={"rct": Noul("Is this an RCT?")})
result.answers["rct"].value # probability of yes
result.cost, result.input_tokens
Cost model¶
The state is billed once per request, and each extra short question adds only about 18 tokens (plus its own text). The same abstract cost 1,128 tokens with one question and 1,254 tokens with eight. So:
- Ask all questions about an item in one request. The decision reviewers always do.
- One question per criterion is nearly free, and gives you a probability for every criterion.
- Long criteria repeated in several questions cost more than the abstract itself; keep criteria concise.
Confidence-gated hybrid workflows¶
Jev is cheap enough to screen everything, and an LLM can then review only the articles Jev is unsure about. Filters are plain row functions, so no special workflow is needed:
from lattereview.providers import SystemOneProvider, OpenAIProvider
from lattereview.agents import DecisionTitleAbstractReviewer, TitleAbstractReviewer
from lattereview.workflows import ReviewWorkflow
jev = DecisionTitleAbstractReviewer(provider=SystemOneProvider(), name="Jev",
inclusion_criteria=inclusion, exclusion_criteria=exclusion)
llm = TitleAbstractReviewer(provider=OpenAIProvider(model="gpt-6-luna"), name="LLM",
inclusion_criteria=inclusion_text, exclusion_criteria=exclusion_text)
workflow = ReviewWorkflow(workflow_schema=[
{"round": "A", "reviewers": [jev], "text_inputs": ["title", "abstract"]},
{"round": "B", "reviewers": [llm], "text_inputs": ["title", "abstract"],
"filter": lambda row: 0.1 <= row["round-A_Jev_include_probability"] < 0.9},
])
Jev's probabilities are mostly close to 0 or 1, so a band of [0.1, 0.9) routed about 20% of articles to the LLM in our evaluation (5% to 43% depending on the dataset). Replacing Jev's decision with the LLM's inside the band gave almost the same recall and precision as reviewing everything with the LLM, for about one fifth of the LLM calls. The LLM is most useful on criteria that need an extra reasoning step, such as telling a segmentation study from a classifier study.
See the hybrid tutorial for a complete pipeline with measurements.
Choosing a threshold¶
A probability cutoff fitted on one dataset (or one backend) does not carry over to another. With a labeled sample of
your own data, suggest_threshold finds the highest cutoff that still reaches a target recall:
from lattereview.utils import suggest_threshold
suggest_threshold(labeled_df, "round-A_Jev_include_probability", "label", target_recall=0.95)
# {'threshold': 0.08, 'recall': 0.96, 'precision': 0.41, 'n_included': 212, 'wss': 0.52}
Articles close to the cutoff can land on either side of it when re-run, because uncertain answers vary slightly between identical requests.
wss is the work saved over sampling: the fraction of articles you would not need to read at that cutoff, minus the
recall given up.
Evaluation¶
We ran DecisionTitleAbstractReviewer (Jev 1.13) on all 11,793 articles of LatteReview's v1 evaluation: three
custom searches over 978 radiology AI articles and six SYNERGY systematic-review datasets. The criteria were passed
exactly as the v1 LLM reviewers got them. The comparison is the stored v1 LLM decision (gemini-2.5-flash and
gpt-4o-mini, with gpt-4o resolving disagreements). The full run took 14 minutes and cost $0.71. The notebook and
per-article results are in
evaluation/decision_evaluation.ipynb.
| Dataset | Articles | Included | Jev AUC | LLM AUC | Jev WSS@95 | LLM WSS@95 | Jev recall at p ≥ 0.5 | LLM recall at ≥ 3 |
|---|---|---|---|---|---|---|---|---|
| search1 | 978 | 697 | 0.957 | 0.937 | 0.25 | 0.00 | 0.92 | 0.92 |
| search2 | 978 | 367 | 0.875 | 0.818 | 0.23 | 0.25 | 0.85 | 0.84 |
| search3 | 978 | 55 | 0.865 | 0.789 | 0.32 | 0.08 | 0.56 | 0.78 |
| Appenzeller-Herzog 2019 | 2,873 | 26 | 0.899 | 0.847 | 0.64 | 0.00 | 0.08 | 0.31 |
| Donners 2021 | 258 | 15 | 0.923 | 0.772 | 0.66 | 0.07 | 0.67 | 0.73 |
| Jeyaraman 2020 | 1,175 | 96 | 0.795 | 0.708 | 0.00 | 0.00 | 0.00 | 0.02 |
| Meijboom 2021 | 882 | 37 | 0.941 | 0.902 | 0.76 | 0.54 | 0.14 | 0.84 |
| Muthu 2021 | 2,719 | 336 | 0.735 | 0.728 | 0.17 | 0.27 | 0.08 | 0.24 |
| Oud 2018 | 952 | 20 | 0.969 | 0.955 | 0.87 | 0.80 | 0.70 | 0.80 |
| Mean | 0.884 | 0.828 | 0.43 | 0.22 | 0.44 | 0.61 |
AUC measures how well a score ranks included articles above excluded ones. WSS@95 is the work saved over sampling at 95% recall: the fraction of articles a reviewer would not need to read at the cutoff that still finds 95% of the included ones, minus the 5% recall given up.
What this shows:
- Jev ranks articles better than the v1 LLM reviewers on every dataset (Muthu is a tie), and saves about twice as much screening work at 95% recall.
- A cutoff of 0.5 is too strict for long, multi-part criteria. On the SYNERGY datasets, whose criteria are
paragraphs with many conditions, Jev's
include_probabilityrarely exceeds 0.5 for included articles. The cutoffs that reach 95% recall ranged from 0.01 to 0.23. Rank by the probability, or fit a cutoff withsuggest_thresholdon labeled data, instead of using 0.5. Splitting long criteria into separate items (a list or numbered dict) also gives you one probability per criterion. - Use the probabilities, not the
evaluationinteger, for ranking. The 1-5evaluationcolumn is kept for compatibility with LLM filters, but it is coarse (mean AUC 0.79). - Other signals rank about as well as
include_probability: the expected evaluation score (mean AUC 0.889) and the product of the per-criterion probabilities (0.894, best on 4 of 9 datasets). None was consistently best, soinclude_probabilityis the recommended signal. - The labels of the custom searches come from structured metadata and contain some noise: several articles labeled "CT" are PET/CT studies that both Jev and the LLMs excluded.
Self-hosting: OpenJev and Kev¶
Open models that serve the same /v1/systemone API can run on your own hardware. Point base_url at them:
-
OpenJev 27B (openjev), an independent open model. The weights are licensed CC BY-NC 4.0: research and non-commercial use only. The Apache-2.0 server
openjev-serverruns it with vLLM or, on Apple silicon, MLX. On a 32 GB Mac, use the 4-bit MLX build (about 15 GB): -
Kev (jaredpalmer/kev, Apache-2.0): Qwen base models with a LoRA adapter, from 0.8B to 27B parameters.
python -m kev.serve --run jaredpalmer/kev-4b --port 8009, thenSystemOneProvider(base_url="http://127.0.0.1:8009", api_key="local"). - openjev-sglang serves any SGLang model by reading first-token probabilities, which its authors describe as not calibrated.
Known differences between backends¶
A compatible API is not an equivalent model. Before relying on another backend, check it on labeled data.
We ran the same 200 articles (100 each from custom searches 2 and 3, stratified by label) through three backends
with DecisionTitleAbstractReviewer:
| Jev 1.13 (TypeSafe) | OpenJev 27B, 4-bit MLX (openjev-server 0.2.0) |
Kev-4B (bf16, MLX) | |
|---|---|---|---|
| AUC, search 2 / search 3 | 0.873 / 0.838 | 0.872 / 0.826 | 0.854 / 0.814 |
| Same include decision as Jev (p ≥ 0.5) | - | 98% / 92% | 77% / 51% |
| Articles with 0.05 < p < 0.95 | 51% / 66% | 34% / 64% | 91% / 99% |
| Median confidence (evaluation score) | 0.93 / 0.77 | 0.93 / 0.68 | 0.36 / 0.50 |
| Time per article (6-8 questions) | about 0.2 s; about 16 articles/s at the default pace | 9.5-13.5 s | 1.2-2.2 s |
| Cost and license | $0.042 per million input tokens | free on your hardware; weights CC BY-NC 4.0 | free on your hardware; Apache-2.0 |
The local models ran on an Apple M5 with 32 GB of memory.
- OpenJev ranks and decides almost like Jev, but on a laptop it is about 50 times slower. A short one-question request took about 0.9 s, and each question with an abstract-length state added about 0.7 s.
- Kev-4B ranks almost as well, but its probabilities are much less extreme and its confidence values are lower. A cutoff or an uncertain band tuned on Jev would route almost every article to the LLM with Kev. Fit thresholds per backend.
- Answers are not fully deterministic on any backend we tried. Confident answers (probabilities near 0 or 1) repeat, but uncertain ones vary between identical requests: six identical requests for one borderline article returned an include probability of 0.62-0.84 on TypeSafe and 0.60-0.88 on OpenRouter. Save your results instead of re-running, and expect articles near a cutoff to land on either side of it.
- Response details differ. OpenRouter reports the cost in the response. OpenJev returns probabilities with four decimals and
counts tokens differently (its prompt template makes requests look 2-4 times larger). TypeSafe requires a
model; local servers use their own default when it is omitted. Missing fields (a missingconfidence,legendorchoice) are handled bySystemOneProvider, so reviewers see the same answer format everywhere. - Jev limits a choice question to 255 options; OpenJev evaluates up to 52 options in one pass. Validation errors from the server are shown in the error message.