Skip to content

Decision Models (Jev)

Since v1.3.0, LatteReview can review with System One decision models such as TypeSafe's Jev, in addition to LLMs. This page explains what they are, how they differ from LLM reviewers, and how to use them well.

What is a decision model?

An LLM reviewer reads a prompt and generates text: LatteReview asks it for JSON with a score and a written reasoning. A decision model generates nothing. It reads a text state (here, an article's title and abstract) and a set of typed questions, and returns a probability for every allowed answer. Three question types exist:

Question You define The answer
Noul (yes/no) an instruction, optionally what "true" and "false" mean the probability of "yes"
Choice named options (up to 255 with Jev), optionally with descriptions the chosen option, a probability per option, and a confidence
Score 2-10 ordered levels the expected level, the most likely level, a probability per level, and a confidence

Jev was released by TypeSafe AI in September 2026. It bills $0.042 per million input tokens (output is free) and answers in about 0.2 seconds. Screening 1,000 abstracts costs about $0.05-0.07: roughly half the cost of a small LLM such as gpt-6-luna (about $0.13), and one to two orders of magnitude less than larger models.

Jev vs LLM reviewers

LLM reviewers (TitleAbstractReviewer, ...) Decision reviewers (DecisionTitleAbstractReviewer, ...)
Output JSON generated by the model probabilities; answers can never be off-schema
Written reasoning yes (reasoning="brief" or "cot") no. The reasoning column is generated by LatteReview from the probabilities
Certainty the model's self-reported number a probability, which can be thresholded for a target recall
Inputs text and images text only (English-optimized, up to 32k tokens for the state plus the longest question)
Multi-step reasoning yes no multi-hop reasoning; keep hard, multi-step criteria for an LLM
Prompting prompt templates, examples, model_args typed questions only; examples, reasoning, model_args and prompt templates raise an error
Cost per 1,000 abstracts about $0.13 (gpt-6-luna) to several dollars (gpt-6-sol, gpt-6-astra) about $0.05-0.07
Latency per article one to several seconds about 0.2 seconds (about 16 articles per second at the default pace)
Repeatability answers can change between runs confident answers repeat; uncertain ones vary between identical requests (we saw 0.62-0.84 for one article)

Decision reviewers are BasicReviewer subclasses, so they work in any ReviewWorkflow, next to LLM reviewers, with the usual round-{R}_{name}_{key} columns, filters, cost tracking and memory. backstory is accepted but ignored.

Backends

SystemOneProvider speaks the /v1/systemone protocol, so every backend works the same way:

from lattereview.providers import SystemOneProvider

provider = SystemOneProvider()                                    # TypeSafe, jev-latest, reads TYPESAFE_API_KEY
provider = SystemOneProvider(backend="openrouter")                # OpenRouter, ~typesafe/jev-latest, reads OPENROUTER_API_KEY
provider = SystemOneProvider(model="jev-1.13.0")                  # pin a version instead of the -latest alias
provider = SystemOneProvider(base_url="http://localhost:3000")    # a self-hosted server (see below)
provider = SystemOneProvider(
    base_url="https://your-gateway.example/v1/systemone",         # any other /v1/systemone gateway
    api_key="...",
    model="...",
    input_price_per_million=0.042,                                # for cost tracking (custom URLs default to 0)
)
Backend How to get access Default model Key variable
typesafe (default) an API key from the TypeSafe console jev-latest TYPESAFE_API_KEY
openrouter an OpenRouter API key; Jev is billed through your OpenRouter credits ~typesafe/jev-latest OPENROUTER_API_KEY
custom base_url your own server or gateway; the URL may end with or without /v1/systemone the server's default none (pass api_key= if needed)

Other options: timeout (default 60 s), max_retries (default 3; retries 429, 529, 5xx and timeouts with exponential backoff, honoring Retry-After), and requests_per_minute. Jev allows 1,200 requests per minute, so the two named backends pace requests to 1,000 per minute by default; custom URLs are not paced. Pass requests_per_minute=0 to turn pacing off.

On OpenRouter, the cost comes from the response; elsewhere it is computed from the reported input tokens.

Screening titles and abstracts

from lattereview.agents import DecisionTitleAbstractReviewer

reviewer = DecisionTitleAbstractReviewer(
    provider=SystemOneProvider(),
    name="Jev",
    inclusion_criteria={1: "The study must involve CT scans.", 2: "The study must use deep learning."},
    exclusion_criteria={1: "The study must not include PET scans."},
)

Criteria can be a string (one criterion), a list, or a numbered dict. For each article, the reviewer sends one request with a 5-level evaluation score (mirroring the 1-5 rubric of TitleAbstractReviewer), an overall include question, and one yes/no question per criterion. It returns these columns:

Column Meaning
evaluation int 1-5, the most likely rubric level; existing filters such as evaluation >= 4 keep working
include_probability the probability that the article meets all inclusion and no exclusion criteria
confidence Jev's confidence in the evaluation score (None if the backend does not report one)
criteria {"inclusion": {criterion: p(met)}, "exclusion": {criterion: p(applies)}}
reasoning a sentence generated by LatteReview from the probabilities, e.g., P(include)=0.03; rated 'absolutely exclude'. Likely fails inclusion 1 (p=0.02: 'The study must involve CT scans.').

The raw normalized answers (per-level probabilities, confidence) are kept under _answers in the output column.

Scoring and extraction

DecisionScoringReviewer is the counterpart of ScoringReviewer. score_descriptions become the answer levels, and certainty is the confidence x 100, on the same 0-100 scale:

from lattereview.agents import DecisionScoringReviewer

reviewer = DecisionScoringReviewer(
    provider=SystemOneProvider(),
    name="Validation",
    scoring_task="How strong is the validation of the model in this study?",
    scoring_set=[1, 2, 3],
    score_descriptions={1: "no validation", 2: "internal validation only", 3: "external validation"},
)
# Columns: score (a value from scoring_set), certainty (0-100), probabilities ({score: p})

For your own questions, use DecisionReviewer. Each question becomes a column: the chosen option for Choice, the probability of "yes" for Noul, and the expected 0-based level for Score:

from lattereview.agents import DecisionReviewer
from lattereview.providers import Noul, Choice, Score

reviewer = DecisionReviewer(
    provider=SystemOneProvider(),
    name="Extractor",
    questions={
        "modality": Choice("What is the main imaging modality?", ["CT", "MRI", "X-ray", "ultrasound"]),
        "deep_learning": Noul("Does the study use deep learning?"),
        "quality": Score("How rigorous is the evaluation?", ["weak", "moderate", "strong"]),
    },
)

This covers categorical extraction. Free-text extraction (e.g., "list the reported complications") still needs an LLM AbstractionReviewer.

You can also call the provider directly:

result = await provider.decide(state="Title: ...\nAbstract: ...", questions={"rct": Noul("Is this an RCT?")})
result.answers["rct"].value   # probability of yes
result.cost, result.input_tokens

Cost model

The state is billed once per request, and each extra short question adds only about 18 tokens (plus its own text). The same abstract cost 1,128 tokens with one question and 1,254 tokens with eight. So:

  • Ask all questions about an item in one request. The decision reviewers always do.
  • One question per criterion is nearly free, and gives you a probability for every criterion.
  • Long criteria repeated in several questions cost more than the abstract itself; keep criteria concise.

Confidence-gated hybrid workflows

Jev is cheap enough to screen everything, and an LLM can then review only the articles Jev is unsure about. Filters are plain row functions, so no special workflow is needed:

from lattereview.providers import SystemOneProvider, OpenAIProvider
from lattereview.agents import DecisionTitleAbstractReviewer, TitleAbstractReviewer
from lattereview.workflows import ReviewWorkflow

jev = DecisionTitleAbstractReviewer(provider=SystemOneProvider(), name="Jev",
                                    inclusion_criteria=inclusion, exclusion_criteria=exclusion)
llm = TitleAbstractReviewer(provider=OpenAIProvider(model="gpt-6-luna"), name="LLM",
                            inclusion_criteria=inclusion_text, exclusion_criteria=exclusion_text)

workflow = ReviewWorkflow(workflow_schema=[
    {"round": "A", "reviewers": [jev], "text_inputs": ["title", "abstract"]},
    {"round": "B", "reviewers": [llm], "text_inputs": ["title", "abstract"],
     "filter": lambda row: 0.1 <= row["round-A_Jev_include_probability"] < 0.9},
])

Jev's probabilities are mostly close to 0 or 1, so a band of [0.1, 0.9) routed about 20% of articles to the LLM in our evaluation (5% to 43% depending on the dataset). Replacing Jev's decision with the LLM's inside the band gave almost the same recall and precision as reviewing everything with the LLM, for about one fifth of the LLM calls. The LLM is most useful on criteria that need an extra reasoning step, such as telling a segmentation study from a classifier study.

See the hybrid tutorial for a complete pipeline with measurements.

Choosing a threshold

A probability cutoff fitted on one dataset (or one backend) does not carry over to another. With a labeled sample of your own data, suggest_threshold finds the highest cutoff that still reaches a target recall:

from lattereview.utils import suggest_threshold

suggest_threshold(labeled_df, "round-A_Jev_include_probability", "label", target_recall=0.95)
# {'threshold': 0.08, 'recall': 0.96, 'precision': 0.41, 'n_included': 212, 'wss': 0.52}

Articles close to the cutoff can land on either side of it when re-run, because uncertain answers vary slightly between identical requests.

wss is the work saved over sampling: the fraction of articles you would not need to read at that cutoff, minus the recall given up.

Evaluation

We ran DecisionTitleAbstractReviewer (Jev 1.13) on all 11,793 articles of LatteReview's v1 evaluation: three custom searches over 978 radiology AI articles and six SYNERGY systematic-review datasets. The criteria were passed exactly as the v1 LLM reviewers got them. The comparison is the stored v1 LLM decision (gemini-2.5-flash and gpt-4o-mini, with gpt-4o resolving disagreements). The full run took 14 minutes and cost $0.71. The notebook and per-article results are in evaluation/decision_evaluation.ipynb.

Dataset Articles Included Jev AUC LLM AUC Jev WSS@95 LLM WSS@95 Jev recall at p ≥ 0.5 LLM recall at ≥ 3
search1 978 697 0.957 0.937 0.25 0.00 0.92 0.92
search2 978 367 0.875 0.818 0.23 0.25 0.85 0.84
search3 978 55 0.865 0.789 0.32 0.08 0.56 0.78
Appenzeller-Herzog 2019 2,873 26 0.899 0.847 0.64 0.00 0.08 0.31
Donners 2021 258 15 0.923 0.772 0.66 0.07 0.67 0.73
Jeyaraman 2020 1,175 96 0.795 0.708 0.00 0.00 0.00 0.02
Meijboom 2021 882 37 0.941 0.902 0.76 0.54 0.14 0.84
Muthu 2021 2,719 336 0.735 0.728 0.17 0.27 0.08 0.24
Oud 2018 952 20 0.969 0.955 0.87 0.80 0.70 0.80
Mean 0.884 0.828 0.43 0.22 0.44 0.61

AUC measures how well a score ranks included articles above excluded ones. WSS@95 is the work saved over sampling at 95% recall: the fraction of articles a reviewer would not need to read at the cutoff that still finds 95% of the included ones, minus the 5% recall given up.

What this shows:

  • Jev ranks articles better than the v1 LLM reviewers on every dataset (Muthu is a tie), and saves about twice as much screening work at 95% recall.
  • A cutoff of 0.5 is too strict for long, multi-part criteria. On the SYNERGY datasets, whose criteria are paragraphs with many conditions, Jev's include_probability rarely exceeds 0.5 for included articles. The cutoffs that reach 95% recall ranged from 0.01 to 0.23. Rank by the probability, or fit a cutoff with suggest_threshold on labeled data, instead of using 0.5. Splitting long criteria into separate items (a list or numbered dict) also gives you one probability per criterion.
  • Use the probabilities, not the evaluation integer, for ranking. The 1-5 evaluation column is kept for compatibility with LLM filters, but it is coarse (mean AUC 0.79).
  • Other signals rank about as well as include_probability: the expected evaluation score (mean AUC 0.889) and the product of the per-criterion probabilities (0.894, best on 4 of 9 datasets). None was consistently best, so include_probability is the recommended signal.
  • The labels of the custom searches come from structured metadata and contain some noise: several articles labeled "CT" are PET/CT studies that both Jev and the LLMs excluded.

Self-hosting: OpenJev and Kev

Open models that serve the same /v1/systemone API can run on your own hardware. Point base_url at them:

  • OpenJev 27B (openjev), an independent open model. The weights are licensed CC BY-NC 4.0: research and non-commercial use only. The Apache-2.0 server openjev-server runs it with vLLM or, on Apple silicon, MLX. On a 32 GB Mac, use the 4-bit MLX build (about 15 GB):

    pip install "openjev-server[mlx] @ git+https://github.com/abhishekgahlot2/openjev-server"
    hf download openjev/openjev-MLX-4bit --local-dir openjev-MLX-4bit
    openjev serve --backend mlx --model openjev-MLX-4bit --profile openjev --port 3000
    
    provider = SystemOneProvider(base_url="http://localhost:3000")
    
  • Kev (jaredpalmer/kev, Apache-2.0): Qwen base models with a LoRA adapter, from 0.8B to 27B parameters. python -m kev.serve --run jaredpalmer/kev-4b --port 8009, then SystemOneProvider(base_url="http://127.0.0.1:8009", api_key="local").

  • openjev-sglang serves any SGLang model by reading first-token probabilities, which its authors describe as not calibrated.

Known differences between backends

A compatible API is not an equivalent model. Before relying on another backend, check it on labeled data.

We ran the same 200 articles (100 each from custom searches 2 and 3, stratified by label) through three backends with DecisionTitleAbstractReviewer:

Jev 1.13 (TypeSafe) OpenJev 27B, 4-bit MLX (openjev-server 0.2.0) Kev-4B (bf16, MLX)
AUC, search 2 / search 3 0.873 / 0.838 0.872 / 0.826 0.854 / 0.814
Same include decision as Jev (p ≥ 0.5) - 98% / 92% 77% / 51%
Articles with 0.05 < p < 0.95 51% / 66% 34% / 64% 91% / 99%
Median confidence (evaluation score) 0.93 / 0.77 0.93 / 0.68 0.36 / 0.50
Time per article (6-8 questions) about 0.2 s; about 16 articles/s at the default pace 9.5-13.5 s 1.2-2.2 s
Cost and license $0.042 per million input tokens free on your hardware; weights CC BY-NC 4.0 free on your hardware; Apache-2.0

The local models ran on an Apple M5 with 32 GB of memory.

  • OpenJev ranks and decides almost like Jev, but on a laptop it is about 50 times slower. A short one-question request took about 0.9 s, and each question with an abstract-length state added about 0.7 s.
  • Kev-4B ranks almost as well, but its probabilities are much less extreme and its confidence values are lower. A cutoff or an uncertain band tuned on Jev would route almost every article to the LLM with Kev. Fit thresholds per backend.
  • Answers are not fully deterministic on any backend we tried. Confident answers (probabilities near 0 or 1) repeat, but uncertain ones vary between identical requests: six identical requests for one borderline article returned an include probability of 0.62-0.84 on TypeSafe and 0.60-0.88 on OpenRouter. Save your results instead of re-running, and expect articles near a cutoff to land on either side of it.
  • Response details differ. OpenRouter reports the cost in the response. OpenJev returns probabilities with four decimals and counts tokens differently (its prompt template makes requests look 2-4 times larger). TypeSafe requires a model; local servers use their own default when it is omitted. Missing fields (a missing confidence, legend or choice) are handled by SystemOneProvider, so reviewers see the same answer format everywhere.
  • Jev limits a choice question to 255 options; OpenJev evaluates up to 52 options in one pass. Validation errors from the server are shown in the error message.