LatteReview π€β¶

A framework for multi-agent review workflows using large language models.
π¨ NEW in v1.3.0: Screen with decision models such as TypeSafe's Jev: probabilities instead of generated text, in about 0.2 seconds per article, next to your LLM reviewers in the same workflow. See What's New below and Decision Models.
π What's New in v1.3.0¶
- Decision reviewers with Jev: LatteReview can now review with System One decision models such as TypeSafe's Jev, which answer typed questions with probabilities instead of generating text.
DecisionTitleAbstractReviewer,DecisionScoringReviewerand the genericDecisionReviewerwork in anyReviewWorkflow, next to LLM reviewers. - One provider, any backend:
SystemOneProviderworks with TypeSafe and OpenRouter, and with any other/v1/systemoneserver viabase_url, including a self-hosted OpenJev model. - Evaluated at full scale: on all 11,793 articles of LatteReview's evaluation datasets, Jev ranked articles better than the v1 LLM reviewers on every dataset (mean AUC 0.88 vs 0.83), for about $0.06 per 1,000 articles. See the evaluation, including where Jev falls short.
- Thresholds for a target recall:
suggest_thresholdfits a probability cutoff on labeled data. - Hybrid workflows: let Jev screen everything and send only uncertain articles to an LLM.
Nothing changes for existing LLM reviewers. Read Decision Models for how Jev differs from LLM reviewers and how to use it well.
Try it in the notebooks: Screening, scoring and extraction with Jev Β· Hybrid Jev + LLM review with measurements.
What Was New in v1.2.0¶
- Current models: tested with OpenAI GPT-6 (
gpt-6-astra,gpt-6-sol,gpt-6-luna) and GPT-5.x, Anthropic Claude Opus 5.5, Sonnet 5, Haiku 4.5 and Fable 5.1, and Google Gemini 3.x (gemini-3.8-flash,gemini-3.5-flash-lite). Older models such asgpt-4o-miniandgemini-2.5-flashkeep working. - No more rejected-parameter errors: if a model rejects a setting in
model_args(e.g.,temperatureon GPT-6 or Claude 5, ormax_tokenson OpenAI reasoning models), LatteReview drops or renames it with a one-time warning and retries. If a reasoning model runs out of tokens before finishing its answer, the call is retried without the limit. - New default models:
OpenAIProviderandLiteLLMProviderdefault togpt-6-luna,GoogleProvidertogemini-3.8-flash, andOllamaProvidertoqwen3.8:27b. Passmodel=to choose another. - Better local models:
OllamaProviderconstrains answers to the reviewer's JSON schema, mapsreasoning_effortto Ollama'sthinksetting, passes othermodel_args(e.g.,top_p) as model options instead of failing, andclose()works again. Tested withqwen3.8:27bon a 32 GB Apple Silicon Mac. - More accurate costs: computed from the token usage each API reports, including hidden reasoning tokens.
- Python 3.10 or later is now required. On Python 3.9,
pipinstalls 1.1.1.
See the CHANGELOG for the full list.
Overview¶
LatteReview is a powerful Python package designed to automate academic literature review processes through AI-powered agents. Just like enjoying a cup of latte β, reviewing numerous research articles should be a pleasant, efficient experience that doesn't consume your entire day!
Features¶
- Multi-agent review system with customizable roles and expertise levels for each reviewer
- Support for multiple review rounds with hierarchical decision-making workflows
- Review diverse content types including article titles, abstracts, custom texts, and even images using LLM-powered reviewer agents
- Define reviewer agents with specialized backgrounds and distinct evaluation capabilities (e.g., scoring or concept abstraction or custom reviewers of your own preferance)
- Create flexible review workflows where multiple agents operate in parallel or sequential arrangements
- Enable reviewer agents to analyze peer feedback, cast votes, and propose corrections to other reviewers' assessments
- Enhance reviews with item-specific context integration, supporting use cases like Retrieval Augmented Generation (RAG)
- Broad compatibility with LLM providers through LiteLLM, including OpenAI and Ollama
- Model-agnostic integration supporting OpenAI, Gemini, Claude, Groq, DeepSeek, OpenRouter, and local models via Ollama
- High-performance asynchronous processing for efficient batch reviews
- Standardized output format featuring detailed scoring metrics and reasoning transparency
- Robust cost tracking and memory management systems
- Extensible architecture supporting custom review workflow implementation
- NEW: Support for RIS (Research Information Systems) file format for academic literature review
- NEW: Decision-model reviewers (TypeSafe's Jev, or a self-hosted OpenJev) that return probabilities for fast, cheap screening
Quick Links¶
License¶
This work is licensed under a Creative Commons Attribution-NonCommercial-NoDerivatives 4.0 International License - see the LICENSE file for details.
π¨βπ» Authors¶
|
Pouria Rouzrokh, MD, MPH, MHPE Medical Practitioner and Machine Learning Engineer Incoming Radiology Resident @Yale University Former Data Scientist @Mayo Clinic AI Lab |
Support LatteReview¶
If you find LatteReview helpful in your research or work, consider supporting its continued development. Since we're already sharing a virtual coffee break while reviewing papers, maybe you'd like to treat me to a real one? β π
Ways to Support:¶
- Become my sponsor on GitHub
- Treat me to a cup of coffee on Ko-fi β
- Star the repository to help others discover the project
- Submit bug reports, feature requests, or contribute code
- Share your experience using LatteReview in your research
Acknowledgement¶
I would like to express my heartfelt gratitude to Moein Shariatnia for his invaluable support and contributions to this project.
π Citation¶
If you use LatteReview in your research, please cite our paper: