Skip to content

LatteReview πŸ€–β˜•

PyPI version License: CC BY-NC-ND 4.0 Python 3.10+ Code style: black Maintained: yes View on arXiv Sponsor me on GitHub Support me on Ko-fi

A framework for multi-agent review workflows using large language models.

🚨 NEW in v1.3.0: Screen with decision models such as TypeSafe's Jev: probabilities instead of generated text, in about 0.2 seconds per article, next to your LLM reviewers in the same workflow. See What's New below and Decision Models.

πŸ†• What's New in v1.3.0

  • Decision reviewers with Jev: LatteReview can now review with System One decision models such as TypeSafe's Jev, which answer typed questions with probabilities instead of generating text. DecisionTitleAbstractReviewer, DecisionScoringReviewer and the generic DecisionReviewer work in any ReviewWorkflow, next to LLM reviewers.
  • One provider, any backend: SystemOneProvider works with TypeSafe and OpenRouter, and with any other /v1/systemone server via base_url, including a self-hosted OpenJev model.
  • Evaluated at full scale: on all 11,793 articles of LatteReview's evaluation datasets, Jev ranked articles better than the v1 LLM reviewers on every dataset (mean AUC 0.88 vs 0.83), for about $0.06 per 1,000 articles. See the evaluation, including where Jev falls short.
  • Thresholds for a target recall: suggest_threshold fits a probability cutoff on labeled data.
  • Hybrid workflows: let Jev screen everything and send only uncertain articles to an LLM.

Nothing changes for existing LLM reviewers. Read Decision Models for how Jev differs from LLM reviewers and how to use it well.

Try it in the notebooks: Screening, scoring and extraction with Jev Β· Hybrid Jev + LLM review with measurements.

What Was New in v1.2.0

  • Current models: tested with OpenAI GPT-6 (gpt-6-astra, gpt-6-sol, gpt-6-luna) and GPT-5.x, Anthropic Claude Opus 5.5, Sonnet 5, Haiku 4.5 and Fable 5.1, and Google Gemini 3.x (gemini-3.8-flash, gemini-3.5-flash-lite). Older models such as gpt-4o-mini and gemini-2.5-flash keep working.
  • No more rejected-parameter errors: if a model rejects a setting in model_args (e.g., temperature on GPT-6 or Claude 5, or max_tokens on OpenAI reasoning models), LatteReview drops or renames it with a one-time warning and retries. If a reasoning model runs out of tokens before finishing its answer, the call is retried without the limit.
  • New default models: OpenAIProvider and LiteLLMProvider default to gpt-6-luna, GoogleProvider to gemini-3.8-flash, and OllamaProvider to qwen3.8:27b. Pass model= to choose another.
  • Better local models: OllamaProvider constrains answers to the reviewer's JSON schema, maps reasoning_effort to Ollama's think setting, passes other model_args (e.g., top_p) as model options instead of failing, and close() works again. Tested with qwen3.8:27b on a 32 GB Apple Silicon Mac.
  • More accurate costs: computed from the token usage each API reports, including hidden reasoning tokens.
  • Python 3.10 or later is now required. On Python 3.9, pip installs 1.1.1.

See the CHANGELOG for the full list.

Overview

LatteReview is a powerful Python package designed to automate academic literature review processes through AI-powered agents. Just like enjoying a cup of latte β˜•, reviewing numerous research articles should be a pleasant, efficient experience that doesn't consume your entire day!

Features

  • Multi-agent review system with customizable roles and expertise levels for each reviewer
  • Support for multiple review rounds with hierarchical decision-making workflows
  • Review diverse content types including article titles, abstracts, custom texts, and even images using LLM-powered reviewer agents
  • Define reviewer agents with specialized backgrounds and distinct evaluation capabilities (e.g., scoring or concept abstraction or custom reviewers of your own preferance)
  • Create flexible review workflows where multiple agents operate in parallel or sequential arrangements
  • Enable reviewer agents to analyze peer feedback, cast votes, and propose corrections to other reviewers' assessments
  • Enhance reviews with item-specific context integration, supporting use cases like Retrieval Augmented Generation (RAG)
  • Broad compatibility with LLM providers through LiteLLM, including OpenAI and Ollama
  • Model-agnostic integration supporting OpenAI, Gemini, Claude, Groq, DeepSeek, OpenRouter, and local models via Ollama
  • High-performance asynchronous processing for efficient batch reviews
  • Standardized output format featuring detailed scoring metrics and reasoning transparency
  • Robust cost tracking and memory management systems
  • Extensible architecture supporting custom review workflow implementation
  • NEW: Support for RIS (Research Information Systems) file format for academic literature review
  • NEW: Decision-model reviewers (TypeSafe's Jev, or a self-hosted OpenJev) that return probabilities for fast, cheap screening

License

This work is licensed under a Creative Commons Attribution-NonCommercial-NoDerivatives 4.0 International License - see the LICENSE file for details.

πŸ‘¨β€πŸ’» Authors

Pouria Rouzrokh Pouria Rouzrokh, MD, MPH, MHPE
Medical Practitioner and Machine Learning Engineer
Incoming Radiology Resident @Yale University
Former Data Scientist @Mayo Clinic AI Lab
Twitter Follow LinkedIn Google Scholar GitHub Email

Support LatteReview

If you find LatteReview helpful in your research or work, consider supporting its continued development. Since we're already sharing a virtual coffee break while reviewing papers, maybe you'd like to treat me to a real one? β˜• 😊

Ways to Support:

Acknowledgement

I would like to express my heartfelt gratitude to Moein Shariatnia for his invaluable support and contributions to this project.

πŸ“š Citation

If you use LatteReview in your research, please cite our paper:

@misc{rouzrokh2025lattereview,
    title={LatteReview: A Multi-Agent Framework for Systematic Review Automation Using Large Language Models},
    author={Pouria Rouzrokh and Moein Shariatnia},
    year={2025},
    eprint={2501.05468},
    archivePrefix={arXiv},
    primaryClass={cs.CL}
}