Providers Module Documentation¶
This module provides different language model provider implementations for the LatteReview package. It handles interactions with various LLM APIs in a consistent, type-safe manner and now supports image-based inputs in addition to text inputs.
Overview¶
The providers module includes:
BaseProvider: Abstract base class defining the provider interfaceOpenAIProvider: Implementation for OpenAI API (including GPT models)OllamaProvider: Implementation for local Ollama modelsLiteLLMProvider: Implementation using LiteLLM for unified API accessGoogleProvider: Implementation for Google's Gemini API (including Gemini Pro and Flash models)SystemOneProvider(new in v1.3.0): Client for System One decision models such as TypeSafe's Jev, which answer typed questions with probabilities instead of generating text. It is used by the decision reviewers; see Decision Models.
** You can use any of the models offered by the providers above as far as they support structured outputs. **
Model compatibility¶
Newer models, especially reasoning models, reject some request parameters that older models accept:
| Models | Parameters they reject |
|---|---|
OpenAI GPT-5 and GPT-6 families and o-series (e.g., gpt-6-sol, gpt-5.5, o4-mini) |
max_tokens (renamed to max_completion_tokens); most also reject temperature other than 1 and top_p |
Anthropic Claude Opus 4.7+, Claude 5 and Fable (e.g., claude-opus-5-5, claude-sonnet-5, claude-fable-5-1) |
temperature, top_p, top_k |
Google Gemini 3.x (e.g., gemini-3.8-flash) |
None yet, but Google has deprecated temperature, top_p and top_k and recommends leaving them at their defaults |
Since v1.2.0, LatteReview handles these differences for you:
- Rejected parameters are dropped. If the API rejects a parameter in
model_args, the provider drops it (or sendsmax_completion_tokensinstead ofmax_tokens), prints a one-time warning, and retries. The adjustment applies to the rest of that provider's calls. Models that accept the parameter are called exactly as before, so code such asmodel_args={"max_tokens": 200, "temperature": 0.1}keeps working with new models; the rejected values are simply not applied. - Token limits don't cut off answers. Reasoning models count hidden reasoning tokens against
max_tokens/max_completion_tokens/max_output_tokens, so a small limit can leave no room for the answer. If a structured answer is cut off by the limit, the call is retried once without it (with a warning), and the limit is not sent again for that provider. - Claude returns structured output natively. Through
LiteLLMProvider, Claude models use Anthropic's native structured outputs. If a model rejects the forced tool call that older LiteLLM releases use instead (Claude Opus 5.5 and Fable 5.1 do), LatteReview falls back to JSON mode.
Recommendations:
- For reasoning models, leave out
max_tokensandtemperature. Usereasoning_effortto trade depth for speed and cost (e.g.,"low"for screening,"high"for an expert reviewer). It works withLiteLLMProviderfor OpenAI, Claude and Gemini models, and withOpenAIProviderfor OpenAI models and Gemini (through Google's OpenAI-compatible endpoint). Models without reasoning (e.g.,gpt-4o-mini) reject it, and it is dropped with a warning. - With
GoogleProvider, use Gemini's ownthinking_configinstead:{"thinking_level": "low"}on Gemini 3.x or{"thinking_budget": 0}on Gemini 2.5 Flash.GoogleProviderpasses onlytemperature,top_p,top_k,max_output_tokens,safety_settingsandthinking_configto Gemini and ignores other keys (such asmax_tokensandreasoning_effort). - Older models (e.g.,
gpt-4o-mini,gemini-2.5-flash,claude-haiku-4-5) keep acceptingtemperatureandmax_tokens.
Models tested with real API calls for v1.2.0:
| Provider path | Models |
|---|---|
OpenAIProvider |
gpt-6-astra, gpt-6-sol, gpt-6-luna, gpt-5.5, gpt-4o-mini |
OpenAIProvider (Gemini endpoint) |
gemini-3.8-flash |
GoogleProvider |
gemini-3.8-flash, gemini-3.5-flash-lite, gemini-3.1-pro-preview, gemini-2.5-pro, gemini-2.5-flash |
OllamaProvider (local, Apple M5, 32 GB) |
qwen3.8:27b |
LiteLLMProvider |
gpt-6-astra, gpt-6-luna, gpt-4o-mini, o4-mini, anthropic/claude-opus-5-5, anthropic/claude-opus-5, anthropic/claude-sonnet-5, anthropic/claude-haiku-4-5, anthropic/claude-fable-5-1, gemini/gemini-3.8-flash, groq/openai/gpt-oss-120b, deepseek/deepseek-flash, openrouter/qwen/qwen3.8-27b, ollama_chat/qwen3.8:27b |
BaseProvider¶
Description¶
The BaseProvider class serves as the foundation for all LLM provider implementations. It provides a consistent interface and error handling for interacting with language models.
Class Definition¶
class BaseProvider(pydantic.BaseModel):
provider: str = "DefaultProvider"
client: Optional[Any] = None
api_key: Optional[str] = None
model: str = "default-model"
system_prompt: str = "You are a helpful assistant."
response_format: Optional[Any] = None
last_response: Optional[Any] = None
calculate_cost: bool = True # Controls whether to calculate token costs
param_adjustments: Dict[str, Optional[str]] = {} # Parameters the model rejected (see Model compatibility)
Error Types¶
class ProviderError(Exception): pass
class ClientCreationError(ProviderError): pass
class ResponseError(ProviderError): pass
class InvalidResponseFormatError(ProviderError): pass
class ClientNotInitializedError(ProviderError): pass
Core Methods¶
create_client()¶
Abstract method for initializing the provider's client.
def create_client(self) -> Any:
"""Create and initialize the client for the provider."""
raise NotImplementedError
get_response()¶
Get a text or image response from the model.
async def get_response(
self,
input_prompt: str,
image_path_list: List[str],
message_list: Optional[List[Dict[str, str]]] = None,
system_message: Optional[str] = None,
) -> tuple[Any, Dict[str, float]]:
"""Get a response from the provider."""
raise NotImplementedError
get_json_response()¶
Get a JSON-formatted response from the model.
async def get_json_response(
self,
input_prompt: str,
image_path_list: List[str],
message_list: Optional[List[Dict[str, str]]] = None,
system_message: Optional[str] = None,
) -> tuple[Any, Dict[str, float]]:
"""Get a JSON-formatted response from the provider."""
raise NotImplementedError
Cost Calculation¶
All providers include built-in cost calculation for both input and output tokens. Costs are computed from the token usage the API reports, which includes hidden reasoning and thinking tokens, priced with LiteLLM's model price map. When the usage or the price is not available, the cost is estimated from the text with the tokencost package. A model missing from both price maps (e.g., one released after your LiteLLM version) is reported as costing 0, with a warning, instead of failing the review. This can be controlled using the calculate_cost parameter:
# Initialize provider with cost calculation disabled
provider = BaseProvider(calculate_cost=False) # All costs will be 0
# Initialize provider with cost calculation enabled (default)
provider = BaseProvider(calculate_cost=True) # Costs will be calculated based on token usage
The cost calculation returns a dictionary with three values:
input_cost: Cost of the input tokensoutput_cost: Cost of the output tokenstotal_cost: Sum of input and output costs
When calculate_cost is False, all costs will be 0.
OpenAIProvider¶
Description¶
Implementation for OpenAI's API, supporting OpenAI models, Gemini models through a compatible endpoint, and OpenRouter's model marketplace. The updated version supports processing text and image inputs.
Class Definition¶
class OpenAIProvider(BaseProvider):
provider: str = "OpenAI"
api_key: str = None
base_url: str = None
model: str = "gpt-6-luna"
response_format_class: Optional[Any] = None
Key Features¶
- Automatic API key handling from environment variables
- Support for OpenAI and Gemini models
- Custom base URL support for alternative endpoints
- OpenRouter integration for access to multiple model providers
- JSON response validation
- Image input handling
- Comprehensive error handling
Usage Example¶
from lattereview.providers import OpenAIProvider
# Initialize with OpenAI model
provider = OpenAIProvider(model="gpt-6-sol")
# Initialize with Gemini model (no "gemini/" prefix here; that prefix is only for LiteLLMProvider)
provider = OpenAIProvider(model="gemini-3.8-flash")
# Initialize with OpenRouter
provider = OpenAIProvider(
model="anthropic/claude-sonnet-4-6",
api_key="your_openrouter_key",
base_url="https://openrouter.ai/api/v1"
)
# Get a response
response, cost = await provider.get_response("What is the capital of the country shown in this map?", ["path/to/image.jpg"])
# Get JSON response
provider.set_response_format({"key": str, "value": int})
response, cost = await provider.get_json_response("Format this as JSON", [])
Using OpenRouter¶
OpenRouter provides access to a wide variety of language models through a unified API endpoint. To use OpenRouter with the OpenAIProvider:
- Sign up at OpenRouter to get an API key
- Set the following configuration:
base_url: "https://openrouter.ai/api/v1"api_key: Your OpenRouter API keymodel: Any model available on OpenRouter (e.g., "anthropic/claude-sonnet-4-6", "meta-llama/llama-3.3-70b-instruct", etc.)
This gives you access to models from:
- Anthropic (Claude models)
- Meta (Llama models)
- Mistral
- Google (Gemini models)
- And many more
Example:
provider = OpenAIProvider(
model="mistralai/mistral-large-2512",
api_key="your_openrouter_key",
base_url="https://openrouter.ai/api/v1"
)
Alternatively, LiteLLMProvider (recommended) reaches OpenRouter models directly with an OPENROUTER_API_KEY environment variable and an openrouter/ model prefix, e.g. LiteLLMProvider(model="openrouter/openai/gpt-4o-mini").
GoogleProvider¶
Description¶
Implementation for Google's Gemini API, supporting the Gemini 3.x and 2.5 models (e.g., gemini-3.8-flash, gemini-3.5-flash-lite, gemini-2.5-pro). This provider enables direct integration with Google's AI models, handling both text and multimodal inputs, as well as structured JSON responses.
Class Definition¶
class GoogleProvider(BaseProvider):
provider: str = "Google"
api_key: Optional[str] = None
model: str = "gemini-3.8-flash"
response_format: Optional[Dict[str, Any]] = None
response_format_class: Optional[Any] = None
last_response: Optional[Any] = None
Key Features¶
- Native Gemini model support via Google's official
genailibrary - Automatic API key handling from environment variables (
GEMINI_API_KEY) - Support for structured JSON responses
- Processing of text and image inputs (multimodal capabilities)
- Cost calculation from the token usage Gemini reports, including thinking tokens
- Gemini's
thinking_configcan be passed throughmodel_args(see Model compatibility) - Comprehensive error handling with detailed error messages
Usage Example¶
from lattereview.providers import GoogleProvider
# Initialize with the default model (gemini-3.8-flash)
provider = GoogleProvider()
# Or specify a different model
provider = GoogleProvider(model="gemini-3.5-flash-lite")
# Get a text response
response, cost = await provider.get_response("What is the capital of France?")
# Get a response with an image
response, cost = await provider.get_response("What's in this image?", ["path/to/image.jpg"])
# Get a structured JSON response
provider.set_response_format({"name": str, "population": int, "landmarks": [str]})
response, cost = await provider.get_json_response("Give me information about Paris.")
Generation Parameters¶
Reviewers pass their model_args to the provider. GoogleProvider copies temperature, top_p, top_k, max_output_tokens, safety_settings and thinking_config into Gemini's generation config and ignores other keys:
from lattereview.agents import ScoringReviewer
reviewer = ScoringReviewer(
provider=GoogleProvider(model="gemini-3.8-flash"),
name="Gemini",
scoring_task="Does the study use deep learning?",
model_args={"thinking_config": {"thinking_level": "low"}}, # Gemini 2.5: {"thinking_budget": 0}
)
Response Format¶
The GoogleProvider supports structured JSON responses through the set_response_format method. It accepts either a Pydantic model or a Python dictionary that specifies the expected schema:
# Using a dictionary schema
provider.set_response_format({"name": str, "age": int, "hobbies": [str]})
# Using a Pydantic model
from pydantic import BaseModel
class Person(BaseModel):
name: str
age: int
hobbies: List[str]
provider.set_response_format(Person)
OllamaProvider¶
Description¶
Implementation for local Ollama models, supporting both chat and streaming responses and image inputs. The default model, qwen3.8:27b, is a 27B multimodal reasoning model (an 18 GB download) that runs well on an Apple Silicon Mac with 32 GB of memory. Pull it first with ollama pull qwen3.8:27b, or pass any other model you have pulled.
Class Definition¶
class OllamaProvider(BaseProvider):
provider: str = "Ollama"
client: Optional[AsyncClient] = None
model: str = "qwen3.8:27b"
response_format_class: Optional[Any] = None
invalid_keywords: List[str] = ["temperature", "max_tokens"]
host: str = "http://localhost:11434"
Key Features¶
- Local model support
- Streaming capability
- Free cost tracking (local models)
- Image input processing
- Structured JSON output constrained to the reviewer's schema
model_argsmapping:reasoning_effortbecomes Ollama'sthinksetting ("none"turns thinking off;"low","medium"and"high"set its level), keywords thatAsyncClient.chataccepts (think,format,options,keep_alive, ...) are passed as they are, and other keys (e.g.,top_p,seed,num_ctx) go into the modeloptions.temperatureandmax_tokensare ignored (seeinvalid_keywords).- Connection management
Usage Example¶
from lattereview.providers import OllamaProvider
# Initialize provider
provider = OllamaProvider(
model="qwen3.8:27b",
host="http://localhost:11434"
)
# Get response
response, cost = await provider.get_response("What is the capital of the country shown in this map?", ["path/to/image1.png"])
# Get JSON response with schema
provider.set_response_format({"answer": str, "confidence": float})
response, cost = await provider.get_json_response("What is the capital of France?", [])
# Stream response
async for chunk in await provider.get_response("Tell me a story", [], stream=True):
print(chunk, end="")
# Clean up
await provider.close()
LiteLLMProvider¶
Description¶
A unified provider implementation using LiteLLM, enabling access to multiple LLM providers through a single interface. It now supports text and image-based interactions.
Class Definition¶
class LiteLLMProvider(BaseProvider):
provider: str = "LiteLLM"
model: str = "gpt-6-luna"
custom_llm_provider: Optional[str] = None
response_format_class: Optional[Any] = None
Key Features¶
- Support for multiple model providers
- JSON schema validation
- Image input handling
- Cost tracking integration
- Tool calls handling
Usage Example¶
from lattereview.providers import LiteLLMProvider
# Initialize with different models
provider = LiteLLMProvider(model="gpt-6-luna")
provider = LiteLLMProvider(model="anthropic/claude-sonnet-5") # Needs ANTHROPIC_API_KEY
provider = LiteLLMProvider(model="gemini/gemini-3.8-flash") # Needs GEMINI_API_KEY
# Get response
response, cost = await provider.get_response("What is the capital of the country shown in this map?", ["path/to/image1.png"])
# Get JSON response with schema
provider.set_response_format({"answer": str, "confidence": float})
response, cost = await provider.get_json_response("What is the capital of France?", [])
SystemOneProvider¶
Description¶
SystemOneProvider (new in v1.3.0) talks to System One decision models over the /v1/systemone protocol. Unlike the LLM providers above, it does not send prompts or parse JSON text: it sends a text state and typed questions in one request and returns a probability for every allowed answer. It is not a BaseProvider subclass, and it is used by DecisionReviewer, DecisionTitleAbstractReviewer and DecisionScoringReviewer. See Decision Models for the concepts, evaluation results and self-hosting.
Class Definition¶
class SystemOneProvider(pydantic.BaseModel):
backend: Literal["typesafe", "openrouter"] = "typesafe"
base_url: Optional[str] = None # any other /v1/systemone server; overrides backend
api_key: Optional[str] = None # else TYPESAFE_API_KEY / OPENROUTER_API_KEY; optional for base_url
model: Optional[str] = None # else "jev-latest" / "~typesafe/jev-latest"; the server default for base_url
timeout: float = 60.0
max_retries: int = 3 # retries 429, 529, 5xx and timeouts, honoring Retry-After
requests_per_minute: Optional[float] = None # 1,000 for the named backends, unlimited for base_url, 0 = off
input_price_per_million: Optional[float] = None # 0.042 for the named backends, 0 for base_url
async def decide(self, state, questions) -> DecisionResult: ...
async def aclose(self) -> None: ...
Questions and answers¶
Noul(instructions, true=None, false=None): a yes/no question.Choice(instructions, options):optionsis a dict of option → description, or a list of option names.Score(instructions, levels): 2-10 ordered levels, lowest first.
instructions, option descriptions and levels may also be JSON structures. The three classes are exported from lattereview.providers.
decide() returns a DecisionResult with answers (question ID → Answer), model (as reported by the server), input_tokens, cost (reported by OpenRouter, otherwise input tokens × price) and raw (the untouched response). Each Answer is normalized the same way for every backend:
| Field | noul | choice | score |
|---|---|---|---|
value |
probability of yes | chosen option | expected level (0-based, may be fractional) |
label |
None | chosen option | label of the most likely level |
level |
None | None | most likely level (0-based) |
probabilities |
None | option → probability | level label → probability, in level order |
confidence |
None (the probability itself shows certainty) | as reported, else None | as reported, else None |
Usage Example¶
from lattereview.providers import SystemOneProvider, Noul, Choice, Score
provider = SystemOneProvider() # TypeSafe; needs TYPESAFE_API_KEY
provider = SystemOneProvider(backend="openrouter") # needs OPENROUTER_API_KEY
provider = SystemOneProvider(base_url="http://localhost:3000") # e.g., a local OpenJev server
result = await provider.decide(
"Title: Deep learning for pneumonia detection on chest radiographs.\nAbstract: ...",
{
"deep_learning": Noul("Does the study use deep learning?"),
"modality": Choice("Which imaging modality is studied?", ["CT", "MRI", "X-ray"]),
"validation": Score("How strong is the validation?", ["none", "internal only", "external"]),
},
)
result.answers["modality"].value # "X-ray"
result.answers["validation"].label # "external"
result.cost, result.input_tokens
Ask all questions about an item in one request: the state is billed once, and each extra question adds only a few tokens.
Errors raise SystemOneResponseError (a ResponseError) with the server's message, the HTTP status_code, and retryable (False for errors a retry cannot fix, such as 401 or 422). A missing API key raises ClientCreationError.
Error Handling¶
Common error scenarios:
- API key errors (missing or invalid keys)
- Unsupported model configurations
- Models not supporting structured outputs or JSON responses
- Invalid image file paths
Best Practices¶
- For all online APIs, prefer using LiteLLMProvider class as it provides unified access and error handling.
- For local APIs, use OllamaProvider directly (rather than through LiteLLMProvider) for better performance and control.
- Set API keys at the environment level using the python-dotenv package and .env files for better security.
Limitations¶
- Requires async/await syntax for all operations
- Depends on external LLM providers' availability and stability
- Rate limits and quotas depend on provider capabilities