Pricing per million tokens, context window, capabilities — pulled from each provider's public docs. All 2 are available via the same AIgateway OpenAI-compatible endpoint; flip the model string to switch.
Gemini 3.5 Flash-Lite is a low-latency, cost-effective multimodal model optimized for high-throughput, low-cost execution for subagent tasks and document parsing.
o3 is OpenAI’s general-purpose reasoning model, balancing strong analytical performance with reasonable latency and cost.
from openai import OpenAI
client = OpenAI(
base_url="https://api.aigateway.sh/v1",
api_key="sk-aig-...",
)
# Gemini 3.5 Flash-Lite
client.chat.completions.create(
model="google/gemini-3.5-flash-lite",
messages=[{"role":"user","content":"hello"}],
)
# O3
client.chat.completions.create(
model="openai/o3",
messages=[{"role":"user","content":"hello"}],
)