Pricing per million tokens, context window, capabilities — pulled from each provider's public docs. All 2 are available via the same AIgateway OpenAI-compatible endpoint; flip the model string to switch.
Gemini 3.5 Flash-Lite is a low-latency, cost-effective multimodal model optimized for high-throughput, low-cost execution for subagent tasks and document parsing.
Kimi K2.6 is a frontier-scale open-source 1T parameter model with a 262.1k context window, multi-turn tool calling, vision inputs, and structured outputs for agentic workloads.
Source: each provider's published benchmarks. Higher is better. Run an eval to compare on your own data.
from openai import OpenAI
client = OpenAI(
base_url="https://api.aigateway.sh/v1",
api_key="sk-aig-...",
)
# Gemini 3.5 Flash-Lite
client.chat.completions.create(
model="google/gemini-3.5-flash-lite",
messages=[{"role":"user","content":"hello"}],
)
# Kimi-K2.6
client.chat.completions.create(
model="moonshot/kimi-k2.6",
messages=[{"role":"user","content":"hello"}],
)