Models

Find the right model for your next idea.

Prices per 1M tokens

Gemini 3.8 Flash

google/gemini-3.8-flash

Official Google Gemini 3.8 Flash model routed through Vertex AI with native thinking levels. Current catalog pricing reflects Google's introductory Standard PayGo pricing through December 31, 2026.

Context 1.0MInput $0.7500Output $3.7500
View Model
Technical details
EndpointsMax output 66K tokens

Gemini 3.7 Flash

google/gemini-3.7-flash

Official Google Gemini 3.7 Flash model routed through Vertex AI with native thinking levels. Current catalog pricing reflects Google's introductory Standard PayGo pricing through December 31, 2026.

Context 1.0MInput $0.7500Output $3.7500
View Model
Technical details
EndpointsMax output 66K tokens

Gemini 3.6 Flash

google/gemini-3.6-flash

Official Google Gemini 3.6 Flash model routed through Vertex AI with native thinking levels. Current catalog pricing reflects Google's introductory Standard PayGo pricing through December 31, 2026.

Context 1.0MInput $0.7500Output $3.7500
View Model
Technical details
EndpointsMax output 66K tokens

Gemini 3.5 Flash

google/gemini-3.5-flash

Official Google Gemini 3.5 Flash model routed through Vertex AI when project access is enabled.

Context 1.0MInput $1.5000Output $9.0000
View Model
Technical details
EndpointsMax output 66K tokens

Gemini 3 Flash Preview

google/gemini-3-flash-preview

Official Google Gemini 3 Flash preview model routed through Vertex AI when project access is enabled.

Context 128KInput $0.5000Output $3.0000
View Model
Technical details
EndpointsMax output 8K tokens

Gemini 2.5 Flash Lite

google/gemini-2.5-flash-lite
Recommended

Official Google Gemini 2.5 Flash-Lite model routed through Vertex AI.

Context 1.0MInput $0.1000Output $0.4000
View Model
Technical details
EndpointsMax output 8K tokens

Gemini 3.1 Flash Lite

google/gemini-3.1-flash-lite
Recommended

Official Google Gemini 3.1 Flash-Lite model routed through Vertex AI when project access is enabled.

Context 1.0MInput $0.2500Output $1.5000
View Model
Technical details
EndpointsMax output 8K tokens

Gemini 2.5 Flash

google/gemini-2.5-flash

Official Google Gemini 2.5 Flash model routed through Vertex AI.

Context 1.0MInput $0.3000Output $2.5000
View Model
Technical details
EndpointsMax output 8K tokens

Gemini 2.5 Pro

google/gemini-2.5-pro

Official Google Gemini 2.5 Pro model routed through Vertex AI.

Context 1.0MInput $1.2500Output $10.0000
View Model
Technical details
EndpointsMax output 8K tokens

Gemini 3.1 Flash Image

google/gemini-3.1-flash-image

Nano Banana 2 image generation and editing through /v1/images/generations. Image endpoint uses per-image billing, not chat token estimates.

Context 131KInput $0.5000Output $60.0000
View Model
Technical details
EndpointsMax output 33K tokens

Gemini 3.5 Flash Lite

google/gemini-3.5-flash-lite
Recommended

Gemini 3.5 Flash-Lite, Standard global pricing; no automatic cross-model fallback.

Context 1.0MInput $0.3000Output $2.5000
View Model
Technical details
EndpointsMax output 66K tokens

GPT 6 Luna

openai/gpt-6-luna

LINE route via official OpenAI Responses API, Standard processing. No subscription access; fallback follows the Router's configured policy. Cache writes billed separately from ordinary input.

Context 128KInput $0.1000Output $0.5000
View Model
Technical details
EndpointsMax output 33K tokens

Glm 5.3 Flash

z-ai/glm-5.3-flash

Opt-in LINE comparison pilot; fallback follows the Router's configured policy. Catalog prices are list prices, not a promise of free credits. Qwen cache rate is for implicit cache; explicit-cache creation/read is not enabled by this route.

Context 128KInput $0.1500Output $0.5000
View Model
Technical details
EndpointsMax output 33K tokens

Glm 5.3

z-ai/glm-5.3

Opt-in LINE comparison pilot; fallback follows the Router's configured policy. Catalog prices are list prices, not a promise of free credits. Qwen cache rate is for implicit cache; explicit-cache creation/read is not enabled by this route.

Context 128KInput $1.4000Output $4.4000
View Model
Technical details
EndpointsMax output 33K tokens

Deepseek V4.1 Flash

deepseek/deepseek-v4.1-flash

Opt-in LINE comparison pilot; fallback follows the Router's configured policy. Catalog prices are list prices, not a promise of free credits. Qwen cache rate is for implicit cache; explicit-cache creation/read is not enabled by this route.

Context 128KInput $0.2200Output $0.6600
View Model
Technical details
EndpointsMax output 33K tokens

Qwen3.8 Max 0902

qwen/qwen3.8-max-0902

Opt-in LINE comparison pilot; fallback follows the Router's configured policy. Catalog prices are list prices, not a promise of free credits. Qwen cache rate is for implicit cache; explicit-cache creation/read is not enabled by this route.

Context 128KInput $2.0000Output $6.0000
View Model
Technical details
EndpointsMax output 33K tokens

Catalog model prices. Service fees are shown separately in your usage.