Google models
Selected API model profiles with sources, verification dates, and explicit pricing assumptions.
API specifications do not determine model availability, limits, or billing in a chat interface, coding agent, or IDE. Check the actual access method.
| Model / API ID | Context / input limit | Max output | Input / output | Verified |
|---|---|---|---|---|
Gemini 3.8 Flashgemini-3.8-flash | 1.048576M Input limit | 65.536K | $0.75 / $3.75 | Pricing source |
Gemini 3.5 Flash-Litegemini-3.5-flash-lite | 1.048576M Input limit | 65.536K | $0.3 / $2.5 | Pricing source |
Compare profiles on your task
Use the same task and acceptance checks. Compare confirmed defects, corrections, elapsed time, and usage. Provider names and similarly named reasoning levels are not a shared performance scale.
- Gemini 3.8 Flash
Evaluate a current Flash profile for your coding or extraction workload.
Reasoning control: thinking level (low, medium, high).
Standard paid-tier token rates through 2026-12-31. Scheduled rates from 2027-01-01: input $1.50, output $7.50. Grounding and caching have separate charges.
- Google lists an input token limit separately from the output limit.
- The minimal thinking level is unsupported. Availability in Gemini and Gemini CLI is separate from API availability.
- Gemini 3.5 Flash-Lite
Evaluate for simple extraction or other tasks where cost and latency matter.
Reasoning control: supported; see configuration.
Standard paid-tier rates. Output includes thinking tokens; grounding, caching and other processing tiers are separate.
- An API capability does not imply that an account exposes the same option in Gemini or Gemini CLI.