Google models

Selected API model profiles with sources, verification dates, and explicit pricing assumptions.

API specifications do not determine model availability, limits, or billing in a chat interface, coding agent, or IDE. Check the actual access method.

Model / API IDContext / input limitMax outputInput / outputVerified
Gemini 3.8 Flash
gemini-3.8-flash
1.048576M
Input limit
65.536K$0.75 / $3.75
Pricing source
Gemini 3.5 Flash-Lite
gemini-3.5-flash-lite
1.048576M
Input limit
65.536K$0.3 / $2.5
Pricing source

Compare profiles on your task

Use the same task and acceptance checks. Compare confirmed defects, corrections, elapsed time, and usage. Provider names and similarly named reasoning levels are not a shared performance scale.

Gemini 3.8 Flash

Evaluate a current Flash profile for your coding or extraction workload.

Reasoning control: thinking level (low, medium, high).

Standard paid-tier token rates through 2026-12-31. Scheduled rates from 2027-01-01: input $1.50, output $7.50. Grounding and caching have separate charges.

  • Google lists an input token limit separately from the output limit.
  • The minimal thinking level is unsupported. Availability in Gemini and Gemini CLI is separate from API availability.
Gemini 3.5 Flash-Lite

Evaluate for simple extraction or other tasks where cost and latency matter.

Reasoning control: supported; see configuration.

Standard paid-tier rates. Output includes thinking tokens; grounding, caching and other processing tiers are separate.

  • An API capability does not imply that an account exposes the same option in Gemini or Gemini CLI.

Access, limits, and other providers