Back
GoogleGemini API
Introduced the new Flex and Priority inference tiers, offering more options for optimizing cost or latency
AI summary
Written by AI from the official notes. Check them for exact details.Google introduced Flex and Priority inference tiers for better cost and latency optimization.
- New Flex inference tier for cost optimization
- New Priority inference tier for latency optimization
- More options for users to customize performance
Why it matters: Users looking to optimize their AI tool costs or performance should consider these new tiers.
Full release notes
- Introduced the new Flex and Priority inference tiers, offering more options for optimizing cost or latency.