Google unveils Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber with better token efficiency

Google recently launched its newest AI models focusing on higher token efficiency and lower latency for agentic workflows. This includes Gemini 3.6 Flash, Gemini 3.5 Flash-Lite, and the specialized Gemini 3.5 Flash Cyber, which are going live starting today across Google AI Studio, Android Studio, Google Antigravity, Google Search, and the Gemini app.

What has been improved exactly?

Gemini 3.6 Flash delivers notable upgrades in coding and multimodal tasks while consuming 17% fewer output tokens overall compared to 3.5 Flash, with savings reaching up to 65% in specific benchmarks like DeepSWE. It also introduces computer use as a built-in client side tool feature via the Gemini API and Gemini Enterprise, alongside enhanced Frontier Safety safeguards against chemical, biological, radiological, nuclear, and cyber risks. 

For high-volume workloads, Gemini 3.5 Flash-Lite executes up to 350 output tokens per second, with Malaysia price details starting at $0.30 (~RM1.23) per million input tokens and $2.50 (~RM10.24) per million output tokens, compared to $1.50 (~RM6.14) input and $7.50 (~RM30.71) output for 3.6 Flash. 

Are you looking forward to Gemini AI models that are more token efficient and therefore won’t use up so much money to utilize? Let us know in the comments below, and stay tuned to TechNave.com for the latest tech news updates.