Gemini 3.6 Flash Launch Draws Harsh Early Reviews as Gemini 3.5 Pro Is Delayed

Google DeepMind has released three new Gemini models: **Gemini 3.6 Flash**, **Gemini 3.5 Flash-Lite**, and **Gemini 3.5 Flash Cyber**. What did not arrive was the model many developers had been waiting for: **Gemini 3.5 Pro**. Instead, Google said 3.5 Pro is still being tested with partners and will be made broadly available when it is ready. At the same time, the company confirmed that work has already started on Gemini 4 through what it describes as its most ambitious pre-training run yet. *Go

发布于 2026年7月25日generalGEO 评分: 010 次阅读
图片以深色背景为底,左侧有蓝色星形图案,右侧大写“Gemini”字样。中间突出显示“Gemini 3.6 Flash”及“Reviews · Benchmarks · Pricing”字样。下方有图表,显示3.6 Flash、3.5 Pro、Competitor 1、Competitor 2的数值,3.6 Flash数值最高。右侧有一个蓝色价格标签图标,下方有“$”符号。底部有“3.5 Pro Delay”字样,配有紫色时钟图标。该图与文档中介绍Gemini 3.6 Flash相关内容相呼应,突出其在各方面的情况。

Gemini 3.6 Flash Launch Draws Harsh Early Reviews as Gemini 3.5 Pro Is Delayed

Introduction

Google DeepMind has released three new Gemini models: Gemini 3.6 Flash, Gemini 3.5 Flash-Lite, and Gemini 3.5 Flash Cyber.

What did not arrive was the model many developers had been waiting for: Gemini 3.5 Pro.

Instead, Google said 3.5 Pro is still being tested with partners and will be made broadly available when it is ready. At the same time, the company confirmed that work has already started on Gemini 4 through what it describes as its most ambitious pre-training run yet.

图片为Google DeepMind官方推文,介绍了其推出的三款新模型。Gemini 3.6 Flash相比3.5 Flash使用更少token,以同等成本交付更高质量成果;Gemini 3.5 Flash-Lite为文档处理和智能搜索等日常任务提供快速、高性价比选择;Gemini 3.5 Flash Cyber专为发现和修复关键软件漏洞而构建的网络安全模型。该图片与上下文紧密相关,是对上文提到的Google DeepMind发布三款新模型的官方说明,为读者提供了具体信息。

Google introduced Gemini 3.6 Flash, Gemini 3.5 Flash-Lite, and Gemini 3.5 Flash Cyber.

The launch quickly became controversial in developer communities. Several early users criticized Gemini 3.6 Flash for weak frontend output, inconsistent spatial reasoning, and disappointing coding results compared with competing frontier models.

That framing is much stronger than the available evidence supports. Early hands-on complaints are real, but Google’s own benchmark results and independent Artificial Analysis testing show a more complicated picture: Gemini 3.6 Flash improves on 3.5 Flash in several coding, computer-use, long-context, and knowledge-work evaluations while keeping the same overall Artificial Analysis Intelligence Index score.

This article follows the original report’s structure while separating official benchmark data from community reactions.

Gemini 3.6 Flash Arrives to Immediate Criticism

Google positions Gemini 3.6 Flash as a workhorse model for production AI agents.

The model is generally available and supports text, image, video, audio, and PDF input, with a context window of 1,048,576 input tokens and up to 65,536 output tokens. It supports capabilities including code execution, computer use, function calling, file search, structured output, search grounding, URL context, and thinking.

Google says its main goal is not simply higher benchmark intelligence. The emphasis is on completing agentic tasks with fewer output tokens, fewer reasoning steps, lower latency, and a lower output-token price than Gemini 3.5 Flash.

Still, some of the first public tests were far less positive.

Frontend Generation Received Harsh Early Feedback

The original report highlights a developer test in which Gemini 3.6 Flash was given a two-shot frontend-generation task.

The developer described the result as unusually poor, criticizing the generated application for broken component structure, inconsistent logic, and unreliable interactions.

该图片展示了一款名为Aetheria的像素风格3D场景,属于Gemini 3.6 Flash用于生成交互式前端的典型成果呈现。场景以深色星空为背景,包含灰色城堡、山地、水域、浮空飞碟、浮空发光体等像素化元素,搭配紫色树木、多彩小建筑等细节。图片左侧设有场景参数调整面板,包含时间与光照、电影视角、世界效果三类设置选项,用于调节呈现效果。该图片对应文档中提及的社区测试内容,直观呈现了Gemini 3.6 Flash生成前端的界面成果,用于佐证开发者对该模型前端生成质量的相关评价。

One early community test used Gemini 3.6 Flash to generate an interactive frontend experience.

Other developers reported similar problems after trying the model on existing codebases. Complaints included ignored design-system constraints, unexpected UI changes, and results that looked weaker than interfaces produced by earlier Gemini Pro models.

These observations matter because frontend coding combines several capabilities at once:

  • Understanding an existing repository
  • Following design constraints
  • Producing valid code
  • Preserving existing behavior
  • Creating visually coherent layouts
  • Handling multi-step edits without damaging working components

A model can score well on a programming benchmark and still feel unreliable in real frontend work.

At the same time, individual demonstrations are not controlled benchmarks. Prompt quality, agent harness, repository structure, thinking level, tool access, and the number of allowed iterations can significantly change the result.

Frontend Arena Results Do Not Yet Support a Simple “Best” Claim

The source report says Gemini 3.6 Flash underperformed models such as Claude Fable 5, GPT-5.6 Sol, and Meta Muse Spark 1.1 in frontend coding.

The live Arena WebDev leaderboard is continuously updated, so rankings can change as more votes arrive. At the time this publication draft was prepared, Kimi K3 remained the preliminary leader, followed by Claude Fable 5 and GPT-5.6 Sol.

Because Gemini 3.6 Flash was only released on July 21, static screenshots and early leaderboard snapshots should not be treated as permanent rankings.

The more useful conclusion is that Google’s new Flash model did not immediately establish itself as the clear frontend leader, despite Google describing it as stronger for coding.

Spatial Reasoning Also Drew Negative Reactions

The original report next turns to spatial reasoning.

One developer recorded tests of Gemini 3.6 Flash on basic positional and directional relationships and reported more errors than with Gemini 3.5 Flash.

The user initially wondered whether the weaker output was caused by not selecting a high thinking level. Other users then tried higher reasoning settings and still reported mixed first impressions.

The complaints included incorrect spatial relationships and visual-scene inconsistencies.

These tests again represent community observations rather than standardized evaluations. They are useful as examples of real-world friction, but they should not be generalized into a claim that the model has universally regressed in spatial reasoning.

Google’s official documentation describes Gemini 3.6 Flash as being designed for multimodal and spatial tasks. Independent vision testing has also found strengths in several image and video categories, although not every visual task improves over 3.5 Flash.

CursorBench Shows Improvement Over 3.5 Flash, but Not Frontier Leadership

The original report also points to CursorBench, a benchmark for real software-engineering work inside the Cursor environment.

A widely shared screenshot compared Gemini 3.6 Flash with models including Cursor Composer 2.5, Claude Fable 5, GPT-5.6 Sol, Grok 4.5, and Claude Sonnet 5.

图片是一张推文截图,显示了CursorBench 3.2分数的图表。图表中,不同模型的分数以不同颜色的点表示,如Grok 4.5、GPT-5.6 Sol、Opus 4.0等。Gemini 3.6 Flash在图表中以蓝色点呈现,其分数低于其他模型。推文下方有用户评论“意思就是 Gemini 3.6 Flash 还不如 Composer 2.5 更需提 Grok 4.5 呗”,并指出Gemini 3.6 Flash现可在Cursor中可用。该图片与文档上下文相关,用于说明CursorBench测试结果,表明Gemini 3.6 Flash在该测试中表现不如部分模型。

CursorBench placed Gemini 3.6 Flash behind several higher-scoring coding models in the snapshot shared by the source.

The current CursorBench page shows several reasoning levels for Gemini 3.6 Flash. In the latest published results available during preparation:

  • Gemini 3.6 Flash Medium scored 51.2%
  • Gemini 3.5 Flash scored 48.8%
  • Gemini 3.6 Flash Low scored 47.4%

That means the medium-thinking configuration improved over Gemini 3.5 Flash in this benchmark, even though it remained behind many stronger coding models.

This is a good example of why model evaluations need context.

The statement “Gemini 3.6 Flash is worse than the competition” can be true for a specific benchmark and comparison set, while “Gemini 3.6 Flash improves over 3.5 Flash” can also be true at the same time.

Google’s Own Benchmarks Show Real Gains

Google’s published evaluation paints a noticeably more positive picture than the harshest community reactions.

The company reports improvements over Gemini 3.5 Flash across coding, machine-learning engineering, computer use, knowledge work, chart reasoning, and long-context tasks.

这是2026年7月的Gemini 3.6 Flash基准测试结果对比表,用于展示该模型与其他AI模型的各项性能指标数据。表格中重点对比了Gemini 3.6 Flash与Gemini 3.5 Flash,清晰呈现两款模型在各评测项的数值差异,核心指标包含输入输出价格、SWE-Bench Pro、DeepSWE v1.1、Terminal-Bench 2.1等多维度的评测结果,还涵盖GPT-4o-Luma、Grok 1.5、Claude Sonnet 5等其他竞品模型的相关数据,整体直观呈现了Google公开的该模型的基准测试表现,与文档中介绍的Gemini 3.6 Flash评测相关内容相呼应,可辅助说明文档中提及的模型性能提升情况。

Google’s published July 2026 benchmark table compares Gemini 3.6 Flash with several competing models.

Selected Google-reported results include:

Benchmark Gemini 3.6 Flash Gemini 3.5 Flash
SWE-Bench Pro 58.7% 55.1%
DeepSWE v1.1 49% 37%
Terminal-Bench 2.1 78.0% 76.2%
MLE-Bench 63.9% 49.7%
GDPval-AA v2 1421 1349
OSWorld-Verified 83.0% 78.4%
CharXiv Reasoning, no tools 85.2% 84.2%
GDM-MRCR v2, 128K 91.8% 77.3%
GDM-MRCR v2, 1M 54.0% 26.6%

Google also says Gemini 3.6 Flash consumes 17% fewer output tokens than Gemini 3.5 Flash on the Artificial Analysis Index and can reduce token use by as much as 65% in some DeepSWE workloads.

The new model is priced at:

  • $1.50 per 1 million input tokens
  • $7.50 per 1 million output tokens

Gemini 3.5 Flash had the same listed input price but a higher $9.00 output price in Google’s launch comparison.

For agentic workloads, that distinction matters because cost depends not only on benchmark quality but also on how many reasoning tokens and tool loops the model uses to finish the job.

Artificial Analysis Gives 3.6 Flash the Same Intelligence Score as 3.5 Flash

Third-party testing provides a more restrained result.

Artificial Analysis gives Gemini 3.6 Flash at high reasoning an Intelligence Index score of 50.

That is the same score as Gemini 3.5 Flash.

图片为Twitter用户@synthwavedd发布的推文,内容是关于Gemini 3.6 Flash在Artificial Analysis上的得分情况。推文上方文字显示Gemini 3.6 Flash在Artificial Analysis上的得分与3.5 Flash相同,且表现不如Meta Spark 1.1等模型。下方图表展示了不同模型在Artificial Analysis中的得分情况,Gemini 3.6 Flash得分在图表中处于较低位置。该图片与文档中“Artificial Analysis给Gemini 3.6 Flash的Intelligence Index得分与3.5 Flash相同”的内容相关,直观呈现了得分情况。

Artificial Analysis found no improvement in its overall Intelligence Index score from 3.5 Flash to 3.6 Flash.

The result explains part of the disappointment.

Developers looking only at the model number may expect a clear jump in overall intelligence. Instead, Artificial Analysis found roughly the same aggregate intelligence level while measuring improvements in efficiency and task completion time.

Its testing shows Gemini 3.6 Flash at approximately 251 output tokens per second, making it one of the fastest models in the comparison.

Artificial Analysis also reports that both Gemini 3.6 Flash and Gemini 3.5 Flash-Lite roughly halve time per task compared with their predecessors.

So the main improvement is better described as:

Similar aggregate intelligence, faster completion, fewer output tokens, and lower cost per completed task in many workloads.

That is less dramatic than a major intelligence breakthrough, but it is still meaningful for high-volume production agents.

Price-to-Performance Is More Nuanced Than the Viral Complaints Suggest

Some users criticized Gemini 3.6 Flash for being more expensive than certain competing models that score higher on intelligence benchmarks.

图片是一张Twitter截图,显示了Jean P.D. Meijer关于Gemini 3.6 Flash价格与性能问题的疑问。图中展示了Intelligence Index vs. Cost per Intelligence Index Task的图表,横轴为每任务成本(美元),纵轴为Intelligence Index。图表中以不同颜色标识了多个模型,如GPT 5.6 Sol High、Grok 4.5等。Gemini 3.6 Flash在图表中位置较高,成本相对较高。该图片与上下文紧密相关,直观呈现了Gemini 3.6 Flash在成本与性能方面的表现,呼应了文档中关于其价格与性能争议的内容。

The source highlighted concerns about Gemini 3.6 Flash’s position on cost versus intelligence.

Raw per-token pricing does not tell the whole story.

A model that uses fewer reasoning steps or fewer output tokens can sometimes complete a task more cheaply even when its listed token price is not the lowest.

Conversely, a cheaper model can become more expensive in practice if it requires repeated retries or longer generations.

For developers, the better metric is usually cost per successful task, measured on their own workload.

Claims About the Knowledge Cutoff Need More Evidence

The source article also cites a user who claimed Gemini 3.6 Flash did not know many events from 2025 and 2026 despite apparently presenting a newer knowledge date.

The user concluded that the model’s practical knowledge appeared to stop around January 2025.

That test should be treated cautiously.

A language model failing to recall an event does not, by itself, prove a specific training-data cutoff. Models can fail because of retrieval behavior, reasoning errors, prompt wording, safety filtering, incomplete training coverage, or uncertainty.

Google’s public model documentation does not provide enough information to independently confirm the specific cutoff date claimed in the social-media test.

For current information, developers should use Google Search grounding or another verified retrieval source rather than assuming a model’s internal memory is complete through a particular date.

Gemini 3.5 Flash-Lite Focuses on Speed

Google also released Gemini 3.5 Flash-Lite, which is designed for high-volume, low-latency agent workloads.

Artificial Analysis measured the model at approximately 350 output tokens per second.

Google lists the API price at:

  • $0.30 per 1 million input tokens
  • $2.50 per 1 million output tokens

That makes it much cheaper than Gemini 3.6 Flash.

The model is intended for workloads such as:

  • Agentic search
  • Document processing
  • High-volume extraction
  • Translation
  • Summarization
  • Subagent execution

Google says Flash-Lite also supports configurable thinking levels and built-in computer use.

Its official comparison shows substantial gains over Gemini 3.1 Flash-Lite, including:

  • Terminal-Bench 2.1: 54% vs. 31%
  • GDM-MRCR v2: 72.2% vs. 60.1%
  • GDPval-AA v2: 1140 vs. 642

It even exceeds the older Gemini 3 Flash on several agentic benchmarks.

The original report dismisses speed as unhelpful if the answer is wrong. That is fair as a general principle, but the benchmark evidence does not support treating Flash-Lite as simply “faster and worse.”

Google’s stated goal is different: use a smaller model for high-throughput workloads where full frontier intelligence is unnecessary.

Gemini 3.5 Flash Cyber Targets Vulnerability Research

The third new model is Gemini 3.5 Flash Cyber.

It is built on Gemini 3.5 Flash and fine-tuned for cybersecurity work such as:

  • Finding software vulnerabilities
  • Validating security issues
  • Producing patches
  • Supporting automated defensive analysis

Google says the model is used inside CodeMender, where multiple Flash Cyber agents can work together and combine their findings into one report.

Unlike Gemini 3.6 Flash and Flash-Lite, Flash Cyber is not being released as a general public model.

Google says it will be available to governments and trusted partners through a limited-access CodeMender pilot because of the dual-use risks associated with automated vulnerability discovery.

The separate Google DeepMind announcement says the model is intended to make high-volume vulnerability discovery and patching less expensive than using larger frontier models.

Gemini 3.5 Pro Is Still Not Ready

For many developers, the biggest announcement was the model that did not launch.

Google said:

Gemini 3.5 Pro is currently testing with partners and will be made broadly available as soon as it is ready.

That confirms that the flagship model has been delayed relative to earlier expectations.

Reports from Reuters and other outlets say Gemini 3.5 Pro had previously been expected in June and that coding performance was one of the areas contributing to the delay.

Google itself did not provide a new public release date in the July 21 announcement.

The company instead emphasized that it wants to release the model when it meets its quality bar.

This is important context for the 3.6 Flash launch. The new Flash model is not a replacement for Gemini 3.5 Pro. It occupies a different position in the lineup: a production-oriented model optimized for cost, speed, multimodal work, and agentic execution.

Gemini 4 Has Already Started Its Largest Pre-Training Run

At the same time, Google confirmed that it has begun training the next Gemini generation.

Google described the Gemini 4 effort as its most ambitious pre-training run yet.

图片为Twitter用户Logan Kilpatrick发布的推文,发布于2026年7月21日晚11:50。推文内容为“我们已经启动了迄今为止最雄心勃勃的预训练项目——Gemini 4,并对目前的进展感到兴奋:)”,并配有蓝色勾勾认证标志和谷歌图标。该推文与文档中Google确认已开始训练下一代Gemini相关,体现了其对Gemini 4预训练项目的重视。

Google said its Gemini 4 pre-training run is already underway.

This creates an unusual product timeline:

  1. Gemini 3.6 Flash is now generally available.
  2. Gemini 3.5 Flash-Lite is now generally available.
  3. Gemini 3.5 Flash Cyber is entering limited deployment.
  4. Gemini 3.5 Pro is still testing with partners.
  5. Gemini 4 pre-training has already begun.

The source article interprets this as evidence that Google is racing through version numbers while its flagship Pro model remains unfinished.

That interpretation is speculative.

Large AI organizations often train future model generations before every product from the current generation has completed its release process. Pre-training, post-training, safety evaluation, product integration, and serving preparation can proceed on overlapping timelines.

Still, the delayed Pro release clearly increases pressure on Google because coding and agentic workflows have become major competitive areas for OpenAI, Anthropic, Moonshot, xAI, Meta, and other model providers.

Is Gemini 3.6 Flash Really Google’s “Worst” Model?

The available evidence does not justify that conclusion.

There are three separate stories happening at once.

1. Some early developers genuinely disliked the model

The negative frontend and spatial-reasoning tests are real user reports. They are worth paying attention to because production quality is not fully captured by benchmarks.

2. Google’s own data shows meaningful improvement over 3.5 Flash

The model improves on DeepSWE, MLE-Bench, OSWorld, GDPval, and long-context evaluations while using fewer output tokens.

3. Third-party aggregate intelligence barely moved

Artificial Analysis gives 3.6 Flash the same Intelligence Index score as 3.5 Flash. That supports the argument that this is primarily an efficiency release rather than a major leap in raw intelligence.

For developers, the right question is not whether the model deserves a dramatic label.

The better question is whether the combination of speed, cost, quality, and reliability works for a specific application.

Gemini 3.6 Flash may be a poor choice for one frontend workflow and a strong choice for a high-throughput multimodal agent. Those outcomes are not mutually exclusive.

常见问题

What is Gemini 3.6 Flash?

Gemini 3.6 Flash is Google’s production-oriented multimodal model for coding, knowledge work, computer use, and agentic workflows. It supports a context window of more than 1 million input tokens and is generally available through the Gemini API and other Google products.

Is Gemini 3.6 Flash worse than Gemini 3.5 Flash?

Not overall. Artificial Analysis gives both models the same Intelligence Index score, but Google reports that 3.6 Flash improves on several coding, computer-use, knowledge-work, and long-context benchmarks while using fewer output tokens.

Why are developers criticizing Gemini 3.6 Flash?

Some early users reported weak frontend generation, broken UI edits, and inconsistent spatial reasoning. These are individual tests rather than controlled benchmarks, so results may vary with prompts, repositories, tools, and thinking levels.

How much does Gemini 3.6 Flash cost?

Google lists Gemini 3.6 Flash at $1.50 per million input tokens and $7.50 per million output tokens. Pricing can change, so developers should verify the current rate in the official Gemini API documentation.

How fast is Gemini 3.5 Flash-Lite?

Artificial Analysis measured Gemini 3.5 Flash-Lite at roughly 350 output tokens per second. Google designed it for high-volume, latency-sensitive tasks such as document processing and agentic search.

What is Gemini 3.5 Flash Cyber?

Gemini 3.5 Flash Cyber is a cybersecurity-focused model built on Gemini 3.5 Flash. Google plans to make it available to governments and trusted partners through CodeMender rather than as a general public API model.

When will Gemini 3.5 Pro be released?

Google has not announced a new public release date. The company says Gemini 3.5 Pro is testing with partners and will be broadly released when it is ready.

Is Google already training Gemini 4?

Yes. Google says it has started its most ambitious pre-training run yet for Gemini 4. The company has not announced a public Gemini 4 release date.

相关工具

  • Google AI Studio: Google’s browser-based environment for testing Gemini models and building with the Gemini API.
  • Gemini API: Official developer documentation for Gemini 3.6 Flash capabilities, limits, and model identifiers.
  • Gemini: Google’s consumer interface for using current Gemini models.
  • Vertex AI: Google Cloud’s enterprise platform for deploying and managing Gemini models.
  • CursorBench: Cursor’s benchmark for codebase understanding, editing, debugging, and agentic software-engineering tasks.
  • Artificial Analysis: Independent model testing across intelligence, speed, latency, pricing, and token usage.
  • CodeMender: Google DeepMind’s security-agent environment using Gemini 3.5 Flash Cyber.

Related Links

Summary

Google’s July Gemini release introduced Gemini 3.6 Flash, Gemini 3.5 Flash-Lite, and Gemini 3.5 Flash Cyber while leaving Gemini 3.5 Pro in partner testing.

Early developer reactions to 3.6 Flash were unusually negative in frontend and spatial-reasoning tests, and the model does not lead the strongest coding benchmarks. However, Google’s published results show clear gains over 3.5 Flash in several important workloads, while Artificial Analysis finds the same aggregate intelligence score with better speed and efficiency.

Gemini 3.5 Flash-Lite targets high-throughput work, Flash Cyber focuses on defensive security, and Google has already begun the Gemini 4 pre-training run.

The most accurate reading is not that Gemini 3.6 Flash is Google’s “worst model,” but that it is an efficiency-focused release whose early real-world reception is much more mixed than Google’s launch benchmarks suggest.