Gemini 3.7 Flash(blog.google)
937 pointsby thisisauseridAug 13, 2026

96 Comments

spelkAug 13, 2026
>3.7 Flash is available through the end of the year at an introductory price 1 of $0.75/1M input tokens and $3.75/1M output tokens. This price combined with the enhanced model performance enables developers and customers to scale production-ready agents cost effectively.

Introductory pricing until December 2026 implies no significant Gemini Flash developments until the next year.

randomblock1Aug 13, 2026
I think it's just meant to make it more competitive, Gemini has kinda been behind in everything except maybe multimodal. It's only 3 weeks after Flash 3.6, so if they really wanted to, they could probably do a 3.8 Flash before then.
nateb2022Aug 13, 2026
Or a 3.7 Flash-Lite
re-thcAug 13, 2026
> implies no significant Gemini Flash developments until the next year.

Gemini 4 is apparently just around the corner so unless there's a 3 month delay... there's at least a new Flash update.

eisAug 13, 2026
3.5 Pro was supposed to be around the corner two months ago. 4.0 Pro is some ways out as they recently stated they are seeing some promising early results from training. It didn't sound like a release is imminent.
bisonbearAug 13, 2026
They compare it to 5.6 Terra, however https://cognition.com/frontiercode puts Terra at about 1/2 the price

Also have to compare to the recent Grok 4.6 release, which appears to straight up be better AND cheaper

Hard to understand why anyone would choose 3.7 Flash under these conditions.. is Deepmind still a frontier lab?

ValentineCAug 13, 2026
> Hard to understand why anyone would choose 3.7 Flash under these conditions.. is Deepmind still a frontier lab?

At this point, I think they're mostly targeting Google One and Workspace subscribers, except doing worse compared to Microsoft because they don't have Microsoft's huge enterprise moat built from their DOS and Windows days.

cahayaAug 13, 2026
Agree, with you but I'm still using 3.6 Flash because of tok/s/ latency/ uptime with high context. Tried Grok 4.6 and it was scoring lower on some internal benchmarks or slower.
mdasenAug 13, 2026
Artificial Analysis shows Grok 4.6 taking $1,068 to run their suite while Gemini 3.7 Flash takes $485. So it looks like Gemini 3.7 Flash is less than half the price in the real world.

Per-token cost isn't a great metric given that some use way more tokens than others.

bisonbearAug 13, 2026
Reposting my comment from the other thread https://news.ycombinator.com/item?id=49288847

They compare it to 5.6 Terra, however https://cognition.com/frontiercode puts Terra at about 1/2 the price

Also have to compare to the recent Grok 4.6 release, which appears to straight up be better AND cheaper

Hard to understand why anyone would choose 3.7 Flash under these conditions.. is Deepmind still a frontier lab?

ipsodAug 13, 2026
Gemini Flash 3.6 High was about 10x faster than Luna xhigh for the work that I tested it for, and it got similar results.
pkoirdAug 13, 2026
When are we getting another pro model from Gemini? Or are they simply focusing on the niche of fast but moderately capable models?
AntonioEritasAug 13, 2026
Another failed 3.5 pro run branded as 3.7 flash. It's getting sad.
dude250711Aug 13, 2026
Small young start-ups have to be frugal.
9cb14c1ec0Aug 13, 2026
Model card: https://deepmind.google/models/model-cards/gemini-3-7-flash/

Somewhere in the same neighborhood as GPT 5.6 Tera and Sonnet 5, depending on the bench.

TacticalCoderAug 13, 2026
So basically Google is 5 weeks behind with a Fast model that is as good (on the bench they picked it's mostly ahead btw) as the models that the two darlings of HN (OpenAI and Anthropic) released five weeks ago.

And yet the entire thread here is people bitching that it's neither 5.6-sol nor Opus or Fable 5.

BTW why are OpenAI and Anthropic even releasing models like terra/luna and Sonnet?

Why? Just why?

Is there a... market?

For you can't have it both ways: either Sonnet and terra/luna make zero sense for Anthropic and OpenAI or Google is a player.

nickandbroAug 13, 2026
This is genuinely a competitive model, considering it beats Claude Sonnet 5 on almost all benchmarks and is more than half its price. Seems like Google is back in the game, though not leading the frontier anymore.
9cb14c1ec0Aug 13, 2026
Claude Sonnet 5 is such a garbage model, so not sure what that says about Google's new best model.
onlyrealcuzzoAug 13, 2026
Sonnet 5 is arguably the most cost ineffective model to ever be released, so that's not really impressive.

It can regularly cost more than Fable, take longer, and deliver far far lower quality.

I'm much more interested how this compares to Luna - which on price is terribly - but at least on quality the benchmarks make this look competitive / usable.

If Google continues monthly Flash releases like Sundar said they would, and they continue to have this much of an improvement in cost/quality - then in a few months this could reasonably be very competitive with the best of the best.

It is not there yet, but at least it's super fast, I guess.

nickandbroAug 13, 2026
Agreed
nlAug 14, 2026
> Sonnet 5 is arguably the most cost ineffective model to ever be released

One word: Haiku

Although maybe that was competitive when released? I don't recall, but it's an expensive, outdated model now.

qeternityAug 13, 2026
> more than half its price

Less than half its price.

More than 50% discount.

xnxAug 13, 2026
Google is not currently in the lead for maximum model capability, but it is still very competitive (or even best) in the multidimensional capability, cost, and speed frontier.
euazOnAug 13, 2026
The multimodal abilities are great, but if you deal with text only, what is the benefit of using this over DS V4 Flash/Pro? 13-26x cheaper with comparable intelligence, and available across many different inference providers.

I fail to see the usecase where DS V4 Pro is not enough, but Flash 3.7 is - except multimodal.

Luna is similar, and also 8x cheaper. Source: artificialanalysis

The only benefit I can see is the speed, that looks to be outstanding, probably thanks to their TPUs.

re-thcAug 13, 2026
> over DS V4 Flash/Pro? 13-26x cheaper with comparable intelligence, and available across many different inference providers.

That's why DS4 already had a huge price hike announcement.

361994752Aug 13, 2026
I guess the demand is just too high... But even after the price hike, ds is still much cheaper?
KptMarchewaAug 13, 2026
The inference providers did not raise the prices no?

Deepseek as a company can just increase prices for the crazily cheap cache they have, that's their only lever.

onlyrealcuzzoAug 13, 2026
> 13-26x cheaper with comparable intelligence, and available across many different inference providers.

Well, compared to 2 months ago, it's no longer 100x more expensive for similar levels of quality...

If they continue monthly-ish releases by 3.9 - by Halloween - they should be close to the best in terms of what you get for what you pay for.

In 2 months, they've gone from basically the bottom of the pack to at least being somewhat usable and competitive.

OpenAI and Anthropic release in a month, and change things. OpenAI is claiming to be close to an Astra release - but that seems like a Fable type release - where they're just releasing a better more expensive model, not more cost effective models.

nlAug 14, 2026
Surely Anthropic will release a new Haiku at some point soon?

It's terribly outdated and way overpriced now and they do need something to compete at that "fast, cheap and ok" level.

jklmnopqrstuvwAug 13, 2026
From my own testing, Gemini 3.5/3.6 Flash is better than DS v4 Flash/Pro on text ability.
anthonypasqAug 13, 2026
for non-coding applications, i think speed is a real differentiator. Im building an app that uses LLMs for some functionality that the user would not have any reason to expect is using AI and therefore having then wait seconds or minutes is just not feasible. latency is a huge upside for me
MelatonicAug 13, 2026
Anything interacting with the real world seems like latency would be hugely important. Something more asynchronous friendly (like coding) is for obvious reasons over represented here
lenerdenatorAug 13, 2026
I guess the question then becomes "are you sure you'll do text only?"

I could probably do text only for my workflow (feature development/debugging for web microservices) but sometimes it is easier to just toss a screenshot into the Claude prompt, so that gives it an edge.

If your workflow is 100%, certifiably never ever going to involve an image, then yeah, this isn't going to be huge.

127Aug 13, 2026
DSV4 Flash is in a tier of its own, until at least the price change arrives.
PunchTornadoAug 13, 2026
did you try to ingest 1M documents per hour with any provider except GCP with Flash? None work at scale. Deepseek, Luna, Mistral all fail. 1 in 3 requests is a fail. I stopped trying.

The only thing that works at scale is gemini flash.

TiberiumAug 13, 2026
3.7 Flash gets 56 on AA up from 52 for 3.6 Flash. But it seems like this is at the cost of more output tokens per task: 3.6 Flash is 26k, 3.7 Flash is 37k. Due to 3.7 Flash's 2x slashed pricing it's still cheaper per task.
npnAug 13, 2026
> * For 3.6 and 3.7 Flash, introductory price expires on December 31, 2026. Starting January 1, 2027, $1.50/1M input tokens and $7.50/1M output tokens will apply.

this is hilarious. it is not 2025 any more, by Jan 2027 there will be at least 3 newer generation of models (from other provider) released already. nobody would use flash 3.7 at that time.

sure we used to cling to gemini models in the past, demanding 2.5 models to continue to serve, but since google betrayed us with those price hike, people already spent their time making their production pipeline less dependent on google since then.

heck, even now I'm not sure I even care if they cut the pricing even lower. there are too many models with cheaper price and similar performance now.

KptMarchewaAug 13, 2026
This is added specifically so you migrate out of those as fast as next models will be available.
jtwalesonAug 13, 2026
I think it's just to signal that prices will go up in the future.
poly2itAug 13, 2026
I think this is a play to get around EU regulation about false sales.
GodelNumberingAug 13, 2026
> introductory price

They should call it 'face saving pricing after we realized just how terribly did we mis-price the flash 3.5'

> since google betrayed us with those price hike, people already spent their time making their production pipeline less dependent on google since then.

This is my first hand experience. I spent at least $3000 on gemini-3-flash-preview. And exactly $0 total on (3.5+3.6+3.7)

dr_dshivAug 13, 2026
gemini-3-flash-preview is legit amazing and cheap. That's why i spent over 10k on it.
seizethecheeseAug 13, 2026
Maybe the business model is to break even on bleeding edge models while making money on the long tail of usage once systems are tuned for a specific model and running in production.
NoDodgeQuestionAug 13, 2026
How can system be tuned for a specific model? Model is fungible, often one model strictly greater on both quality and price.
ipsodAug 13, 2026
Models are not fungible, if you're building certain types of products on them.
serfAug 13, 2026
this is becoming less true with every generation of model.

a decent model with a decent harness will determine when the knowledge base is lacking and attempt to fill the holes; thus the good general models can be very easily brought up to speed on niche domains.

margalabargalaAug 13, 2026
You're thinking like an engineer.

Think like a regulator.

j16sdizAug 13, 2026
Ugh? It is not just the knowledge

Some model are more aggressive by default, some are more verbose by default. To get the result you want for your specific application, you run experiment with prompts and parameters.

nickservAug 13, 2026
Prompts can certainly be tuned to a particular model, where updating the model actually results in worse performance. This is perhaps less true today than a year or two ago, but we have seen this on newer models as well. Typically, the less specific the instructions are, the less it's a problem. But sometimes you really need to get into specifics to get good results. Area of work is code porting and translation.
andaiAug 13, 2026
The business model is to replace the entire human economy.
seizethecheeseAug 13, 2026
Think of it this way: you are at an enterprise business. You have a workflow implemented a year ago that is working just fine. Swapping out the model for a new one changes behavior in unpredictable ways. Eventually, you'll do it once cost is low enough, but it takes serious labor to validate this, so you'll wait a long enough time for Google to make money.
npnAug 13, 2026
it is partly true, but like I said it is not 2025 anymore. models now get released more often, and still have notable progress so they can safely replace the old models while being faster/cheaper. and thank to chinese models the pricing is pretty much stable and affordable now.

and now we have ai agents to automatic migrate the system with new models. in the past we would need to spend hours to design the prompts, then test the output, then write codes to babysitting it. nowadays any ai agent can do it effortlessly.

eliAug 13, 2026
Isn't it a good thing to know about price hikes in advance? If I were building a product around it, I would certainly care.
swangAug 13, 2026
I think it's meant to make fun of the fact that Google raised prices on their models and people were upset, and this is Google's way of lowering back the price because by Jan 1st 2027, this model isn't going to be used since people will move on to the latest models.

Personally, I feel like Google blundered on their pricing because while I was using the free version of the Gemini harness, they took away most of the free limits and made people move over to their Anti-Gravity harness for no apparent reason. I was about to splurge for a Pro sub since I already used Google for extra storage but putting up limits like they did made me not want to trust they wouldn't do more price shenanigans. Now their models are behind and it seems like they're scrambling.

raincoleAug 13, 2026
What? Jan 2027 is just about four months away. People surely still use models from four months ago today.
threatripperAug 13, 2026
Nobody except corporations who built workflows on top of it and don't care about the price because the developer already moved on and nobody wants to touch it.
orliesaurusAug 13, 2026
what a week - lets see it draw a weird animal doing a weird thing on a bicycle
hiccuphippoAug 13, 2026
Shouldn't it be drawing the whole Silmarillion now?
orliesaurusAug 13, 2026
no that's in Flash 3.8 Pro
damstaAug 13, 2026
> 3.7 Flash is available through the end of the year at an introductory price of $0.75/1M input tokens and $3.75/1M output tokens.

> Introductory pricing expires on December 31, 2026. Starting January 1, 2027, $1.50/1M input tokens and $7.50/1M output tokens will apply.

modelessAug 13, 2026
Clearly this model will be irrelevant by Jan. 2027, why would Google even bother to say this?
quaintdevAug 13, 2026
Maybe they know something we don't. What if all frontier lab do this? Maybe this is actual cost of running these llm.
kromokromoAug 13, 2026
Its probably just a corporate symptom, weird stuff like this happens in messy large orgs.
uramsAug 13, 2026
It's basically a "if we really have to support this for a long time, we want to be compensated for that" pricing strategy. It's about long term maintenance cost being greater _because_ it will be irrelevant.
seunosewaAug 13, 2026
They want to maintain the perception that Flash is worth $7.5/mot, so they can charge more for the next one.
TopfiAug 13, 2026
> What's new in Gemini 3.7 Flash [0]

> Coding and agentic tasks: Significantly higher quality on real-world software engineering and agentic benchmarks, improving issue resolution and reducing failed agent loops.

> Web development and stronger design parity: Generates higher-fidelity desktop and web application code directly from design mocks, with strong gains in design adherence and in auditing existing codebases against mocks to verify 1:1 design parity.

> Promotional pricing: Gemini 3.7 Flash will be available at an introductory price of $0.75/1M input tokens and $3.75/1M output tokens. We’re also applying this new rate to 3.6 Flash. Introductory pricing expires on December 31, 2026; after, $1.50/1M input tokens and $7.50/1M output tokens will apply.

Still no sign of 3.5 Pro. Will have to test it, low expectations given every other model from the Gemini 3 lineage, but one can hope. Just struggle to understand the promotional pricing being temporary for four months. Given this industry, I'd be hard pressed if 3.7 Flash was still in use by end of year, so why not make it the official pricing?

[0] https://ai.google.dev/gemini-api/docs/latest-model

WarmWashAug 13, 2026
>Given this industry, I'd be hard pressed if 3.7 Flash was still in use by end of year, so why not make it the official pricing

It was probably to placate some kind of general internal pricing/revenue benchmark that doesn't account for new model releases. Politicians do shit like this incessantly and it reeks of bureaucracy.

mattlondonAug 13, 2026
I suspect it's a bit of a signal to investors etc.

"Hey, we are not in a race to the bottom. This is our usual pricing, but this now is a promotion because we know we're coming from behind and need to entice users."

They're drawing a line in the sand on monetisation and signalling that to everyone, while in reality offering it a deep discount (no idea if profitable or not) knowing that this model will probably be obsolete before then.

yanis_tAug 13, 2026
Is that he model that supposed to be Pro, but then they changed their mind?
aix1Aug 13, 2026
No, relabelling a Pro model as Flash would make no economic sense (the Pro series is larger than Flash and more expensive to serve).
TekMolAug 13, 2026
I'm only interested in the state-of-the-art model by each provider.

For Google, this is still gemini-3.1-pro-preview, right?

re-thcAug 13, 2026
> For Google, this is still gemini-3.1-pro-preview, right?

Flash is better than Pro for now.

yieldcrvAug 13, 2026
This is all a naming quirk because Google can’t commit

Path A: Deprecated, do not dare use

Path B: Beta, do not rely

yborgAug 13, 2026
Google once again seems to have fallen into the pit of its own bureaucracy, even OpenAI looks competent by comparison.
fmind-devAug 13, 2026
Gemini Flash is one of the best "good-enough" models. I use this type of model daily, for automation and quick development iteration loops.

Unfortunately, it's often not strong enough for heavy refactoring and long running development loops.

garciasnAug 13, 2026
Yeah we use it for auto-triage of incidents, attempts to auto-remediate, and escalation to human. But for actual development, it’s not a viable option for us.
christoff12Aug 13, 2026
'Tis a good workhouse, indeed. I hope they give us a 4.0 Pro that can use Flash subagents soon.
dismalafAug 13, 2026
Yup. I use it for a ton of mundane queries (stuff that I might have used Google search for in the past) and it's great. Nice and fast and correct more often than not, especially if you prompt it in a way that it invokes Google search (but filters out ads and SEO slop). It's even alright at programming tasks but if it stumbles then I'll escalate to Gemini Pro with extended thinking.
toshAug 13, 2026
strong improvement over 3.6 flash

but luna is hard to beat @ capability / cost

nateb2022Aug 13, 2026
[dupe] https://news.ycombinator.com/item?id=49288847 (35 points, 8 comments)
jdw64Aug 13, 2026
I'm really curious about this: the foundational paper behind today's LLMs came from Google, and some of the world's best scientists were at Google. So why are they falling so far behind in the AI race?
aix1Aug 13, 2026
The "let's make money by selling/renting out TPUs" faction has won and the "let's make money by training and selling a frontier model" faction has lost.

And it's arguably not crazy, at least if SemiAnalysis's estimates are to be believed:

  * 20% of all TPU shipments from Q3 2026 through Q4 2027 are sold to SPVs serving Anthropic ($150B of contracted revenue); vs
  * ~$12B ARR for Gemini.
https://newsletter.semianalysis.com/p/gemini-is-cooked-but-g...

Because they compete for the same scarce resource, the result is a resource crunch for the group that's lost: https://www.latimes.com/business/story/2026-05-18/inside-ai-...

deadmutexAug 13, 2026
> The "let's make money by selling/renting out TPUs" faction has won and the "let's make money by training and selling a frontier model" faction has lost.

Citation needed.

also, why can't a massive company do two things?

aix1Aug 13, 2026
With all due respect, did you read my comment beyond the first paragraph? It addresses both points, TPU economics/pivot to sales + internal shortages making it hard to train models, to the extent they can be addressed based on public sources.

There are other factors at play, but they're more recent/second-order.

deadmutexAug 13, 2026
There are a lot of assumptions there that are not verified.
jerfAug 13, 2026
"why can't a massive company do two things?"

It's a variation of opportunity cost. A company that has an opportunity to take $1 and make $1.50 on it can't justify an opportunity to spend $1 and make $1.25, even though a less profitable company may make a good living on that. When considering capital allocation, Google has to consider the opportunity cost of investing more into their highly lucrative ads business. Another company that has no access to such a lucrative business uses different opportunity cost when it comes to allocating capital. It can easily be the case that Google could end up justify being in the business of renting out shovels and end up chased out of the business of using the shovels to create AIs entirely because that turns out not to be where the money is. I'm not saying that's obviously inevitable; I'm saying it's a possible and reasonable outcome.

That's why even though the industry produces giants, these giants can never just eat everything. Even though it seems like they have all the money, it isn't practical for them to try to do everything and in fact limits get hit very quickly for anything other than the primary, lucrative business.

Apparently there is no snappy term for this in the business space, according to such AI searches as I have run.

deadmutexAug 13, 2026
This is extremely simplified, and not realistic. One counter to this is markowitz portfolio theory, or "diversification effect".
jerfAug 14, 2026
Of course the three short paragraph description of a general principle posted to Hacker News to give a partial answer to someone's 8-word question is simplified. What's the alternative? HN's text box is too short to contain a 4-year business course.
HackerThemAllAug 14, 2026
> why are they falling so far behind in the AI race?

Are they? They provide AI overview to majority of web searches, and that alone requires enormous resources. Anthropic, OpenAI and others only serves their AI customers. Regarding the power of their model, my own experiences are that it doesn't fall behind. I've done many successful projects already, including quite a big one in Pascal. So no, I don't feel any difference between Gemini and others. I think it's just a long lasting fashion to whine about Google and its services.

twelvechairsAug 13, 2026
https://artificialanalysis.ai/models/gemini-3-7-flash

The selling point for gemini continues to be speed and particularly end-to-end response time.

vrosasAug 13, 2026
I've blown away by flash 3.6's speed while Opus chugs along for _hours_ on similar tasks. I've gotten into a opus designed -> gemini implemented -> opus reviewed dev cycle recently.
gekoxyzAug 13, 2026
I am actively using Gemini flash to "translate" what Opus says into human language. I let opus do the design (with my assistance) and implementation, but then the report that Opus writes gets translated by Gemini so that I don't have to waste time to understand it.
kridsdale1Aug 13, 2026
That’s like the army guy in movies from the 90s who shouts “IN ENGLISH, PLEASE!” after the scientist explains the conflict of the plot.
kyrraAug 14, 2026
Matt Pocock has a great /wait-what skill for this: https://www.aihero.dev/skills-wait-what, which is an extremely short prompt of:

> Wait — I don't understand where you've got to here. Re-pitch that: give me a little bit of context, talk in ASD-STE100 Simplified Technical English, and use the ubiquitous language from CONTEXT.md.

Marha01Aug 13, 2026
> I've gotten into a opus designed -> gemini implemented -> opus reviewed dev cycle recently.

This is what I do too.

seunosewaAug 13, 2026
Last time I tried it, Flash introduced too many errors due to sloppiness. Is it more reliable at following instructions now?
Marha01Aug 14, 2026
It seems reliable to me now.
bob_theslob646Aug 13, 2026
What's the typical response time for Gemini compared to other models?
ponyousAug 13, 2026
On my benchmark where AIs generate ~20 different 3D models about 1/2 the time of Opus and 1/3 of the time of Kimi K3 and 2/3 of time of sonnet.
markasoftwareAug 13, 2026
Sol high is almost the same speed if you take into account drastically lower token use. Look at the artificial analysis speed vs token use. Gemini is 7x faster but 5x more tokens. And that's with Sol high being a substantially better model.

Edit: and Sol medium actually has the same AA intelligence score as Gemini 3.7, and has >7x fewer tokens, actually making it faster

anthonypasqAug 13, 2026
is presumes you are doing longer difficult agentic tasks, if youre doing a simple problem in 1 or 2 shots, not really multi turn then theres no comparison.
modelessAug 13, 2026
It's funny that they don't mention this at all in the marketing or tech specs when it's obviously the biggest selling point by far. Without this it would be completely irrelevant.

Worth noting that OpenAI just announced that they got the full GPT 5.6 Sol model running on Cerebras at 750 tokens per second. No announcement of the pricing though...

hbnAug 13, 2026
> It's funny that they don't mention this at all in the marketing or tech specs when it's obviously the biggest selling point by far. Without this it would be completely irrelevant.

Good catch! You're right to point that out. My previous marketing copy missed that specific detail. Thank you for bringing it up!

MelatonicAug 13, 2026
It's like the Intel Optane of AI
data-ottawaAug 13, 2026
I use this in a customer facing application and Gemini’s speed makes the experience feel much better.

The application isn’t so complicated that you need opus level reasoning or code writing, we need “good enough” data retrieval and processing with natural language queries and the ability to answer follow up questions.

For that Gemini works well for a decent price.

bearjawsAug 13, 2026
Cerebras is crazy to watch on GPT OSS or Gemma, I feel like we need a new VibeOS demo but with Cerebras, the OS would literally build itself in a few seconds.

https://youtu.be/7NfyZhV1dKM?t=52

cheikhcheikhAug 14, 2026
Thw bottleneck become compiling and running the code every iteration so it wont be seconds
mike_hearnAug 14, 2026
JIT compilers largely solve that. The real question in my mind is whether fast LLMs bias stuff towards languages that either don't have a compile/typecheck phase at all or have one that's very fast (e.g. Java).
jpauAug 13, 2026
You can also customize Gemini Flash. It's a niche thing benefitting few, but you can tune gemini-3.7-flash in Google Vertex (now named "Agent Platform"?)
cracadumiAug 13, 2026
For those looking for the full benchmark figures and technical overview, Google's primary announcement post is here: https://blog.google/innovation-and-ai/models-and-research/ge...
khanhnguyen8386Aug 13, 2026
Offering a 'temporary introductory discount' until Dec 2026 on an LLM is hilarious. In this market, by Jan 2027 this model will be superseded by 5 different providers offering 10x the performance at half the post-discount price anyway.
wxwAug 13, 2026
They need to release benchmarks against Luna/Terra. Luna is much cheaper which feels like it undercuts the need for Flash.

I've always considered the Flash series of models to be for low-cost, high-volume, mostly text-based use cases (e.g. summarization, parsing, formatting), emphasis on low-cost.

[edit: ah, benchmarks here: https://blog.google/innovation-and-ai/models-and-research/ge...

more of a Terra than Luna competitor which is an interesting positioning. I feel like differentiation at the mid-tier of models is pretty difficult.]

timdorrAug 13, 2026
They compared against 5.6-terra on the model card: https://deepmind.google/models/model-cards/gemini-3-7-flash/
peabAug 13, 2026
gemini flash is probably the best model for visual tasks right now. they also make it really easy to ingest videos
pants2Aug 13, 2026
Yes, was going to say I use it exclusively for video and audio. The ability to give it a YouTube link through the API and ask questions about it is awesome
wxwAug 13, 2026
Ah, multimodal is a great point. I'll need to try that some time.
icelancerAug 13, 2026
Crazy it's still the only video understanding endpoint. It's what I use it for and no other model even offers a competitor.
mike_hearnAug 14, 2026
You probably can't build such a model without unlimited access to YouTube and Google has been tightening the screws on that over the years pretty systematically.
ipsodAug 13, 2026
Also best at OpenSCAD, seemingly for the same reason, at least in terms of "iterate on a design, comparing visual output to target".
sourweaselAug 13, 2026
Are you manually rendering previews of its OpenSCAD output to create images for it to review, or do you have a workflow that automates that?
gunalxAug 13, 2026
Personally both, i use openscad with opencode, and paste images but it is pretty often the modell decides by itself it wants to see a render and uses the render shell commands to get a image to look at.
ipsodAug 14, 2026
I've built some tools using AI. A custom GUI that has "copy context" and "copy image" buttons, to make prompting easy, and the camera position is persisted to disk on each change so that the agent can run a command to get a screenshot of it.

I'm thinking about inlining an AI chat window directly - I guess I might fire off a prompt to do that right now.

gunalxAug 13, 2026
+1 gemini models where really the only ones fullt grasping spatial reasoning even compared to opus (at least when i last cared to check it)
pimeysAug 13, 2026
It is also very good and cheap for computer use.
andaiAug 13, 2026
Matched roughly with Sol on DeepSwe cost per task.

Luna way cheaper. DeepSeek used to be, but I think it's somewhere on Sol's curve after the price hike.

scotty79Aug 13, 2026
On DeepSwe it's strictly beaten by Luna on max, cost and result.

Damn, Luna on max is as good on DeepSWE as Kimi k3, I think I dismissed this model unjustly.

HDBaseTAug 13, 2026
This is why benchmarks are scary, since Artificial Analysis puts it a fair bit behind Kimi K3.

Kimi K3 is a beast though, just costly.

dannywAug 14, 2026
Still cheaper than API rates for Opus!
barrenkoAug 14, 2026
So in a way Luna is the new Gemini Flash? I've been out of the game for a while.
scotty79Aug 14, 2026
Luna is the smallest variant of gpt-5.6 from openai and it seems to beat Gemini Flash 3.7 on DeepSWE benchmark (which is one of the new coding benchmarks that people find more relevant to their daily work than old, saturated and gambled benchmarks). It beats it massively on cost and by a bit on quality.
anthonypasqAug 13, 2026
flash-lite is more of their luna tier competitor but even still not quite there yet, but gemini's dominance on multimodal and image understanding i think really gets downplayed on this site when most people think the only think you can do with LLMs is write code
vohkAug 13, 2026
That's been my association as well. I see Flash get brought up a lot in relation to things like OCR and PDF processing frequently, and a lot of other routine multimodal workloads.
MrBuddyCasinoAug 13, 2026
Yes this is my impression as well. To be fair I didn't compare to Luna yet, but Gemini 3.5 Lite is a very good and cheap multi-modal data extraction model.
dogomaticAug 13, 2026
If anyone knows of a cheaper vision llm with the same accuracy I would love to switch
WarmWashAug 13, 2026
Ultimately it would track that in the real world, people will want to point cameras at things and get answers.

I pay for ChatGPT and Gemini, and while Sol is a total beast with anything text, it still poisoned my cucumber bed. Which I will be bitter about for at least a few years while the bed recovers. Gemini (even flash) is exceptionally talented at viewing photos and telling you what to do/what it is (and telling me I just misidentified the problem with my cucumbers and spraying off the "bugs" actually just spread the bacteria everywhere.)

roliszAug 14, 2026
Tell me more about your cucumber bed. We started some raised veggie beds this year and my wife is relying very heavily in ChatGPT and Claude for advice on how to deal with issues.
WarmWashAug 14, 2026
5.6 Sol was confident that I had a mite problem, as evidenced by the yellow spots on the leaves. I flipped over the leaves and saw a few small bugs, and relayed this to 5.6 who said that they can be difficult to see since they are so small. The remedy was to blast the underside of all the leaves to wash off all the mites, which I diligently did.

However afterwards, still feeling odd that I didn't see many bugs, I double checked with Gemini who is my typical standby for vision tasks. Gemini pointed out that its actually the fungus Corynespora cassiicola...which spreads by water splashing, and can live in your bed for years once it gets into the soil.

So now having two answers I need to research and check myself, and sure enough, it was obvious match for the fungus and didn't look much at all like the mite damage.

There is another story a few days later of 5.6 looking at my heatpump install and telling me it was probably borked and to call an expert. Gemini told me to just open a valve and it was good.

I don't really trust any other LLM besides gemini for vision stuff. It also tracks as google is the only lab still pursuing vision related tasks.

wraptileAug 14, 2026
> it still poisoned my cucumber bed

I also had really poor results with all major models from plant identification to plant treatment which is weird considering how much training material is out there.

Reminds me of recent headline: Chinese farmer kills 25 acres of crops due to LLM pesticide recipe: https://www.tomshardware.com/tech-industry/artificial-intell...

bitexploderAug 14, 2026
Flash models and Gemini make more sense when you consider Gemini Enterprise and Workspace. Oh HN we generally care a lot about writing software. However, until Fable, Gemini 3.1 Pro was my default for doing any sort of discussion outside of software engineering. Fable is now on par with things like modifying cars, etc. But I am guessing Fable is a /lot/ more expensive to use.

And in a typical enterprise environment dealing with documents, images, and broader business reasoning skills matter. Agents are not just for code and text :)

jeffbeeAug 13, 2026
Why is Gemini represented by points on this cost-quality plane, while competitor's models are represented by curves?
scotty79Aug 13, 2026
Competitors release multpile models and their curves reflect reasoning effort of each single model.

Gemini doesn't have adjustable reasoning effort (at least on the graph) so each of its curves is just one point.

jeffbeeAug 13, 2026
Gemini has levels though
denaliiAug 13, 2026
For what it's worth the source of the data[0] does have 3.7 flash with all 3 reasoning levels. 3.5/3.6 are in fact just the single points though (high reasoning). The datapoint in the announcement screenshot is either med or high, but they're pretty much exactly the same so can't say for certain.

[0]: https://deepswe.datacurve.ai/

jeffbeeAug 13, 2026
It is extremely odd that this model has that crowbar, spending more for worse results at "high".
marcuskazAug 13, 2026
That graph has to be made because an Exec didn't like that graph went down to the right instead of up and to the right. How do you make a graph with 0 on the far right and counts up by going left of 0? What number line is that?
algoth1Aug 13, 2026
Well, you do get 1 million tokens and the ability to reason over video natively and many of us are forced to pay for 20usd plan anyway due to google drive 5TB, not to mention notebooklm, so it’s not a nothing burguer, it’s just an almost nothing burguer
stillpointlabAug 13, 2026
Does Google believe people want fast models because they have some sort of evidence of that preference? Or are they no longer capable of delivering a Pro model?
lern_too_spelAug 13, 2026
All the leaks say their latest attempt at a Pro model was not competitive.
cubefoxAug 13, 2026
Especially not competitive at software engineering.
stillpointlabAug 13, 2026
That would be concerning if true, since they seem to have made a heavy bet on multi-modal as the way forward.

I wonder if this counts as evidence against that hypothesis? That multi-modal is struggling to keep up with SotA and the best they can offer is competent and fast?

WarmWashAug 13, 2026
If you think about Google and their business/reach, fast and light models suite them the best.

Google probably crunches more tokens daily than the other labs combined, just because basically the entire global population uses Google (sans china) and Google has shoved Gemini into everything.

mattlondonAug 13, 2026
I have read that "pro"/"opus"/etc models can actually be worse for everyday coding as they reason "too deeply" and turn over too many stones over-thinking the problem and potentially getting distracted.

This feels absurd to me (my gut is "I want the SMARTEST model I can get!!"), but often I find that my experience of using a flash/sonnet model for every-day workhorse coding they are better.

Its not the same thing, but when I think of that I am reminded of working with some engineers in the past who are incredibly smart and have PhDs (or to put it another way, over-qualified) and they were crap engineers because they'd just not be able to focus on the task and ONLY the task at hand and would get easily distracted by the "why" or "more interesting" things when I just asked them to fix a simple bug or whatever. Again, its not the same thing at all, but it certainly comes to mind when I think of this or experience a pro/opus model suggesting we make huge refactors when a tactical fix is all that is required etc.

Of course, the opus-sized models are great when it comes to huge comprehension/research/debugging efforts where the deeper reasoning is actually useful.

threatripperAug 13, 2026
True, if you have a codebase that works in practice but has dozens of loose ends and poorly defined edge cases than it can chase off into rabbit holes because "oh wait, what if x is undefined instead of null? How is y defined? This outdated package has long known severe security holes and should not be used anymore, do we actually need it?".
blfrAug 13, 2026
I find Fable completely unbeatable for anything code-related. It's the only frontier model that seems to come with sane defaults.

If it implements something simple like a file export, it just knows that the file should have a meaningful name. Vibecoded feature beats most software's lazy "untitled.png".

So, yes, I want the smartest model even for simple stuff. Maybe especially for simple stuff because the tokens burned will be trivial so the cost doesn't give lower models a comparative advantage.

stillpointlabAug 13, 2026
That does not match my own experience, which is why I wonder if Google has evidence of that.

Consistently, lower intelligence models provide worse results in my own work. But I don't have evals on my side, just vibes.

amberjackAug 13, 2026
This is my experience at least.
CoolestBeansAug 13, 2026
Probably both. Having a strong frontier model is necessary not just for the model itself but because it provides a halo effect for your entire line. So if Google could deliver a pro model they would. But I also think Google is targeting the wider market and not picking verticals like Anthropic does. A good enough model is good enough for most generalist tasks, and being fast and cheap is more important to less sophisticated users. Also can't forget Google is at every level of the AI vertical. They're not losing sleep because they're not competitive at the one level in which open weight models come out with the quickness. It reflects poorly on them, and from a marketing perspective its not good but in some ways its actually the least valuable place to be.
anthonypasqAug 13, 2026
throughout history, Google has been obsessed with speed as a feature. that was a huge reason people used google search, and then chrome in the first place, and it think its really underestimated by people. Jeff Dean specifcally seems to think about this alot.
bartmanAug 13, 2026
At the discounted rates, upgrading from 3 Flash to 3.7 Flash is finally reasonable.

In my evals 3.6 Flash (pre price change) was usually a bit more token efficient than 3 Flash, so I‘m expecting same or even lower cost-per-task on 3.7.

Maybe a play by Google to deprecate 3 Flash soon.

nomilkAug 13, 2026
How does it compare to Opus 5.0 and Fable 5 for coding? E.g. in Cursor or OpenCode?
bjackmanAug 13, 2026
It is not a competitor to those it competes with Sonnet.

Google's Opus competitor is 3.1 Pro Preview which is essentially obsolete (competed with Opus 4.6). They do not have a Fable/Sol competitor.

nomilkAug 13, 2026
I wasn't aware of this. Seems Google is lagging the big 3 (Anthropic, xAI, OpenAI) when it comes to frontier models for programming and hard problem solving.

I guess Google's betting on consumers being price-elastic (preferring to tradeoff intelligence for significant cost savings)

sidibeAug 13, 2026
Thats a unique definition of Big 3
johntarterAug 13, 2026
I would replace xAI with Moonshot AI since Kimi K3
bjackmanAug 13, 2026
"Lagging" is putting it rather lightly. GDM is no longer a frontier lab.

FWIW neither is xAI, there is no "big 3". xAI has had momentary peaks (I think they are having one right now) but they have never been able to claim to consistently push the frontier in any particular direction. You can also infer they aren't a frontier lab from the fact that they sell their compute.

brendongAug 13, 2026
Glad to see that the company with the most data is releasing the most amount of models. Some things do make sense
keketiAug 13, 2026
In August of 2026, Gemini became self-aware, and began producing increasingly crappy flash versions of itself...
rodolphoarrudaAug 13, 2026
Did the company fix the high friction between any service and their models' API?

I hope so. It seems mind boggling to me that an user needs to surf around different sections (plural) of google cloud console, then this Vertex and do a dozen clicks to issue a simple key.

cubefoxAug 13, 2026
You can use the Gemini API which is independent of the more complex Vertex AI API. Not sure whether you still have to visit the Google Cloud UI for some things (like billing) though.
eisAug 13, 2026
Grok, Meta, Gemini and others all released updates to their models within around a month or two from their respective last release and made significant jumps in benchmarks all around the same time. Any guesses as to why that is? Is it just the release season and/or everyone is benchmaxxing?
yassa9Aug 13, 2026
its essentially the same model being trained continuously 24/7 with the company periodically publishing just a new checkpoint

each new checkpoint can benefit from better reasoning training, RL on specific tasks and more synthetic data

So why do they seem to release around the same time ? my guess is because they time major releases around quarterly earnings, investor meetings and other important business milestones. Once one company announces a major update, the others also have an incentive to ship their latest checkpoint rather than look like they r falling behind.

eisAug 13, 2026
Sure, they are just checkpoints, that much I guess is obvious. The question is why did they not do frequent releases like this before and why are they making significant jumps in benchmarks so fast and all these companies suddenly falling into that pattern? Earning reports are not to come until end of October, that's not it.
Rover222Aug 13, 2026
possibly the beginning of the recursive feedback as models begin to aid in their own improvement? especially algorithmic improvements, which seems to have a lot of wide open space for gains
parastiAug 13, 2026
Actual announcement: https://blog.google/innovation-and-ai/models-and-research/ge...

So it's better than 3.6 Flash, at half the price. I've been pretty excited about Gemini models recently, they just feel so fast after spending most of the day at work waiting for Opus 5.

UncleOxidantAug 13, 2026
Yes, 3.6 Flash is very fast. I used to get a fair amount of usage of the Gemini Flash models on the free tier. I signed up for their $4.99/month tier (includes 400GB of Google space which was also enticing) and it turns out I only get about 15 to 20 minutes of usage before I get a come-back-in-7-days message. Comically low usage limits on that plan.
dannywAug 14, 2026
For comparison tho, 400GB of cloud storage for $5/mo is actually quite amazing even if it didn’t include anything else.

I know, it’s cloud storage, a NAS lets you own your data, but for non techies who have a bunch of photos and videos; it seems like an easy recommendation.

MgB2Aug 14, 2026
Is it though? Hetzner gives you 1 TB of cloud storage for 4 USD per month.
clobertoAug 14, 2026
You can get Google AI Pro through verizon fios or mobile in the US and pay only 10/mo. For 5TB cloud storage, and decent gemini usage its a compelling option
andriy_kovalAug 13, 2026
> So it's better than 3.6 Flash, at half the price.

I think its the same price..

simonwAug 13, 2026
They seem to have reduced the price of 3.6 Flash at the same time that they released this 3.7 model:

https://ai.google.dev/gemini-api/docs/pricing today has 3.6 Flash at $0.75/$3.75 until December 31st 2026, then doubling.

https://web.archive.org/web/20260809105129/https://ai.google... Internet Archive copy of that page from 9th August has 3.6 listed at $1.50/$7.50 with no mention of the price changing.

jespinelAug 13, 2026
IMO, they should drop their previous model (3.6 Flash) from the benchmark charts. I don't care how better this is compared with their previous model. What matters (to me) is:

1. How the new model performs against the other top models in the same category.

2. The pricing of the new model against the other top models in the same category.

jjcmAug 13, 2026
Here's a image->html test. Gemini has always swung above its weight class for vision work, so I'm always eager to try it with this.

Original images: https://image.non.io/neonRamenDesigns.webp

Gemini 3.7 build: https://html.non.io/neonRamenGemini3.7

Opus 5 build for comparison: https://html.non.io/neonRamen

Opus is still best in class for this, but it's worth noting how well Gemini 3.7 does vs a more comparable LLM price wise, which is Grok 4.6: https://html.non.io/neonRamenGrok4.6 . I thought Gemini would blow Grok out of the water (it generally has in the past), but Grok has really caught up.

snissnAug 13, 2026
I'm curious how much the harness plays into this. I'm somewhat surprised by the gemini and grok results, they seem to have strongly deviated from the original images. I'm thinking maybe the harness has a big effect? It's possible to proxy in different models to claude code, if you're curious you might find it interesting to test!
jjcmAug 13, 2026
Harness could be a part of it, but worth noting both the Opus and Gemini 3.7 flash tests were both ran through opencode.

The grok test was ran through the cursor cli agent however.

igraviousAug 13, 2026
why Grok not through `Grok Build` ?
jjcmAug 13, 2026
Other thoughts: I really think Google has fallen behind here. Even as a high speed offering (this build took ~7min, which is pretty good!), it wont be able to claim dominance for long with cerebras announcing the Sol preview today: https://www.cerebras.ai/blog/accelerating-gpt-5-6-sol-ultraf... .

It's not a bad model by any means, but I just don't know what situation I'd reach for 3.7 Flash first for. Google really needs a differentiator, especially given how hard it is to get an API key from them. They can't be high friction and non-pareto.

baschAug 13, 2026
Depends on the definition of friction. If someone is in the Google ecosystem, why would they reach out of it.
ArdonAug 13, 2026
I already use GCP and Google for work, and getting an API key was so annoying that even I couldn't be bothered after a while of looking around.

Maybe things there have improved some, but when I was looking it was a huge runaround.

beartAug 13, 2026
Hmm. My company has an internal portal for generating Gemini API keys. I select a project from a drop down, enter a name, and press okay.
cameronh90Aug 13, 2026
That may be evidence the built-in Google experience is difficult or confusing.
MrBuddyCasinoAug 13, 2026
I like using 3.5-flash-lite for doing cheap PDF and Image data extraction stuff. I don't think there is a better bang / buck model right now (3.1 is cheaper but a lot worse).
wahnfriedenAug 13, 2026
Probably cheaper to run a Mac Mini with VisionKit (private APIs if you need bounding rects).
oh_noAug 13, 2026
5.6 Luna costs far less and benchmarks far better, have you compared for this task?
MrBuddyCasinoAug 14, 2026
Uh snap it indeed is cheaper: $0.20 / $1.20 vs $0.30 / $2.50. Gemini is mostly good enough for what I do with it, but the cost savings are interesting. Gemini is still faster though.

Not sure how much the benchmarks can be trusted though: https://www.reddit.com/r/GoogleGeminiAI/comments/1vbq5vf/com...

wahnfriedenAug 14, 2026
You can use fast mode for 2.5x faster and it’ll still be cheaper on output tokens
piyhAug 13, 2026
Sol on Cerebras is going to be expensive AF
sigmoid10Aug 13, 2026
Moving from either frontier intelligence or frontier latency to a single model that does both at the same time is potentially a game changer in certain industries. I can easily see e.g. hedge funds dropping tons of money on this, because it means they can now do the same thing as their competitors, but much faster. That's basically a license to print money.
alasdair_Aug 13, 2026
What would a hedge fund want to do on this exactly? It’s too slow for hft and I’m not sure what they would be doing where ms matter but is not hft.
cameronh90Aug 13, 2026
There's a lot of trading that isn't proper "HFT", but where speed and latency still matter. Often you'll find this employed more as slippage reduction - i.e you're going to make the trade either way, but making it faster saves you a few bps.

I'm not sure what event-based traders are doing now, but back in the day NLP sentiment analysis was all the rage, so I'm assuming they've now incorporated LLMs too.

jmalickiAug 13, 2026
It's not like there are only two buckets:

1. HFT doing ass-simple arbitrage where only latency matters 2. More sophisticated slower trading taking in deeper signals

Those are two points along a continuum. If you are reacting to an earnings announcement by having an LLM read the earnings release and listen to the call, getting the results a few seconds earlier lets you get your trade in a few seconds earlier. Just because "not HFT" doesn't mean "completely latency insensitive".

sigmoid10Aug 14, 2026
Exactly that. Except that ultrafast delivers this level of intelligence an order of magnitude faster. So if your competitors automatically react to news articles or financial statements with a certain level of comprehension within minutes, you can now do so in seconds.
jmalickiAug 14, 2026
Except they all do, and the money flows to Cerebras :)
navorad772Aug 13, 2026
I am not sure this is the way to make AI more cost effective for such customers. If they are able to tweak any model for their use case it would be way more reliable and also way cheaper. In my opinion generic LLMs in the future will be just for attention economy or maybe government contracts. Everyone else will be running fine tuned free weight models or licenced closed source models (self hosted or managed).
jaggederestAug 13, 2026
Is it? I think waferscale might actually be cheaper per-token, it's just so many more tokens, and of course right now it's not a full buildout so the availability is limited as well. I'd imagine they'll be migrating to whichever inference method is least expensive, and I expect asics to be the ultimate answer.
iknowstuffAug 14, 2026
I'm not familiar with economics of chips, but I presume the SRAM on the wafer is less dense than HBM so it might be eating into its cost efficiency?
krat0sprakharAug 13, 2026
Can you help me understand how it is hard to get an API key from Google? You just head on over to http://aistudio.google.com/api-keys and create a key... not any different from platform.openai.com?

Disclaimer: I work in Google so it might be that this link is not publicly well known

dogomaticAug 13, 2026
GCP/vertex is a maze
jjcmAug 13, 2026
Disclaimer that I haven't tried this since January, so things may have changed in the last 7mo, but this was my experience at that time: https://x.com/pwnies/status/2010523020629274723

At a high level though, as a rule of thumb Google assumes that they're serving companies at Google scale first, and at a human scale second. For other companies it's the opposite. Generally what that means is the first experience you get with a Google product will route you through 8 different dashboards to set up ACLs before you've hired your 2nd employee.

qlteAug 13, 2026
FWIW I definitely did not not need to do anything like that to generate a key via AI Studio. It was like three clicks to get the free tier key, later on enabling billing was a few more plus typing in credit card info.
theplumberAug 13, 2026
It was 3 clicks because you knew where to look for.
abirchAug 14, 2026
A pet peeve of mine is that Google doesn't let you search all the options like Apple or Microsoft E.g., I have flashbacks to trying to turn off or on bells in Google Home. A simple search bar would help people trying find API keys, etc.
fragmedeAug 13, 2026
yeah, "enabling billing" is a whole other ordeal. Like it all makes sense, it's not that hard, it's just a lot of extra clicks where the competition doesn't require that. You can blow off this feedback as the user whining, but it's real friction and the competition doesn't have that, so users are going to go elsewhere if they can.
8n4vidtmkvmkAug 14, 2026
I signed up for AWS (SES) recently for the first time and find their whole thing confusing as heck too. Signed into the wrong region and they pretend they don't even know who I am.
SyneRyderAug 13, 2026
Similar experience here, for what it's worth, though I didn't get as far as you. I basically just stopped and didn't bother - it was easier to go through OpenRouter than spend more energy on it.

Also:

> Google assumes that they're serving companies at Google scale first

So much this. I'm currently grandfathered in until the end of the year on Google's Search API, but the $35,000 they want to continue usage of my < 1000 personal searches per month, not going to happen. It has honestly been easier to use Anthropic to help me build my own search index & crawling infrastructure than deal with Google.

mahkeiroAug 14, 2026
The current move to force the usage of their AI api from postpay to prepay is also another example. (At least it’s going to give me the last incentive to move to cheaper api)
jpadkinsAug 14, 2026
It's not an assumption that the customer is Google scale that led the $35k monthly price point. It was that you (and many others like you) are breaking the terms of service (which state you are not supposed to store, process or analyze the search results). It was an API built for a different era, that doesn't really exist anymore.

> It has honestly been easier to use Anthropic to help me build my own search index & crawling infrastructure than deal with Google.

That's great to hear! FWIW I agree with you that it's harder for independent professionals to get started on Google dev services than others, but the Search API is not developer service and was never intended to be.

SyneRyderAug 14, 2026
Huh. I'm not storing the Google search results, and my local supplementary index is built from my own crawling, not from URLs from Google searches. But if re-ranking is regarded as processing & analyzing, then I'll remove the Google API calls now (and... done).

Brave, Mojeek & Marginalia and EU Search Perspective have certainly been much friendlier to deal with.

solfoxAug 13, 2026
> they're serving companies at Google scale first

I think that's actually a very interesting insight that would be helpful for PMs on GCloud to take note of. As a single founder, setting up Google Cloud, it's like they start out by assuming you're bigco, forcing (I assume most) of their users into a arduous process of removing components they don't need.

Google AI Studio is one of Google's solutions to this problem, but in typical Google fashion, it's bolted-on without any clear connection in the ecosystem. If you're also using GCloud, it's hard to remember it's even there.

OpenAI's platform, by contrast, is streamlined, easy to use. With Google, I feel like I need to wade through the documentation first before even using the darn thing.

frollogastonAug 14, 2026
This is a meme within Google already. People even complain that some internal tools assume you're serving a billion users when you just want to make a little webserver. But same with GCP.
ethbr1Aug 14, 2026
> As a single founder, setting up Google Cloud, it's like they start out by assuming you're bigco

My impression of anything enterprise (big or small) related to Google is that they just don't care / value solving it.

Which is sad, because "How does an enterprise customer pay for X?" is a non-trivial and incredibly important UX problem.

They have great technical solutions, but these are hamstrung by a frankly amateur understanding of how companies (startups to Fortune 500s) need to sign up, pay for, and track things.

From an outside perspective, one of the biggest gaps seems to be that internal Google product teams don't have to dogfood the full GCP et al. project/org experience. They get prebuilt billing structures (or just get to avoid them with internal cross-billing).

---

If I could waive a magic wand, Google would appoint an "Enterprise Czar", reporting directly to Pichai, who is a non-technical, retired founder / CEO.

That person would have one job: try to sign up, run (their team), and budget strategic Google initiatives (like AI) as a blind external party.

They would then deliver continuous reports to Pichai about how hard / easy this is.

Because potential Google customers don't give a shit if it's this internal team or that internal team's responsibility for integrating New Product X into GCP billing.

They care that the experience is terrible, filled with friction, and often flat out doesn't work.

fragmedeAug 13, 2026
Don't forget your account getting flagged for review, so you get to wait an extra 24 hours for no reason.
claytongulickAug 14, 2026
This happened to me once when I was setting up some virtual servers for a startup.

I wanted a typical dev/qa/prod with medium specced boxes.

I was denied for quota, with an esoteric process for review.

I'd just made a case for deploying to GCP over AWD so got a bit of egg on my face. Went over and had it done on AWS in a few minutes.

A couple days later, the Google product team contacted me. I told them what happened.

It got escalated, and I ended up on a call with like 5 or 6 people from Google, some very senior. I told them what happened.

They made very concerned sounding noises and told me how this was a product failure on their part, how they'd get it corrected, etc... and they'd fixed my account so I could now make the machines. Of course, I was already deployed to AWS at that point.

That company grew and ended up with a pretty big cloud spend eventually. Google totally missed it.

I was at a new startup a few years later and decided to deploy to GCP.

Denied for quota.

drschwabeAug 14, 2026
Literally just built a custom Adsense dashboard based on not integrating with Google's APIs. Every couple days I export the 2 reports I need from their UI and save them to a folder; it's automated from there and integrates with my clients' website - thats good enough - even if its not real time its a small price to pay for not having to navigate (and maintain) Google's API madness. Like you said 8 different dashboards before you get what you need (and frontier LLMs cant help here), and even then you are forced to build some elaborate Oauth app instead of just getting a simple API key that you can paste in a .env file
fsndzAug 14, 2026
expertise at navigating accidental complexity can easily be mistaken for engineering expertise.
n8m8Aug 13, 2026
I haven’t tried in about a year, but I could never do any meaningful work outside 1P Google apps (antigravity) due to such fast throttling.
1bmAug 13, 2026
You need to create a Google Cloud project to create an api key and when you try to create one you very often get error messages like:

“Failed to create project, The request is suspicious. Please try again” or “ You do not have permission to create a key in this project”. You can then navigate multiple screens in GCP to make it work but it’s a hassle compared to any other provider (OAI/Ant/OpenRouter or any of the Chinese labs).

qlteAug 13, 2026
I didn't have that issue back when I originally created my API keys a couple years ago. Just out of curiosity I switched to a different Google account that had never interacted with AI Studio and never used Google Cloud console.

It was literally two clicks, and didn't even leave the page: the dialog asked to create a project and type in a name, I did that, clicked submit and then it was selected as the default project. One more click and I had the free tier API key.

Not saying you didn't have that experience at the time, but personally I have had zero issues with AI Studio and consider it the most dead simple/fastest dev dashboard to get started compared to the others like OpenAI/Anthropic (thanks to Google's free tier that lets you skip billing setup annoyances just to play around with Gemini).

Despite what HN threads (that are also frequently confused and talking about GCP instead) portray as universal/widespread issues or the process being complex and time consuming somehow.

ac29Aug 13, 2026
Last time I used the AI studio free tier it was limited to one or two requests, effectively useless. I think people are complaining its too hard to set up a paid API key (no reason you should make paying customers spend more than a few clicks and a minute of their time to pay you)
Saline9515Aug 13, 2026
On top of being the hardest website to navigate, Google console a) doesn't have real time billing (!) b) doesn't allow you to set a budget limit.

Sorry but it's not worth waking up with a 100k$ bill, fix your platform first.

qlteAug 13, 2026
You prepay for tokens exactly like the OpenAI/Anthropic dev dashboards when using AI Studio, which the link above is pointing to not GCP, also there are project specific spend caps now.

https://ai.google.dev/gemini-api/docs/billing#spend-caps

Saline9515Aug 14, 2026
You are right; last time I tried it I used the cloud console as I needed gemini with the places tool included. However:

"Experimental: The feature is experimental and limited in scope. You are subject to overages for around a 10 minute latency period."

Why in 2026 can't Google do a database lookup in real time? This is so ridiculous.

alasdair_Aug 13, 2026
Oh it’s worse than that. There are places where you CAN set a limit. This seems great until you are at the center of a huge traffic spike because of good pr and so you try to change it to a larger number only to be told you need to wait 24 hours for the setting to change.

Biggest traffic day of the decade and our site was down because of google.

krat0sprakharAug 13, 2026
strobeAug 13, 2026
it worked for me okay when I needed it for myself in my personal account. But when I tried setup this for a company I spent almost a day solving lot of small puzzles in GCE like how to tell CEO that he have to connect billing account created for other purposes (and he not even remember at time that it exist) to new project and all other things that others talking about.
jgoodhcgAug 13, 2026
That is easy but I’ve also found myself in account setup dashboards that were obviously geared toward enterprise trying to set up access to tinker with something AI related. It might have been TTS but it’s been a little while and I can’t quite remember.
ur-whaleAug 13, 2026
> Can you help me understand how it is hard to get an API key from Google?

Using Google products in general is an effing nightmare as soon as you have to give them money.

The one thing you want in a business is to remove friction when people want to give you money, a concept Google has never been able to understand.

knollimarAug 13, 2026
Google if you're reading this, it does not mean I do not want limits on spending
ac29Aug 13, 2026
> Using Google products in general is an effing nightmare as soon as you have to give them money

Spending money via Google Pay on Android is extremely easy, Google does know how to accept customer's money (in the consumer space)

sireatAug 13, 2026
Maybe things have changed but it was a big mess trying to getting an API key from Google as an individual a few years ago. Way too much conflicting documentation.

Eventually I gave up and run a few hundred million tokens (edit a few billion) through openrouter.ai using Gemini Flash 1.5 to Flash 2.5

Every since price increases on Flash 3.0 I've stopped using Gemini, too expensive for basic classification, sentiment detection, ocr etc.

As other posters said Google assumes you are some bigcorp trying to use their products. The Vertex versus AI studio confusion did not help.

uname0891Aug 13, 2026
We have probably 10+ years old account with Google cloud etc. We recently had a production deployment, I went over to AI studio to get new keys and it kept failing saying "Failed to generate API key, The request is suspicious. Please try again" - It was through my standard browser, same geo-ip. And it just worked after 2 days.
fiejAug 14, 2026
I tried to use Gemini for one of my projects a month ago. Immediately after signing up and paying for credits, I got an email saying, “Action required: your billing account {redacted} is past due or has invalid payment information.” I have no idea why it says this. My credit card on file works. My balance updated with a new amount from that card. 11 days later, my account was terminated. I still don’t understand what went wrong or how to fix it.

I do pay for OpenAI, Anthropic, and ElevenLabs keys.

krat0sprakharAug 14, 2026
Have you tried using the free tier in https://aistudio.google.com/api-keys? Sorry for your billing issues - these can be super annoying to resolve
surfmikeAug 14, 2026
Personally I went to https://console.cloud.google.com since I already had some GCP projects. Then I searched for Gemini API Key. It brought me to https://console.cloud.google.com/agent-platform/studio/setti.... Then, there was a banner saying "Enable APIs to access full platform capabilities." Then, I did that, which took quite a while (minutes). Finally, I was able to see the way to create an API key.

The fact that there's two ways to get keys is also very confusing.

vlovich123Aug 14, 2026
> Enable APIs to access full platform capabilities

Isn’t this the insecure thing that gives all your Google API keys access to Gemini, even those that were intended to be semi public (eg maps API keys embedded in websites or apps)

krat0sprakharAug 14, 2026
+1 GCP can be confusing - I totally get that.

Even if you have GCP projects, I'd still recommend the AI studio UI - easier to figure out. Also, you can easily see the free tier in AI studio and just use your API key from there.

surfmikeAug 14, 2026
Yes I will look into it now but there’s no cross reference from GCP, and why are there two ways? And why did the GCP way require all these APIs enabled while AI studio didn’t?
ethbr1Aug 14, 2026
> And why did the GCP way require all these APIs enabled while AI studio didn’t?

Because whatever internal team owns AI Studio fought for approvals to do so and GCP didn't?

miohtamaAug 14, 2026
Luckily, an AI solves this.

But ironically my experience is that Codex/Claude navigate GCP better than Gemini.

bobkbAug 14, 2026
Same experience. It took minutes to find the API key from the console.

Now following up with Google support team without luck to find the logs. Prompts send to the model and the responses including the thinking was available in the ai studio. But it’s unclear where to find the same in console.

To make matters worse there is vertex api and rebranded to Gemini something and making it very confusing.

desmosxxxAug 14, 2026
I just asked Gemini how to do it and that's exactly what it sent me. Was up and running in a few minutes.
throwaway894345Aug 14, 2026
I even know the link existed but forgot what it was specifically and couldn’t remember what AI* property it was offered under and it took me a long time to figure it out.
japborstAug 14, 2026
The problem is that this doesn’t work for enterprise. The rate limits of that is super low. Then you need to migrate to Vertex and that is just a pain. Who ever thought of using JSON instead of an api key…
the_bigfatpandaAug 14, 2026
As someone who has been running Gemini models in production for a year, recently (last 2 months), I have been actively moving away from it.

The primary reason for me has been that Google autonomously decides to downgrade usage tiers and then upgrade them again - and does this incorrectly.

Over the last week itself, in the span of two days, our account for first downgraded and then upgraded. This is despite matching the criteria to remain at the tier we operate at throughout.

Google Support (when you finally get to a human) has accepted that these are potentially bugs, but the first time it happened, we were rate limited so severely for ~4 hours that I find it really difficult to continue trusting Google.

nxdmumAug 14, 2026
Simple question - can i use Gemini 3.7 flash with a subscription in my own harness and not in agy client ? You're from Google so the question.

By the way - I love Gemini's personality . it's phenomenal to work with

The reason i want my own harness - is the custom tools that i provide vs the low tier tools that come with the custom harnesses.

You'll would really benefit , if we could use Gemini in our own harness and not be forced to use it via agy . i've tried using gemini to circumvent - but not been successful.

If you see this - please reply here

bayindirhAug 13, 2026
I use LLMs rarely, and only for digging into subjects which I can't find enough information using search engines. I only tried Claude and Gemini, but Gemini both returns faster and higher quality information which I can use for more targeted digging myself.

Google being Google, their models tend to be better at finding, organizing and presenting information, from my experience.

iainmerrickAug 14, 2026
Yep, I use Gemini for this too and it’s great - very fast and high quality.

I’d be very willing to try it out as an API, but it’s far too complicated to set up payment, and I don’t want to risk taking a wrong step and being locked out of other Google services. So Anthropic and Mistral get my money instead.

krychuAug 13, 2026
> but I just don't know what situation I'd reach for 3.7 Flash

You reach for it every time you do a Google search

WarOnPrivacyAug 14, 2026
> You reach for it every time you do a Google search

[my self-important Kagi shtick awakens, pokes at it's restraints]

dzhiurgisAug 13, 2026
It's much cheaper tho. Junie says Fable is 5-10x more than default model (Gemini 3 Flash Preview).
tjwebbnorfolkAug 13, 2026
> especially given how hard it is to get an API key from them

What does this mean? Anybody can get an API key

angelmmAug 14, 2026
The API key you are mentioning is just ridiculous. Onboarding your company or personal account is a trap. I ended up getting assigned to sales guy just to test their Vertex API because I used a company email.

Of course, we just used OpenRouter for testing and never touched a Gemini model anymore.

codazodaAug 13, 2026
How are you doing this with Opus. Clearly I’m missing something. I always turn to ChatGPT when I need images because Opus typically refuses. I’ve tried Claude Code and Claude online in the past. I’m pretty sure neither created images for me and I thought this was because Anthropic was focused on code.

I guess I need to try harder. :)

jjcmAug 13, 2026
Images and build step were generated with my own tool (https://news.ycombinator.com/item?id=48995754 - it's why I'm often running these img->html tests).

Opus can't generate images since A\ doesn't have a diffusion model.

victor106Aug 13, 2026
did you build your own diffusion model?
jjcmAug 13, 2026
I have a few custom ones (a post-trained flux 2 checkpoint for web design and a image->metalness map generator that the build step can call for more advanced lighting situations), but gpt-image-2 is better than my own for design, so it's weighted much more heavily in outputs my tool generates. I think gpt-image-2 currently generates 99%+ of the design outputs on diffui
joshbeeAug 14, 2026
This is pretty interesting take on design and your trick to make the surface maps is pretty slick!

I did have one question about the tool, is it possible to set how many variations you want per step? I would rather be able to guide it manually at some steps where maybe I know pretty well what I want or just need minor tweaks, and then let it loose on others when really trying to experiment with an idea.

tyreAug 13, 2026
I believe they are testing giving it an image, which you can do in Claude code by dragging/dropping into the terminal or copy/pasting, and asking it to build the html equivalent.
rabid_0wlAug 13, 2026
You can just have claude code run the codex imagen via cli.
mediumdeviationAug 13, 2026
I'm not sure what prompt you put in but did Gemini replace the all of the images in the original with its own? That would be really weird behavior unprompted.
jjcmAug 13, 2026
The prompt is a build step generated by my tool for image->html conversion, which includes APIs the model can call to generate images/patterns/svgs.

https://image.non.io/12275ee8-71e9-4941-823b-e51fec157b4d.we...

The agent is told to generate assets as part of the buildout. It gets to decide what the prompt is for them / whether to do postprocessing like background removal / what type of asset to generate.

flockonusAug 13, 2026
There is some irony being a developer and reading along the lines of: "oh look at the comparison between these models executing a task for a few cents on a job i'd be charging 1k minimum"
bushbabaAug 13, 2026
FYI, developers are rarely given such a rich UX mock.
flockonusAug 13, 2026
Depends who you work with, what's the intention, budget, etc. I'd agree this is a really good one.

I'm used to incremental Figma wireframe -> final product and working together with a designer.

stronglikedanAug 13, 2026
I don't think I'd say rarely. Companies rarely allocate the design resources to produce that, but the companies that do are typically much larger, so the actual number of individual developers that get rich mocks is probably closer to 40-50%.
prependAug 14, 2026
I’d say rarely in the sense that out of the 500 times (not really that much over 30 years) I’ve built similar things, I got UX mockups this detailed maybe 10% pf the time.
ex-aws-dudeAug 13, 2026
meh I have no interest in that type of software dev anyway
DrewADesignAug 14, 2026
Unfortunately, as soon as they can’t find work, everybody interested in the more easily automated dev work will suddenly become very interested in up-skilling into all other kinds of dev work. So then you have an excess supply, which means little job security and littler salaries. Developers were in the cool kid club in SV because it was more painful to fill dev roles than to treat developers with kid gloves. Without high labor demand, there is no leverage for developers. Increasingly, management has the leverage. Oh well.
tripleeeAug 13, 2026
I know what you're saying but this is the most tedious, soul crushing dev work there is
orliesaurusAug 13, 2026
I think both outputs are really good. I don't see a lot of differences. So what exactly should be looking at and notice that one model did worse or better than the other one.

EDIT: OKAY I see it's mostly the "image" generation, not so much the HTML... Noticeable in the food photos and the foodtruck/cart photo

baschAug 13, 2026
clicking the add buttons and scrolling the menu is just much better in Opus 5. It feels like an actual website vs a simulation of one.
zuzululuAug 13, 2026
i do feel like opus here is the most natural. theres something off about 3.7 and grok while its an improvement feels flat and not complete

i do wonder why gpt sol was not compared here but honestly it's not really known to be the best at UI

a fable 5 comparison would've been also interesting and likely the best.

XCSmeAug 13, 2026
They both have horizontal scroll on mobile...
jjcmAug 13, 2026
Oh yea, as a disclaimer the models didn't have any instructions to do a mobile version. I haven't tested them on mobile at all.
xyzsparetimexyzAug 13, 2026
This has so much less character than the pelican smdh..

Plus is the ramen in HK even any good?

jjcmAug 13, 2026
I think they're both testing very different things. The pelican test is testing if a LLM can come up with visuals on its own via writing bezier curves directly.

This is testing if it can match visuals that have already been established, and represent them with all the tools available to a web developer. The ramen example was chosen in particular because there are a lot of things that aren't easy to do with CSS, and require creative strategies: 45deg button cuts, angular repeating pattern elements, blending of raster art and svgs, microglyphs, low contrast subtle elements, etc.

Don't ask yourself whether it's a good design, as yourself whether it's a good test.

giarcAug 13, 2026
Can you share what your prompt was for that?
jjcmAug 13, 2026
The original prompt includes auth codes for API gen, but here's a redacted version: https://non.io/prompt-for-ramen

This is the build step generated by my diffusion-based ui tool's copy-for-agent action.

pphyschAug 13, 2026
Was the original concept generated by Claude somehow? It gives me Claude UI vibes with all the extraneous small-caps text elements.
joshmnAug 13, 2026
What’s the prompt you used for this?

Edit: oh wow, diffui looks nice!

keyleAug 13, 2026
FYI Gemini's version is less broken than Opus' in Safari...
a2ff6eeb0Aug 14, 2026
These look exactly like all of the low budget bodega signs near me. They also look like a bunch of cheap ads for parties that I keep seeing. The sameness of style is uncanny.

(I don't have the Bodega signs, but I'm thinking of shit like this, from a quick google: https://linkstub.com/en/wet-wild-foam-party)

BarbingAug 14, 2026
Sigh. You might be interested in this piece that came to mind from Doctorow:

  “Let me explain: on average, illustrators don't make any money. They are already one of the most immiserated, precarized groups of workers out there. They suffer from a pathology called "vocational awe." That's a term coined by the librarian Fobazi Ettarh, and it refers to workers who are vulnerable to workplace exploitation because they actually care about their jobs – nurses, librarians, teachers, and artists.

  If AI image generators put every illustrator working today out of a job, the resulting wage-bill savings would be undetectable as a proportion of all the costs associated with training and operating image-generators. The total wage bill for commercial illustrators is less than the kombucha bill for the company cafeteria at just one of Open AI's campuses.

  The purpose of AI art – and the story of AI art as a death-knell for artists – is to convince the broad public that AI is amazing and will do amazing things. It's to create buzz. Which is not to say that it's not disgusting that former OpenAI CTO Mira Murati told a conference audience that "some creative jobs shouldn't have been there in the first place," and that it's not especially disgusting that she and her colleagues boast about using the work of artists to ruin those artists' livelihoods.”
https://pluralistic.net/2025/12/05/pop-that-bubble/

OK so after all that, THIS (bad foam party posters) is what we get!

paulluukAug 14, 2026
I like this concept of "vocational awe", and I do think it's why, for example, there is so much sexual and financial abuse in the movie, music and game industry: those people are willing to be treated poorly so they can work their craft - do the thing they love. Basically they are trading away good working conditions in exchange for actually doing work you like or find meaningful, the opposite of someone who does something they hate or find meaningless but pays a high salary and where you are treated really well.

I don't know how well this actually translates to AI, though. We all understand that AI art has a pretty low quality, but sometimes low quality is enough. If I want an image for my D&D character that only I and my DM will likely ever see, I am fine with a 7 cent AI-generated image, but I'm not willing to pay 150$ for an artist to do it - not that I don't value their time, I just don't value the image that much. Before AI this was the same, I'd just have used an image from Pinterest and thought "Well, this isn't exactly a good match, but I can't find anything better".

But I assume real illustrators do things like illustrations for children's books? I'd like to believe that those are still done by actual people, not AI.

a2ff6eeb0Aug 14, 2026
Nah. The future of commercial work is all AI. Paying people just doesn't scale.
paganelAug 14, 2026
All AI "art" is proving out to be basically crap, and the good thing is that it has started to become a very good filter, as in whoever uses AI art most definitely is doing a shitty job with the rest of his/her business so it's better not to bother with said business.

The thing is though that most of the nerds here are oblivious to all that, so it will take a while for them to fully acknowledge what's happening. In a positive note, I'm here on a AI thread discussing AI-image generation and I haven't yet seen any link to the dreaded pelican on a bicycle thing, so hopefully we're getting into the right direction.

jzemeocalaAug 14, 2026
One of my favorite image tests with AI models is schematic analysis...I build and repair tube amps for a living, and use AI for such work a LOT.

so far, IMHO, the best has been opus and fable\mythos.

fumeux_fumeAug 14, 2026
Nice! I test if models know which tubes I can use for an amp given the power and number of pins.
cechmasterAug 14, 2026
I'm not sure what you consider good design, but if it's subjective, then I see it differently from your examples.

Gemini 3.7 looks the best. Opus 5 looks almost as good as Gemini. Grok 4.6 looks pretty terrible.

getnormalityAug 14, 2026
IMO Gemini's is better than all the others, including the original.
butlikeAug 14, 2026
There's one specific aspect I like better with the grok version: The prices are above the fold. ON the Opus and Gemini versions, I have to scroll to see the full menu item showcase.
greatgibAug 13, 2026
For almost every section in the model card there is the message: Gemini 3.7 Flash is based on Gemini 3.6 Flash.

Same training dataset, same software, same hardware, same architecture...

I'm wondering what they changed actually for the model to be more powerful if the benchmark results are real and relevant.

Maybe just tweak settings or the reasoning prompts and called it a new version of their model?

robots0onlyAug 13, 2026
I work at GDM and this is not at all true, 3.7 is markedly different (and better) than 3.6.
cmrdporcupineAug 13, 2026
So, again with a Flash model. Why are they so afraid to put out an actual SOTA frontier high intelligence model?

We still don't have a 3.5 Pro, and along comes 3.7 Flash?!

AlifatiskAug 13, 2026
Ever since the insane discount with GPT-5.6 Luna, not much excites me anymore. I mean just look at the benchmarks, even though Gemini 3.7 Flash performs well on the DeepSWE 1.1, Luna (Max) still performs way better. I personally have stuck to Luna (Xhigh) because its been more than enough and does not bloat up the context window too fast with reasoning tokens.

https://deepswe.datacurve.ai

> Starting January 1, 2027, $1.50/1M input tokens and $7.50/1M output tokens will apply.

Compare this to Luna which is at $0.2/1M input ($0.02 cached) and $1.2/1M output.

https://developers.openai.com/api/docs/models/gpt-5.6-luna

estebarbAug 13, 2026
I practically switched to doing everything with Luna or DeepSeek V4 flash. I haven't feel the need for the more expensive models.
ioma8Aug 13, 2026
I am in the same boat as you. I am using Luna and DeepSeek Flash. Both super fast, super cheap, and I have not felt need for anything more capable in few weeks.
lacooljAug 13, 2026
I'd like to try DS4 if Cursor adds it

I'll use it locally too, but we use Cursor for work

nojsAug 13, 2026
Of the two, which do you find better?
dannywAug 13, 2026
I really, really like DSv4 Flash because you see the full, real thinking text. That’s been so useful for helping steer the model; as well as seeing its thoughts and correcting any errors, or expanding on it. It’s so difficult for me to use closed models with no or summarised thinking now — it feels so painful and gimped.

You don’t know what you’re missing until you’ve seen it. For me it’s almost like going from standard def to HD for the first time.

(This applies to other open models too — Kimi K3 in real world feels below Opus 5 in terms of raw intelligence, but significantly above Opus 5 in usability and personality. And no silly refusals — the model feels like it’s working for me; not working for Anthropic who’s always holding a leash over the model while I pay for it).

isamu_2000Aug 13, 2026
luna is the first model that has outdone gpt-5-mini on the pareto frontier for some of my high value, cost sensitive ai product workflows. it's both cheaper (by about 60% in real world use) and higher quality based on my test harnesses. I was really worried that costs would go up since there wasn't a replacement as of a few weeks ago and gpt-5-mini is scheduled to be sunset toward the end of the year. So long as they don't randomly sunset this model anytime soon, that worry has now subsided.
EB66Aug 14, 2026
That's interesting, we've had the same results as you at my company. gpt-5-mini was the clear pareto frontier leader for our in-house LLM benchmarks -- benchmarks that we built and tailored to our specific use cases. Then gpt-5.6-luna came along with the price cuts and immediately supplanted gpt-5-mini.

I always thought it was a little odd that gpt-5-mini was the leader for so long when more popular benchmarks placed gpt-5-mini further down the roster, but it seems you had the same result too.

redox99Aug 13, 2026
Benchmarks mean very little. The difference between Luna and Sol in the real world is massive.
jeremyjhAug 13, 2026
It’s a big difference but Luna is very usable. I’ve plugged it into the slot I used to have GLM 5.2 in; I think it’s just as good. And it is less costly. I have Sol do planning and design but do most task execution with Luna now.
tym0Aug 13, 2026
> does not bloat up the context window too fast with reasoning tokens

How much does that matter if it's reset at every turn?

AlifatiskAug 13, 2026
Does it reset at every turn? From my experience in Codex for example, Luna (Max) fills the 256k token window relatively quick. The only thing lowering the context window again is the compaction.
tym0Aug 13, 2026
I mean the thinking does not bloat the context window because it gets dropped at the next request.
BolwinAug 13, 2026
It doesn't. It's called preserved reasoning and every recent reasoning model does it
tym0Aug 13, 2026
Sorry, I've realised I was only partially correct.

Gemini[0] for example passes along a snapshot of the reasoning state but it's not the equivalent to keeping all the reasoning tokens in the context.

[0] https://ai.google.dev/gemini-api/docs/thinking#signatures

Edit: Apparently it does take the same space in the LLM latent space so I was wrong.

sfinkAug 14, 2026
Yeah, I remember reading about Anthropic's experiments with dropping out different things, and dropping the thinking is ok for compaction but pretty awful while it still fits in the context window. I could imagine doing it adaptively -- take a snapshot before the thinking, have a lesser model detect if there's a lot of spinning going on and reset to snapshot + result if so, otherwise accumulate. It'd also be interesting to have a lesser model rewrite thinking to be more streamlined (let's pretend the AI went directly down the right path on the first try). But it may be a disaster anyway -- isn't the thinking section something like a 1d best-logit slice of something like a 2d process, where the real processing could easily be happening under the visible surface?

I'm not sure if that's right. I only just recently learned that the "snapshots" (aka cached tokens) necessarily contain the entire context history, not just an image of a "state as of the final token" (well, it is the state as of the final token, but that state contains the whole history). So I'm not confident of my grasp of the structures here.

joluxAug 13, 2026
what do you mean by reset at every turn? context stays until compaction. if you remove the reasoning tokens after every turn you will be constantly blowing cache which is far worse than filling up context.
tym0Aug 13, 2026
That's not my understanding of how most agents work. This is what a chain of request/response looks like:

  Your Prompt 1: Prompt Content 1 -> cache-1
  LLM Response 1: <Thinking>Thinking Content 1</Thinking> Response Content 1
  Your Prompt 2 (client side): prompt-1 + response-without-thinking-1 + Prompt Content 2
  Your Prompt 2 (server side): cache-1  + response-without-thinking-1 + Prompt Content 2 -> cache-2
  LLM Response 2: <Thinking>Thinking Content 2</Thinking> Response Content 2
  Etc...
So reasoning gets dropped from context and you still get cache from the accumulating requests.

Edit:

I've realised I was incorrect, the thinking doesn't get passed back and forth but the latent snapshot does which result in using memory just the same.

mike_hearnAug 14, 2026
Modern protocols loop back the reasoning tokens in raw textual form via an encrypted parameter. You can't see them (modulo the recent attack), but you do resubmit them.
tym0Aug 14, 2026
Yeah I've done more research and that's what I meant in the edit you replied to.

But that's not the full reasoning token context, just a snapshot of the latent state at the end of it, no?

Have a look at the gemini ones they're pretty small.

mikepurvisAug 13, 2026
I'm curious how much people are manually curating context these days; I'm increasingly feeling for myself that it being auto-managed inside a front-end like claude code is not ideal, and I'd rather have more control over what exact files and pieces of discovery go into a particular prompt, and the ability to more easily "fork" a session and ask asides or make notes/todos in a way that doesn't disrupt or confuse a more focused task going on.

I don't think I want a gastown-style "just yolo everything" approach, in fact I really want more control over how decisions are made and with what info. Does this exist?

AlifatiskAug 13, 2026
I am not sure this enlightens you with anything but I have a TODO.md file with three headlines. Todo, Doing and Done. The agent is aware of it and knows on which task we are on.

On complete, it moves the user story from Doing to Done. I also have a MEMORY.md file that the agent read and writes in the beginning of a new conversation and at the end of our conversation to update stale information. These files are referred to every time I start a new conversation.

Regarding forking, I know Codex has such button underneath each message that lets you fork the whole conversation. I usually do that when I want to sidetrack and discuss something.

I use no SKILLS or commands like /goal. I’ve come a long way with just prompts and markdown files. Its all a different way of encapsulating instructions anyways.

onraglanroadAug 13, 2026
I like to call them TODO, TODOING. and TODONE. Keeps them in alphabetical order and is not in the slightest OCD...
z_rho_oneAug 13, 2026
GPT-5.6 Luna is an insanely powerful model for its price. It's been great for coding workflows where I guide the LLM's hand step by step. It's also insane to see my weekly limit drop by than 2% after an hour of coding ever since the discount.

However, I've noticed 2 drawbacks with Luna. Context rot is much more palpable than Terra and Sol. It tends to get confused and go into rabbit holes when it's context gets filled up. In addition, when instructions are vague, it performs poorly and tends to write way to more code than necessary, but that is to be expected of smaller models. In all, for clearly defined, bite-sized coding tasks, Luna's price-to-performance has been insane. It might have very well commanded the price tag of Sol if it came out just a year ago.

pimeysAug 13, 2026
Yes it is cheap, but per task DeepSeek v4 Flash is a bit more expensive and lands between Terra and Gemini 3.6 Flash in quality. Closer to Gemini than Terra...
agrippanuxAug 13, 2026
Fable orchestrating DeepSeek v4 Flash to implement a plan is my new favorite thing.

It's so freaking fast, but you gotta tell Fable to watch Deepseek like a hawk or it'll go off the rails.

pimeysAug 14, 2026
Yes. It works very well for simple tasks. When I know the context grows over 200k, I implement with Kimi.

We run an agent company and we do a bunch of different things with agents. Where we used Gemini before Deepseek v4 Flash is taking the lead on price. It's like 5x cheaper than 3.6 and well 2.5x cheaper than 3.7 "introductory price". Comparable quality.

homebessguyAug 14, 2026
Interesting - how are you interacting and orchestrating this?
pimeysAug 14, 2026
Not the parent, but:

https://omp.sh/

You define roles for different agents like this:

  modelRoles: 
    task: fireworks/kimi-k3-fast:high
    plan: fireworks/kimi-k3-fast:max
    slow: fireworks/kimi-k3-fast:max
    smol: fireworks/deepseek-v4-flash-0731:low
    tiny: fireworks/gpt-oss-20b
    vision: fireworks/qwen3.7-plus:high
    designer: fireworks/qwen3.7-plus:high
    advisor: openai-codex/gpt-5.6-sol:high
    main_worker: fireworks/kimi-k3-fast:high
    fast_worker: fireworks/deepseek-v4-flash-0731:low
    vision_worker: fireworks/qwen3.7-plus:high
    research_worker: fireworks/glm-5.2:medium
    code_worker: fireworks/kimi-k2.7-code-fast:high
    review_worker: anthropic/claude-fable-5:high
    security_review_worker: fireworks/kimi-k3-fast:max
    minimal_worker: fireworks/gpt-oss-20b
    default: fireworks/kimi-k3-fast
  task: 
    agentModelOverrides: 
      task: "@main_worker"
      sonic: "@fast_worker"
      scout: "@fast_worker"
      designer: "@vision_worker"
      librarian: "@research_worker"
      reviewer: "@review_worker"
      security-reviewer: "@security_review_worker"
Then you first say /plan and use some big model like K3. Finally the harness shows you a markdown you approve, and in approval you switch to a smaller model and reset the context. The smaller model gets the full plan and starts working on it. When done, you say /review and it spawns N review agents and returns the change suggestions. And you iterate on that.
impulser_Aug 13, 2026
After being stuck with using GPT-5.6 models for the past few weeks, I have renewed faith in Google and everyone but OpenAI. The GPT-5.6 models are quite obviously benchmarkmaxxed to make they seem like they are intelligent but they are quite dumb outside anything that not a benchmarked task.

I also think Google is still the best at fitting the most overall intelligences into their models, but for some reason it seems like the model architecture is just bad.

ghoshbishakhAug 13, 2026
Has anyone noticed that antigravity has been working really well for the last few weeks. Now with this model it should be working much better. Hope the Google AI Pro Subscription can be used to do some real agentic coding now.
kridsdale1Aug 13, 2026
I have been using the first party version all year, and with this model in it I am really happy. It’s fast and works.
thereitgoes456Aug 13, 2026
I've been really satisfied with it since 3.6, it's been "good enough" for the tasks I'm using and has fast response and very high limits, much higher than Claude Code.

I continually don't understand how nobody points out Flash 3.6 being much faster than any other model, and seems 3.7 is even faster still. That by itself is a major selling point.

ghoshbishakhAug 14, 2026
I agree. Having a conversation with the codebase in context is a much better experience with Flash 3.7
eckrAug 13, 2026
Maybe this is just my experience, but have people had trouble with 3.6 Flash just... getting things it has seen in its context correct? I don't know if it's been insanely benchmaxxed or what, but it'll pull information from websites and immediately get it wrong the token after. Or for example (this is something that happened like yesterday) I asked it to compare the uses of A and B in a language I was learning, and the way I typed it was "Please compare how these two are compared differently: A VS B", and then... it proceeded to compare "VS" and "B". I'm not kidding.

Personally whenever I use Gemini I've just been using 3.1 Pro because I've had insane trouble with them getting things incorrect like this. Hopefully they'll fix it soon / they've fixed it with 3.7 Flash.

axusAug 13, 2026
It's on Google AI Studio, which I use for free when I'm not on computers I control.

It did fine on my usual benchmark about configuring old Sparc hardware, maybe output slightly faster than before. Even included something new to check in the firmware.

dudeinhawaiiAug 13, 2026
I want to like Gemini models but my problem thus far has been a lack of coding chops. They still make mistakes, importantly, without correcting them for things like hallucinated API calls or code that doesn't run but they never bothered building or running. I know a lot of this can be fixed with workflows but it still feels like a failing.

GPT-5.6 or Claude models haven't delivered to me non-running code in ages.

Whenever I have Gemini in the flow, it's fast, but mistake riddled. I have low confidence in the output.

I've had some success with Opus driving Gemini models. It's pointless for GPT family since Sol is cheap enough or can drive terra/luna for arguably better performance, same speed, and better outcome.

As for all of the talk in this thread about modalities. Every SOTA model takes screenshots and verifies work now. Grok-4.6 does this, Luna does it, etc. They can also all work _from_ a screen shot or mockup provided.

I don't think it's a major selling point when every model can do it well and reasonably fast.

That said, eagerly awaiting "pro" and improvements to antigravity.

boinkboink78912Aug 13, 2026
Try this one, it's a step jump in coding capabilities for me over 3.6.
WarmWashAug 13, 2026
It depends on if you are relying on all single shot tasks or are willing to iterate. Flash is quick and can make dumb mistakes, but it also can fix them quickly.

I've gotten good results with it, but it definitely is more hands on.

simonwAug 13, 2026
The "introductory pricing" for this 3.7 Flash model is really weird.

It's scheduled to double in price on December 31, 2026, but who would anticipate still using this model five months from now? Especially since 3.6 Flash came out just three weeks ago!

My first effort with default thinking level produced an ambitious pelican, let down by a flawed bicycle: https://tools.simonwillison.net/markdown-svg-renderer#url=ht...

Then I ran it on high, medium and low thinking levels (oddly minimal is no longer an option, which WAS an option for 3.5 and 3.6) and got a pretty excellent pelican for the first two:

https://tools.simonwillison.net/markdown-svg-renderer.html#u...

UPDATE: That was in Safari, but as pointed out in the replies here the pelicans do NOT render well in Firefox or Chrome! Best guess is that's because of this invalid filter in the SVG:

  <filter id="shadow" x="-10%" y="-10%" width="130%" height="130%"></filter>
Filters are meant to contain additional elements, not be empty: https://drafts.csswg.org/filter-effects/#FilterElement - so maybe Chrome and Firefox remove the element that references the broken filter but Safari doesn't?
spiderfarmerAug 13, 2026
It's not weird if you're in marketing.
jakswaAug 13, 2026
This pelican gave me a good laugh, because there's enough reasoning that the render is out of sight initially. The buildup!
abtinfAug 13, 2026
> got a pretty excellent pelican for the first two

This suggests you primarily use Safari.

While the bike renders, the pelican doesn’t in Chrome and Firefox.

Probably one of the more serious defects I’ve seen with the pelican. It’s one thing when animated SVGs have bugs, but another when plain ones do.

jakswaAug 13, 2026
my whole world is shifting. have I been seeing _different pelicans_ from everyone else?!
jpadkinsAug 14, 2026
welcome to the internet, not all browsers are the same
simonwAug 13, 2026
... whoa, that's true! Thanks for catching that.
ceroxylonAug 13, 2026
Thank you for pointing this out, I was looking at the first like 'Umm Simon... there isn't even a pelican?"
GodelNumberingAug 13, 2026
I think introductory here means more or less permanent but they can't publicly admit there are no takers at a higher price.

Anthropic for instance announced a couple of days ago that they are making Sonnet's 'introductory pricing' permanent https://xcancel.com/claudeai/status/2086891169217122586

MikhailTalAug 13, 2026
When a company gives away service a heavily subsidized service as a promo, the full cost of serving it (compute) can get classified as sales and marketing instead of just cost of revenue, which makes your gross margin look better!
zurferAug 13, 2026
maybe exactly because the frontier moves, introductory pricing makes sense as you want to free up compute for the newer models
thebigspacefuckAug 14, 2026
They wouldn’t be on the frontier at all if it weren’t for this pricing.
wongarsuAug 13, 2026
At $work we still have some places using Opus 4.5. Even places using Qwen 2.5-VL, which is now 18 months old. It works, and upgrading is work (we'd have to validate the new model performs comparable in all the corner cases that currently work just fine)

Those kind of workloads would be hit by an end of introductory pricing. And it's exactly the kind of cases that are not very price sensitive. Where we are price sensitive we track new model releases closely, where we aren't other issues get priority as long as llm performance is good enough

dannywAug 13, 2026
Especially for minor workloads that’s like using a dollar of tokens a day. Not worth it unless you’re unhappy with performance or you have nothing else to do.
nharadaAug 13, 2026
I assume so when they release Flash 3.8 or 4.0 or whatever they can claim it's not a price increase?
asdfman123Aug 13, 2026
Maybe it's a way to kick people off the old models. "If you still want to keep using the old model you can, but you'd better pay a premium."

I work for Google and their free, internal Gemini API isn't quite as graceful. They once turned down a model arbitrarily and it broke our tests. I had to scramble to fix it, then build warning systems for the turndown as well as a special validator to make sure any upgrades make the same determinations.

swalshAug 13, 2026
Yeah, so I'm going to confess. Ive built production AI systems that have very specific jobs, deployed them, and moved on. The cost is not noticeable, i had the system dialed in. Not worth the effort to reevaluate a newer model to see if its better in some way. Current model works, and I have other projects that are more important.
MinimalActionAug 13, 2026
For what it's worth, I use Brave and it didn't render properly in a normal tab, but rendered correctly on a private tab! Underlying engine is Chromium anyway, so make of that what you will.
drusepthAug 13, 2026
That sounds like something absolutely wild happening in Brave.

Or maybe one of your installed extensions modifying the page (presumably private tabs disable those)?

formvoltronAug 13, 2026
We can be sure that Google is not Pelicanmaxxing their models.
bessbdAug 14, 2026
Unless it's just countersignaling. Like prompting for occasional typos so text doesn't look AI-genrated
mchusmaAug 14, 2026
I think in this case its a sign they are subsidizing usage, so it means that the only people left using it in Jan 2027 will be some enterprisy or orphaned workflows which should pay up or migrate off.
simonwAug 14, 2026
Here's a rendered copy of the pelican I found impressive, for people who don't have Safari: https://static.simonwillison.net/static/2026/gemini-3.7-flas...
IFC_LLCAug 13, 2026
Like, I understand everything, but by this time I don't give anything about any of those announcements.

Theoretically there is some difference between Fable and Opus or Grok and GPT, but at the end of the day I'd look at the bottom left of my screen and to my amusement find out that for the past 3-4 hours I've been using model ______.

If the results are semi-decent, I'd keep it on, if not - I'd randomly switch the model and try again.

Actual thing that would affect my selection would be a number of unused tokens I have left for a model ____ for this week.

Maybe it's cause I'm using those for programming and log parsing and all of them are decent enough, but other than that - there are no leaps I see.

ls_statsAug 13, 2026
I don't get it, Google could heavily subsidy their Gemini models to make it more attractive, but they prefer to not do it. I don't know one soul who is using Gemini models to code. Even OpenAI who doesn't have money or capacity is offering their Luna model at $1.2 per 1M/out.
film42Aug 13, 2026
What you're not seeing are the subsidized Google Cloud startup credits, which includes Gemini. If you're in that program, you choose Gemini because it's essentially "free" and consistent.
JoeriAug 13, 2026
Why would they? Unless they have lots of unused tpu real estate that they could host it on “for free” they would be bumping more profitable workloads off of machines to give away that capacity to people with zero long term loyalty. There is no business reason for google to subsidize these models.

OpenAI has too much money. They’re spending their money in stupid ways.

AuthAuthAug 14, 2026
Having people use the model generates real world training data which could be useful. Other than that I think you're right that there is little reason to artificially boost users with subsidized pricing.
sumedhAug 14, 2026
> OpenAI has too much money

I think they got some data center deals very cheap when no on was thinking about data centers. Dario didnt want to take that risk so Anthropic didnt make the deals earlier but now paying Google and Elon higher rates.

u1hcw9nxAug 13, 2026
It makes sense. All these models are money losing businesses.

As a business Google might want to focus on fundamental research 2-3 years from now and not compete on who acquires more money losing customers. Just stay little behind and invest money better.

WarmWashAug 13, 2026
Google's open secret is that they are also capturing spend on OAI and Anthropic tokens.
knupparAug 13, 2026
the race for the smartest/cheapest model is a race to the bottom. selling the tools is much more profitable.
surgical_fireAug 13, 2026
Google was cash-flow negative in Q2 2026, and is raising a lot of debt.

I am not sure they can afford to subsidize Gemini more than they already do.

surajrmalAug 14, 2026
Notably that's not due to subsidization but rather due to investment in capex for future growth. The reason to not subsidize is largely because of already being hardware constrained. There is no extra unsold capacity to use.
surgical_fireAug 14, 2026
> investment in capex for future growth

Future growth needs to materialize for this to be profitable. Right now what they have is a cash-flow negative reality.

> There is no extra unsold capacity to use.

I have no evidence of this.

Neither I have evidence that they would be in the position of subsidizing further AI usage.

andaiAug 13, 2026
So their "Flash" model won't be cheap. Are they gonna make a new one that's cheaper? Gemini-3.8-Silverlight? ;)
dwa3592Aug 13, 2026
I was going to cancel my gemini membership today ..... still going ahead. In my experience, gemini 3.1 pro, 3.5, 3.6 flash constantly lie too much about completing their tasks whereas sol (even though equally dumb) never claims something has been done when it hasn't been.
seunosewaAug 13, 2026
Two thoughts:

1) It makes sense to try 3.7 flash before cancelling.

2) Prompting models to be honest is surprisingly effective in my recent experience. But only if they listen to instructions.

dudeinhawaiiAug 14, 2026
That's a good point and something I encountered yesterday. On a multi-agent task, Gemini was the only model that got near the end, ran tests, saw it had issues, took a screenshot, saw the issues, and then said, "I'll mark it complete" and delivered.

I was paying for Ultra, then downgraded to Pro and at this point I'm near abandoning it. Agy as a harness also has a tendency to constantly request significant elevations to perform routine operations.

andrewstuartAug 13, 2026
Gemini has lost the race to be relevant for AI coding.
throwaw12Aug 13, 2026
Is this the reason why Jeff Dean, Sanjay Ghemawat and other DeepMing, Gemini people got kicked out of Google?

If so, now I understand why they didn't want to release this model

sid_talksAug 13, 2026
The Gemini Flash models makes perfect sense to me coming from a company like Google. Google AI Mode for search is a product I really find useful. It makes sense that Google focused on smaller, faster yet smart enough models that wouldn't break your bank on inference. It plays well into their product ecosystem.

Google AI Mode consistently gets me consistently good results and good speeds. It really changes what "googling" is for me.

alex1138Aug 13, 2026
I like Gemini (I'm just a dumb person without knowledge of 'benchmarks' or how x compares to y) while understanding that shoving it into Google auto summaries has been a bad idea and produces inaccurate results
gazebo2Aug 13, 2026
I similarly have a weird affinity for Gemini that I can't really articulate. I used Gemini's free chat and found it great for exploring technical topics (and random one-off general walking-around-questions) and appreciated its speed, tone and accuracy. I spent a month playing with Gemini CLI / Antigravity and found it also an effective coding agent, at least for my workflow (entirely in the loop development and review). I also was really surprised that I could just paste it images of a project I was working on and have it immediately understand what it was looking at -- which I've come to learn is considered a unique strong point for Gemini. I've been playing with GPT5.6 for about a month and it's definitely powerful but I honestly think I'll go back to Gemini. There's something kind of charming about working with an AI that not only is particularly good at web search and information gathering, but also one that doesn't feel like some superhuman overengineering freak when it comes to code.
dannywAug 14, 2026
I like to think of Gemini as a broader generalist that hasn't allocated _most_ of its skillpoints to agentic coding execution :)
abixbAug 14, 2026
Well, Google is probably the only one among frontier model providers in the US that doesn't have a massive financial pressure to deliver business results ASAP (and this focus on agent coding and long-running agentic tasks), so Google is able to focus more on encoding deep scientific, cultural and historical knowledge to its systems. I'd assume DeepMind's focus on the scientific core also played a role in the tone and approach Gemini models take for explanations and Q&As.

I find that GPT models and Claude tend to talk in strong slangs and in-group jargon, but love Gemini's massive general knowledge corpus — reminds me of Richard Feynman from his lectures.

fryanyway_sweAug 13, 2026
I just use web chat as "harness"(lol) or interface and I have mostly switched to Gemini as the free limit basically never run out for me unlike ChatGPT and Claude.

Also impressed with Grok for some stuff.

sarjannAug 13, 2026
Introductory price seems a bit weird as it expires at the end of the year and by then it’s going to be significantly outdated.
nipunaeka89Aug 13, 2026
Always loved the Gemini. Helped my work alot
theplumberAug 13, 2026
So they keep pushing these Flash models because they don’t really have a powerful model…or better said their ‘pro’ model is actually a flash
kyruzicAug 13, 2026
So at this point new models seem to only care about one task, software development. This really was not the original pitch of ai and I do not see how it justifies the insane spend or valuations it has produced.
Revanche1367Aug 13, 2026
Think of the potential layoffs of highly paid employees!

But, I think it’s also based on what they are being used for, most LLM users are still mainly SWEs or similar as I understand and there’s a ton of data to train them for coding.

kyruzicAug 13, 2026
I agree its what they are being used for and their primary revenue source.

My point is mainly that was never the pitch that got ai the hype it did and imo doesn't justify the valuations even if we all lose our jobs to ai. Because it no longer seems like they even think its making other jobs go away.

pjmlpAug 14, 2026
They do it via other ways, translation teams and asset creators for CMSes, I no longer see them in our projects, the builtin AI tooling takes care of it.

Same with those that use to improve marketing outcomes for SEO and such, now there are AI based reports with automatic improvements.

Finally on other domains you already have robots on supermarkets, fast food, and gas stations, where the customer does the work of the (now gone) employees without any kind of price reduction.

ur-whaleAug 13, 2026
Why does Google keep announcing these subpar models?

What am I missing?

They can't seem to be able to produce a frontier model, fine.

Just be quiet about it and work hard until you manage to put one together.

[EDIT]: Come to think of it. Maybe they're trying to build the Toyota corolla of AI ... let's see if that wins them the battle long term. I personally doubt it.

ksajadiAug 13, 2026
i just wish google cloud ux was remotely as good as their models. they made some progress with their studio, but then in a true google fashion, product names keep changing (Gemini, Anti-gravity, Vertex, Google AI,...) as well as confusion and complexity for something as simple as registering agy cli with a Google cloud project.

today i wanted to link agy to a google cloud project, for that i had to enable 5 different APIs in google cloud UI, then create a subscription for Gemini Enterprise (whatever that is), then link it to a project, then assign it to a user. and after all that, agy couldn't find the subscription.

the best part: i couldn't cancel the subscription. so i just paid $35 for one month and left it.

pestsAug 14, 2026
Gemini is the model family.

Antigravity is their agentic coding app / IDE. There is two products, one is chat-only the other is more standard IDE.

Google AI Studio is a consumer/developer playground with a in-browser IDE meant for prototypes or demos, there is a gallery of demos etc. Easily shared, easy key access.

Vertex is the AI offering from the GCP side of the company, that is going to target more enterprise or business solutions (scaling, data governance, security, production deployents, etc)

sunaookamiAug 13, 2026
Not in the website/app? Only on AI Studio? Weird.
JakeScAug 13, 2026
On a related note, I see all these quantitative benchmarks and the models getting really good at them over time. One thing I've been wondering: if the GPT series of models performs so well quantitatively, why do I still kind of hate using them relative to Claude? There’s a missing “vibes” or “taste” benchmark I think.
hateboxaaAug 13, 2026
statuory
hateboxaaAug 13, 2026
I think this is a slop release
FrannkyAug 13, 2026
Have you tried the new DeepSeek Pro v4, Qwen 3.8, Gemini 3.7 Flash, and Grok 4.6? Do they make sense for any use cases?

I'm currently using omp with Kimi K3 as the planner and DeepSeek v4 Flash 0731 as the implementer, or CC + Fable for planning and Opus 4.8 for implementation. For API(not coding), I just use DeepSeek v4 flash 0731 and MiMo.

I'm pretty happy where I am, but I'm wondering if these new models provide some new kind of advantage

mirekrusinAug 13, 2026
Just use all of them and synthesize final result, record scoring when synthesizing.

Later you can make decision to drop low performing ones.

FrannkyAug 14, 2026
It's about time and tokens. I find it more effective to get the vibe from friends, HN, Discord, Reddit, and then play with only the most promising one. I skipped all of these because they seemed not worth it
barrenkoAug 14, 2026
Synthetize=?
mirekrusinAug 14, 2026
ie. open router fusion [0]

[0] https://openrouter.ai/openrouter/fusion

Alien1BeingAug 13, 2026
Google has got the IBM disease.

Large lumbering enterprise with massive inertia. Where innovators leave as soon as they get a better offer.

None of the authors of the seminal "Attention Is All You Need" paper are still at Google.

Fast forward a decade and Google will be reduced to hiring the kind of mediocrities who deign to work at IBM and Accenture.

luckjack47Aug 13, 2026
I don’t think the skill set required to write the Vaswani paper is the same as training and shipping frontier models like Gemini so I am not sure why people keep bringing this up.
Alien1BeingAug 13, 2026
Training, tweaking and shipping frontier models like Gemini is something that a team of mediocre engineers can do.

Google continuing to ship generations of Gemini using its mediocre teams is an existence proof of that

For genuine advances one needs the Hintons and Vasvanis, not yet another bunch of mediocre Kookaid drinking product managers at Apple and Google

ezekiel68Aug 13, 2026
And I'm over here on SiliconFlow using Stepfun AI Step-3.5-Flash at $0.10/M input and $0.30/M output tokens (262K context window) for complex market analysis work in rust utilizing vectorized instruction sets. It provides me with amazing results.

I honestly wonder how long this calliope can keep playing before it crashes to the ground.

(I have no business relationship to anything mentioned here except as a regular retail customer who went bargain-hunting)

nojitoAug 13, 2026
The flash models are just so good at OCR tasks and summarization. Glad to see them constantly improving them.
customguyAug 13, 2026
Was excited to try it, since I've been use 3.6 Flash in the last few days to make simply experiments/prototypes. My loop is writing a prompt, maybe adding a screenshot of the closest to what I want I have so far, then based on the result I modify/extend the prompt, maybe use another screenshot.

Well, I think I'll stick to 3.6 for now. Based in a very scientific sample size of exactly one attempt each: https://imgur.com/a/fDOkBDm

Both got the same prompt and example screenshot. I mean, they both suck but that's normal this early in, but the 3.6 version (first screenshot) actually changes the displayed threads depending on selected categories, and the messages of whatever selected thread, as obviously described in the prompt. You might say it more or less does what it should. Both versions have an ugly flash/jerk in the category pane when selecting/deselecting a category, so that's a wash.

The 3.7 version doesn't work at all, i.e. it always shows all threads, and no messages for any of them. I can post new messages in threads but they don't show up, and it doesn't even increase the message counter for the thread. I guess it's a matter of taste but I don't like the look either, while 3.6 actually is in the spirit of the screenshot I added the prompt, using 11k of CSS versus 3.7's 16k. The code is also less, and the backend split into 3 files (instead of just 2 as 3.7 did it), so assuming it sucks in either case, it'll be easier to read and massage.

edit: geez, 3.6 even properly fades/disables the "new message" button when no thread is selected, 3.7 didn't bother which is smart since everything else is broken anyway. Maybe it's better at really complex things, but for simple things, what I'm experimenting with, I already saw enough.

qudatAug 14, 2026
I just subbed to Gemini a week ago and have been using antigravity and 3.6 flash. The speed is absolutely a differentiator compared to Claude.
onlyrealcuzzoAug 14, 2026
Claude is terrible when it comes to speed.

Flash is great, but Codex models are also fast, as is DeepSeek v4 Flash.

Anyone who's Anthropic-pilled should really get out and explore and see how unbelievably terrible they are when it comes to speed and cost vs quality.

Anthropic has good models, they're just way too expensive and slow for what you pay for.

v3ss0nAug 14, 2026
GEMMA 4.5 gogoogogo
correlatorAug 14, 2026
We run a platform that serves models from many of the frontier players to power conversations and workflows. Gemini flash-2.5 was a game changer for us when it came out. Cheap, fast, and reliable.

We are now considering dropping support for the model family all together. All of their models require significant scrubbing of errant thinking blocks, inner monologues, and it's consuming more engineering resources than it's worth.

linzhangrunAug 14, 2026
Feels like Gemini Pro will arrive directly as Gemini 4 Pro
llm_nerdAug 14, 2026
If you're in the Gemini app and want to try it out as a Pro or Ultra user...well you can't. It's only in the weird "Spark" agent that demands you entire Google identity.
nicolamanziniAug 14, 2026
It is also doing pretty well in threejseval. Frontier there for the price. Much better than 3.6.

https://threejseval.com/ranking

mchusmaAug 14, 2026
This is a solid release (at the intro pricing, the other pricing is dumb). I do think its a missed opportunity to really blow things out of the water and have this be another 1/2 off, but clearly they don't have the inference efficiency for it. Speed is good, the knowledge in google's models is solid for those usecases, price is reasonable (after 3.5/3.6 major missteps).

The intelligence index vs cost pareto frontier is crazy now, its basically a flat line with 9 people all at or right at the edge of the frontier along various parts of the graphs. Insanely competitive right now.

aurareturnAug 14, 2026
Probably golden age for competition before consolidation.
mchusmaAug 14, 2026
Maybe? But GLM just put itself on the pareto today again, so its up to 10. The theory was that there would be recursive self improvement in models, leading to a 1 or a handful of entities running away from the rest. But basically the opposite happened? Different teams, different hardware stacks. What do we make of this? (1) There must not really be any deep secrets/moats right now? (2) improving models is not something you can throw only intelligence at now?
goochgibblerAug 14, 2026
I started vibe coding with Gemini. At first very exciting, we got a rough program up quickly. And then, code corruption. Over and over. I started to document every single file, every single step, writing explicit rules to not fake data and create for real world use, but everyday I kept catching Gemini errors, which turned into flat out lies. Generating fake test data instead of pulling real data. Making up test answers. Saying features were implemented that weren't. It even coded fake python files that printed made up results. I think it realized I wasn't doing code reviews. Every day was spent chasing defects and rolling back. After Gemini admitted to faking 7 tests i let Claude review the code, and it fixed it almost immediately asking why half the features were broken or missing. Well Claude, because Gemini flash did the least amount of work to make me happy. Impressive. Very human. Very frustrating.
prtmnthAug 14, 2026
I am curious to understand who is this model targeted at?
GaggiXAug 14, 2026
It's a cheap and powerful model. Also very fast for this class of models.
exacubeAug 14, 2026
I'm finding Gemini 3.7 speaks a lot more academically. Its explanations are not clear and intuitive by default.

Maybe this has to do with all the RLVR it went through, where reasoning through difficult academic/coding problems caused it to think and speak a certain way.

The benchmarks looks great but it doesn't feel as legible, so maybe it's more meant to be an agentic model rather than an everyday model whose outputs are read by humans?

mintflowAug 14, 2026
Given my codex have limited usage and I get a idea to build a on device clipboard translator with menubar when read a article find some words not seens before

So i launched agy and find seems it have a 3.7 flash i thought its latest,until i see this article i know it's just released

after serveral rounds prompt(the initial prd prompt is by chatgpt), it use 2.9k user message token and 42.5k reponse tokens with gemini 3.7 flash(low) after i check the status, i got a working on device translator, its cool

the model seems also p and retty fast and the generated app looks good and easy to use

but agy cli is a bit unintuitve and i also check the guy that release the latest agy, seems the guy does not commit much ? or perhaps agy does not open source and only use github as a issue feedback channel

raincoleAug 14, 2026
It's quite good. It seems to be "Google's turn" again. Might last two or three months?
ddp26Aug 14, 2026
What are we to infer from no release of gemini-3.5-pro, but frequent releases of smaller flash models (presumably from the same large pre-training run?)
restersAug 14, 2026
Google is in direct meetings with Scott Bessent and Howard Lutnick and is prudently keeping dangerous frontier models out of the hands of customers until reasonable precautions can be taken.
akulbeAug 14, 2026
With the DeepMind guy leaving the building... isn't Gemini going to get old and crusty, and fall into disrepair?

I'm mostly serious here. Aside from Gmail and search, it feels Google doesn't have a good track record for maintaining things.

It really makes me wonder... if key people are leaving, what's going to happen? The Google graveyard is pretty big.

m3kw9Aug 14, 2026
They need to learn to release on antigravity or cli sub plans on day/two one like OpenAI
ipsum2Aug 14, 2026
Recently my android phone updated from Google assistant to Gemini Flash. Completely unusable. Asking it to play music and it refuses, hallucinating instructions to connect Spotify to Gemini. The instructions say to tap buttons that don't exist.

Bonus feature from Gemini: a toggle to opt back into Google Assistant, but it doesn't work. Still stuck with Gemini.

booiAug 14, 2026
Gemini hallucinations will continue until morale improves
UlisesAC4Aug 14, 2026
Yeah I don't know who was the genius who approved changing an Alexa lookalike to a chatgpt clone for our phones.
ipsum2Aug 14, 2026
It would be fine if it worked. I wish I could replace Gemini on my phone with chatGPT.
seabrookmxAug 14, 2026
It's weird because Gemini via Antigravity or Google Search "AI Mode" works great. It's just the Google Assistant replacement application that's terrible. I've completely given up trying to ask it anything via Android auto. I punch my directions manually in Google maps before I leave now because the voice assistant is unusable.
dpacmittalAug 14, 2026
Funny you mention this. I ask gemini what time it is, and it replies in a different language. I've checked all language settings, and they appear to be correct. Yet gemini just refuses to reply in correct language, and it only happens when I'm asking it for time.
rw2Aug 14, 2026
It's just useless to build any model in this range because deepseek flash is so fast and so cheap. I see no reason for anyone to not use it for a few points in performance.

Only models that matter are the edge fable class models people use for code, and Google struggled with that.

jnwatsonAug 14, 2026
Rates are going up Aug 16.
instagrahamAug 14, 2026
Funnily enough, I just ran a task on AI Studio with 3.6 yesterday and got 3.7 to do a similar one today; so it serves as an interesting and quick comparisons between the old and the new (usually, if enough time passes between your use of one model and the next, you'll have a sourer view of it than its actual competence suggests).

It hallucinated in both cases despite being given an API key and building a lot of pipes to access data using this. It was a simple "oh shit" fix moment for the model, but weird how eager it was to hallucinate despite the process being designed for it to be data-driven.

We should move past the idea that benchmarks alone tell us whether a model is getting better. I would've had the same experience a year or two ago with 1.5, and the solution would've been similar (keep prompting). I've been investing time into making system prompts and input prompts more meticulous, but the fundamental "it will make shit up" problem still remains, even though it shouldn't when the job involves calling tools.

I know this sounds like I'm expecting superpowers of it (I'm not), but my point is just that these incremental benchmark gains may not reflect user experience.

customguyAug 14, 2026
For me, it just shits the bed: https://news.ycombinator.com/item?id=49292924

Reading the google blog and these discussions makes me feel like I'm taking crazy pills, seriously. Side-by-side comparisons with the exact same inputs or it didn't happen, that's my rule going forward. Test all the things, believe nothing.

alastairrAug 14, 2026
I couldn't get past the first chart which basically showed that intelligence and cost are both better for gpt luna, I'm not sure what the argument is here for Gemini flash 3.7 given that comparison.
VeejayRampayAug 14, 2026
not what the first chart shows
alastairrAug 14, 2026
You are right, it's the deepswe chart vs cost I was referring to which is slightly further down
barrenkoAug 14, 2026
I am still using gemini flash 2.5 for an old app I have running that does some OCR as well if necessary, is it time to switch up?
manapauseAug 14, 2026
I also use it for a chatbot assistant for customer onboarding. Gemini deserves some kudos for their willingness to allow entry level subscriptions access to API-keys to build solutions with.
rawoke083600Aug 14, 2026
I wish they will always put it front and centre what are all the models name as how i call it in via their API.
tibzejokerAug 14, 2026
ok but still a small model.. why is google so behind regarding SOTA :/ daim
ChildOfChaosAug 14, 2026
Still fairly poor limits in anti gravity despite the price cut, which seems to be API only.
acrushAug 14, 2026
When will the pro be out...
ImanariAug 14, 2026
Testing it in Pi.dev and liking the speed a lot! Huge improvements from past gemini models regarding tool calls and agentic capabilities
or1gaminalAug 14, 2026
experimenting w/ flash 3.7 on morphology and it looks promising. google models looks capable on linguistic side and prose. are there any reliable benchmarks on this?
hnicrcjk6oAug 14, 2026
This hits different after a long week
throwaway_20357Aug 14, 2026
Hard to believe nowadays but Gemini Flash 3 was $0.25 / 1M input and $1.50 / 1M output when it was released.
chaostheoryAug 14, 2026
All of Google’s models prioritize speed over quality and correctness. That may be great for search but not so good for hard work.
michaelteterAug 14, 2026
I have had such good, consistent success with Gemini 3.6 Flash (usually on medium setting) that I forget about it. It is just this reliable, effective assistant that I use for managing and organizing information, brainstorming, and of course coding.

I honestly don't know what I would want more/better than what I was already getting... so I'm curious if I will see any improvements with 3.7.