News: 0184962057

  ARM Give a man a fire and he's warm for a day, but set fire to him and he's warm for the rest of his life (Terry Pratchett, Jingo)

Google's Gemini 3.7 Flash Targets Coding and Agents With a 50% Price Cut

(Thursday August 13, 2026 @11:30PM (BeauHD) from the new-and-improved dept.)


Google has [1]released Gemini 3.7 Flash just three weeks [2]after 3.6 Flash , focusing on better coding, agentic workflows, and enterprise automation while [3]temporarily cutting API prices in half through the end of 2026 . VentureBeat reports:

> For enterprise developers, the more consequential story may be the combination of those intelligence gains with lower inference costs: through the end of 2026, Gemini 3.7 Flash costs $0.75 per million input tokens and $3.75 per million output tokens.

>

> Starting Jan. 1, 2027, pricing rises to $1.50 per million input tokens and $7.50 per million output tokens. That means the current discount is temporary, but it gives teams deploying high-volume coding and business agents several months to evaluate whether Google's claimed reductions in retries and manual oversight translate into lower total operating costs. The launch also underscores Google's rapid iteration on its Flash line while its next flagship Pro model remains absent. [...]

>

> Google describes Gemini 3.7 Flash as its "most intelligent workhorse model yet for coding and agents." The company says the model is better at adapting when it encounters roadblocks, clarifying intent when necessary and following instructions with greater fidelity. [...] Google's benchmarks show a large generational improvement in several software engineering tests. [...] Google's own results do not show 3.7 Flash universally displacing higher-priced competitors. They instead suggest a model that has become substantially more competitive in coding and agent workloads while occupying a lower price tier.



[1] https://blog.google/innovation-and-ai/models-and-research/gemini-models/introducing-gemini-3-7-flash/

[2] https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-3-6-flash-3-5-flash-lite-3-5-flash-cyber/

[3] https://venturebeat.com/technology/googles-gemini-3-7-flash-targets-coding-and-agents-with-a-50-introductory-price-cut



Begun, the AI price wars have... (Score:2)

by barc0001 ( 173002 )

Be interesting to see how many of these companies bite the dust if the bigger players with deeper pockets like Google keep access pricing down longer to drive adoption and build their market share. I know many companies are already freaking out about some of the jumps in pricing that still don't pay the 'real' cost of using AI and have already aggressively cut back so the industry as a whole is in for a rough time not properly factoring in exactly how miserly their potential customer base is, but this wrin

Re: (Score:3)

by barc0001 ( 173002 )

You'd be astonished at how many companies will be absolutely fine with you driving that Yugo as long as it's cheaper.

The company I work at has top down shifted which model is tied into all the devs IDEs 4 times in 6 weeks, always seeking cheaper usage even at the cost of suitability. And the company is a multinational with 25K employees worldwide, not some little shop... If my company's doing it, others are doing it too.

Re: Begun, the AI price wars have... (Score:2)

by Kycin ( 6650188 )

We use llmlite, so we have centralized api key access and each user gets an allowance of usage per month. We provide guidance on what model to use went and what it's good for. We also have a small cluster of 7900xtx for local inference when you go over your allowance but itvonky supports low volume.

Re: Begun, the AI price wars have... (Score:4, Insightful)

by barc0001 ( 173002 )

> We also have a small cluster of 7900xtx for local inference when you go over your allowance

That's actually one of the conversations I've been hearing at work is that 'free' LLMs we can self host at one of our datacenters is the long term plan once they're Good Enough(tm). I think a lot of companies are planning the same. Medium to big corporations with large tech departments remember vividly how Broadcom screwed them on VMWare pricing and they're not going to get caught in the same bear trap with AI. They're already planning to be mostly to fully self sufficient - which again is going to cause a lot of those "AI will be worth XXXXXXXX" predictions to be completely and utterly wrong and kill a lot of those companies.

Re: (Score:2)

by Rei ( 128717 )

Also, the big AI companies get a number of advantages, such as large batch sizes and being able to balance on-peak and off-peak workloads better.

Re: (Score:1)

by Anonyrnous ( 10465021 )

> In a few years that $30k compute card is worth $300. Are you or your company willing to completely replace a multimillion dollar infrastructure every 2 or 3 years?

At the moment all of the big AI providers building data centers are paying 10x the normal price for GPUs and RAM which is why companies like NVidia and Micron are making out like bandits. Keep in mind the AI providers have to recoup that spending at some point. Of course the rest of us are also paying 10x normal prices at the moment if we attempt to build our own servers.

NVidia and Micron (TSMC, Samsung etc) must have been massively ramping up production capacity to meet the demand.

Not only that but every

Re: (Score:2)

by Rei ( 128717 )

Is most of your daily driving "a race"?

This is a midrange model. This is competitive midrange pricing. For my midrange work I've been using GLM-5.2 - Gemini 3.7 flash is a bit cheaper and a bit better on the benchmarks, so yeah, I'll probably switch so long as the prices stay low.

Sometimes you need a high-end model (Claude Fable, GPT 5.6 Sol, Kimi K3, etc). Sometimes a low-end model is fine (GPT 5.6 Luna, DeepSeek 4 Flash 0731, etc). Usually a midrange model is the best balance.

Re: (Score:2)

by ceoyoyo ( 59147 )

You're need to get to work, get groceries, maybe go for an occasional Sunday drive. Do you buy a Porsche or a Honda?

Re: (Score:2)

by allo ( 1728082 )

Google cuts the price, because more powerful Chinese models were cheaper. When a Chinese company releases a good model for free, you instantly have a lot of companies (many American) that host it. As they don't have to pay for training they can operate it for hosting price plus some profit margin and undercut everyone else.

First hit is free (Score:2)

by liqu1d ( 4349325 )

Or at least heavily subsidised.

Gemini, what is google's electricity cost (Score:2)

by blue trane ( 110704 )

At hyperscale data center rates (approx. $0.06 รข" $0.08 per kWh in North America/global hubs), the raw power cost to generate 1 million tokens is only $0.008 to $0.032 (less than 3 cents).

Re: (Score:2)

by allo ( 1728082 )

The other question is, how efficient is the hosting? In case of Gemini it is highly optimized. They split for example different tasks in the pipeline over different servers so a single server can be optimized to do a single task highly parallel.

Measuring the environmental impact of delivering AI at Google Scale: [1]https://arxiv.org/abs/2508.157... [arxiv.org]

[1] https://arxiv.org/abs/2508.15734v1

"That means the current discount is temporary" (Score:2)

by memory_register ( 6248354 )

Let's get you hooked at prices you can afford, then hope you stay when it becomes unaffordable.

Remember, there's a big difference between kneeling down and bending over.
-- Frank Zappa