Gemini 3.7 Flash Goes Half-Price as Alibaba Open-Sources a 2.4T Model
August 25, 2026
Gemini 3.7 Flash's half-price launch and Alibaba's open 2.4T Qwen3.8 look like a value-tier win, but the discount expires January 1 and the open model needs 72 GPUs. Read the catches.

Two moves in mid-August reshaped the cheap end of the AI market, and both come with a catch worth reading before you re-plan your stack. Google shipped Gemini 3.7 Flash at half price on August 13, and Alibaba open-sourced Qwen3.8, a 2.4-trillion-parameter model.
Gemini 3.7 Flash: half price, with an expiry date
Google priced 3.7 Flash at $0.75 per million input tokens and $3.75 per million output, and called it half price. Read the asterisk: those rates run only through December 31, 2026, then double to $1.50 and $7.50 on January 1, 2027. The "half" is measured against the workhorse tier's standard rate, which is exactly where it lands again next year. It is a time-boxed discount, not a permanent cut, and standardizing on it now risks meeting doubled rates in January.
The capability gain looks real, though. The model scored 65.3% on the DeepSWE v1.1 coding benchmark, up from 49.0% for the prior Flash, and it carries a 1M-token context. Google did not release open weights.
Qwen3.8 is open, but open is not local
Alibaba went the other way and published the Qwen3.8 checkpoint on August 12: 2.4 trillion parameters, free to download. Before you plan to self-host it, note the hardware. Nvidia's day-zero serving test ran it on a rack of 72 Blackwell GPUs. Open weights at this scale are a licensing headline, not a home-lab option, and the open checkpoint is text-only. The managed sibling, Qwen3.8-Max (2.4T mixture-of-experts, roughly 95B active, 1M context, with vision), reached general availability on August 3 at $2 and $6 per million tokens, undercutting Gemini Flash's standard $7.50 output rate.
What it means for builders
The value tier is now a genuine fight, and Google's cut was itself a response to cheap Chinese models. OpenAI recently cut GPT-5.6 Luna by up to 80%, and models like Moonshot's Kimi K3, GLM 5.2, and Qwen3.8-Max are delivering near-frontier quality at a fraction of the cost. That competitive pressure, not generosity, is what put Flash on sale.
Two takeaways. The cheaper winner depends entirely on your input-to-output token mix, so test both on your actual workload instead of the headline rate. That is the same measure-first discipline behind the AI productivity proof gap, and it applies to model choice too. And whether you rent a closed model by the token or download open weights and rent the GPUs, at this scale you are still renting the rack.
Join the discussion on Gemini 3.7 Flash Goes Half-Price as Alibaba Open-Sources a 2.4T Model
Likes, comments, and replies are available for authenticated readers with verified email addresses.
Comments (0)
Loading discussion...More news

Two Critical Next.js RCE Bugs Are Patched in 16.3.3 and 15.5.24
Vercel shipped an emergency Next.js release fixing two critical unauthenticated RCE bugs, a Windows path traversal and an AVIF image flaw. Self-hosters, upgrade to 16.3.3 or 15.5.24 now.

Everyone Feels Faster With AI. Almost Nobody Can Prove It.
84% of developers feel more productive with AI. Only 20% of their organizations measure whether it's true. That gap between feeling and evidence is becoming a business problem.

AI Agents Just Had a Rough Patch Tuesday
Two critical Microsoft agent CVEs and a Spring AI prompt-injection bug landed in the same window, all pointing to the same lesson: what you tell an agent it can't do isn't what stops it.