← Back to NewsNEWSAI Coding

Gemini 3.7 Flash Goes Half-Price as Alibaba Open-Sources a 2.4T Model

August 25, 2026

Gemini 3.7 Flash's half-price launch and Alibaba's open 2.4T Qwen3.8 look like a value-tier win, but the discount expires January 1 and the open model needs 72 GPUs. Read the catches.

S
Shubham Sharma
Aug 29, 2026
❤️ 0 likes💬 0 comments
ai-modelsAI codingaillm
Gemini 3.7 Flash Goes Half-Price as Alibaba Open-Sources a 2.4T Model

Two moves in mid-August reshaped the cheap end of the AI market, and both come with a catch worth reading before you re-plan your stack. Google shipped Gemini 3.7 Flash at half price on August 13, and Alibaba open-sourced Qwen3.8, a 2.4-trillion-parameter model.

Gemini 3.7 Flash: half price, with an expiry date

Google priced 3.7 Flash at $0.75 per million input tokens and $3.75 per million output, and called it half price. Read the asterisk: those rates run only through December 31, 2026, then double to $1.50 and $7.50 on January 1, 2027. The "half" is measured against the workhorse tier's standard rate, which is exactly where it lands again next year. It is a time-boxed discount, not a permanent cut, and standardizing on it now risks meeting doubled rates in January.

The capability gain looks real, though. The model scored 65.3% on the DeepSWE v1.1 coding benchmark, up from 49.0% for the prior Flash, and it carries a 1M-token context. Google did not release open weights.

Qwen3.8 is open, but open is not local

Alibaba went the other way and published the Qwen3.8 checkpoint on August 12: 2.4 trillion parameters, free to download. Before you plan to self-host it, note the hardware. Nvidia's day-zero serving test ran it on a rack of 72 Blackwell GPUs. Open weights at this scale are a licensing headline, not a home-lab option, and the open checkpoint is text-only. The managed sibling, Qwen3.8-Max (2.4T mixture-of-experts, roughly 95B active, 1M context, with vision), reached general availability on August 3 at $2 and $6 per million tokens, undercutting Gemini Flash's standard $7.50 output rate.

What it means for builders

The value tier is now a genuine fight, and Google's cut was itself a response to cheap Chinese models. OpenAI recently cut GPT-5.6 Luna by up to 80%, and models like Moonshot's Kimi K3, GLM 5.2, and Qwen3.8-Max are delivering near-frontier quality at a fraction of the cost. That competitive pressure, not generosity, is what put Flash on sale.

Two takeaways. The cheaper winner depends entirely on your input-to-output token mix, so test both on your actual workload instead of the headline rate. That is the same measure-first discipline behind the AI productivity proof gap, and it applies to model choice too. And whether you rent a closed model by the token or download open weights and rent the GPUs, at this scale you are still renting the rack.

Join the discussion on Gemini 3.7 Flash Goes Half-Price as Alibaba Open-Sources a 2.4T Model

Likes, comments, and replies are available for authenticated readers with verified email addresses.

Comments (0)

Loading discussion...

More news