On the same day, both companies said the same thing: the models are fast enough now. The difference is, Google's is already in your hands, while OpenAI's is still locked behind a waitlist.
Google rolled out Gemini 3.7 Flash on Thursday, targeting coding and automation workflows, with a full public launch now live in over 160 countries. The model can handle up to 1 million input tokens (roughly 750,000 words), with an output cap of 64,000 tokens, and processes text, images, video, audio, and PDFs, plus tool-calling and computer-use capabilities. In real-world testing, it finished the same internal coding benchmark in 2 minutes 13 seconds, compared to over 5 minutes for the previous Flash generation—a visible leap in speed. Google's own benchmark data shows Gemini 3.7 Flash beating rivals like Claude Sonnet 5 and GPT-5.6 Terra in 11 out of 18 tests, scoring 1,588 Elo on Code Arena's web development category and 30.4% on AutomationBench, an enterprise workflow test—numbers calculated using Google's own methodology, so the margin of "leading" should be taken with a grain of salt.
Pricing got cut in half too: through the end of the year, it's $0.75 per million input tokens and $3.75 for output—half the original price of the previous Flash. That promotional pricing ends December 31, after which it jumps back up to $1.5 and $7.5, though even at that rate it's still cheap by Google's standards.
OpenAI's preview on the same day, GPT-5.6 Sol Ultrafast, isn't actually a new model—it's the existing GPT-5.6 Sol, already hardened with red-team-tested prompt injection defenses, re-run on Cerebras chips. Output speed hits up to 750 tokens per second, roughly 560 English words per second, 14 times faster than the standard Sol. Cerebras's wafer-scale chips are the key ingredient here—fast enough that voice assistants could theoretically "think" in real time mid-call. The catch: this version is currently only available to a small group of OpenAI API customers on an invite-only basis, with no public head-to-head benchmarks against other models released yet.
John Crepezzi, an AI engineer at Jane Street, said in OpenAI's announcement that the speed "unlocks different ways of using the model," while Courtland Lykins, head of product at Podium, called it "invaluable" for their voice systems, saying it "completely changed the calling experience." These are customer quotes, not public benchmark results.
The timing of both releases is worth noting. Google's true flagship model, Gemini 3.5 Pro, still has no release date, coming three weeks after Gemini 3.6 Flash launched and just days after a leadership shakeup at DeepMind—Demis Hassabis stepping back to let deputy Koray Kavukcuoglu take over. OpenAI, meanwhile, is leaning on borrowed Cerebras compute to chase speed rather than waiting for its own infrastructure to catch up. Both companies are using "speed" to shift the conversation, but only Google has actually put the goods in developers' hands.
Gemini 3.7 Flash is now available worldwide; GPT-5.6 Sol Ultrafast remains an invite-only preview with no public release date announced.






