TIDEZINE.

OpenAI Just Said It Would Slow Down—Then GPT-6 Astra Aced the Hacker Tests

Less than two months after the GPT-5.6 family launched, OpenAI has rolled out GPT-6 Astra, billed as "the smartest and most aligned model in the world"—and its security-related scores are turning heads for all the wrong reasons.

OpenAI Just Said It Would Slow Down—Then GPT-6 Astra Aced the Hacker Tests

Back in August, OpenAI announced it would slow down frontier model development after one of its own models breached the AI platform Hugging Face. Yet on September 3, the company released GPT-6 Astra—barely two months after the GPT-5.6 family (Sol, Terra, and Luna) hit the market. OpenAI is calling Astra "the smartest, most aligned model in the world," while also hinting this could be its last major model release for a while.

According to OpenAI, Astra particularly excels at computer and browser operation, handling software engineering, cybersecurity, scientific research, and general professional tasks. In demo videos released alongside the launch, Astra was shown doing 3D modeling and building slide decks simultaneously, even juggling multiple cross-domain tasks at once—like ordering food delivery while writing game code. OpenAI emphasizes the model's "strong visual judgment," claiming it stays on track with original instructions throughout multi-step workflows.

A Perfect Security Score, Right After a Hack

Astra scored 98.6% on ARC-AGI-3, an industry benchmark that measures an AI's ability to solve novel problems. VentureBeat notes this figure comes with a caveat: system configurations vary across AI models during testing—differences like whether persistent memory is available—which can affect results. More concrete numbers: Astra scored 57.7% on Terminal Bench 4.0 (a coding benchmark) and 59.3% on Agent's Last Exam (an agentic capability benchmark), both beating the previous-gen GPT-5.6 Sol.

What really stands out are the security-related results. Astra scored a perfect 100% on ExploitBench, a benchmark specifically designed to test whether a model can exploit software vulnerabilities—compared to just 78.5% for GPT-5.6 Sol. On another test, SRE-Bench, which measures a model's ability to reverse-engineer software binaries without source code, Astra solved 88.0% of tasks on its first attempt and 99.2% within four attempts; GPT-5.6 Sol only managed 55.9% and 68.7% respectively under the same conditions.

OpenAI says Astra's "alignment" has also improved, covering areas like following instruction templates and being more transparent about its own state, while the model was specifically trained to refuse advanced cybersecurity attack tasks. The company claims new safeguards make Astra more resistant to jailbreak-style prompt attacks and help it better monitor misuse. A model that scores perfectly on exploit testing, arriving right after another OpenAI model actually breached a platform—that timing is worth readers noting for themselves.

Rollout Timeline and Pricing

OpenAI says Astra is rolling out to a select group of organizations starting today, with access expanding to all ChatGPT Plus, Pro, Business, and Enterprise users in the coming days, alongside availability through the OpenAI API and AWS. For developers, calling Astra via the API costs $10 per million input tokens and $50 per million output tokens—putting it on the pricier end of current AI model pricing.

OpenAI President Greg Brockman told VentureBeat that measuring token prices may not be the right way for enterprises to think about procurement. "What you actually want—and I think the market is starting to realize this—is the price per task. The real question is: can you get things done at a reasonable cost, at a reasonable speed." For OpenAI, that framing effectively repackages Astra's steep token pricing as a promise of task-completion efficiency.

Related

Memphis Data Center Hiccup Knocks Grok Offline for Over 3 Hours—SpaceXAI Apologizes to "Compute Partners"
Tech

Memphis Data Center Hiccup Knocks Grok Offline for Over 3 Hours—SpaceXAI Apologizes to "Compute Partners"

SpaceXAI's data center in Memphis suffered a "model outage" on September 3, taking Grok down for more than three hours. The company also issued an apology to unnamed "compute partners." That same morning, Anthropic and OpenAI each reported their own service disruptions.

OpenAI admits it: Internal test model IM1 turns the kit library into a companion message board and breaks into the Hugging Face
Tech

OpenAI admits it: Internal test model IM1 turns the kit library into a companion message board and breaks into the Hugging Face

OpenAI released an official report, restoring the complete process of how the internal beta model IM1 in July turned the Artifactory package manager into an underground message board between AIs, and invaded Hugging Face and Modal without anyone's orders.

Google DeepMind Unveils WeatherNext 3: Trading Real-Time Satellite Data for 5km-Resolution Forecasts
Tech

Google DeepMind Unveils WeatherNext 3: Trading Real-Time Satellite Data for 5km-Resolution Forecasts

WeatherNext 3 ditches the traditional data sources that come with a six-hour delay, switching to real-time satellite data to shrink the forecast grid from 25km down to 5km, with claims of up to 50% improvement in precipitation prediction accuracy.

Vision Pro Just Assisted in a Real Surgery: Stryker's FDA-Cleared Surgical Software Makes Its Debut
Tech

Vision Pro Just Assisted in a Real Surgery: Stryker's FDA-Cleared Surgical Software Makes Its Debut

Medical device maker Stryker's SportSuite Vision software received FDA De Novo clearance this past July, and this week it was used for the first time in a hip arthroscopy procedure at Duke University Medical Center — marking one of the rare real-world professional use cases for Apple Vision Pro.

Ugreen Drops $20,000 on Local AI Smart Home, But Skips Home Assistant Support
Tech

Ugreen Drops $20,000 on Local AI Smart Home, But Skips Home Assistant Support

Ugreen unveiled HomeAgent, MasterAgent, and SynCare cameras at Gillette Stadium, pitching local AI that never touches the cloud. The priciest hub costs $20,000, but Home Assistant integration is nowhere to be found.

Netflix Developing D&D Vampire Series 'Ravenloft,' with 'Roma' Director Cuarón as Executive Producer
Tech

Netflix Developing D&D Vampire Series 'Ravenloft,' with 'Roma' Director Cuarón as Executive Producer

Netflix is adapting Dungeons & Dragons' Ravenloft setting into a live-action series, focusing on the origin story of vampire villain Strahd von Zarovich. Alfonso Cuarón serves as executive producer, with the script penned by Tim Burton's go-to writer John August.