THE AI RANKINGS

The Daily Debrief

Astra goes critical, and the DOJ sides with OpenAI

OpenAI says Astra is the first of its models to cross the Critical cybersecurity threshold in its own Preparedness Framework, and the Justice Department filed late on Tuesday arguing that training large language models on copyrighted text is fair use.

Issue 15 · Thursday, 3 September 2026

Astra is the first OpenAI model rated a critical cyber risk OpenAI said on Wednesday that Astra crosses the “Critical” cybersecurity threshold in its own Preparedness Framework, scoring full marks on ExploitBench and finding two zero-day vulnerabilities unaided in a separate evaluation against more recently disclosed flaws. It says the safeguards now “sufficiently minimize the risk of severe harm for release”, the model is coming “soon”, and the cyber capabilities go first to vetted testers through its Daybreak Blue program. Source

The Justice Department told a court AI training is fair use The government filed a statement of interest late on Tuesday in the New York Times’ copyright case against OpenAI and Microsoft in Manhattan federal court, arguing that training large language models on copyrighted text is transformative and that AI development is a national security interest. The filing also argues that requiring licences would hand the largest technology companies an oligopoly on model training; the Times said the administration is “siding with a handful of trillion-dollar AI companies at the expense of the countless American creators whose work they stole”. Source

Google shipped Gemini 3.8 Flash and a cybersecurity version on Wednesday Google held the introductory price at $0.75 per million input tokens and $3.75 per million output until 31 December, and says the model beats Gemini 3.7 Flash and other frontier models on finance and legal agent benchmarks while scoring 54.9 per cent on HLE-Verified. Gemini 3.8 Flash Cyber goes to vetted defenders through a new Fairwind Program, so both frontier labs spent this week deciding who is allowed a cyber model rather than whether to build one. Source

New York City banned AI for students through eighth grade Mayor Zohran Mamdani announced a one-year moratorium on Wednesday covering all student-facing generative AI software up to eighth grade across the largest school district in the United States, which the city calls the most expansive such prohibition in the country, with companion chatbots barred at every grade. High school students instead get twice-yearly AI critical-thinking modules and supervised pilots in a small number of classrooms. Source

A third of Perplexity’s citations don’t contain the figure cited Haus Research put 310 factual questions about 210 technology companies to Perplexity’s two search models on Wednesday, fetched every source they cited, and found that 34.7 per cent of the 1,826 citations attached to a sentence stating a figure pointed at a page that would not open or carried none of that sentence’s numbers. Scored per claim rather than per citation the failure rate is 14.4 per cent of 872 claims, and one cited URL in six was behind a gate the reader cannot pass. Source

Uber is cutting 3,300 jobs to fund robotaxis Bloomberg reported on Wednesday that the cuts are roughly 10 per cent of Uber’s workforce and come with 20 per cent fewer managers, taking headcount back to about its 2021 level and making this the company’s largest reduction since the pandemic. Dara Khosrowshahi told staff the company had accumulated “more layers, more coordination, more fragmented ownership”, and the money goes to ride-sharing, delivery and more than $10bn of robotaxi partnerships. Source

Alibaba’s Qwen3.8-Max update took first place on CodeArena The 0902 snapshot of Qwen3.8-Max landed on Wednesday post-trained for coding and agent work, lifting the model’s front-end CodeArena score 22 points to 1,691 and putting it top of that leaderboard. Alibaba left the version number alone, along with the 2.4 trillion parameters, the one-million-token context window and the $2 and $6 per million input and output tokens. Source

Anthropic’s watermark detection API is in private preview Anthropic’s Fable and Mythos 5.1 announcement lists the EU AI Act detection tool as available to regulators, law enforcement, media, fact-checkers, independent researchers, educational organisations and EU civil society groups, plus enterprises carrying their own compliance obligation. General availability has no date. Source