BGAD Consulting BGAD Consulting STRATEGIES. DELIVERED.

BGAD News Flash

BGAD News Flash

Daily briefing: digital, tech and AI

10 September 2026

OpenAI says it cracked Navier-Stokes, but has published no proof

OpenAI reported a solution to the Navier-Stokes Millennium Prize Problem using a swarm of agents running on a next generation model described as significantly more capable than GPT-6 Astra. Ethan Knight put the effort at roughly 10,000 agents, trained over the past year with multi-agent reinforcement learning and self organising across large amounts of unstructured parallel test time compute. No theorem statement, preprint, proof sketch or formal verification artifact has been made public, so mathematical acceptance remains unresolved. Latent Space also warns that the widely repeated 88 hour figure traces back to a satirical post and should not be treated as confirmed. The AI Daily Brief carried the same story as a disputed headline.

Source: Latent Space, 9 September 2026, AINews on the OpenAI Navier-Stokes report, newsletter post; and The AI Daily Brief, 9 September 2026, AI Model Month Is Off to a Blistering Start

Meta launches Muse, an always on consumer agent with a separate security layer

Meta has shipped Muse, a personal AI agent that is always on, connects to apps, drives a browser and is distributed through Meta's own properties. Each Muse runs in its own persistent isolated Linux virtual machine, actions are mediated by a separate component called Sentinel, secrets are never directly exposed to the agent, and sensitive actions require approval, with a bug bounty of up to 300,000 dollars. Connectors span Gmail, Calendar, Outlook, Plaid, OpenTable, Docs, Spotify and Peloton alongside Instagram, Messenger, Facebook and Marketplace, with commerce running on Stripe Link and an agentic payment protection guarantee. Meta says day one usage came in at ten times its internal projections.

Source: Latent Space, 9 September 2026, AINews on the OpenAI Navier-Stokes report, newsletter post

Model selection becomes a skill in its own right as September floods the market

Nathaniel Whittemore's argument this week is that the sheer pace of releases has made choosing the right model a distinct competence rather than an afterthought. The first ten days of September alone brought Gemini 3.8 Flash, Meta's Muse Spark 1.3, the Muse consumer agent and ChatGPT Images 2.5. His point for operators is that faster, cheaper and more specialised models mean the default of routing everything to one frontier model is now leaving both money and quality on the table.

Source: The AI Daily Brief, 9 September 2026, AI Model Month Is Off to a Blistering Start

Artificial Analysis rewrote its benchmark index over a weekend to place Astra

GPT-6 Astra initially scored 61 on the Artificial Analysis Intelligence Index, identical to GPT-5.6 Sol, five points behind Fable 5.1 and one point behind Meta's Muse Spark. Artificial Analysis then pushed out version 4.2 over the weekend, adding agentic weighting through its AA Briefcase test, after which Astra ranked second behind Fable 5.1. Ethan Mollick criticised the decision to change all index criteria at once. For anyone using public leaderboards to inform procurement, this is a reminder that the ranking is a moving artefact as much as a measurement.

Source: The AI Daily Brief, 8 September 2026, Why GPT-6 Astra Is So Significant and So Confounding

Astra scores 100 per cent on ExploitBench at every effort level

The clearest step change in the Astra benchmark set is security. The model reached 100 per cent on ExploitBench at every effort setting, and 39 per cent on an internal benchmark of recently disclosed vulnerabilities against 5.5 per cent for GPT-5.6 Sol. That is a roughly sevenfold jump on real, recent vulnerabilities in a single generation. The practical read for security teams is that the offensive capability now sits inside a generally available product rather than a research preview.

Source: The AI Daily Brief, 8 September 2026, Why GPT-6 Astra Is So Significant and So Confounding

Practitioners say coding has plateaued while computer use jumped

Martin Casado of a16z reported a clear step up in computer use but no meaningful step in coding for his own work, and questioned whether it still makes economic sense to push models on high skill development tasks. Armin Ronacher said Astra's Python goes strange one step removed from conventional code and that its unit tests are poor. Several practitioners including Kris Puckett and Dan Dris reported weak front end design output despite strong results on 3D, spatial reasoning, mathematics and agents, with "make a beautiful website" cited as possibly the single biggest remaining reason to use Claude. Dan Shipper of Every called Astra the best writing model he has tried and its computer use incredible, but said it overcomplicates interfaces and still trails Fable 5.1 on the largest tasks.

Source: The AI Daily Brief, 8 September 2026, Why GPT-6 Astra Is So Significant and So Confounding

3D generation is the breakout use case, and the unit economics are startling

The demo wave following Astra's release has centred on 3D rather than text or code: the launch demo house rebuilt in Unreal Engine 5, a Craig Federighi parkour clip recreated in Blender across 900 hand rendered frames, a Tesla Model X decomposed into 334 modelled pieces, and 3D house tours generated from a Zillow listing. The cost figures are the story. Theo built a browser game in one shot for under 30 dollars, and one builder produced Sonic in Godot in 53 minutes using 4 per cent of a weekly usage allowance, or 25 minutes and 1 per cent on a medium setting. Adjacent uses are already appearing in explanation and training, including interactive 3D explainers and a one session interactive 3D atlas of a user's own ankle.

Source: The AI Daily Brief, 8 September 2026, Why GPT-6 Astra Is So Significant and So Confounding

Astra's staged rollout drew an apology from Altman before reaching everyone

Astra was announced with access initially limited to partners in OpenAI's cybersecurity focused Daybreak programme, with the delay attributed to safety and alignment standards. Sam Altman apologised for the messy rollout, and paying subscribers were offered banked usage resets for the days they went without access, with general availability arriving by the Friday night. OpenAI has since confirmed full rollout to Plus, Pro, Business and Enterprise users across Codex and ChatGPT Work. The three minute launch video passed 132 million views and 100,000 saves by the Tuesday morning, and every use case it shows is hands free, with a foreground task and a background admin task running at once.

Source: The AI Daily Brief, 8 September 2026, Why GPT-6 Astra Is So Significant and So Confounding; and Latent Space, 9 September 2026, AINews, newsletter post

Cognition reaches a 48 billion dollar valuation as capital keeps arriving

Cognition has raised more than 2 billion dollars at a 48 billion dollar valuation, saying run rate revenue has grown from 492 million to close to 900 million dollars since May. Mistral raised a round of roughly 24 billion dollars, and ElevenLabs is reported to be preparing for an initial public offering. Whatever the debate about model progress, the funding market for the applications and infrastructure layers has not cooled.

Source: The AI Daily Brief, 9 September 2026, AI Model Month Is Off to a Blistering Start; and Latent Space, 9 September 2026, AINews, newsletter post

Harvey shows harness design nearly triples accuracy, and post-training beats it again

Harvey and Baseten published results from a recursive language model harness for mergers and acquisitions diligence, in which a root agent searches a data room, delegates document review to sub-agents and aggregates across corpora of up to 80 million tokens. On the LAB Diligence benchmark, moving from a standard tool loop to the recursive harness lifted the mean rubric pass rate from 23 to 62 per cent across models. Post-training then went further: self distilled supervised fine tuning on GLM-5.2 lifted pass rate from 46 to 60 per cent, and GRPO on Qwen3.5-122B-A10B lifted it from 30 to 63 per cent on held out rooms while raising document coverage from 62 to 96 per cent. For enterprises weighing model choice against engineering effort, this is evidence the effort side still pays.

Source: Latent Space, 9 September 2026, AINews on the OpenAI Navier-Stokes report, newsletter post

Dwarkesh Patel finds most pretraining progress came from data, not model design

In new research with Jerry Han, Dwarkesh Patel measured where compute efficiency gains between 2019 and 2025 actually came from. At a fixed budget of 1e19 FLOP, data improvements delivered 12.0 times compute efficiency gains against 3.7 times from model recipe improvements, a ratio of 3.24 in data's favour, with the two largely additive. Their measured year on year gain of 1.57 times jointly sits well below Epoch's estimate of roughly 3 times. Their argument is that model research's real contribution was making larger compute usable rather than raw efficiency, which matters for anyone forecasting how much headroom remains.

Source: Dwarkesh Podcast, 8 September 2026, Pretraining progress is mostly coming from data, written post

Students using AI chatbots to do the work score a year behind, says the OECD

Prof G Markets took up the OECD's global education report card with the Financial Times chief data reporter John Burn-Murdoch, on the finding that results are the worst on record. The underlying PISA 2025 study of 760,000 fifteen year olds across 91 countries found reading, maths and science at their lowest since records began in 2000, with reading falling from a 2012 high of 501 to 466. Nearly half of students now use AI chatbots regularly for learning, and those using them to draft essays, summarise texts and research topics score about 20 points lower in science, which the OECD equates to roughly a year of learning. OECD head Mathias Cormann said more screen time and less reading for pleasure have gone hand in hand with weaker results, with scores starting to fall above five hours of device use a day.

Source: Prof G Markets, 9 September 2026, Canadian Economist: Trump's Tariffs Are A Gift To Mark Carney

Taiwan's drone ambitions run into an annual budget and a software gap

ChinaTalk's analysis of Taiwan's drone industry finds procurement targets escalating far faster than delivery. A 2022 target of 3,200 military drones by mid-2024 produced roughly 1,000 units, yet the goal was raised to 48,750 military drones by 2027 and then to 200,000 units in January 2026. In August the legislature passed a headline 240 billion New Taiwan dollar authorisation but stripped out the multi-year procurement commitment, so funding must now be re-approved annually, and made the economics ministry lead overseer while defence still does the buying. The piece argues Taiwan's realistic play is supplying components such as airframes, gimbals and actuators into other countries' integrators rather than exporting finished drones, with Western software dependency and constrained flight testing as the named bottlenecks.

Source: ChinaTalk, 9 September 2026, Can Taiwan Build Drones, newsletter post

DNA synthesis screening emerges as the one AI safety measure Washington and Beijing might both want

ChinaTalk makes the case that mandatory screening of nucleic acid orders is a rare AI risk control both the United States and China have independent reasons to adopt. A June open letter signed by Sam Altman, Dario Amodei and Demis Hassabis called for the United States to mandate that synthesis providers screen DNA and RNA orders before fulfilment, and two bills are live in Congress after the previous framework was cancelled in May 2025 without replacement. Firms representing about 80 per cent of global synthesis capacity already screen voluntarily, so a mandate mostly raises the floor for the rest and protects compliant majors from being undercut. With China accounting for roughly 34 per cent of the world's synthesis providers, the author argues that open weight releases give Beijing more reason than Washington to secure the physical chokepoint, since safeguards at the model layer disappear once weights are public.

Source: ChinaTalk, 8 September 2026, US-China Biorisk Cooperation: Yes, It's Possible, newsletter post

Subscribe to our updates
Past daily issues Weekly newsletter