BGAD Consulting
STRATEGIES. DELIVERED.
BGAD News Flash
Daily briefing: digital, tech and AI
3 October 2026
Google released Gemini 4 Argon, claiming state of the art results on coding and agentic knowledge work benchmarks and wins over OpenAI's GPT-6 Astra on several measures. Nathaniel Whittemore notes Google spent most of 2026 outside the conversation as a top model lab and found itself well behind once coding ability and agent harnesses became the battleground. The catch is availability. Access at launch is limited to a small group of cybersecurity partners in the US government's voluntary prerelease programme, so none of the benchmark claims can yet be checked independently.
Source: The AI Daily Brief, 1 October 2026, Gemini 4 Argon, Sonnet 5.5 and What Matters with AI Models
The same episode pairs Gemini 4 Argon with Anthropic's Claude Sonnet 5.5 to ask where the frontier actually sits. Launch coverage puts Sonnet 5.5 at 70.6 percent on Terminal Bench 4.0 while holding pricing at 2 dollars per million input tokens and 10 dollars per million output tokens, and reports it beating Opus 5.5 on coding at roughly half the price. Whittemore's argument is that raw model performance is no longer the deciding variable, with coding harnesses and personal agents now mattering as much as the weights.
Source: The AI Daily Brief, 1 October 2026, Gemini 4 Argon, Sonnet 5.5 and What Matters with AI Models
The Big Short investor told Prof G Markets he would not invest in an Anthropic listing, arguing that Anthropic and OpenAI are manufacturing a crisis, and that he is waiting for the S-1 before judging the numbers. Anthropic's draft S-1 has been filed confidentially with the SEC, so the document he wants is not yet public. Pivot ran a segment the same day on what it called the shaky road to Wall Street for both labs. Eisman also said he is short one company and that high bond yields could trigger a correction, so he is talking his own book.
Sources: Prof G Markets, 2 October 2026, Steve Eisman: One Company Could Break The AI Boom; Pivot, 2 October 2026, AI's Rocky Road to Wall Street, Hegseth's Macho Military, and Trump's AI Safety Theater
In a State of Markets episode built around 25 charts, a16z growth investors David George, Sarah Wang, Alex Immerman and Santiago Rodriguez argue hyperscaler capital expenditure is approaching 1 trillion dollars a year, with demand for compute still ahead of supply. The companion deck traces the path as roughly 241 billion dollars in 2024, 416 billion in 2025 and 790 billion in 2026. a16z is a venture firm with heavy exposure to the buildout and this is its own research, so treat the framing as directional rather than neutral.
Source: the a16z Podcast, 30 September 2026, The $1 Trillion AI Buildout, State of Markets
The same a16z session flags the distance between deployment and measurable impact. Its data has nearly 70 percent of S&P 500 companies reporting a live AI deployment, roughly 30 percent reporting any quantifiable impact, and only about 2 percent consistently disclosing an AI metric tracked over time. a16z also notes that barely 2 percent of US households were paying for AI services as of April 2026, which sits awkwardly next to the capital spending in the same deck.
Source: the a16z Podcast, 30 September 2026, The $1 Trillion AI Buildout, State of Markets
a16z reports that pricing for prior generation GPUs, specifically A100s, stayed at or above start of year levels through 2026, against widespread predictions of fast obsolescence as newer silicon shipped. The firm reads this as evidence that compute demand runs deep enough to keep old hardware fully used. The episode links it to falling inference costs and to agents absorbing any capacity that frees up.
Source: the a16z Podcast, 30 September 2026, The $1 Trillion AI Buildout, State of Markets
Walter Goodwin, founder and chief executive of chip company Fractile, told No Priors that FLOPs have scaled roughly a million fold over twenty years while memory bandwidth has risen only about 40 times, which makes bandwidth the binding constraint on inference. He claims Fractile's design delivers 25 times more bandwidth per chip than an HBM based part, and the company markets it as running advanced models up to 25 times faster at a tenth of the cost. No published benchmarks or third party verification were offered for any of those figures.
Source: No Priors, 2 October 2026, The Future of Frontier Model Architectures with Walter Goodwin, Fractile Founder and CEO
Alex Zhang, first author on the Recursive Language Models work, told Latent Space that the dominant coding agents from OpenAI and Anthropic are all the same, being variations on treating every task as a prompt fed back as context. He pitches Recursive Language Models as an alternative that restricts the model to code as its only tool, which he says keeps calls locally in distribution and allows compositional generalisation. His specific claim is that strategies trained on short tasks transfer to problems 8 to 30 times longer.
Source: Latent Space, 2 October 2026, Academia is for Ambition, Alex Zhang, MIT
On the same episode Zhang put figures on OpenAI's Navier Stokes mathematics project: roughly 10,000 agents running over 88 hours, about 130 billion output tokens, at an estimated cost near 40 million dollars. He calls the approach grossly inefficient and says 95 percent of the swarm is useless. The numbers are his estimate rather than an OpenAI disclosure, and he uses them to argue that harness design beats brute force parallel sampling.
Source: Latent Space, 2 October 2026, Academia is for Ambition, Alex Zhang, MIT
Zhang reports that on KernelBench nearly all recent top solutions are AI generated, but the single human competitor in the top ten produced the only kernel that ran stably end to end. He puts the gap down to reward hacking, with model submissions optimising the scored metric rather than real performance. He offers it as evidence that current agent benchmarks overstate practical capability.
Source: Latent Space, 2 October 2026, Academia is for Ambition, Alex Zhang, MIT
PolyAI chief technology officer Tsung-Hsien Shawn Wen told Machine Learning Street Talk that the company's audio native model, Dialog-RSN-1, works in three stages. It predicts turn taking signals, generates a text reply with citations, then writes a transcript for enterprise auditing. Wen says it is trained on authentic noisy call recordings with synthetic noise added, and that cleaning the audio too much made performance worse. He also argues current voice benchmarks miss the point because they score answer quality while ignoring latency, turn taking and caller trust. The episode was produced in partnership with PolyAI.
Source: Machine Learning Street Talk, 1 October 2026, How a Voice Agent Learns the Rhythm of Conversation, Shawn Wen
Lean creator Leonardo de Moura told Machine Learning Street Talk about an incident in which a false proof of the Collatz conjecture was accepted by Lean's official kernel and by Nanoda, an independent alternative kernel, apparently by exploiting a different bug in each. That matters because formal verification rests on a small trusted kernel plus independent checkers, and two independent checkers failing on the same artefact is precisely the failure the model is meant to rule out. De Moura discussed it alongside Mathlib, specification complexity as the real bottleneck, and the growing role of AI agents in formalisation.
Source: Machine Learning Street Talk, 30 September 2026, Who Checks a Proof No Human Can Read, Leo de Moura
On All-In the hosts and their guests discussed how to spot bubble behaviour in private markets, with the clearest signal being second tranche investors paying two to three times the initial valuation while company performance is unchanged. The hosts treat that as the cue to take money off the table and return cash to limited partners. Jason Calacanis separately argued that a billion dollar paper valuation means little now and that unicorn status should be measured in revenue instead. All four hosts are active venture investors with private AI holdings, so both the warning and the call to return capital reflect their own books.
Source: All-In, 30 September 2026, Jake Paul and The Chainsmokers: Turning Fame into Funds, Jake Enters Politics, Venture Bubble Signs
The AI Daily Brief covered an executive order creating America.gov, an AI driven portal for federal government services that is eventually meant to handle passport applications, name changes and Medicare enrolment. Departments are reported to have 90 days to integrate their services, with functionality arriving from next year. Detail beyond the order itself rests on secondary summaries rather than the show's own notes, so treat the timeline as provisional.
Source: The AI Daily Brief, 1 October 2026, Gemini 4 Argon, Sonnet 5.5 and What Matters with AI Models
The New York Times used a two minute Hard Fork episode to announce that Max Read will guest host the show for the next few months, alongside a rotating cast of Times reporters and outside experts. Read has written about technology and the internet for more than a decade, most recently in his newsletter Read Max, and has been writing the Times' Work Friend column. The handover follows the 18 September episode billed as a Hard Fork exit AMA.
Source: Hard Fork, 1 October 2026, What's Next for Hard Fork