How are they Losing so Bad
Dan PetersonSynopsis — AI-drafted from Dan's notes
An 11-minute rant from ThePrimeagen arguing that Google is fumbling AI on a generational scale. It had every advantage: its first TPU in service in 2015, eleven years before OpenAI’s first chip; the 2017 transformer paper, written by eight of its researchers; and spending he puts at almost $500 million a day. Yet he finds it eighth on the Artificial Analysis Intelligence Index, behind much smaller labs: Z.ai, which he says has 800 to 1,100 staff, Moonshot with about 300, and xAI’s Grok, a joke at coding in March and now near the top.
Beyond the ranking he has three exhibits. The Gemini site gives a free user Gemini 3.1 Pro, a February model, and stops him short of Google’s best. He gave Gemini 3.8 Flash, inside Cursor, a trivial bug and came back 40 minutes later to find it re-reading one file in a loop: 330 million tokens, $118. Other tasks went “not horrible”, but he calls the loop a bug people have reported for eight months and says he is now wary of leaving the model unattended. And in August, as an anonymous model called Ox Alpha drew attention on OpenRouter, Google AI Studio staff posted what read as teasers. It turned out to be Z.ai’s GLM-5.3 Flash, and Google’s Logan Kilpatrick called it an “unfortunate whirlwind of timing” around the Gemini 3.7 Flash launch.
The history is right and one detail is not. Google ran TPUs internally from 2015 and announced them in 2016; OpenAI and Broadcom unveiled Jalapeño, OpenAI’s first chip, in June 2026. Alphabet’s raised 2026 capital spending guidance of $195 to $205 billion comes to roughly $530 to $560 million a day, though that covers all its capital spending and not only AI. The transformer paper’s authors were at Google Brain and Google Research, not DeepMind, which merged with Brain to form Google DeepMind only in 2023. Ox Alpha ran on OpenRouter from 20 August and was revealed on the 26th; the Google posts went up on the 22nd, and one of them did name Ox Alpha.
The ranking is a snapshot, and it moved both ways. A week after the video OfficeChai reported Google tenth among labs on the index, with Gemini 3.8 Flash its best model at fourteenth. In February, Gemini 3.1 Pro Preview had led the same index, and on 30 September Gemini 4 Argon put Google back among the top three labs. The reading loop is real and not only Cursor’s: users on Google’s own developer forum report Gemini 3.7 and 3.8 re-reading the same lines in Google’s Antigravity tool, with a similar loop report there from March. His $118 run is his own and can’t be checked.
Links
- https://www.youtube.com/watch?v=d3Gjq-BffuI
- https://en.wikipedia.org/wiki/Tensor_Processing_Unit
- https://en.wikipedia.org/wiki/Attention_Is_All_You_Need
- https://techcrunch.com/2026/06/24/openai-unveils-its-first-custom-chip-built-by-broadcom/
- https://w.media/ai-infrastructure-demand-pushes-alphabets-2026-capex-guidance-to-us-205-billion/
- https://officechai.com/ai/google-slips-to-10th-place-among-ai-labs-on-the-artificial-analysis-intelligence-index/
- https://artificialanalysis.ai/articles/gemini-3-1-pro-preview-new-leader-in-ai
- https://artificialanalysis.ai/articles/gemini-4-argon-google-top-three-labs
- https://wavect.io/blog/ox-alpha-free-ai-model-guide-2026/
- https://explainx.ai/blog/gemini-team-ox-alpha-timing-backlash-timeline-august-2026
- https://discuss.ai.google.dev/t/antigravity-gemini-models-3-8-and-3-7-get-stuck-on-reading-repeated-files-nonstop/183779