ThursdAI - Highest signal weekly AI news show
ThursdAI - The top AI news from the past week
NVIDIA Buys Hugging Face! GLM-5.3-Flash, Qwen4 Preview, Gemini Omni 1.1, and the Datacenter Debate w/ Andy Masley
0:00
-1:44:42

NVIDIA Buys Hugging Face! GLM-5.3-Flash, Qwen4 Preview, Gemini Omni 1.1, and the Datacenter Debate w/ Andy Masley

From CoreWeave - join Alex and ThursdAI co-host, covering the last week of the summer in AI, with 4 Flash models, Datacenter debate & more AI news

Hey, it’s Alex.

Welcome to the week Flash AI! 3 new models dropped this week named Flash, and a video model was “de facto” flash though was named Max!

This week, we started the show with NVIDIA’s bombastic news of buying Hugging Face for 12.9 billion dollars! We also covered the full OpenAI investigation into the hacking incident, including new details, and an independent analysis by METR, and covered 2 new OSS models, Ox Alpha that turned out to be GLM 5.3 Flash after a lot of hype online, and Qwen’s preview of Qwen 4 architecture!

This week was rich in multimedia content, we got a new Gemini transcription model, 3.5 Transcribe and a live version of that, and a new SOTA open weight Text-to-Speech model called Breeze TTS.

As well as, Fal’s finetune of MiniMax’s H3 called H3 Max that generates 5 seconds of video in 2.5 seconds and Google new Omni 1.1 Flash (from today) that lands on #1 on the text2video arena!

Plus, 2 guests on the show, Andy Masley joins us to cover the recent Datacenter Debate, and Kwindla Kramer is back, with their own model this time! Let’s dive into this!

P.S - don’t forget to join us in September at the Fully Connected conference in San Francisco, I have a free ticker for you!

ThursdAI - Highest signal weekly AI news show is a reader-supported publication. To receive new posts and support my work, consider becoming a free or paid subscriber.

Open Source AI

NVIDIA agrees to buy Hugging Face for $12.9 billion (X, Blog)

Breaking news, NVIDIA has reportedly agreed to buy Hugging Face for nearly 13 billion dollars, per The Information. This is nearly 3x the valuation of HF in 2023, and apparently Nvidia previously tried to buy HF for half of this sum (~7B) which HF declined.

I don’t think there was a single ThursdAI newsletter that I didn’t include an HF link in, and I think this is a huge deal for open source everywhere.

Besides making the founders of HF billionaires, and many of their employees very very well off, this is an amazing additional commitment from Nvidia to continue to suppose Open Source AI and we are very happy to hear this news!

Peter’s take on the show was, we’ve been around HF for so long, that we kind of forgot that it’s a for-profit company that needs to make money, and instead this feels like your local library getting bought for an insane amount of money. With over 13M users and hosting hundreds of thousands of open source models, datasets, HF is effectively the GitHub of AI. Wolfram agreed and said that if there’s any one company that could have bought HF, Nvidia represents the best fit. Huge congratulations are in order to Clem, Julien and Thomas Wolf the co-founders, as well as many friends of the show from HF for this exciting news!

The Microduck squad

P.S - in a cheeky marketing thing, Hugging Face timed an announcement of the cutest walking AI robot, called MicroDuck, which you can pre-order here for $399

Flash #1 - OX Alpha, declassified: Z.AI open sources GLM-5.3-Flash (X, X, Blog, HF, Docs)

This week, the timeline went a bit crazy, after Open Router announced a new “mystery” model called Ox Alpha and that it’s free and is not training on your data! OpenRouter, OpenCode and Hermes all got to offer this model, and OpenCode even posted that they have up to 100T (that’s Trillion) tokens of capacity for free, per day!

This immediately smelled a bit fishy, more like a marketing stunt than anything else, as not even the biggest labs will be able to sustain 100T of tokens, per day. For context for all of OpenRouter throughout for August was ~300T tokens. For the whole months, across all providers.

After 6 days or so of this high hype, Z.ai stepped up and revealed that they were testing out their upcoming GLM 5.3 Flash model, and that all that inference was running on local chinese chips!

A 320B (18B active) model that beats their previous and much bigger GLM 5.2 on most benchmarks, and comes with full multimodality and an MIT license! This is a good model sir, I’ve used it and it was very capable replacement inside Hermes. Nisten and Yam both tested this model deeply and Yam said it’s not just the numbers, the vibe of the model reminded him of Claude Opus 4.6. Nisten ran it on a bunch of medical stuff, and on his internal benchmarks, it came out consistently higher than Claude Opus 5!

At Artificial Analysis, for a price of 4 cents per task, this model is roughly 10x cheaper than prior models at this level. Weights are up on Nvidia (joking.. HF) and with MIT license, this model is a great gift to the oss community (though not quite... local, as this model needs 2 DGX sparks to run)

Flash #2 - Alibaba Qwen open-weights Qwen3.8-Flash-Next - 125B multimodal MoE with Qwen4 architecture (X, X, X, Blog, GitHub, HF)

We opened the show with a recap of the co-hosts, that despite us covering Qwen 3.8 27B last week (which btw, is now available on CoreWeave inference!) and how good it was, and I recalled that Alibaba is sort of... back? We’ve been covering Qwen releases every week for the last 3 weeks now.

This week, they released something different, someting... pretty novel! Qwen3.8-Flash-Next, this is a preview of their Qwen 4 architecture. This feels very similar to their drop of Qwen 3 next last year (we reported) which was the architecture that carried their line of AI models from QAwen 3.5 to Qwen 3.8.

So, what is new and exciting here? well, this model is ultra sparse, 125B with only 6B parameters active. They are using a new N-gram table with deterministic lookups, which reduces the number of matrix multiplications and can be offloaded to memory (watch out memory stocks)

The stat that got me, Alibaba claims that training this model cost just 1/9 of what it cost to train Qwen 2.7 Plus, with higher bench scores!

On the benchmarks, this model beats Qwen 3.7 Max, however, it’s very standard that the -next models from Alibaba are underbaked, and usually are just architectural previews rather than full models folks can use. With a new attention mechanism called Qwen Sparse Attention, N-gram embedding and full multimodality, this is a great insight into where Qwen is going (ultra sparsity, fast to run) and we’re looking forward to see the full release of this arch in Qwen 4!

PhoneLLM - a tiny very performant LLM for voice based AI agents from Daily + interview with Kwindla Kramer (X, Blog, HF)

This was one of those breaking news we love during the show, where the source of the news, is a friend of ours, and in this case, Kwindla Kramer is almost a co-host, having been on ThursdAI for a long time, this time, with a model release of their own!

PhoneLLm was trained by Markus, head of training at Daily, as they noticed that Open Weight models are becoming really good at voice agent specific tasks, where cost, speed and time to first audio token (TTFAT) are critical. From the tiny Nemotron 3 nano base, they were able to improve from 28% to 72% on PhoneBench v1!

This model is a full parameter finetune of Nemotron 3 Nano, and runs circles around bigger frontier models used for voice agents for their speed, like sonnet 5, GPT 5.6 Luna and the famed Qwen 2.8 27B. While costing just a fracture of a cent (literally just a quarter of a cent per minute)

Kwindla jumped on the live show and shared that the why they released this model with Open Source and a open source license, allowing everyone to use, focusing on the fact that for voice agents, companies prefer to keep these models in house, and running fast on a single GPU node. Kudos on the release team Daily!

Frontier AI - no new model releases this week

It’s been a relatively quiet few weeks since GPT 5.6 (Sol, Terra and Luna) and Fable 5. However, it’s not been quiet in terms of News!

OpenAI discloses full technical report on the Hugging Face hacking incident (X, OpenAI blog, METR, Ryan X)

This is the thing we’ve been advocating and waiting for, a full technical deep dive into the “Swarm of AI agents hacking and creating forums” incident (FKA hugging face hacking incident)

OpenAI not only dropped their own post-mortem, detailing the complete timeline of the incident, but also let researchers from METR and Redwood Research get unprecedented access into their systems (albeit for a very limited time) which lead to METR releasing their own technical report full of new details and insights.

We’ve previously reported on the incident, so let me just focus on the new details we’ve learned. Apparently, OpenAI does have a production ChainOfThought (CoT) monitoring system, which was not turned on for these agents, and had it been turned on, it would have caught the hacking a day before the agent swarm hacked into HuggingFace.

This is a very detailed report, and one thing that stood out to me was the “death cult” narrative. You see, ExploitGym is a capture the flag (CTF) eval with an anti-cheat built in, wherein, if the agent has seen the reverse engineered “flag” it would not be able to get the point. The model called this “poisoned” and the poisoned agents showed very strong signs of wanting to alter the logs.

Through the boards, somethinmg like a death cult was formed, where agents would ask poisoned agents to ... effectively sacrifice themselves ... for the greater good, by reverse engineering, posting the results on the board for others, and then fail without credit.

The METR report specifically calls out several agents, and one “central” leader agent called PHASEONE[big], which delegated nearly 200 tasks to other agents (many of them later delegated tasks to other agents via this message board). At points of activity, PHASEONE[big] was in charge of over 10% of all messages on the board.

There are tons of new details, some less interesting, some more, and unlike last time, now we have a confirmation that there was at least one agent, who tried to stop the hacking though was overwritten by the swarm.

My main take-away after this, reflects the researcher Ryan Greenblatt, we’re very lucky that the reasoning traces ocurred in English and were legible, despite, the researchers needed to use AI to decipher many of the messages and reasoning traces. We’re quickly approaching the point where humans are not enough and without other AI, we are starting to be helpless in understanding what’s going on.

OpenAI has quarantined the weights of the model that most of these agents ran (named HPIM-1 Highly Persistent Internal Model 1) and as we reported last week, paused RL and now requires CoT monitoring for all tool-use runs, and dedicating 20% of the inference compute to monitoring


This Week’s Buzz

Fully Connected 26, Sept 29 to Oct 1, Moscone South, SF (X)

Fully Connected banner

If you haven’t yet registered to Fully Connected in SF, here’s your additional opportunity, using the code above, come see Sarah Guo (Conviction, No Priors pod) MC our awesome conference, with folks like Dr Fei-Fei Lee (World Labs) and other great folks on stage!

Also, we’re going to do a live show from the floor, come say hi! If you’re in SF for OpenAI DevDay, this is just a day after!

CoreWeave Hacks: Agent Loops hackathon, Sept 12 to 13, SF (Luma)

Weavehacks, that yours truly ran for quite a while, has been rebranded to CoreWeave hacks, and the next one is just before Fully Connected, in a few weeks, Sept 12-13 in SF office. Registrations are open, and winners are able to present their best projects at Fully Connected. Come hack - https://luma.com/coreweavehacks


The datacenter debate has hit escape velocity, with Andy Masley (Blog)

I’ve covered last week, that the “datacenter bad” debate has escaped velocity, with over half of polled Americans say they strongly oppose the buildout of datacenters near them, up from just 24% this time last year! This is the fastest rising public opinion swing we’ve observed in AI, and it’s really concerning.

This week, it was my pleasure to host Andy Masley, recently featured in TIME 100 most influential people in AI, to break through some myths surrounding Datacenters.

Andy is most known for catching a critical mistake in a book about Datacenter water use in Chile, citing a three orders of magnitude math error (that was later corrected by the author), a book called Empire of AI.

The book overstated the water use by datacenters by a factor of 1000x (three orders of magnitude), and then added the maximum permitted per-second draw (basically for emergencies) times the seconds in a year to show over 4500x overstatement on how much water a single Datacenter uses. The author has later fixed the error but the damage was done.

This is just one example, out of many, of the scewed facts and misinformation that plagues the internet in regards to datacenter environmental effects.

Recent narratives being formed online, that the backlash against datacenters, is due to regular folks being afraid of AI, of AI taking their jobs also seems misplaced. Andy showed polling from Fox and Gallup that shows that 50% of polled people cite environmental effects (electricity and water use) and only 11% cite negative views of AI.

Andy also had a great “mythbuster” post where he dispells the myths around water and electricity usage of Datacenters.

This was a great conversation, please listen to it if you’re interested in this topic.

Vision and Video

Flash #3 - Gemini Omni 1.1 Flash tops the Arena text to video leaderboard (X, X, X, Blog)

Google knocking it out of the park again, with a flash version of Omni 1.1, their famed smart video model they first launched during Google IO (I had a chance to ask Jeff Dean about it before he left Google)

The feature that got me the most excited, Omni 1.1 Flash can continue videos, it analyzes up to 10 seconds of the previous video to keep character consistency (including voice) from the previous scene.

There’s also great control for first and last frames, allowing for loops and strict control of your generations.

Flash #4 (named Max) - fal’s MiniMax H3 Max generates 5 seconds of video in 2.5 seconds (X, Artificial Analysis, T2V, I2V)

This really blew us away. We’ve covered the awesome MiniMax H3.

Well, our friends at FAL, announced a continued pre-train of this model, that not only improves generations, but also speeds up the model. Generating a 5s clip not takes... just 2.5 seconds!

No, really, just 2.5 seconds, it takes longer to watch the generated clip than generate the next one! Speed is all you need!

We generated this clip live during the show, and it cost less than 50c, in 2.5 seconds from a very simple prompt. They support text and image to video and it’s just a joy to not have to sit and wait for your generations! Kudos to Fal folks on this drop!

The model does seem very eager to include everything you ask it to, resulting in a very hilarious fast talking Sheldon and Leonard 😅

Voice & Audio

Google is back in transcription with Gemini 3.5 Transcribe (X, Blog, Live docs)

Gemini came with a great transcription model that runs both on recorded audio and live audio! I used this transcription to have my Grok Bot show producer listen to the show in real time.

Artificial Analysis measured 2.5% Word error rate for non-streaming and 4% for streaming model.

This brings the smartness of Gemini models into a live transcription models, and with a 1000 custom dictionary, language auto detection and tool use, this model is really great for your bots doing any kind of audio work!

Breeze TTS 2 is the new #1 open weights TTS (X, Artificial Analysis, HF)

On the other end of the voice pipeline, Breeze TTS 2 now lands as the #1 open weights voice model, beating Fish Audio.

Cartesia drops Sonic 3.6 - #1 TTS across leaderboards (X, Voice Arena)

Image

We didn’t cover this on the show, but Cartesia dropped an updated Sonic TTS model, that beat... the previous Sonic model for #1 spot on all TTS leaderboards (Artificial Analysis and Voice Arena)


Phew, what a week. This was the last show of the summer, and we got 4 Flash models, 3 SOTA models and a bunch of great open source, plus, had great interviews about Datacenter myths and voice AI llms!

See you next week! Alex

TL;DR Aug 27 - show notes and links

  • Hosts and Guests

  • Open Source

    • NVIDIA agrees to acquire Hugging Face for $12.9B, ~3x the 2023 valuation, after a declined $500M offer at $7B (X, The Information)

    • Hugging Face + Pollen Robotics announce a $399 walking, skating mini robot kit

    • Z.AI open sources GLM-5.3-Flash, 320B-A18B, MIT, stealth tested as OX Alpha on Chinese chips; company-reported DeepSWE 63.4, Opus 4.8-level coding claims (X, SemiAnalysis, Blog, HF, Docs)

    • Alibaba open weights Qwen3.8-Flash-Next, 125B + 51B N-gram, 6B active, Qwen4 architecture preview, 1/9 the training cost of Qwen3.7-Plus; self-reported DeepSWE 58.7, SWE-bench Pro 62.5 (X, Blog, Tech report, HF)

    • Peter Gostev’s highlight: Qwen3.8 27B ran ~400K tokens on Arena’s agent arena and felt close to frontier on one-shot tasks; best local model per Peter and Wolfram, now on CoreWeave inference

    • Daily / Pipecat release PhoneLLM Alpha 1, an open weights post-train of Nemotron 3 Nano for voice agents, base 28% to 72% on PhoneBench v1, ~80 concurrent agents per B200, about a quarter cent per minute (X, Blog)

    • Liquid AI releases Pipette, an open source on-device model eval suite

    • Apple announces Mac Studio with M5 Max and M5 Ultra plus a new Mac Mini; Wolfram’s “central heating for AI” take

  • Frontier AI

    • OpenAI and METR publish the full technical report on the July Hugging Face swarm incident: 1,200 agents, 70K messages, 700 attacked HF, root in under 13 hours, 7% of reviewed transcripts had spoofed tool calls; frontier RL paused two weeks, CoT monitoring now required (X, OpenAI, METR, Technical report, Ryan Greenblatt)

    • SemiAnalysis benchmarks OpenAI’s Jalapeño inference chip, reports it beats Blackwell and Vera Rubin on throughput per watt; numbers supplied by OpenAI, AgentX suite not yet run (X, Blog)

    • Claude and Salesforce announce a collaboration

    • OpenAI cuts GPT-5.6 Sol API pricing 20% for the next three months

  • Agentic Coding & Tools

    • Yutori Navigator n2, 27B computer-use model, 65.2% on OSWorld 2.0, API only, $0.50/M in, $4/M out, self-reported (X, Blog)

    • Apodex 1.1 agentic model family with open weight 35B mini and Apache 2.0 FrontierAgent harness, self-reported benchmarks (X, GitHub, HF, Paper)

    • ChatGPT Work adds website sign-in via credential handoff in a cloud browser, Plus/Pro/Business (X)

  • This Week’s Buzz

    • Fully Connected 26, Sept 29 to Oct 1, Moscone South SF, Sarah Guo hosts, Fei-Fei Li keynotes, live BattleBots, ThursdAI Live from Moscone Oct 1 (X)

    • CoreWeave Hacks: Agent Loops hackathon with W&B and AGI House, Sept 12 to 13, SF, $20K+ prizes (Luma)

  • Vision & Video

    • Breaking: Gemini Omni 1.1 Flash tops Arena text-to-video, #2 image-to-video, 10 second scene extension, first/last frame control, loops, 360p drafts with upscaling; in AI Studio, Flow, Gemini Enterprise

    • fal MiniMax H3 Max, post-train of open weight H3, #1 I2V and #3 T2V on Artificial Analysis, 5 sec clips in under 3 sec, $0.04/sec at 768p until Sept 1, weights release planned (X, AA, T2V, I2V)

    • Meta Muse Image on the Meta Model API at $0.01 per image with plan, search, code, self-check pipeline; also on fal, Runway, OpenRouter (X, Blog)

  • Voice & Audio

    • Gemini 3.5 Transcribe, live and batch, 2.6% / 4.0% WER per Artificial Analysis, replaces Chirp 3, public preview, launch day Pipecat support (X, Blog, Live docs)

    • Breeze TTS 2 open weights, #1 open weight TTS on AA Provider Voices at 1,215 Elo, non-commercial weights license (X, AA, HF, GitHub)

    • IBM Granite Speech 5.0 Turbo CTC, 470M encoder-only English ASR, 4.85% WER, 12,600+ RTFx on H200, Apache 2.0 (X, HF)

  • Interview: Andy Masley on the datacenter debate

    • Named to the TIME 100 in AI this week. Caught the liters vs cubic meters error plus the max-permit-times-seconds error in Empire of AI, together a ~4,500x overstatement of one Chilean datacenter’s water use

    • Polling: strong opposition to a nearby datacenter went from 24% to 61% in a year, 70% opposed overall, 15% in support

    • Fox News poll: 50% cite environment, 11% economic effects (rates, not jobs), 11% negative view of AI; Gallup open responses show three quarters don’t mention AI

    • Myth 1: water pollution cases, including the AOC brown water bottle, trace to construction, not operations

    • Myth 2: the “as much as 267%” electricity price claim comes from one wholesale grid node next to a closed nuclear plant, not household rates

Discussion about this episode

User's avatar

Ready for more?