Hey everyone, Alex here 👋
Summer is over. Wolfram said it in the first minute of the show and he was right. In 48 hours Anthropic shipped Fable 5.1, Meta’s Muse Spark 1.3 caught up to Fable 5 on the Artificial Analysis index at a fifth of the price, Google shipped another Flash, 3.8 this time, Z.ai put the full GLM-5.3 weights out, and three labs shipped world models that run in real time. It seems that they all tried to send their best work before Astra drops.
This week’s ThursdAI was so long that I decided to split it into two episodes. This is the regular format you know and love. And OpenAI Astra is so good, it deserves its own episode, which you can find at thursdai.news/astra.
By the way, as you guys know, I test these models continuously on my own stuff, and this week I was able to build a live studio for the show, with real-time transcription and an agent producer, in about four hours with Fable 5.1. More on that in the Fable section.
Joining me: Wolfram Ravenwolf, Nisten Tahiraj, LDJ, Yam Peleg and Peter Gostev. Plus, Ryan Carson hopped back to chat about Astra in the second part! Let’s get into it.
Frontier AI: the .1 week
It looks like all the frontier labs tried to ship something before OpenAI dropped Astra.
Fable 5.1, the SOTA LLM until a few hours ago, and it fixes the jargon douche problem (X, Blog, System card, EFS)
This was the story of the week until noon on Thursday, and it’s still my favorite model to use. Fable 5.1 and Mythos are the same weights, Fable is the one we actually have access to. OpenAI, and from this week Google, seem to converge on the same strategy.
Anthropic’s numbers: Terminal-Bench 4.0 goes to 55.8% from 42.0 for Fable 5, Terminal-Bench Science more than doubles to 52.6%, and SWE-bench Pro lands at 81.2.
Price stays at $10 and $50 per million, and the number that matters if you build agents is cache reads down 75% to $0.25 per million. Anthropic says that makes typical workloads about 25% cheaper and heavy agentic ones up to 45%, but that wasn’t proven, and folks complained about draining quotas!
Peter’s counterpoint from actually running it: his front-end generations on Code Arena cost $40 to $60 each where Sol cost $3 to $10, and the Max version still came in first on Code Arena by a large margin. His point, and mine: with a model like this we need to imagine bigger and be more ambitious. More on that in a second.
Mannered prose, finally acknowledged
We finally have acknowledgment from Anthropic that this was a problem. For months I called the way Opus 5 speaks “jargon douche” (my post on it): everything was load-bearing, everything was a control plane, every problem was a pain point. Not only did they fix it with Fable 5.1, they gave it a name. Anthropic’s prompting guide (Writing density) calls it mannered prose, and it comes with a fix: add it to your personalized settings, or just ask Claude to not use mannered prose.
I said on the show that Fable 5.1 is the best writer I have used. It’s still AI writing, you can feel it a little, but it’s concise in a way no earlier Claude was, and the jargon is gone when you ask. The one thing to watch is that it’s trigger-happy: ask it to plan something big and it will, then ask a simple follow-up and it answers with the same intensity, writes scripts, runs them. You have to tell it when you’re just making a comment between colleagues.
We’ve been testing the Mars mass driver launch on every model for over three years, and this was by far the best one we’ve seen. Two prompts, and it built more than just Mars: the whole solar system, a textured Earth, a mission planner, an autopilot, and we could land the thing! It was mind-blowing.
How I built thursdai.news/live in one sitting
As these models get more capable, we talked on the show about needing to be more ambitious. The day before the show I was playing around with Muse Voice Transcribe, the new model I’ll mention below, and Fable 5.1, and I wanted to do something very ambitious. So I asked GrokBot: how long would it take to build a live page for you guys to watch our stream, so that GrokBot could be our producer, put up chyrons and highlight the topics we’ve covered? GrokBot said it’s going to take a while. So I just YOLOed into Claude Design with Fable 5.1 and built a design for this, then went to Claude Code, entered plan mode, built a plan, and handed it off to three agents in Cursor.
I never wrote a line of code, and the whole setup is significantly more than a Three.js demo. This is a real working three-part system: a website, streaming video on Cloudflare, and streaming transcription that gets read by a bot, which can control our show. I think I’ve hit around 400 million tokens, if not more. Yam asked me on the show how I did this, so I decided to tell you guys here. I am mind-blown that this was possible, and after four hours I was able to go to thursdai.news/live and actually see it working.
Meta Muse Spark 1.3 catches Fable 5 at a fifth of the price (Zuck, AA analysis, AA model page)
As we say on the show, don’t bet against Zuck. The MSL folks have been on a tear lately, and this is the fourth Spark .1 version in around five months.
More than how this one model performs, look at the jumps in capabilities from version to version. This is the first time that MSL is showing up as a frontier lab, because an unreleased version of Spark with max reasoning beats GPT-5.6, Grok 4.6 and company, and lands around Fable-level capability. On the AA index the version you can use today scores 61, the max preview scores 62, Fable 5.1 sits at 66. Now, it doesn’t mean this model is that good, but there are a few more things here.
The gains are mostly agentic: banking-style tool use, terminal work, GDPval. The asterisks are that it thinks more, so cost per task went up, and AA’s own long-context test regressed a bit. Meta’s own chart looks rosier than AA’s
Then I asked the panel who’s using it. Nobody raised a hand.
Wolfram plans to put a bot on the contributor tier for open source work. That’s $0.10 in and $0.20 out, if you’re fine with Meta training on your prompts. Nisten wants it as a cheap verifier for the medical datasets he builds, because he needs something that isn’t Fable or a Chinese model trained on Fable. LDJ tried it on interface building and creative writing and called it pretty good, with its own taste.
The exciting part: Open weights and a model codenamed with a 🍉 are “coming soon,” and nobody knows what the watermelon is but it’s very exciting!
Gemini 3.8 Flash and 3.8 Flash Cyber: another Flash, and the price doubles in January (X, Cyber thread, Fairwind, Pricing)
Google’s turn. 3.8 Flash lands three weeks after 3.7 Flash. HLE-Verified 54.9, 1M in and 64K out, same $0.75 and $3.75 as 3.7, live in AI Studio, Antigravity and the Gemini app.
The underreported line is on Google’s own pricing page. On January 1, 2027, both 3.7 and 3.8 Flash go to $1.50 and $7.50. That’s double.
The WSJ reported that Google scrapped its 3.5 Pro checkpoints because Flash kept overtaking them, and Gemini 4 is still in post-training. Wolfram, our resident Gemini user, put the update straight into his home assistant and still asked the question everyone asks: where’s the Pro?
3.8 Flash Cyber is Google’s version of the Mythos split. CWE-Bench 47.2% at $3.64 per rollout, against Fable 5’s 47.8% at $10.27 (Artificial Analysis ran it), and 2.6x more valid patches for the Chrome team. It’s only available through the Fairwind Program, 650-plus vetted partners, governments and critical infrastructure, background check included.
Qwen3.8-Max-0902 claims the Code Arena crown (X, Arena, QwenCloud)
Alibaba updated its API-only Max model: 2.4T MoE, 1M context, post-trained on coding and “cowork,” number one overall on Code Arena with a WebDev Elo of 1691, at $2 and $6. A third-party DeepSWE run puts it at 56.6 behind Sol’s 73, so the number one is a front-end number one, not an agentic coding one. I asked the panel if they know anyone using Qwen Max through the API. Nisten knows one IT guy running OpenClaw on it and some people generating datasets. That’s the honest read on where it sits outside China.
Also from the frontier: Elon says Grok 4.7 lands next week, which makes xAI the one lab that didn’t ship before Astra.
Open Source LLMs
Wolfram’s correction when I called this a quiet open source week: we are so spoiled. He’s right.
Z.ai opens the full GLM-5.3 weights (custom license, not MIT) (X, HF, Blog)
We covered GLM-5.3-Flash last week as the OX Alpha mystery model. This week the full 753B model with 40B active got its weights on Hugging Face, under a custom “glm-5.3” license rather than the MIT the Flash version shipped with, so read it before you call it fully open. LDJ’s correction on air: the model itself isn’t new, we covered it, the open weights are the news, and that is a big deal because people can run it on their own rigs now. Z.ai’s own numbers: CyberGym 84.5%, above Fable 5 and Sol, ExploitBench 54.4 (Fable 5 is at 78), Terminal Bench 3.0 up to 28.3 from 5.2’s 4.6, and a claim of 2,436 real vulnerabilities found across 269 open source projects, the oldest from 1981, 53 disclosed so far.
Also on this base: we interviewed the co-founder of Abliteration AI, the folks who went viral by providing a product where they took GLM-5.3 and removed the refusals for anything besides CSAM and self-harm. We actually had this person, who asked to remain anonymous, as a guest on the show. Definitely check out that conversation, it’s very interesting. More on that below.
Tencent Hy4 preview: 770B, Apache 2.0, and a quant that fits it in 214 GB (X, Sherry, HF, Blog)
A 770B MoE with 49B active, 1M context, Apache 2.0, at $0.834 and $2.501 per million on the API. I wouldn’t put Tencent in the top tier of Chinese labs with DeepSeek, Z.ai, Alibaba and Moonshot yet, and their benchmarks are image-only charts, so the only number I’ll quote is theirs: a blind eval by 163 internal experts rated it “slightly ahead of GLM 5.3 and Kimi K3” on 203 engineering tasks.
The interesting part is Sherry, their quantization that takes the 1.5 TB of weights to 214 GB at 2.38 bits per weight, running at 205 tokens per second prefill and 20 decode on eight H20s. Nisten, who does one-bit models at Prism ML, gave the necessary asterisk: two-bit on a model this big keeps something useful, and it tends to drop things like multilingual ability that the headline benchmarks don’t measure. Test it for your use case and expect losses elsewhere.
OpenAI pulls its models from Cursor, and Wolfram’s case for open harnesses (OpenAI)
Wolfram raised this in the open source segment on purpose. OpenAI no longer allows its models in Cursor, now that Cursor belongs to SpaceX. Anthropic set the precedent when it pulled Claude from Windsurf during the OpenAI acquisition rumors, but Cursor’s whole pitch was that no model lab owned it, so you could use every model in one place.
Wolfram called it a bad precedent. If providers can decide “I don’t like you, you don’t get the model,” you want open weights you can host anywhere and an open harness that can swap models. My read on air was that the battle lines are being drawn: Anthropic buys GPU capacity from SpaceX, OpenAI is aligned with Microsoft, and Jensen is aligned with everyone, since he just made the Hugging Face acquisition official at $12,930,300,000, which we covered last week (last issue).
This Week’s Buzz 🐝: Kimi K3 on CoreWeave, Fully Connected, CoreWeave Hacks (Kimi K3, Deploy docs, Fully Connected, CoreWeave Hacks)
Kimi K3 is now live on CoreWeave Dedicated Inference, on GB300 NVL72, and it purrs like a kitten. Check it out
Fully Connected 26 is September 29 to October 1 at Moscone South in San Francisco: three days, 32 sessions, 2,000-plus people, Sarah Guo hosting, Fei-Fei Li keynoting (you’ll see why that’s timely in the world models section), live BattleBots, and ThursdAI broadcasting live from the floor. The regular ticket is $1,299 and early bird is over. On the show I dropped a code for a 100% free ticket for people who follow ThursdAI, and it’s here too: THURSDAIFC2026 (register here)
Before that, CoreWeave Hacks (formerly WeaveHacks) runs September 12 and 13 in SF with Weights & Biases, AGI House and Typesafe AI. The theme is Agent Loops, build agent loops that catch their own mistakes, with $20k-plus in prizes, a robot dog for best loop design, Formula 1 tickets for the most production-ready hack, and a Fully Connected ticket for attending. Apply on Luma and say you’re with ThursdAI, we’ll let you in!
Voice & Audio: two closed ASR models in one week
Meta Muse Voice Transcribe, and why this transcript has names on it (X, Architecture, Zuck)
Meta Superintelligence Labs shipped its first audio model, and I took it for a test drive. It’s pretty incredible. Muse Voice Transcribe does streaming ASR, diarization for 20-plus speakers and endpointing in one model from the Muse Spark family: 80ms audio chunks, one token each, with an adaptive delay so it waits on hard words and commits fast on easy ones. Meta claims 3.1% streaming WER and 17.5% diarization error on Artificial Analysis. 70-plus languages, 25 validated, code-switching. API only, no weights and $0.18 per hour make this model a no brainer!
On air you could watch it work. As I spoke, thursdai.news/live labeled me, then Wolfram when he interjected, then Yam and LDJ, and it got “ThursdAI,” “GPT-5.6 Sol,” “Alex Wang” and “Scale AI” right because we gave it a keyword dictionary (Wolfram asked, and yes, it takes one). The speaker names are a second trick: Fable set up a voiceprint for each co-host the night before, and the site matches the live diarization against them. For three and a half years I labeled speakers by hand in Descript every week. In 2026 you should not be doing that manually, and now I’m not. At 18 cents an hour, the whole show cost under a dollar to transcribe. Wolfram was the only one on the panel as excited as me, because he already runs voice agents and knows that who-spoke-when is the unsolved part.
Microsoft MAI-Transcribe-2, #2 on the leaderboard at less than half the price (Launch, AA thread, Leaderboard)
Microsoft AI announced this the morning of the show, with the claim of the highest quality and cheapest transcription at the fastest speed, 10x faster than GPT-Transcribe, live on Microsoft Foundry. Artificial Analysis had the independent numbers within the hour: second on the word error rate board at 2.0%, about 400x real time, $1.67 per 1,000 minutes, less than half the price of its peers, with diarization and 60 languages.
I tried it after the show and I was blown away. It took the 90-minute Astra episode and transcribed it in 15 seconds, fillers and all . It’s a batch model, not streaming, so it’s a different board from Meta’s, but two labs shipping ASR with diarization in the same week tells you where the agent builders are pushing.
Inworld Realtime TTS-2 goes GA (X)
Sub-100ms time to first byte at $25 per million characters on demand, a Flash variant at 25ms for $15, free-form stage directions instead of preset emotions, one voice identity across 200-plus languages. Inworld’s “#1 on Artificial Analysis” is on the Controlled Voice Arena. On the Provider Voice Arena, TTS-2 Flash sits fourth behind Cartesia Sonic 3.6. Say which board.
Completely uncensored: the founder of Abliteration AI on the refusal-free GLM-5.3 (Launch, Docs, Pricing)
Three labs spent the week telling you the powerful version of their model is for vetted defenders only, and Abliteration AI is the opposite bet: GLM-5.3 with the refusals removed, hosted as a US-based API at $5 per million, with only CSAM and self-harm hard-blocked. The founder joined us anonymously, and it’s a very interesting conversation about who actually buys this (agent red-teaming for banks first, then cyber, then trust and safety teams), why they don’t do KYC, and why they think gating frontier cyber models to big known names leaves every small security shop behind. Nisten and Wolfram pushed back and agreed in equal measure. Go listen to it, it starts at 59:08 in the video, and my take from the show stands: this is inevitable and already happening inside every serious offensive security shop, this founder just did it in public.
AI Coding & Agents
Muse Code is out of beta (Zuck, Pricing, Blog)
Meta’s coding agent went GA on August 31, ahead of Spark 1.3. One-command install , plans at $5, $20 and $50 a month, a TypeScript SDK preview over the Muse Session Protocol, multi-agent workflows, inter-session messaging and rewind. API pricing is the standard $1.25 and $4.25, or the contributor tier at $0.10 in, $0.20 out and a fifth of a cent cached if you let Meta train on your prompts. Cheaper than everyone, and it only runs Meta’s own model. My jest that’s also true: if you have an Instagram account, you already let Zuck train on your data.
OpenClaw 2.0 gets native computer use through Cua (X, Cua, Blog)
OpenClaw’s biggest release, and the part I care about is the Cua integration, first-class computer use through the Cua Driver SDK (we had Francesco on the show when they shipped background computer use, first after OpenAI), plus cloud fleets of Linux desktops an agent can see and click, a rebuilt browser app, and support for pretty much every OS you own. Wolfram asked the audience who still runs it since he left over instability, and the comments said they’d moved to Codex Mobile and Claude’s mobile app during the two-month release gap. I agree that both got a lot better, everything I start on my desktop now shows up in the Claude app, but shout out to the maintainers regardless.
Vision & Video: the world models went real-time
Three labs, three world models, one week. Wolfram called this the Stable Diffusion moment for video, and by the end of the segment I agreed.
World Labs Atlas: bullet time from three phones (X, Blog)
The one that left me speechless on air, and the best world model demo I have seen. Atlas is a multimodal autoregressive diffusion transformer that World Labs pretrained from scratch on text, images, video, camera poses and depth. Give it one to six reference images and it generates up to a minute of 1440p video with exact camera control. Give it more (over a hundred in one spatial context) and it reconstructs the scene into frames, depth maps, point clouds or Gaussian splats you can walk through. Marble rendered splats, Atlas generates them.
The demo that got me: a watermelon smashed in front of three ordinary phone cameras on tripods, and Atlas replays it from any angle, every drop, the Matrix bullet-time shot with no rig. It also builds Real-to-Sim robot training environments from about 24 phone frames. Partner early access only, no weights, no price, no parameter count.
I said ont he show that I’m excited that dr Fei-Fei is keynoting Fully Connected in four weeks and I get to hear her talk about this on stage.
Runway Solaris: a world model for interfaces (X, Cristóbal, Blog)
No HTML, no CSS. Solaris is Gen-4.5 distilled into a real-time autoregressive frame generator, and the frames are the interface: a photo of a living room where clicking the lamp turns it on, a guy whose shoes you can drag onto him, ingredients on a table you cook by dragging.
Runway’s own 250-person study preferred it to a coded Claude Opus 5 result 61 to 24 on following instructions and 71 to 21 on natural behavior, with the limits listed: unreliable text, drift in long sessions, no accessibility APIs. Wolfram thinks this is where all interfaces go, generated on the fly and changed by asking. I think it’s one of the most important things this week because it’s how people learn, by touching things. Cristóbal, please let people play with it. If the problem is GPUs, talk to us.
Runway GWM Worlds 2 dropped mid-show, with sound (X, Research)
LDJ broke this one live, three days after Solaris: a general world model that generates a steerable, open-ended simulation at continuous 720p, 24 frames per second, with 48 kHz audio and no fixed clip length. You define the world, then talk to any subject in it or move the camera, and it continues from your input.
We played the speech demo, asking a generated stranger for directions to the transit station and getting an answer. LDJ’s note: most world models with audio so far had low-resolution, obviously synthetic sound, and this is a real jump. Nisten wanted her to reverse a binary tree. Research preview, early-access form, and Runway’s own caveat that real-time still trades fidelity for speed. Two Runway drops in three days, and I’m still not allowed to touch either.
fal’s H3 Max week: banned twice, built its own streaming site, shipped Turbo at a cent a second (Turbo, fal.live, Rehan, FastH3)
Last week we told you fal’s post-trained MiniMax H3 Max generates video faster than it plays. This week fal piped it into a Twitch stream as “infinite interdimensional cable,” an endless Rick and Morty channel steered by chat, and Twitch banned it for copyright within an hour, then Kick did the same. So fal built its own streaming site over the weekend. fal.live went up on August 31 on a checkpoint tuned for continuous generation, with channels (anime, sitcom, chaos) where viewers vote on the next scene, and it’s the reason I thought I could build a live site in a night too. Then on Tuesday they shipped H3 Max Turbo: 2x the speed at half the cost, targeting the 97th percentile of H3 Max quality on fal’s own evals, at a promo price of one cent per second of 768p video. Rehan’s demo is a five-second clip in 1.4 seconds.
I generated a Big Bang Theory scene live on Turbo and it came back with generic actors, and on air I said fal pulled a fast one on us. A correction, which I posted after the show: that was MiniMax’s prompt expansion doing it, not fal, and Batuhan from fal set me straight . The speed is the real story, a 15-second video generated in nine seconds, which is ridiculous. Wolfram admitted it’s so addictive you keep regenerating, and he has spent a lot at fal this month. He also runs H3 locally with a fast LoRA.
Wrapping up
I really enjoyed Fable 5.1 this week, and the Atlas demo is the thing I keep replaying. Everything else in this issue happened before noon on Thursday. Then OpenAI shipped GPT-6 Astra while we were live, Peter showed what it can do, Ryan came back to the show for it, and the stream ran to five hours. That’s the other episode, a separate video and newsletter, and if you want the ARC-AGI-3 number and the AGI argument, go there next: thursdai.news/astra.
Next week: Grok 4.7 if Elon keeps his word, the Muse Spark open weights watch, and hopefully Astra in our own hands. Come hack with us at CoreWeave Hacks on the 12th, and grab the free Fully Connected ticket while the code lasts.
If you missed any of it, ThursdAI is a live show at 8:30 AM Pacific every Thursday, a podcast, and a newsletter. Subscribe to one, then go check out the others. And come watch the next one on thursdai.news/live, where the transcript will have your co-hosts’ names on it. See you next week.
TL;DR and show notes
Hosts and Guests
Alex Volkov - AI Evangelist, Weights & Biases & CoreWeave (@altryne)
Co-hosts: @WolframRvnwlf, @nisten, @ldjconfirmed, @yampeleg, @petergostev
The founder of Abliteration AI, who joined anonymously
GPT-6 Astra
Launched mid-show, covered in full in its own episode (thursdai.news/astra)
Frontier AI
Anthropic Claude Fable 5.1 and Mythos 5.1, same weights: Terminal-Bench 4.0 55.8 vs 42.0, cache reads down 75% to $0.25/M, and “mannered prose” gets a name in the prompting guide (X, Blog, System card, EFS, Writing density)
Meta Muse Spark 1.3: xhigh scores 61 on the AA index (ties Sol and Grok 4.6), limited-preview max 62 (ties Fable 5), unchanged $1.25/$4.25, open weights and a watermelon model “coming soon” (Zuck, AA analysis, AA model page)
Google Gemini 3.8 Flash and 3.8 Flash Cyber: HLE-Verified 54.9, $0.75/$3.75 until a doubling on Jan 1, 2027, Cyber is Fairwind-only with CWE-Bench 47.2% (X, Cyber, Fairwind, Pricing)
Alibaba Qwen3.8-Max-0902: 2.4T, 1M ctx, #1 on Code Arena, $2/$6, API only (X, Arena, QwenCloud)
Grok 4.7 lands next week, per Elon
Open Source LLMs
Z.ai releases the full GLM-5.3 weights: 753B/40B active, custom glm-5.3 license, CyberGym 84.5% claimed, 2,436 vulnerabilities found (X, HF, Blog)
Tencent Hy4 preview: 770B/49B active, 1M ctx, Apache 2.0, Sherry quant takes 1.5 TB to 214 GB at 2.38 bpw (X, Sherry, HF, Blog)
Abliteration AI abliterated-model-large-v2: refusal-removed GLM-5.3 as a hosted API, $5/M, only CSAM and self-harm blocked (X, Docs, Pricing)
OpenAI pulls its models from Cursor after the SpaceX acquisition (OpenAI)
NVIDIA makes the Hugging Face acquisition official at $12,930,300,000 (Clem, last week’s issue)
This Week’s Buzz
Kimi K3 (2.8T) on CoreWeave Dedicated Inference on GB300 NVL72 (X, Docs, Dedicated Inference)
Fully Connected 26, Sept 29 to Oct 1, Moscone South SF, Fei-Fei Li keynotes, ThursdAI live from the floor, free ticket code on the show (X, Keynote teaser, Register)
CoreWeave Hacks: Agent Loops, Sept 12 to 13 SF, $20k+ prizes, robot dog, F1 tickets (X, Luma)
Voice & Audio
Meta Muse Voice Transcribe: streaming ASR, diarization and endpointing in one model, 3.1% streaming WER claimed, API only, powers thursdai.news/live (X, Architecture, Zuck)
Microsoft MAI-Transcribe-2: #2 on AA WER at 2.0%, about 400x real time, $1.67 per 1,000 minutes (Launch, AA thread, Leaderboard)
Inworld Realtime TTS-2 GA: sub-100ms, $25/M chars, #1 on AA’s Controlled Voice Arena, #4 on the Provider arena (X)
AI Coding & Agents
Vision & Video
World Labs Atlas: up to 1 minute at 1440p from 1 to 6 images, 3D reconstruction, bullet time from three phones, partner access only (X, Blog)
Runway Solaris: Interface World Model, UIs generated frame by frame, preferred 61 to 24 over Opus 5 in Runway’s own study (X, Cristóbal, Blog)
Runway GWM Worlds 2: real-time 720p, 24 fps world model with 48 kHz audio and open-ended sessions, research preview (X, Research)
fal: infinite Rick and Morty stream banned from Twitch and Kick, fal.live built in a weekend, H3 Max Turbo at $0.01/sec, open FastH3 (Turbo, fal.live, Rehan, FastH3, HF)
H3 World: an open LoRA that turns MiniMax H3 into a walkable world model (Github)
Guest




























