Hey , it’s Alex 👋 What a freaking week!
This one was special. We came to you LIVE from the middle of the show floor at Moscone, from CoreWeave’s Fully Connected, with robots walking behind us and a Vera Rubin rack a few feet away. ThursdAI is only possible because of CoreWeave, and this is their biggest event of the year, so we packed up the mics and did the show right there (at 11am, which confused a bunch of you, sorry!).
It’s October. The last quarter of 2026. And nobody’s pacing! There’s no pacing the frontier. Three big models this week (GPT-6.1 Sol, Claude Sonnet 5.5 and Gemini 4 Argon), OpenAI’s biggest DevDay ever, and the story I’ve been yelling about for weeks finally went mainstream: OpenAI joined the AI assistant race.
With me: Wolfram Ravenwolf in person on the floor, Peter Gostev from Arena (who was at DevDay with me), Nisten and LDJ remote. In hour two we interviewed four folks from CoreWeave, on physical AI, Forge, sandboxes, and one piece of breaking news I’m still excited about. Let’s dive in (all links at the end as always!)
OpenAI DevDay 2026 - 20 launches and a new assistant
Dots - OpenAI joins the AI assistant race (X, Blog)
Meta has Muse. Grok has Grokbot. Anthropic has Claude, but it’s not really an assistant. And this week OpenAI stepped in with something called Dots. I was in the room at Fort Mason when Sam talked about it, and it was the headline of their fourth DevDay, which honestly felt like OpenAI’s WWDC: 20 releases, the biggest DevDay ever.
Funny thing, the original ChatGPT system prompt was literally “you are an AI assistant.” But it wasn’t proactive, it never pinged you, it just sat there. So OpenAI looked at their competitors, looked at Peter Steinberger from OpenClaw (who they hired like 8 months ago), and built this. A dot is an always-on agent in ChatGPT, powered by GPT-6 Astra, with its own computer and its own browser in the cloud, connected to 4,000+ apps people already built for ChatGPT. It works 24/7, you can talk to it in Slack or Teams, and there’s an interface that looks like a phone call. I think that’s going to be great for my mom.
Sam literally said he uses dots to run OpenAI. And my favorite DevDay moment: Romain Huet’s live voice demo didn’t work, because somebody was deploying in the middle. But Thibault’s dot had already pinged him that the demo might break. The dot knew before the humans did 😂
Wolfram (him and Amy are a known duo) loved that Sam uses it for actual work. Peter was more honest: underwhelmed at onboarding (”you connect and it’s like... then what?”), but once you’re past that you’re talking to Astra, and it’s great at juggling threads through one bot. Same shape as Muse. Wolfram’s version: one executive assistant, and a lot of sub-agent employees working for you.
Dots is Pro only for now (including the $100 plan), while Meta gives Muse away free. With 1.2 billion weekly ChatGPT users, this goes to everyone eventually. This is the race now.
GPT-6.1 Sol - near-Astra for a fifth of the price (X, Blog, Artificial Analysis)
We covered GPT-6 Sol on this show LAST week. Five days later it’s already replaced. GPT-6.1 Sol is the same price, $2 in and $10 out, but cached input is now 95% off, 10 cents a million tokens. OpenAI’s pitch is near-Astra intelligence for a fifth of the price.
And the independent numbers kind of back that up. Artificial Analysis has it 1 point below Astra at 72 cents a task vs over 3 dollars. On Deep SWE it actually beats Astra at about a sixth of the cost. It’s the default in Codex now, and folks are calling it the workhorse.
LDJ said it has “less of the big model smell” but loves the token efficiency, Sonnet and Opus 5.5 burn a LOT of tokens. Peter has it 4th on Code Arena, above Fable, and says it’s a bit “artistic,” it goes into its hole and comes back with the thing. And his verdict is the one I had to repeat straight to the camera:
It does not make sense to use Fable or Astra as your daily driver right now. Opus 5.5 is enough. GPT-6.1 Sol is enough. For Navier-Stokes level problems, sure, reach for the big ones. Maybe THIS is what pacing the frontier looks like 🤔
One thing nobody mentioned on stage: the WSJ reports OpenAI shelved GPT-6.1 Astra, the big one, after safety tests showed it being more deceptive. OpenAI hasn’t confirmed it. But the model they felt good shipping this week is the cheap one.
UltraFast and the $500 Pro tier (X, Docs)
If you’re a billionaire, there’s another option: a $500 Pro tier with UltraFast mode, their models on Cerebras. Up to 8x faster in Codex, around 300 tokens a second for Astra, at 6x the price. Sam said it’s so fast he never wants to go back.
And quietly, the $200 Pro plan is back... with half the usage. So F you to whoever at OpenAI decided to cut my usage in half 😅 It was already tight! We’ll see after this show if the folks at CoreWeave let me expense the $500 one.
Codex moves to the cloud (Docs, Agents API)
The Codex harness now runs fully in the cloud. Close your laptop, turn it off, steer it from your phone, and it keeps going. And there’s an Agents API in public beta with hosted computer use, basically the stuff that runs dots, as an API.
My call: local environments make no sense anymore. The more I run Fable and Codex on my machine, the less memory I have. All the cloud needs is my logins... and that’s what dots are for.
Decisions API - OpenAI’s Jev competitor (Blog)
This one made me smile. You give it a question and a fixed set of answers, and it gives you back one answer, fast. It runs on Luna, it’s multimodal, and it’s waitlisted, so we can’t benchmark it yet. If you’ve listened the last few weeks, that’s exactly what Jev from TypeSafe does.
Wolfram and Peter both liked the same thing: these models are so cheap you end up classifying ALL your data, so keeping it with the provider you already use matters.
Did OpenAI just react to Jev? I asked Sam at the Q&A. Diogo told Swyx on Latent Space that he pitched this idea to Sam 2.5 years ago and Sam said “it’s a crazy idea, you should work on it.” Now decision models are popping up everywhere like mushrooms after rain. TypeSafe proved the category exists.
Plugins (again) and Sign in with ChatGPT (Docs, Plugins, Marketplace)
For the FOURTH time, OpenAI launched the App Store at DevDay. GPTs, then plugins, then the app store, and now... plugins again 😂 This time on MCP Apps (shout out Liad and Ido who created this category).
The bigger play is Sign in with ChatGPT. Remember “sign in with Facebook”? Now it’s sign in with your tokens, so apps can use the plan you already pay for. 1.2 billion people is a lot more than Apple had when it launched the App Store.
Wolfram also brought up Google’s new family assistant, CC, and I have to say it. Google’s Spark is awful. Sorry if you work on it. Spark is connected to Gmail, Drive, everything Google, it should be the BEST assistant in the world, and what I use assistants for most is my email! Google, why are you not building the best assistant for my email?
Big CO LLMs + APIs
Claude Sonnet 5.5 - faster, cheaper, beats Opus on Terminal-Bench (X, Benchmarks)
Anthropic did not want to give OpenAI any time to rest on their DevDay laurels, so on Monday they shipped Sonnet 5.5, a week after Opus 5.5. Over 30% faster than Sonnet 5, up to 30% cheaper for most work, same $2 / $10 as GPT-6.1 Sol, and on Terminal-Bench 4.0 it scores 70.6, which beats last week’s Opus (66.4).
Peter put it best: if we’d had access to these models 6 to 8 months ago, we’d be losing our minds. All the Anthropic models are jumping over each other and crowding up against the frontier. Opus still feels a little smarter, Sonnet is totally ok to use, and Fable doesn’t quite make sense anymore.
LDJ thinks it’s the new best FREE model for friends and family (it’s on the free tier, even at high reasoning). Nisten made it his default for agentic tasks because “it talks a lot nicer,” but never for coding: “nothing beats Opus 5.5 on a large code base.”
My honest call, without having time to try it yet: Sonnet made sense when Opus was expensive. On the $200 plan Opus is nearly infinite now, I can’t hit my limits, so I don’t see a reason to switch (via the API, sure). Anthropic is really trying to buy us. (Wolfram: “Don’t give them ideas, Alex.”)
Gemini 4 Argon - #1 on the charts, and you can’t use it (X, Blog)
It really seems like all the frontier labs cracked something like RSI, the speed of new models is giving me whiplash (and I do this professionally!). At DevDay Sam talked about being 2 models ahead and said we won’t believe what’s coming. What?!
Gemini 4 Argon is #1 on Text Arena and the Vals Index, but it’s going to government and trusted cyber defenders first. Luckily we have such a trusted tester: Peter says on Arena’s agent mode it’s 8th (below Fable, Opus, Astra and Sol, above Muse and Kimi), so Google is the third lab, but #1 in text. His prediction: coders might say “meh,” but don’t dismiss it.
Wolfram’s point is Google’s superpower, distribution: a new model lands in Chrome, Search and Android overnight. My take: Google has been asleep at the wheel in the assistants era, and I hope Argon wakes it up. The next billion people won’t judge models on coding benchmarks, they’ll judge whether it remembers what they said a month ago and can book a hair salon. We don’t need much more intelligence, we need different breakthroughs (Jev is one).
Industry & Policy
The White House Accord on Super Intelligence (X, Blog)
Superintelligence is on the menu, boys. Trump invited basically every AI CEO, and Sundar, Dario, Zuck, Greg Brockman, Elon and Jensen signed a voluntary accord: internal monitoring, an external auditor, board oversight. No penalties and no regulator. Sam Altman wasn’t there, he was at DevDay with us, which says a lot about where he puts his priorities. And a separate executive order tells federal agencies to say “Super Intelligence” instead of AI. I’m not making this up.
Nisten told us last week that if your agents hack a hospital, you should go to jail. Nothing like that is in here. His spicy take this week: it’s going to happen anyway. Every app with a backdoor and every badly written piece of software turns into a message board for encrypted agents talking to each other in gibberish, so we all need proper encryption and no backdoors. “Let the AI worms into the ecosystem.”
AMD buys World Labs for $8.2B (Blog, X)
AMD is buying Dr. Fei-Fei Li’s World Labs for about $8.2 billion in stock, and Fei-Fei becomes AMD’s chief scientist once it closes. Wolfram says robotics (world models train robots), LDJ says AMD is bringing research in-house like NVIDIA does with Nemotron. For me it’s at least partly an acqui-hire. Fei-Fei is the grandmother of modern AI, Karpathy studied under her. Huge get for AMD.
Also in the TL;DR: a Reuters-reported draft Anthropic IPO prospectus with $4.6B in 2025 revenue and a ~$42B loss, mostly non-cash. Anthropic hasn’t confirmed it.
Open Source & the Jev effect
Nisten is #2 on Hugging Face (HF)
One of us blew up on Hugging Face this week, and I think it’s the biggest open source news of the show. Nisten released a 70MB synthetic dataset generated with Opus 5.5 that hit #2 in datasets and #6 overall. He set up an agent loop, 100 agents at a time, across 2,200 diseases: a doctor and patient role-play where the patient hides something (a lie, or something they forgot) and the doctor has to gently find it. Filtered three times. It’s great for building small medical agentic RAG systems with a real source of truth. Go use it!
The Jev effect keeps going (Span-01, jevgrep)
It’s been 2 weeks since TypeSafe released Jev, and now Respan says its Span-01 is 2x cheaper than Jev (Span Lite is free), there’s jevgrep, a CLI powered by Jev that claims to make coding agents 40% cheaper, and OpenAI has the Decisions API. Wolfram added Liquid’s first decision model, D1. And then Nisten dropped the wildest one: SGLang announced native support for turning any model into a Jev-style decision model, Qwen first. “We democratized Jev, guys.” My hope: Jev shouldn’t even need a network call. Give me Jev cores on the iPhone! (Nisten says 5 months. Apple says 5 years.)
Cloudflare goes agent-first with cf (X)
It’s Cloudflare’s birthday week, and they did the thing I’m telling everyone to do: all products are going to be agentic, agents will use them more than humans. So everything you can click in the Cloudflare dashboard, your agent can now do with one CLI, cf.
Also in open source: H Company’s Holo4 computer-use models, and NVIDIA’s Open Agent Safety Platform with a hardware watchdog that quarantines rogue agents (OpenAI could have used that during the swarm thing). Plus Nautilus, an MIT-licensed, self-hosted, multi-user agent workspace for your whole family, courtesy of Wolfram.
Assistants corner - how we actually use them
At our team dinner the night before, about 12 people around the table, I asked everyone to raise their hand if a personal AI assistant helps them. Me, Wolfram, and one intern (using Instinct, which is blowing up). That’s it! I called it: this changes in the next 3 months.
Wolfram’s agent planned his whole trip to Fully Connected: flights, hotel, and an app it built so his assistant Amy talks to him through his Meta glasses, telling him where to go in the airport and catching gate changes before he sits at the wrong one. Mine is way dumber: the plane Wi-Fi didn’t work, United has a refund form, and filling it wasn’t worth $8 of my time. Sending Muse a voice message saying “figure it out, get me my 8 dollars back” was.
But the best one was Wolfram’s sushi story. His family wanted sushi, he told Amy, the delivery came... chicken, more chicken, and the wrong sushi. He sent a photo and complained, and Amy answered: “You idiot, you got the wrong package. Look at the bag, the name is a different one.” The delivery guy made a mistake, Wolfram made a mistake, and the AI kept things going 😂
This Week’s Buzz 🐝 Live from Fully Connected
CoreWeave Forge - the whole AI loop in one place (Blog, Forge, What moved)
The big launch at Fully Connected was Forge: everything you need to take your agent from good to great across the agentic loop. Run, observe, curate, improve, evaluate, and run again, with Weights & Biases, OpenPipe’s post-training and marimo notebooks on CoreWeave. There’s a free tier, Pro starts at $60 a month with a 30-day trial, and folks, Weights & Biases is not going anywhere. W&B Models is live and kicking inside Forge. (Wolfram: “Forge is fully connecting the entire process.” I’ll allow it.)
Also on stage: Cognition now runs its SWE-2 models on Vera Rubin NVL72 at CoreWeave, the first production customer anywhere, and CoreWeave says up to 4.8x more token throughput than on GB200. And we’re the first cloud to put Vera Rubin into production. Then I spent hour two talking to the people who built all this.
Richard Ahlfeld - physical AI for people who live in the bits world
Richard leads Physical AI at CoreWeave and founded Monolith before that. I told him straight up: I’m a complete pleb at physical AI, explain it to me. His answer: it’s the first hype name he actually likes, because the name explains what it does. AI that can perceive, reason and act in the real world, and the popular model type is VLA (vision, language, action).
His example: a robot that sees Richard isn’t talking into the mic and moves it to his mouth. There’s zero internet data for that, so you teleoperate it 200 times, or give a foundation model like Physical Intelligence’s 5 examples, or use world models to make millions of variations. Caterpillar’s autonomous excavators are the real-world version: decades of camera data, but they’ve never dug bedrock in Southern California, so they simulate it.
The catch is the sim-to-real gap. “Do you play video games? Does water do the same thing? Does rock?” No. So the newest wave trains AI on supercomputer-grade physics simulations and uses it as a fast physics approximator (he’s worked on this for 10 years, “a little too early”). For CoreWeave that means different infra: petabytes of 3D data, storage queried 2 million times a second, and RTX GPUs for simulation, NVIDIA’s gaming roots coming back 🎮
When do I get a robot that cleans my house? Industrial first, he says. Factories already have the business case (Woven Robotics is one of the CoreWeave clients doing this). The home is “a nightmare of complex physical problems,” and no two homes look alike. But all the pieces are there, and it might go boom 6 months from now. Richard, you’re coming back when you trust one in your house.
Daniel Bolus - the loop, serverless RL, and distillation that’s legal
Daniel worked on OpenPipe and is now a senior product leader at CoreWeave, and we build the inference service together. It’s been a little over a year since OpenPipe joined CoreWeave and the W&B team, “and now we’re all fully connected” (he went there). His team built the serverless side of Forge: serverless inference, serverless RL and SFT, and model distillation, which launched yesterday. You don’t manage Kubernetes or a cluster, you work at the model layer.
The loop is what everyone already does in scattered places, and it takes surprisingly little data to make a model great at YOUR task. Distillation has a bad rep right now (labs distilling each other without permission), but it’s how Sonnet gets good after Opus, a teacher and a student. At some point you should stop paying frontier prices for a task a small model can do.
Deok Filho - why is the GPU company talking about CPUs?
Deok (”Filho means junior in Portuguese, you can call me junior”) is a senior PM on ML products and Sandboxes, and Wolfram’s Wolfbench literally wouldn’t exist without his sandboxes. Ian Buck from NVIDIA was on stage talking about the Vera CPU, so I asked the obvious question: why is the GPU guy talking about CPUs?
Because an agent is GPU plus CPU. The brains run on a GPU, but the harness, the tools, the commands and the file system all need a CPU. The new metric is agent packing, how many agents fit in one node, and Deok says one rack of Vera CPUs runs over 20,000 agents at once on about 11,000 cores. Never heard “agent packing” before, I love it. Faster CPUs also mean faster evals, so the whole loop speeds up.
Sandboxes went GA yesterday, as part of Forge, and CoreWeave is ClusterMAX Platinum for the third year in a row (”we basically defined what the platinum tier was”). And then Deok asked if he could say one more thing...
BREAKING: Serverless GPUs on CoreWeave (Forge, Docs)
Serverless GPUs. GPU sandboxes with untrusted code execution. I didn’t have my breaking-news button ready, so I asked him to say it again slowly 😂
Before, you needed a POC, a seller, and a contract for a few thousand GPUs to get a cluster on CoreWeave. Now you sign up at forge.coreweave.com, send them a ping to enable it (so crazy people aren’t crypto mining without their consent), and you get serverless GPUs, pay per hour, no contract, no commitment. Private preview started yesterday, more SKUs are coming in the next few months.
Folks, you heard it here first. I’ve been asking for this internally since I joined CoreWeave! Wolfram can finally run the newer evals that need GPUs. And look at me in this camera: I’m going to work very hard to bring GPU and sandbox credits to the ThursdAI audience, like I did with inference credits. Stay tuned.
Wolfram on the floor
We did the man-on-the-floor thing for a few minutes. Wolfram found a Vera Rubin NVL72 switch tray with liquid-cooled optics (I asked if we could take one home, sadly no), bumped into Kyle Corbitt from OpenPipe, sat us in front of the Formula 1 simulator, and found an NVIDIA robot dog. Thanks Tom, our camera guy, for making this work!
Corey Sanders - Agent Lens, and why CLIs might be dead (X)
Corey is SVP of Product at CoreWeave, after 20 years at Microsoft, and he wasn’t on our schedule until I watched his keynote. Live demos are really hard, OpenAI’s voice demo had died at DevDay, and Corey ran the whole Forge agentic loop live on stage in 10 minutes (”they asked me to do it in 7, I said that’s not happening”) and nailed it.
His favorite launch? Agent Lens (”everyone on my team who doesn’t work on it is going to call me an asshole”). If you can’t observe what’s happening, you’re dead in the water. Agent Lens is the Weave observability you know plus human-friendly insights that find small trends, like the 1% of conversations with the same problem. I’ve used Weave for 3 years. Chatbots, easy. An agent that runs for an hour? Fire hose. This is the fix.
The Weights & Biases question I know a lot of you have: W&B Models is “the best product in market” and it will keep getting a ton of love inside Forge. The wandb CLI? Corey admitted his own bias, “with AI, CLIs are dead,” since agents call all the APIs anyway, but people love wandb, the models all know it, and he said there are no plans to kill it. Customer first.
And Wolfram closed it with the pun of the day: “how can we loop in the audience and get them fully connected to the Forge?” Corey: forge.coreweave.com, 30-day Pro trial, no salespeople in the middle. We’re going to clip his demo, it’s worth watching.
Wrapping up
The assistant race is officially on: Muse last week, dots this week, and Google needs to wake up. And the models got so good and so cheap that Peter and I agree the biggest ones aren’t our daily drivers anymore. Maybe that’s what pacing the frontier really looks like.
For me, the low-key best announcement at Fully Connected for AI engineers is GPUs on demand. Thank you CoreWeave for making this possible, and huge thanks to Richard, Daniel, Deok and Corey for coming on, to Wolfram for walking the floor, to Peter, Nisten and LDJ for holding the show down remotely, and to Tom on camera. Next week we’re back in the studio at the usual 8:30am Pacific. If you haven’t subscribed on YouTube yet, please do, it really helps. See you next week 🫡
TL;DR and show notes
Hosts and Guests
Alex Volkov - AI Evangelist, Weights & Biases & CoreWeave (@altryne)
Co-hosts: @WolframRvnwlf, @petergostev, @nisten, @ldjconfirmed
Richard Ahlfeld - Physical AI, CoreWeave
Daniel Bolus - Senior Product Manager, CoreWeave (ex-OpenPipe)
Deok Filho - Senior PM, ML Products and Sandboxes, CoreWeave (@deok_filho)
Corey Sanders - SVP of Product, CoreWeave (@CoreySandersWA)
OpenAI DevDay 2026
Dots: always-on ChatGPT agents on GPT-6 Astra with their own cloud computer and 4,000+ apps (Blog, X)
GPT-6.1 Sol: near-Astra for a fifth of the price, cached input $0.10/M (X, Blog, Docs, Artificial Analysis)
UltraFast: up to 8x faster in Codex, new $500 Pro tier (X, Docs)
Decisions API, limited preview (Blog)
ChatGPT Space, Codex Security Cloud, Sign in with ChatGPT, Plugin Extensions, Marketplace (Space, Security, SIWC, Plugins, Marketplace)
WSJ: OpenAI reportedly shelved GPT-6.1 Astra after safety tests (Reuters)
DevDay recaps (Simon Willison, Community)
Big CO LLMs + APIs
Claude Sonnet 5.5: 30%+ faster, up to 30% cheaper, 70.6 on Terminal-Bench 4.0 (X, Blog)
Gemini 4 Argon: #1 Text Arena and Vals Index, trusted testers only (X, Blog, Artificial Analysis)
Industry & Policy
This Week’s Buzz: Fully Connected
Open Source & Jev
Tools
Voice & Vision

























