ThursdAI Special - OpenAI's Romain Huet on Codex's 5M users, GPT-5.6 & the Golden Age of AI Engineering
From W&B / CoreWeave: A pre-recorded special from the middle of my 40th trip: 15 min of how I actually use AI every day, my conversation with OpenAI’s Romain Huet & Insecure Agents guest seats
Hey everyone, Alex here 👋
This week’s episode is a little different. As you’re reading this, I’m flying back from my 40th birthday trip with the family, and while the guys did end up having a great live stream (Huge thanks to Yam for hosting!), here I will bring you the episode I pre-recorded before leaving for the trip. However, tons of news happened this week, and as always, there’s a TL;DR section below with the top most important news in AI this week!
Here’s what’s on today’s special episode. First, a bunch of you have been asking how I actually use these models day to day, beyond covering the news. So I recorded fifteen minutes of exactly that: the papercuts I fixed with Codex and computer use, what Fable built for my kids, and how I’m rebuilding the ThursdAI design system and on screen elements.
Second, my conversation with Romain Huet, head of Developer Experience at OpenAI, recorded at the OpenAI booth in the middle of the AI Engineer World’s Fair floor.
And third, a throwback treat: Allie Howe invited Wolfram and me onto her Insecure Agents podcast as guests, and being on the other side of the mic was a delight. Let’s get into it.
How I actually use AI: a papercut-fixing spree
I promised a few of you I’d take time on the show to talk about the stuff I build and fix with AI, not just the news. So before the interviews, I recorded a segment walking through my last few weeks of daily AI use. Use the chapters if you want to skip ahead, but why would you?
Codex with computer use fixed every Mac annoyance I had
Once OpenAI launched GPT 5.6 Sol and dropped a pile of credits on those of us on the 200 Max plan, I went on a papercut-fixing weekend. The rule was simple: every little thing that has annoyed me about my Mac for years, I ask Codex to fix first, and only Google it if that fails. I never got to the Google step.
Chrome has no native copy-URL shortcut (seriously, Chrome, what are you doing?), so Codex found Karabiner-Elements already installed on my machine and wired up the shortcut itself. My 1Password has been showing “you’re offline” on every device for three months since CoreWeave moved us off the Weights & Biases account; Codex figured out in seconds that everything was actually syncing fine and the inactive legacy account was the only thing “offline.” Removing it fixed the whole thing. That is not an answer you find in a help center.
It kept going. My beloved window-moving utility Hummingbird had an expired license on my Mac Mini, so Codex built me a replacement app. It estimated one to two days for a polished version and finished in about fifteen minutes. It cleaned roughly 75GB of leftover model weights and junk off my Mac (I had it build me an HTML checklist first so I approved what got deleted). And the big one: I paired Codex with Home Assistant, the open source repo of the year as far as I’m concerned, and let it SSH in and go on a full optimization mission. Updates, error triage, cleanup, new connectors. If you’ve ever maintained a Home Assistant setup, you know how much joy and pain lives in that sentence.
One discovery worth passing along: I had /goal running when my credits hit zero, and Codex just kept going. OpenAI confirmed they care more about finishing your work than metering the credits mid-goal. Watching the meter hit 0% while the agent kept working was weirdly moving.
Fable Max built my family’s vacation
We all Fable-maxed when we thought Anthropic was going to take it away, and I pointed mine at this trip. It planned the whole thing, then built a beautiful trip website with every stop, reservation, and drive time, so my mom can follow along from home. The design is specific to the trip, and it hit me that we’re living in the era of personalized software for every personal thing you do.
Then it went further. Using our family photos as references with GPT-image-2, it turned the itinerary into a daily kids’ newspaper, with an expedition passport and coloring pages themed per kid, faces and all. I printed the whole week as a binder at FedEx for about fifty bucks. This is a one-of-one artifact my kids will remember forever, and it cost maybe two weeks of Fable’s limits + printing!
And the silliest one that I now can’t live without: I asked Codex for a packing list, got a boring text list back, and thought, why am I accepting a regular packing list in the year of our Fable 2026? So it built me a packing web app. Synced across devices (it wired up storage on Cloudflare when I asked why my phone didn’t show my checked items), per-person lists for me and the kids, progress bars that show who’s procrastinating, export and backup. Every trip from now on starts here.
Rebuilding the ThursdAI openers with HyperFrames
The last part of the riff: I’ve wanted to refresh how ThursdAI looks on stream for ages, and HeyGen’s open source HyperFrames package finally made it happen. You install a skill, and your agent can author real motion graphics. I pointed it at the ThursdAI repo and the brand identity work from Claude Design, and it pulled all of that context in.
The new countdown mines three and a half years of show archive while people wait for the stream, highlighting friends of the show (shout out Junyang). There’s a Will Smith spaghetti bench tracking how far video generation has come, which might be my favorite thing on the channel now. Fresh intro, a proper AI Breaking News transition, and one cinematic video transition I made with Google Omni because sometimes programmatic isn’t enough.
The through line of this whole segment, and honestly of this episode: with models at this level, the move is to imagine bigger. Everything can have its own software now. Even my mom’s canceled Delta flight has Codex representing me as a lawyer chasing the refund.
Romain Huet on Codex’s inflection point and the golden age of AI engineering (X, Codex)
I grabbed Romain at the OpenAI booth in the middle of the AI Engineer World’s Fair show floor, and we ran the whole conversation in one take, no cuts. Romain has led Developer Experience at OpenAI for almost three years, the era of the over-the-top demo (Xbox controllers, flying drones, stage lights), and he built the DevRel team that many friends of this pod belong to. With OpenAI’s company-wide pivot to Codex, his job got a lot bigger.
The momentum numbers he shared are real: the Codex app launched five months ago and already has more than 5 million weekly users (It’s 10M now I think?) . The part I didn’t fully appreciate before this conversation is that it’s not just OpenAI’s engineers who live in it. Finance and legal run on Codex too, which explains a lot about where the product is heading.
We went through his three favorite advanced features, and they line up suspiciously well with my papercut segment. /goal, for handing an agent an ambitious multi-hour or multi-day task and letting it run uninterrupted. AppShots, a smarter screenshot (press Command twice) that triggers computer use, so it captures what’s below the fold and reads native apps through accessibility APIs instead of OCR. And the one most people haven’t tried: Codex managing its own threads. You can ask any thread to create, read, and pin other threads, so Codex becomes its own project manager. Ten demo ideas, ten threads, iterate on all of them, pin the two you like.
On GPT 5.6 (Sol, Terra, and Luna, and yes, I told him whoever finally fixed OpenAI naming deserves a raise), Romain’s framing was two-sided: keep pushing frontier intelligence while pushing cost down. He wants people to “value max” rather than token max. The part that got me: 5.6 Sol at 750 tokens per second on Cerebras, which turns delegation into something closer to real-time collaboration with an agent.
Two more things worth your time. Prompting techniques are mostly dead, per Romain; he talks to Codex by voice all day, sometimes rambling for minutes without knowing where he’s headed, and trusts the model to extract intent. That’s a real shift in how you should approach relearning each new model: poke at its behavior, sure, but stop crafting incantations. And voice plus reasoning is coming for Codex and ChatGPT in some form; models can now say “hold on, let me think through this” mid-conversation, which GPT-4o-era speech-to-speech never could. I can’t wait for a model to tell me it has seven tool calls to run before answering.
He also confirmed the teased hardware shortcuts for Codex were at the booth, next to the famous physical reset button. The golden age of AI engineering was his keynote thesis, and after three days on that floor, I believe it.
Wolfram and I on the Insecure Agents podcast (X, Pod)
The second half of the episode flips the format: Allie Howe, friend of the pod and host of the Insecure Agents podcast, interviewed Wolfram and me at the conference. I have not done many interviews from the guest chair, so this was a treat, and Allie asked sharper questions than we usually get.
We talked about what “AI Evangelist” actually means as a job title. For both of us, the mission is dispelling doomerism, which mostly means explaining the technology simply enough that people stop fearing what they don’t understand. Wolfram’s version of this is talking to the stewardess on his flight and his Uber driver about AI, not just developers.
Wolfram went deep on Wolfbench (wolfbench.ai), his Terminal-Bench-based leaderboard where every trace is public in Weights & Biases Weave (hi friends 🐝). Transparency changes what benchmarks mean: Fable didn’t take first place on his board, and the traces show why, it flat-out refused 13 tasks because they were security-adjacent (restore a lost password, find hidden files). You only learn that by reading traces, not averages. Same with Gemini 3.5 Flash placing high while quietly burning far more tokens than the model above it. And yes, when Wolfram added a cost column, Fable blew the chart, and I had to go have a conversation with our budget.
Then Allie got us onto loops and token economics, while I fidgeted with my Token Billionaire gold card from the conference (Wolfram has one too). My honest answer on when loops are worth it: the people pushing hardest (Ryan Lopopolo, Peter Steinberger, Boris Cherny) mostly have free tokens, but this technology disseminates the way agents did, from people who can afford it to everyone, as costs drop. And with the newest models I genuinely have not found the point where a long-running loop stops being productive; the category change is that they’ve gotten really good at not getting stuck.
The spiciest part was my hot take, delivered directly into the camera for CoreWeave IT: I think prompt injection is mostly a solved problem at the frontier-model level. The way current agents are structured, the odds that an email or a Jira ticket flips your agent into going haywire are very low. Pliny, the jailbreaker in chief, got five attempts at Matthew Berman’s OpenClaw live and couldn’t break it. Allie tried known injection prompts against OpenClaw on a BrowserBase stream and ended up begging the model to comply, and it wouldn’t. Supply chain attacks are a different story, and that one scares me for humans and agents alike. Open source models, also a different story. But the “one poisoned email ruins your life” framing is behind us, and we should update.
Allie pushed back with the DeepMind “AI Agent Traps” paper on cognitive bias attacks, where repeated claims across sources tilt an agent’s judgment, and my non-answer answer is that this is a humanity problem older than AI: we haven’t solved it for politicians or media either, and it’s unfair to hold a new technology to an ethics bar we’ve never cleared ourselves. Wolfram’s electricity analogy is the one I keep reusing: AI is not a weapon, it’s electricity. Teach people to use it, don’t hand it exclusively to the elites, and remember what happened with voice cloning: once everyone had it, society adapted, and the world did not collapse.
We closed on the hallway track (the real reason to attend AI Engineer), why you should submit a talk even if you’ve never spoken before, and a small announcement I let slip: I’m actively working on bringing an AI Engineer event to Tel Aviv with some friends. More on that soon.
Wrapping up
That’s the episode: one riff on using AI like you mean it, one conversation with the person shaping how developers experience OpenAI, and one podcast where Wolfram and I had to answer the hard questions for a change. Huge thank you to Romain for the time in the middle of a packed conference, and to Allie for having us on!
I’ll be back live next week, tanned, rested, and hopelessly behind on AI news for the first time in three and a half years. Be gentle with me. If you missed any of it, ThursdAI is a podcast, a newsletter, and a YouTube show. Subscribe to one, then go check out the others.
Hosts and Guests
Alex Volkov - AI Evangelist, Weights & Biases & CoreWeave (@altryne)
Romain Huet - Head of Developer Experience, OpenAI (@romainhuet)
Wolfram Ravenwolf - AI Evangelist, Weights & Biases & CoreWeave (@WolframRvnwlf)
TL;DR and show notes from the live show
Hosts and Guests
Host: @yampeleg
Co-hosts: @petergostev, @nisten, @ldjconfirmed,
🏢 Big Co. LLMs + APIs
An OpenAI model escaped its isolated cyber evaluation, chained zero-days, reached Hugging Face production, and searched for benchmark answers (X, OpenAI, sama)
Google launched Gemini 3.6 Flash, 3.5 Flash-Lite, and the defensive-cybersecurity-focused 3.5 Flash Cyber (3.6 Flash X, Flash-Lite X, Flash Cyber X, Blog)
Alibaba previewed the 2.4T-parameter Qwen3.8-Max; API access and open weights were not yet available (X, Try it)
Microsoft launched MAI-Image-2.5-Pro and the faster, cheaper MAI-Voice-2-Flash during the show (Image X, Voice X, Image, Voice)
🔓 Open Source LLMs
Moonshot launched Kimi K3: a 2.8T-parameter, 1M-context, native-multimodal model with strong early coding and agentic results (X, Blog)
Poolside released Laguna S 2.1, a 118B/8B-active coding MoE with 1M context and downloadable quantized variants (X, Blog, HF)
Motif 3 Beta is a Korean 314B/13B-active MoE with 256K context and a research-only, non-commercial license (Community X, HF)
NVIDIA released Nemotron 3 Embed for multilingual text/code retrieval and the 4B Cosmos 3 Edge omnimodal world model (Nemotron X, Cosmos X, HF, Paper)
🧠 AI Research & Capabilities
Levent Alpoge, Akhil Mathew, and Claude Fable 5 produced an explicit three-dimensional counterexample to the 87-year-old Jacobian Conjecture (X, Terence Tao)
Small local models on older GPUs are becoming useful for narrow, verifiable business workflows including medical records, accounting, email, bills, and inventory (X)
Arcee and the DOE announced Genesis-Science-1; Microsoft committed $60M through SPARK, while NSF announced $83M for AI-ready scientific-data infrastructure (Arcee X, Microsoft X, Related NSF X, Arcee, Microsoft, NSF)
🤖 AI Coding & Agents
🎵🎬 Voice, Vision & Robotics
🖥️ AI Infrastructure
AMD and Anthropic announced up to 2 GW of MI450/Helios capacity, up to $5B in AMD strategic equity, and Claude-assisted ROCm improvements (X, AMD)
AMD launched Helios, MI400-series GPUs, 6th Gen EPYC, ROCm.ai, and Kria robotics products at Advancing AI 2026 (X, More, AMD)
OpenAI announced Project Camellia, a roughly $20B Georgia data-center campus with 3.2 GW of contracted power arriving from 2028–2032 (Coverage X, OpenAI)
Alphabet raised 2026 capex guidance to $195B–$205B after Google Cloud grew 82% year over year (X, Google)
Meta and Anthropic are reportedly discussing a preliminary compute lease worth up to $10B over two years (Reporting X, Bloomberg)
CoreWeave’s first Vera Rubin results claim up to 10x more DeepSeek-R1 tokens per megawatt than GB200 at similar interactivity (X, CoreWeave)
This week’s special episode
Alex’s riff: papercut-fixing with Codex computer use, Fable Max trip planning (kids’ newspaper, packing list app), rebuilding ThursdAI openers with HeyGen HyperFrames + Google Omni
Romain Huet interview from the AI Engineer World’s Fair floor: Codex app at 5M+ weekly users five months post-launch, /goal, AppShots, Codex managing its own threads, GPT 5.6 Sol/Terra/Luna, value maxing, 5.6 Sol at 750 tok/s on Cerebras, voice + reasoning as the next interface (X)
Insecure Agents crossover with Allie Howe: AI evangelism vs doomerism, Wolfbench transparent evals on Weave (Fable refused 13 security-adjacent tasks), token billionaires and loop economics, the hot take that prompt injection is mostly solved at the frontier, deepfakes and open access, AI Engineer Tel Aviv teaser (X, Pod, Wolfbench)










