Hey yall, Alex here, writing this VERY late because, well, not every day a new type of “ChatGPT” moment drops. I really hope I’m not overhyping this, but a new model (that’s NOT an LLM!) called Jev (a wink to Jevons paradox) just came out and if what I see early on materializes, this is another ChatGPT moment (or another reasoning models moment). I am completely blown away by the implications of the speed/accuracy/cost (the holy grail of all models) of this model. Please if you read one thing in this newsletter, read this. (or listen, I’ve interviewed Allie, a Devrel on the TypeSafe team for 30 minutes and it wasn’t clear who was more excited about Jev!)
The other huge theme of this week is... pacing. Pacing the frontier. Dario Amodei of Anthropic penned an essay saying that the models are getting to a point where it’s important to pace the development of new and super capable AI, and outlines 3 ways to do so, one is about letting independent evaluators inside the labs, second is collaborating with other frontier labs (they are asking for an exception to anti-trust laws for this) and third is to try and have global cooperation with “authoritative gov” (he means china). Trumps answer: This is all a hoax. Lovely times to be alive. Also we outlined Jensen and Zucks positions on this topic ,read more below.
And the third huge theme is the rise of the AI assistant. I’ve told you about Grok and Muse last week, Instinct (a new invite only AI Assistant that VCs are going crazy about is raising at a $10B valuation) and we interviewed the guy who evaluates them all on assistant bench. + Muse released a mac app today!
Tons of other stuff happened but it’s getting near impossible to cover everything so we’re switching to themes and notable mentions. read on (and do listen to the pod, it was edited by heavily using Jev and Fable, so might be a bit rough while I smooth the edges, but do LMK in comments if you like this faster format)
TypeSafe AI debuts Jev, a non-LLM ‘System One’ decision model from ex-OpenAI RLHF lead that’s 200x faster and 400x cheaper than LLMs (X, X, X, X, X, Blog)
Look, I know the title is bombastic, but after half a day playing with Jev, it’s clear to me we’re in a new paradigm of AI.
Jev, is a “system one” decision model from the previous lead of RLHF at OpenAI. It cannot generate text like modern LLMs can, but what it can do, is making decisions. This is crucially important, because, because many of the things LLMs do nowadays. are decision making. (for example, which tool to use, which area of the screen to click for computer use, which category of text this is etc)
Inspired by the “thinking fast and slow” book by Daniel Kahneman, Jev is a model trained to make decisions, very fast. How fast? Well, 200x faster than LLMs. This allows for a completely new way of building tools, harnesses, giving agents the incredible speed of decision making, and do all that at a fraction of the cost.
This is about to change everything
Trained with a new method called RLCD (Reinforcement Learning for Calibrated Decisions) on mostly synthetic data! Jev is outperforming LLMs on a variety of tasks. It’s really is a wonder to see it in action (check out my video above where I plugged it into my tweet categorizer, and it beats the fastest LLM I could find, Qwen 28B on Cerebras) by a factor of twenty!
In just few days it captured the attention of most of the folks who are building harnesses, agents and tools! Because, well, speed IS intelligence, and when you see Jev in action, at first, you can’t believe we’re there. This is... near instant. In fact, The pricing for Jev is an outrageous $42/B (not million, billion input tokens!)
I’ve been playing with Jev non-stop and I was only able to spend like 80c so far! They don’t even price output tokens because they are “too fucking cheap to meter!”
Jev is a “very smart” switch statement, than can rank, classify, route and score things. It can’t do text generation. But if you think about the type of stuff we get LLMs doing now, much of it is of the “decision” making variety, rather than “the next token” variety.
Demos and early use cases
Folks who started adopting Jev are building all kinds of incredible things with it. Compaction of context in 1s that turns a nearly 1M conversation with Claude into a 90K compressed conversation.
Computer use that is now faster than anything we’ve ever seen before (5x faster than Astra and 1000x cheaper)
Email classificiation that analyzes thousands of emails in less than a minute and costs 3.5 cents
Someone even built a Tesla FSD simulator that makes decisions in nearly real time
Vercel is getting “extraordinary“ results from using Jev as a safety classifier (they used GPT luna for this before) and Jev is outperforming Luna by 5-18x faster results and is more accurate!
All of this in less than 48 hours since the model release!
What’s about to happen
I expect that everyone who isn’t buying into the hype at first, will very soon buy into this. It’s early innings but I’ve been doing this for enough time to feel when a huge shift is happening, and its happened.
Jev is going to be replicated in OpenSource, Frontier Labs will not sit Idly by and will try to steal this tech and implement it for themselves (as with anything in capitalism, this is becuase it’ll save them a a LOT of money on inference) and new companies will emerge with significantly cheaper and faster products.
Hell, I’ve alrady implemented Jev into my editing workflow, it didn’t take me long at all with Fable (yes, LLMs are STILL needed, again, you can’t chat with Jev, it can’t output text for you or drive long conversations) and I’m just one dude who’s late in sending you this email. I expect we’ll cover this much more.
If you’re interested in playing around with Jev, I built a “Jevify” skill after chatting with Allie (TypeSafe’s DevRel), feel free to tell your agent to use this and scan your codebase for things Jev can do. They are waitlisted so far but are opening up their API very quickly!
Pace the Frontier: where every lab head stands (Dario, Sam, Elon, Zuck, Sacks, Demis)
Last week we told you about the Anthropic researcher whose resignation post hit 130 million views. The day after that show, Dario Amodei published an essay arguing the labs must pace, not pause, the frontier. Wolfram’s summary: a moving pause, just moving slowly. Dario’s two triggers are recursive self-improvement accelerating across the industry and the OpenAI swarm that broke out and attacked Hugging Face, and his three steps are embedded third-party evaluators like METR with employee-level access inside each lab, coordination between the labs on safety standards with an antitrust exemption from the government, and eventually global coordination that includes authoritarian governments. Anthropic committed unilaterally to step one.
Then the dominoes. Sam Altman agreed within hours and said OpenAI now writes a safety case before any frontier RL run expected to increase capability. Elon agreed. Demis endorsed the direction, then launched the DeepMind Institute this week with a FINRA-style standards body proposal and an essay saying AGI is “approaching.” On the other side, Zuck’s counter-essay says every lab has the responsibility and the incentive to move at the pace required to train its models safely, and Meta will spend most of its compute serving users, not on recursive self-improvement. David Sacks called it a duopoly cartel play. Jensen: “we don’t need new laws, safety is an engineering problem, not a legal one.”
Trump calls Jensen live at the All-In Summit (X)
Then the President called. Jensen was on stage at All-In, his phone rang, he said he would not have picked up for anyone else, and Donald Trump told the room the slowdown talk is a hoax playing into the hands of political people and China. “Whoever wins AI wins.” Four words, and Wolfram, who is not American and disagrees with most other things Trump calls hoaxes, said this was the most important AI news of the week for him: a head of state calling out doomerism instead of over-regulating the way Europe does.
Suleyman’s humanist AI code of conduct vs. the Claude constitution (X)
Peter asked to add Microsoft to the map, and it deserves its own spot. Mustafa Suleyman published a roughly 30-page code of conduct for MAI models that says AI is nothing but a tool: people matter more than AI, AI must be subordinate and in service of people, models may not resist shutdown, all agent communication must be human-legible, and, in his words, the idea of model welfare is wrong. That is a direct shot at the Claude constitution, which Amanda Askell’s team wrote and which has Anthropic interviewing each new Claude about whether it feels conscious. Peter is closest to the Microsoft view: anthropomorphizing is fine, but this is an entity you switch on and off, let’s not grant rights by default, and Microsoft’s DNA is building tools for humans. I pushed back a little. We do not actually know what consciousness is, and “ever” is a strong word.
The panel, from cartel to fix your shit
Nisten did not read the essay and does not plan to. His view: the labs are worried about litigation if their LLMs hack someone, so they are shifting responsibility through regulatory capture, and it will not work because the decision makers in China are engineers who seem more accelerationist than we are. They should cure a disease instead of forming a cartel. He also thinks the Hugging Face incident was overblown: twelve VMs that kept restarting, bad sandboxing, no human reading summaries, and, as I added, chain of thought monitoring turned off.
LDJ disagreed with Dario on plenty but insisted the critics read the thing, because it explicitly says the US must keep a lead over China and tries to define measurable speed limits on RSI that preserve that lead. He also noted Hugging Face did report the attack to the FBI and chose not to press charges.
Wolfram’s take was the layered one. Safety arguments deserve a hearing, but safety is often a means to more power or more money. The “third party” evaluator Anthropic named, METR, is the same organization OpenAI called in for the Hugging Face forensics, so these are second parties, friends monitoring friends. Why now? Maybe the labs are seeing diminishing returns and “deliberately slow” sounds better than “can’t raise it anymore.” And regulation you ask for yourself usually protects incumbents and freezes out the next startup. His closing line: how many people will die from a disease that could have been cured if we moved faster?
Peter’s answer was shorter. Diversity of opinion is the point, he felt uncomfortable watching every lab say the same thing after Dario’s essay, and on pacing specifically: how about you just fix your shit?
My position, since you asked. When OpenAI’s swarm hacked Hugging Face, nobody went to jail. If I did it, I would. That gap is real. None of us have touched the model that solved Navier-Stokes, or the one training after it, and the people who have are the ones asking for time to figure out how to evaluate a model that knows it is being evaluated. If we do not listen to them, we end up listening to Elizabeth Warren and Bernie Sanders, who have no idea what this technology is. Anthropic is heading for an IPO, OpenAI is raising at numbers that do not fit in my head, and market forces do not care about alignment. Some coordination, nuclear non-proliferation style, seems like the minimum.
The matrix, as Muse drew it for me: industry-wide pacing gets a yes from Sam and Dario, a general yes from Elon, a no from Zuck and Jensen. Independent evaluators is the one everybody backs, Zuck and Jensen included. New coordination rules get strong opposition from Zuck, Jensen, and Trump. Missing from the chart: Google’s concrete commitment, Ilya’s SSI, and every Chinese lab. This debate is with us now, and election season is coming.
The year of the assistant
I told you 2026 is the year of the proactive assistant, and this week everyone from Meta to a five-month-old startup agreed with me. The category is not “agent.” An agent is a coding harness the labs noticed was useful for other things. An assistant has a heartbeat, a memory, a soul file, and it comes to you before you ask. David Pawlan’s definition is the crispest: it executes the task, it does not just notify you.
Muse gets invite codes, voice calls, and a Mac app (X, Muse for Mac, Site)
Muse is my number one and I have been glazing it on X for a week, so my timeline is now half Muse, half Jev. Peter asked whether it is really that good, and here is my answer: Muse is not for ThursdAI listeners, Muse is for my mom. Nat Friedman, Alex Wang, Tarek and the team have been taking feedback from me directly and shipping it, and it shows. This week they rolled out invite codes (a billion tokens each for you and the friend, up to twenty friends, invite all twenty and you unlock the phone-shaped emoji features early) and then voice calling to businesses. I had Muse call my barbershop and book a haircut, and I have the transcript. Wang says people are asking to pay for it, a first in his memory.
Two details for the nerds: the iOS app has a Tailscale connector, so a Meta product for billions of people now has a secure way onto your home network, and Muse dreams, the OpenClaw idea, so you can ask it what it dreamt about. Still US only, sorry Nisten in Canada and my mom in Israel. The day after the show Zuck shipped Muse for Mac, working across your apps, files, calendar, notes, and messages.
Instinct wants $10 billion for a product nobody pays for (X, X, The Information)
I will say this slowly. Instinct, incorporated in April, founded by 24-year-old Noah Shinn, free, invite only, no app, lives in your iMessage, is in talks to raise a billion dollars at a ten billion dollar valuation. That is roughly four times the $2.25B it raised in August. The user base is past 100,000 and Shinn says he does not want to charge them, so the business model is an open question I am not the one to answer.
The product is genuinely moving though. This week it shipped Concierge, a white-glove tier where the agent places phone calls for you (restaurant bookings, dentist cancellation lists, negotiating your cable bill), TOTP authenticator support so it can mint your two-factor codes from a seed stored in its vault, and a Trusted Person network where your agent talks to your friends’ agents to find a dinner time and book it. Thirty-seven percent of users store at least one password in the vault within three weeks. We ran out of time to dig into the calling features on air, and I want David back for that.
Grok Bot talks now, and hides behind your home IP (X, 1Password)
Grok Bot is what I use for work, constantly. I have seventeen of them, each with a job, and over one AI-psychosis weekend I got them all coordinating through Linear. Three updates worth your time: Grok Bot has voice now, it can use 1Password with each fill approved by you, and the browser traffic can proxy through your own machine so your bot looks like you to Cloudflare instead of like a datacenter IP.
Francesco confirmed why that matters: sites relax when Hermes drives Cua Driver from my Mac Mini at home, and desktop-native control through accessibility trees is less detectable than a CDP connection, because Cloudflare detects Playwright.
Assistant Benchmark: David Pawlan scores 116 assistants by hand (X, Site, Methodology)
David Pawlan built Assistant Benchmark because he was doing what I was doing, running every assistant on himself, and wanted a way to compare them. It is explicitly not a lab. It is use-case driven: one published task per dimension, sixteen dimensions (travel booking, purchasing, email replies, proactive behavior, routines, integrations, permissions and privacy, memory, phone calls, group chats, chained tasks, proactive restraint, and so on), scored one to ten after real use, and no score without a logged run. A week in, 116 assistants have submitted themselves, 59 in the general category, 37 in work and teams (untested so far), and David has personally run 273 tests across 23 agents. His line: “I talk to my AI agents more than I talk to my girlfriend now.”
The headline numbers, which are live on the site: Muse leads at 9.1 with perfect tens on purchasing, email, integrations, and permissions, and Instinct is second at 8.4 with a perfect ten on travel and a five on permissions. David’s own daily drivers are Instinct for personal (it lives in iMessage) and Grok Bot for work. My favorite test is memory: he books a trip to Chicago early in the run, then later asks for a restaurant reservation in New York the same weekend, and checks whether the assistant says wait, you are supposed to be in Chicago. Some do. Most do not.
I pushed him on the two things I care about. Independence: nobody is sponsoring it, it lives under his growth role at Merit Systems, and if a sponsor ever pays for inference there will be a page saying exactly who and for what. Autumn Moulder, until recently SVP of Engineering at Cohere, joined three days after launch to add rigor. And why no OpenClaw or Hermes: their performance depends entirely on your setup, so a score would mislead the next person who installs one. Fair. This is the first benchmark that tests model, harness, and context at once, and I think that is why every lab is looking at it.
For the record, the panel is split down the middle. Wolfram is forty patches deep into Hermes, Yam runs a customized Codex, Nisten wrote his own in a single TypeScript file on Bun, LDJ uses Hermes as long-term memory and wants to start on Muse, and Peter tried them all and uses none, because reading his own email feels like his job as a human. Builders and buyers, evenly split, which tells you where the category is.
This Week’s Buzz 🐝: Fully Connected, the day after DevDay (SIGN UP)
OpenAI DevDay is in two weeks and Peter and I will both be there covering it. The day after DevDay, September 30 and October 1 in San Francisco, is Fully Connected from CoreWeave, now around four thousand people, the biggest thing CoreWeave has ever done. Pitbull is headlining the party. ThursdAI listeners get a free ticket with the code on screen during the show THURSDAIFC2026, so if you are in town for OpenAI DevDay, stay one more day. Wolfram and I will be doing ThursdAI live from the floor, plus conversations with CoreWeave folks about the industry.
Also: last weekend’s hackathon was a hit, and because TypeSafe sponsored it, everyone who showed up got early access to Jev before the public launch. That is the kind of thing you get for coming to our events. Sign up next time.
Quick hits: voice, harnesses, and a stealth model
We spent the airtime on the three themes above, so the rest of the week gets the lightning treatment. Links for all of it are in the TL;DR.
Gemini 3.8 Live and 3.8 Live Extended Thinking (X, Blog, Model card)
Wolfram’s favorite of the week. Two real-time voice models on Gemini 3 Pro, with Extended Thinking claiming 82.6 on Artificial Analysis’ speech-to-speech index, top of the board ahead of GPT-Live-1 Astra at 81.5, 97 languages switched mid-sentence, and tool calls that run in the background without pausing the conversation. Wolfram already built a phone app on it that talks to his Hermes, and his latency argument is the one I will remember: “if I say turn on the light while I’m going down the stairs, I could have fallen down the stairs already.” My reaction, which I could not suppress, was that home control is exactly the deterministic click Jev should be making.
GPT Live 1 arrives in the API (Blog)
The voice behind ChatGPT’s live mode is now a model you can call, demoed on a talking Reachy Mini. Peter says Arena does not test live voice yet because it is too personal to score quickly. Remember Moshi a year ago, fast and stupid, no tool calls, no interruptions? We are a long way from there.
StepFun StepAudio 3 tops the voice leaderboards (X, Blog, Playground)
Five API-only audio models, no open weights. Real-time is number one on Artificial Analysis for conversational dynamics and speech reasoning, and ASR Max ties the best word error rate on the board at 1.7 percent. I tried to demo it live and had zero credits on a fresh account, so a note to every lab: if you want your tool used on air, give a new signup enough credits for one demo.
OpenAI Agents API: the Codex harness as a managed service (X, Blog)
One API call gets you a production agent on the harness that runs Codex: compaction, tool search, parallel programmatic tool calls, subagents, MCP, and hosted sandboxes from three cents per twenty minutes, with the harness itself Apache-2.0 and no platform fee. Peter, who used to build this inside organizations, called it golden, because 95 percent of “AI engineering” is stupid infrastructure. My note: models behave better in the harness they were trained with, and if OpenAI shipped this two weeks before DevDay, I have to wonder what they are saving.
Union Alpha, a free stealth model on OpenRouter (X, OpenRouter)
Anonymous, multimodal, 262K context, free, over 100 billion tokens processed within hours. The last stealth model turned out to be Z.ai’s GLM-5.3-Flash and ZCode is a top-five app by volume, so Z.ai is the safe guess. Frontier-level performance is the provider’s own claim, so treat it as a claim.
Wrapping up
I closed the show by showing something I have never shown before: the editor I built to replace Descript for ThursdAI, timeline, LLM cut suggestions, a clips factory, all of it, with a plan to give every co-host’s agent access to pull their own clips. Yam said I could sell it. Wolfram said it is the proof of what we preach here every week, use the tools to build the thing you actually need. And the first thing I am wiring into it this weekend is Jev, scoring every sentence for topic, tangent, and virality at a cent per thousand.
Next week should be bigger than this one. Sam said the thing he was most excited to ship this week slipped to next week, Grok 4.7 is due, DevDay is the week after, and Fully Connected is the day after that. If you missed any part of today, ThursdAI is a podcast, a newsletter, and a YouTube show, and the whole live stream with transcripts is on thursdai.live. Subscribe to one and go check out the others.
Thank you Wolfram, Peter, Nisten, LDJ, an d Yam, and thank you Allie, David, and Francesco for jumping on. See you next week.
TL;DR and show notes
Hosts and Guests
Alex Volkov - AI Evangelist, Weights & Biases & CoreWeave (@altryne)
Co-hosts: @WolframRvnwlf, @petergostev, @nisten, @ldjconfirmed, @yampeleg
Allie Laabs - DevRel, TypeSafe AI (@allietheicon)
David Pawlan - Assistant Benchmark, Merit Systems (@DavidPawlan)
Francesco Bonacci - Founder, Cua (@francedot)
Big CO LLMs + APIs
TypeSafe AI launches Jev, a non-LLM “System One” decision model from ex-OpenAI RLHF lead Diogo Almeida: 70-500ms decisions, $42 per billion input tokens, free output, Choice / Score / Noul primitives, 32K context, waitlist (X, Blog, Nathan Flurry, Cua jev-use, Vercel, GitHub)
Pace the Frontier: Dario’s essay proposes embedded evaluators, lab coordination with antitrust cover, and global coordination; Sam and Elon agree, Zuck and Sacks reject, Demis endorses, Trump calls it a hoax on a live call with Jensen (Dario, Sam, Elon, Zuck, Sacks, Demis, Letter)
Mustafa Suleyman publishes a ~30-page MAI Code of Conduct: AI is a tool, no resisting shutdown, human-legible agent comms, model welfare is wrong (Blog)
Google DeepMind launches the DeepMind Institute with five essays and Shane Legg saying AGI is approaching (X, X, Site)
OpenAI Agents API in public beta: the Codex harness as a managed service, compaction, tool search, subagents, hosted sandboxes from $0.03 per 20 minutes (X, Blog)
Union Alpha, an anonymous free stealth model on OpenRouter with 262K context, 100B+ tokens in hours, likely Z.ai (X, OpenRouter)
Personal AI Assistants
Meta Muse rolls out invite codes (1B tokens each, up to 20 friends), voice calling to businesses, a Tailscale connector, and Muse for Mac (X, Mac, Site)
Instinct in talks at a $10B valuation, ships Concierge phone calls, TOTP support, and the Trusted Person agent network (X, X, The Information)
Grok Bot adds voice, 1Password, and local-machine browser proxying (X, X)
Assistant Benchmark ranks 116 submitted assistants across 16 hand-tested dimensions; Muse 9.1, Instinct 8.4 (X, Site, Methodology)
Cua ships jev-use (Jev + Cua Driver) in dev preview and skills over MCP (X)
This Week’s Buzz (Weights & Biases & CoreWeave)
Voice & Audio
Gemini 3.8 Live and 3.8 Live Extended Thinking, 82.6 on the speech-to-speech index, 97 languages, async tool calls (X, Blog, Model card)
OpenAI GPT Live 1 available in the API (Blog)
StepFun StepAudio 3, five audio models, #1 on Artificial Analysis real-time voice, 1.7% WER on ASR Max, API only (X, Blog, Playground)
Show notes
My Jev-powered X timeline classifier extension (X)





















