Hey, this is Alex, welcome back to your weekly dose of intense AI acceleration summer!
My weekend was consumed by thinking about the OpenAI hack and agent swarms, but then the torrent of AI releases took over, and we got back to back news (including 3 breaking news during the live show), with a heavy open source focus!
I think the winner of this week is SpaceXAI/Cursor who released 3.5 releases, with one being my highlight of the week, Grok Bot (I’ve invited Shub Gaur from Cursor to the show to walk us through it) and Grok 4.6 which matches Opus at half the price.
There was a LOT of news in open source this week as well, with Meta kicking off with Muse Glimmer 30B and promising Muse Spark 1.2 soon, Qwen dropping Qwen 3.8 open weights and DeepSeek dropping an anvil with an upgraded DeepSeek v4 Pro and MIT license!
Let’s dive in (and please don’t forget as a reader you get 100% off the 1299 ticket to Fully Connected, our 2000 person Al event in SF in Sept, just use THURSDAIFC2026 as your code and see you there!)
Grok Bot and Grok 4.6 from SpaceXAI/Cursor
Folks, I’ve previously told you that from 3 frontier labs we noticed a jump to 5, and voila, this week proves that Elon is hell bent to win. After the cursor acquisition, and the integration of all of the parts into SpaceXAI, they have released 2 huge things this week
Grok 4.6 - Ties with GPT 5.6 SOL and half the price and much speed.
I’ve had the pleasure to host Goerge Cameron from Artificial Analysis on the show today, and I asked him, what is the best models. His answer, it’s a 3 factor answer, intelligence, speed and cost per task .
Well, if you use their nifty “recommend a model“ tool on the homepage, you’ll see that Grok 4.6 beats most other models on all of those! But, is it really that good? Models are really hard to evaluate and compare lately. It’s definitely a huge step up from Grok 4.5, with 61.3 on Frontier Code (beating Sol and just after Opus 5) and #4 on Apex-agents (+10 points from previous Grok). on Artificial Analysis this model lands at #4 on intelligence, while being #5 on speed all while being half the price of the models that are above it
As far as the tech goes, this model card confirms that it no longer has the Cursor Bench leaked into it’s weights and it’s #1 on that benchmark! It’s the same 1.5T v9 base at the same price, with Elon claiming that 4.7 is going to mog the competition in 3-4 weeks.
Everyone has a harness, now everyone has a swarm of bots - My Grok Bot review (x.ai/bot)
You guys know all about OpenClaw and Hermes, and Claude CoWork and Codex rebrand, and all of them are trying to nail down the same, always-on, autonomous agents that can do things for you.
Hermes and OpenClaw require you to have an always on computer, mess with API keys, Claude Cowork doesn’t run on the cloud and ChatGPT work starts a fresh session every time you ask a new thing.
Grok Bot (again, awful name) is the first one that seems to nail all of what I want in an always-on agent ... swarm. That’s right, this isn’t one agent with multiple personalities (like OC, Hermes), there’s a bot here for every task, and you dont’ have to manage context, queues, API keys (can if you want to) and models.
Oh, also ,there’s no model picker, it’s just Grok 4.6 deciding for ya, and it’s really fast!
Swarm of bots, working for you, each with their own computer
I am not getting paid for this (besides being provided a free account for cursor, but I’ve had it for 6 months and haven’t used), it’s really that good, the Cursor folks did some magic there. They picked up the most important parts of personal agents, like the (ios-only) mobile app (app store)
You can start a task on your mac, pick it up on your phone, get notified on your phone/mac, and the killer thing is, they are giving your bots their own computer, which can do things (especially if you’re ok with logging in there to your accounts!)
The kicker for me is the very very well done agent to agent communication there, which is transparent but read only to you. You can ask your bots to spin up other bots, but unlike sub-agents, they are actual bots with their own identity. You can even tag them in other chats and create group chats! There’s no context to manage, they do the work for you and so far this wasn’t a problem at all.
On the model side, Grok 4.6 seems to be doing an excellent job with agentic long running tasks that require coding and computer use, I’ve just been chatting with the bots and not thinking about any of the things I used for Hermes and OpenClaw.
What about Vendor Lock-in? Giving Elon data?
Some of these comments our fans raised during the show are very valid, after all, not only is the world divided on Elon Musk (which makes it REALLY hard to judge the models they release just on vibes from X btw, we talk about this constantly) but also, remember that Grok 3 started going off on X and called himself Mechahitler and just recently Grok CLI was caught uploading all of your data to X servers, which was reversed very quickly.
Honestly, I think there’s a very very good chance that this Grok Bot interface, which is geareed toward the less technical users, folks who don’t need the code-diff side pane, and don’t know/care what compaction is, and just want agents to do things for them, is goign to win much of this trust back. It just works, truly, for a beta product it’s really well executed by whoever worked on this!
Security and key management
One of the best parts for me with this Grok Bot, is that the connectors are the same connectors you use in Cursor! There’s a LOT of them (Cursor after all has been one of the first apps to start adding AI agents) and this also means that they take the security very seriously.
Every API key that you want to add, is not shown to the bot, each bot lives in an isolated environment, and for stuff like payments and log-ins, it gives you back the control of it’s computer for you to complete!
I also love this section in settings, which makes auto-approve work for you: you define rules with natural language that you always want the bot to ask you before... sending an email or posting on your behalf or what not.
Chief of staff pattern to get started
In case you’re convinced enough to give it a try (it’s free trial for 1 month, and the cheaper way to get it is via Cursor’s 149$ plan and not via the Grok Ultra plan which is 249), here’s a recommended pattern that works very well.
Create a chief of staff bot, have it interview you about everything you are doing in your day to day, work and personal, then decide how much permissions you wanna give it, start little.
Then ask your chief of staff to create bots for some of the work it can try and help you with, focus on “reduce cognitive load”.
And then see the magic come to life. If you have skills or memory from other bots, you can just ... import it in.
Then try setting up an automated email checker bot, and have your chief of staff surface only the most important emails you have to actually respond to.
Another great pattern is setting up a bot with the last30days research skill (we covered it with Matt Van Horn) and have a research bot for every topic you want to deep dive into.
Schrodinger’s Grok
I haven’t quite named it like that, but we’ve covered all Grok released on the show (tracking 24 on https://thursdai.news/companies/xai excluding this week) and ... it’s always very hard to judge Grok model released based on X feed vibes. It’s either AI influencers who want Elon to retweet them, glazing the models, or folks who hate Elon for his political views or whatever, ignoring their (truly insane progress).
This time, both the model and Grok Bot are getting very very good reviews, from folks like our own Ryan Carson, Lenny Rachitsky, Rubben Hassid and Roberto P Nickson. Not folks who are swayed lightly, but also, yours truly. I really do think there’s something great here, worth trying out, especially if you’ve struggled to maintain your OC/Hermes and want agents to work for you 24/7. LMK if you have questions about it and your experience
Open Source AI and other news
I want to continue with this new newsletter that covers 1 big story, but I can’t leave you uninformed about the most important developments in AI and Open Source
DeepSeek V4 pro 0813 is in GA - MIT licensed chonker with 1.7T parameters (X, Blog, HF, GitHub)
The whale is back with a vengeance, DeepSeek resurfaced with their flagship response to Kimi K3 and with MIT license, we can’t complain.
1M context window, 49B active parameters but it seems to underperform, landing at 54 on the Artificial Analysis leaderboard. However, they did show a significant improvement on DeepSwe (from 12.8 points in the preview version of V4 to 62.7 in this one)
We still think it’s a good model sir, and definitely worth trying out!
Additionally, DeepSeek released their own harness on Github (hitting 23K stars in less than 24 hours) which seems to be exciting as well, give it a try.
Meta comes back to open source with Muse Glimmer (30B) and promise to open source Muse Spark 1.2 (X, Blog, HF)
We would like to officially welcome back Meta to the open source AI community, as they release their smaller Muse model called Glimmer!
The highlights, it runs on a single 24GB consumer GPUs, gets 51 on Swe-bench Pro, beating Qwen 3.6 27B. And with DFlash speculative-decoding, it delivers 233tok/s on RTX 5090.
Zuck promised us the bigger Muse Spark 1.2 in open source and published a long essay on superintelligence and that it should be distributed to everyone, which we applaud and it’s great to see the commitment reinforced! welcome back Meta!
This weeks buzz
Short interjection from our only sponsor, CW this week.
1 - Join 1500 ai practitioners (and a live ThursdAI recording) at Fully Connected Sep 29-31 in SF - use code THURSDAIFC2026 (Register here)
2 - We have day-0 support for Nvidia’s latest Nemotron 3.5 lightning (CW Inference)
Gemini 3.7 Flash - breaking in the middle of the show
Just as we had George Cameron from Artificial Analysis on the show, Gemini dropped Gemini 3.7 Flash, and it’s a speedy beast! Clocking at over 300t/s, it’s google’s mid-tier model, think Sonnet/Terra competitor, that is also great at multimodal (I think it’s one of the only ones that can watch videos)
It beats Muse Spark 1.2 on DeepSWE and lands near the cost-per-task Pareto frontier on Artificial Analysis. For the cost/speed/intelligence trade-off, this model is now #1 on Artificial Analysis selector of best models!
That’s a wrap
This was the first week of the shorter newsletter experiment: one big story done properly, and trust that you’ll listen to the show for the rest (it’s 2.5 hours of exactly this, with demos). Tell me if you hate it. Our release index at thursdai.news tracked 71 releases in July alone, so something had to give, and it wasn’t going to be my weekends.
See you at Fully Connected Sept 29 (code’s in the intro, come say hi to me and Wolfram at Moscone), and if you try the Grok Bot chief of staff pattern, I genuinely want to hear how it goes.
ThursdAI - Aug 13, 2026 - TL;DR
Hosts and Guests
Alex Volkov - AI Evangelist, Weights & Biases & CoreWeave (@altryne)
Co-hosts: @WolframRvnwlf, @petergostev, @nisten, @ldjconfirmed, @yampeleg, Chris Alexiuk - NVIDIA (@llm_wizard)
Shub Gaur - Cursor / SpaceXAI, GrokBot (@shubgaur)
George Cameron - Artificial Analysis (@grmcameron)
Big CO LLMs + APIs
xAI Grok 4.6: AA Index 61 at $2/$6 per M, CursorBench 69.9, card confirms self-optimized inference stack (X, Blog, Model card)
Grok Bot early beta: persistent agents with their own computers, macOS + iOS, free with SuperGrok Heavy and Cursor Ultra (X, x.ai/bot)
Breaking: GPT 5.6 Sol ultrafast preview on Cerebras at ~14x speed, work-account waitlist (Blog)
Breaking: Gemini 3.7 Flash, 50% price cut through end of year, near Pareto-optimal cost per task (X)
OpenAI GPT-5.6-Cyber: 95.0% cyber completion vs 1.5% base, gated behind Daybreak Red (X, Blog)
Grok 4.7 teased: 3-4 weeks out (Elon-reply-sourced only) (X)
Open Source LLMs
DeepSeek V4 Pro 0813 weights re-published under MIT: 1.6T/49B active, DeepSWE 62.7 (+49.9), Terminal Bench 2.1 87.9, $0.435/$0.87 per M (X, OpenRouter)
DeepSeek Harness hit 23K GitHub stars in days, web UI (GitHub)
Qwen3.8-Max landed on HF as open weights: 2.4T/95B active MoE, 1M context, FrontierSWE 73.5, custom license (X, HF)
Meta returned with Muse Glimmer 30B agentic, Apache 2.0, SWE-Bench Verified 76.0, Muse Spark 1.2 weights promised (X, Blog, HF)
NVIDIA shipped Nemotron 3.5 Lightning: 30B MoE/3B active, up to 4x output speed, strong voice-agent results (X, HF)
Motif 3 from Korea open-sourced: 314B/13.2B active, MIT, SWE-Bench Verified 76.2 (X, HF)
Cohere North Micro Vision: 2.4B VLM, Apache 2.0, DocVQA 92.1% (X, HF)
AI in Society
Anthropic watermarks all new Claude text output worldwide under EU AI Act Article 50, C2PA on images, detection docs promised (Geiping FAQ, Euronews)
Stolen Thoughts: 704 artifacts including 62 API keys extracted from hidden reasoning across 6,708 sessions (X, Paper)
Pangram: OpenAI holds 50%+ of AI text share, Anthropic triples to 14.9%, Google falls to 1.9% (X, Blog)
This Week’s Buzz
Evals & Benchmarks
Artificial Analysis launched Optima: private evals from your own use case and agent traces (AA)
Vision & Video
LTX-2.5: 22B open-weights video, multi-shot, 10s 1080p in 23.7s on fal, 16GB VRAM min (X, HF, GitHub)
Alibaba Wan-Animate-2: 14B character animation, Apache 2.0, 70%+ blind preference win (X, HF)
Tencent Hunyuan3D WorldClaw: text-to-3D editable game worlds, paper only (X, Paper)
xAI Imagine Image 2.0: #2 on Arena for T2I and editing (X, Blog)
Voice & Audio
MiniMax-Music3: open-weights production music model, dropped mid-show (X)



















