<?xml version="1.0" encoding="UTF-8"?><rss xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:atom="http://www.w3.org/2005/Atom" version="2.0" xmlns:itunes="http://www.itunes.com/dtds/podcast-1.0.dtd" xmlns:googleplay="http://www.google.com/schemas/play-podcasts/1.0"><channel><title><![CDATA[Towards AI Newsletter]]></title><description><![CDATA[Towards AI's thoughts on the week's biggest AI developments. 
All major AI news, models, tools and papers covered. 
Read by over 130,000 AI Practitioners, Industry Professionals and AI Students.]]></description><link>https://newsletter.towardsai.net</link><image><url>https://substackcdn.com/image/fetch/$s_!ZBHF!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Faea4e29a-6b40-4b9a-9a98-00d0f6550a2e_512x512.png</url><title>Towards AI Newsletter</title><link>https://newsletter.towardsai.net</link></image><generator>Substack</generator><lastBuildDate>Sun, 11 Oct 2026 03:55:12 GMT</lastBuildDate><atom:link href="https://newsletter.towardsai.net/feed" rel="self" type="application/rss+xml"/><copyright><![CDATA[Towards AI, Inc.]]></copyright><language><![CDATA[en]]></language><webMaster><![CDATA[pub@towardsai.net]]></webMaster><itunes:owner><itunes:email><![CDATA[pub@towardsai.net]]></itunes:email><itunes:name><![CDATA[Towards AI]]></itunes:name></itunes:owner><itunes:author><![CDATA[Towards AI]]></itunes:author><googleplay:owner><![CDATA[pub@towardsai.net]]></googleplay:owner><googleplay:email><![CDATA[pub@towardsai.net]]></googleplay:email><googleplay:author><![CDATA[Towards AI]]></googleplay:author><itunes:block><![CDATA[Yes]]></itunes:block><item><title><![CDATA[Why we’re taking more time with the book]]></title><description><![CDATA[More on agents, evaluations, coding workflows, and guidance we think AI engineers will need most in 2026.]]></description><link>https://newsletter.towardsai.net/p/what-were-adding-to-ai-engineering</link><guid isPermaLink="false">https://newsletter.towardsai.net/p/what-were-adding-to-ai-engineering</guid><dc:creator><![CDATA[Louie Peters]]></dc:creator><pubDate>Wed, 07 Oct 2026 16:55:57 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!ts1I!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8cd89b08-6446-44b1-bac9-03bac874f349_2239x1372.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>A quick update on our upcoming book, AI Engineering for Production.</p><blockquote><p>TL;DR: If you missed our first announcement, this is the book we have been working on as the 2026 follow-up to <em>Building LLMs for Production</em>. It is about what it takes to build LLM and agent systems that actually work in production, especially now that models are becoming significantly more capable and, at the same time, just one component of a reliable system.</p></blockquote><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://www.louisbouchard.ai/book/?utm_source=sub-tai&amp;utm_medium=email&amp;utm_campaign=book" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!ts1I!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8cd89b08-6446-44b1-bac9-03bac874f349_2239x1372.png 424w, https://substackcdn.com/image/fetch/$s_!ts1I!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8cd89b08-6446-44b1-bac9-03bac874f349_2239x1372.png 848w, https://substackcdn.com/image/fetch/$s_!ts1I!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8cd89b08-6446-44b1-bac9-03bac874f349_2239x1372.png 1272w, https://substackcdn.com/image/fetch/$s_!ts1I!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8cd89b08-6446-44b1-bac9-03bac874f349_2239x1372.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!ts1I!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8cd89b08-6446-44b1-bac9-03bac874f349_2239x1372.png" width="1456" height="892" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/8cd89b08-6446-44b1-bac9-03bac874f349_2239x1372.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:892,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:451556,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:&quot;https://www.louisbouchard.ai/book/?utm_source=sub-tai&amp;utm_medium=email&amp;utm_campaign=book&quot;,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://newsletter.towardsai.net/i/219292487?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8cd89b08-6446-44b1-bac9-03bac874f349_2239x1372.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!ts1I!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8cd89b08-6446-44b1-bac9-03bac874f349_2239x1372.png 424w, https://substackcdn.com/image/fetch/$s_!ts1I!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8cd89b08-6446-44b1-bac9-03bac874f349_2239x1372.png 848w, https://substackcdn.com/image/fetch/$s_!ts1I!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8cd89b08-6446-44b1-bac9-03bac874f349_2239x1372.png 1272w, https://substackcdn.com/image/fetch/$s_!ts1I!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8cd89b08-6446-44b1-bac9-03bac874f349_2239x1372.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>We first announced the book for <strong>October 20</strong>. But realized it needed more work to become the book we actually want to hand you. In this final push, we want to focus on adding as much practical, production-ready information on agents, evaluations, and the engineering ecosystem as we can.</p><p>We don&#8217;t have a date yet, but it is definitely coming by the end of the year.</p><p>Want to know the moment it&#8217;s out? Get book updates &#8594; <a href="http://louisbouchard.ai/book">louisbouchard.ai/book</a></p><p>Here&#8217;s what we&#8217;ve been working on so far:</p><ol><li><p><strong>An entirely new chapter on coding with AI.</strong> Although this is a book about AI engineering, very few developers in 2026 are writing all their code themselves. So we are adding our best practices, strict boundaries, and the standards we use in enterprise work to show how to use coding agents reliably.</p></li><li><p><strong>An AI tutor made specifically for the book.</strong> It runs as an MCP server that you can plug into your own agent to debug and better understand every topic. We will also show you how to build one yourself and deploy it at scale in the book.</p></li><li><p><strong>Coding sections rewritten to go beyond implementation details.</strong> We are adding more insight into the decisions behind the code, rather than only explaining how to implement it. We are also doing our best to make it fun to read.</p></li><li><p><strong>A new conclusion focused on guardrails when building with AI.</strong> Given where the field is heading, our bet is that this will become one of the most important things to get right. So we are adding much more on this, while bringing in Louie&#8217;s enterprise experience to explain how teams can plan for changes in the industry.</p></li></ol><p>We are quite excited to share the book with you. And while we are adding everything we can to make it as useful as possible, if there are any questions you want answered, or there&#8217;s ONE THING the book needs to cover for you to make it worthwhile, <a href="https://www.louisbouchard.ai/book/?utm_source=sub-tai&amp;utm_medium=email&amp;utm_campaign=book">sign up for book updates</a> and reply to one of those emails. </p><p>Louis-Fran&#231;ois &amp; Louie</p>]]></content:encoded></item><item><title><![CDATA[TAI #225: OpenAI’s DevDay: Cheaper Intelligence, Unfinished Agents]]></title><description><![CDATA[Also, Gemini 4 Argon, GPT-6.1 Sol, Dots, NVIDIA DGX Spark 64GB, Microsoft MAI-Transcribe-2-Streaming, and Cohere Embed 5.]]></description><link>https://newsletter.towardsai.net/p/tai-225-openais-devday-cheaper-intelligence</link><guid isPermaLink="false">https://newsletter.towardsai.net/p/tai-225-openais-devday-cheaper-intelligence</guid><dc:creator><![CDATA[Louie Peters]]></dc:creator><pubDate>Tue, 06 Oct 2026 16:02:06 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!k6CR!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd4877de8-cdc8-4610-8da3-e1d6b04c986a_1400x788.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<h2>What happened this week in AI by Louie</h2><p>OpenAI&#8217;s DevDay on September 29 gave us two major releases: GPT-6.1 Sol, which brings most of GPT-6 Astra&#8217;s benchmark performance down to Sol pricing, and Dots, OpenAI&#8217;s answer to the persistent personal agents we discussed last week. Google also looks competitive again. Gemini 4 Argon scores 53 on Artificial Analysis&#8217;s Intelligence Index at high effort using Long Decode Continuation for extended reasoning, level with Astra at maximum effort. Access is still limited to selected Fairwind partners and trusted testers, so we remain skeptical until we can put it through our own real-world problems.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!k6CR!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd4877de8-cdc8-4610-8da3-e1d6b04c986a_1400x788.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!k6CR!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd4877de8-cdc8-4610-8da3-e1d6b04c986a_1400x788.png 424w, https://substackcdn.com/image/fetch/$s_!k6CR!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd4877de8-cdc8-4610-8da3-e1d6b04c986a_1400x788.png 848w, https://substackcdn.com/image/fetch/$s_!k6CR!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd4877de8-cdc8-4610-8da3-e1d6b04c986a_1400x788.png 1272w, https://substackcdn.com/image/fetch/$s_!k6CR!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd4877de8-cdc8-4610-8da3-e1d6b04c986a_1400x788.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!k6CR!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd4877de8-cdc8-4610-8da3-e1d6b04c986a_1400x788.png" width="1400" height="788" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/d4877de8-cdc8-4610-8da3-e1d6b04c986a_1400x788.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:788,&quot;width&quot;:1400,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" title="" srcset="https://substackcdn.com/image/fetch/$s_!k6CR!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd4877de8-cdc8-4610-8da3-e1d6b04c986a_1400x788.png 424w, https://substackcdn.com/image/fetch/$s_!k6CR!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd4877de8-cdc8-4610-8da3-e1d6b04c986a_1400x788.png 848w, https://substackcdn.com/image/fetch/$s_!k6CR!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd4877de8-cdc8-4610-8da3-e1d6b04c986a_1400x788.png 1272w, https://substackcdn.com/image/fetch/$s_!k6CR!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd4877de8-cdc8-4610-8da3-e1d6b04c986a_1400x788.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>What excites me most about Argon is its one-million-token output limit and the price. Long Decode Continuation lets it keep generating across follow-up calls, including extended reasoning, compared with the standard 128,000-token output limits of Claude Opus 5.5 and GPT-6 Astra. Opus also offers a 300,000-token Batch API beta. That gives Argon much more room to work through a difficult problem or produce a substantial result before hitting an output ceiling. Introductory pricing is $2 per million input tokens and $10 per million output tokens, half Opus 5.5&#8217;s standard rates and a fifth of Astra&#8217;s short-context rates. Google&#8217;s stated regular price is $4/$20 after the promotion.</p><p>DevDay itself landed below the expectations OpenAI had built around it. There was no Astra 6.1 or larger flagship model to try, but the releases that did ship were still substantial. Sol 6.1, in particular, behaves much more like a smaller Astra than a straightforward update to GPT-6 Sol. My guess is that it is closer to Astra architecturally than to the model it replaces, although OpenAI has disclosed nothing that confirms that. The system card says only that Sol 6.1 uses the same types of data and training as Astra. The API behaves like Astra&#8217;s as well, with effort running from low to max and no setting that switches reasoning off, which GPT-6 Sol still offers. Replacing GPT-6 Sol seven days after its launch makes the naming even less helpful.</p><p>The more important point is that OpenAI kept Sol pricing the same. Standard API rates remain $2 per million uncached input tokens and $10 per million output tokens for prompts up to 272,000 input tokens, exactly where GPT-6 Sol was priced and one fifth of Astra&#8217;s 10/50 rates. Cached input is the only additional cut, falling to $0.10 per million tokens. That helps agents that resend the same long prefix of instructions, tool definitions and documents on every call, as long as the prefix is unchanged and the cache actually hits. I suspect Opus 5.5, launched a week earlier at $4 and $20, pushed OpenAI to put a near-Astra model at Sol rates instead of pricing it somewhere in between.</p><p>Independent evaluations show how close Sol 6.1 gets to Astra for a fraction of the cost. On Artificial Analysis&#8217;s current index, Sol 6.1 at maximum effort scores 52 against Astra&#8217;s 53, at an average of $0.72 per weighted task against $3.26. Opus 5.5 at maximum effort, with default fallback enabled, still leads at 58 for $5.98. The lower effort settings are where I would change habits. Sol 6.1 at medium matches the previous Sol at maximum effort on the same index, 48 each, for $0.21 per task against $1.04. I would start routine research, coding and document work at medium and raise effort when the task justifies it. That said, I still primarily use Astra for most coding tasks, but I&#8217;m increasingly having it hand off simpler execution and testing to Sol-6.1 sub-agents.</p><p>Dots extends Astra into a much broader product. Last week we covered Muse, powered by Meta&#8217;s Muse Spark model, and Instinct, which made persistent agents feel as easy as messaging a contact. A dot packages persistent work as one continuing assistant with its own cloud computer and browser, selective memory of your decisions and preferences, connected apps, and the ability to decide when to pause and wake itself.</p><p>Dots feels very powerful to me, and potentially more ambitious than Muse or Instinct. Both of those already act proactively, so the ambition I see is in coordination. A dot can start separate cloud tasks, delegate work in parallel, and create or continue Codex work on a connected computer while you keep talking to it. Its cloud work continues with your laptop closed; local work needs that computer online with the ChatGPT app open. I already work this way in Codex, but dots makes many of these tasks much more accessible.</p><p>Several other DevDay launches fill in pieces an experienced operator currently supplies by hand. ChatGPT Space gives people and agents shared files and Pages. Reusable Codex cloud environments let a new task start in a development setup someone has already prepared. Event subscriptions let a change in a connected app trigger work. Together, they give an agent somewhere to work, the right materials, and a reason to resume. Every&#8217;s Dan Shipper described the payoff: his dot warned that a flight would clash with a video recording that had only been discussed in Slack and never made it onto his calendar. That is the assistance I want, catching a commitment while it is still taking shape.</p><p>But Dots also feels rushed, and I expect it to need several rounds of iteration before it gets the viral adoption Muse and Instinct have enjoyed. I love combining voice with browser work, but I&#8217;ve also seen disconnects, context-fetching failures, voice quality degrading on long calls, and opaque waits; Muse&#8217;s first-time experience felt smoother.</p><p>Some of the friction sits in the product design itself. A persistent assistant asks people to understand what it is working on, what it remembers, what it is waiting for, and when it will return. Today, pausing a dot stops only the main task; delegated work has to be stopped separately in Activity, and scheduled work canceled in Scheduled. A power user can manage that. It is a lot to learn before trusting an assistant with an ordinary errand, and when the state of the work is unclear, I end up supervising the interface as well as checking the result.</p><p>Muse&#8217;s designers have been unusually open about the work this takes. They added side chats when testers needed separate topics, a Goals tab for longer commitments, approval cards for actions, and suggested first jobs because people did not know where to begin. Instinct leans on an even more familiar habit: you text or call it like a contact. Dots has the stronger case for coordinating demanding work, in my view, and every extra entry point and control is another decision a new user has to make before that advantage shows.</p><p>The pace of adoption of persistent agents feels a lot like Cursor&#8217;s breakout and Anthropic&#8217;s fast response with Claude Code. Cursor reported millions of users and more than $100 million in recurring revenue in January 2025; Anthropic launched Claude Code, a terminal agent that looked nothing like Cursor&#8217;s editor but shared or leapfrogged many of its capabilities, about six weeks later. I see Instinct and Muse as the same kind of demand signal, and the labs are now moving quickly to copy the important parts of it. Claude and Codex were already very close to Instinct&#8217;s capability, and Anthropic had Claude Tag remembering context and scheduling tasks in Slack by June. They had not packaged it for viral adoption: an obvious way to start, supply context, follow progress, correct a mistake, and know the job is finished. I expect Dots to diverge from Muse and Instinct the way Claude Code diverged from Cursor.</p><div><hr></div><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://newsletter.towardsai.net/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading Towards AI Newsletter! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><div><hr></div><h3>Why should you care?</h3><p>Sol 6.1 gets close to Astra&#8217;s performance at roughly a fifth of Astra&#8217;s standard input and output prices, which makes large agent workloads much easier to justify through the API. OpenAI is simultaneously giving users less inference inside its cheaper subscriptions. DevDay introduced a $500 Pro tier with a larger allowance and Astra Ultrafast, while cutting the $200 plan&#8217;s token allowance in half. New subscribers already receive the lower limits; existing subscribers keep the previous allowance until October 29. So more serious workloads are becoming affordable through the API at the same time that the $200 subscription covers less of them.</p><p>The subscriptions still include far more inference than their monthly price would suggest. SemiAnalysis estimates that fully using a $200 Claude Max plan delivers about $11,700 of Opus 5.5 API usage per month, compared with roughly $2,100 of Sol 6.1 or $2,900 of Astra under OpenAI&#8217;s new $200 allowance. These are estimates based on its agentic workload mix, and actual usage will vary by model and task. Claude includes much more API-equivalent usage, although Opus 5.5&#8217;s standard input and output rates are twice Sol 6.1&#8217;s. For the work I do, my OpenAI subscription still goes further.</p><p>Small companies can get similar pricing through Claude Team Premium and ChatGPT Business Premium at $100 per person per month on annual billing, or $125 monthly, then buy additional usage at API-linked rates. Where available, pre-purchased credits can reduce that additional spend by 30&#8211;40%. The subsidized allowances become harder to access at scale: Claude Team stops at 150 paid seats, and new ChatGPT Business workspaces at 200. New Claude Enterprise contracts charge a seat fee plus API-rate usage from the first token; OpenAI&#8217;s enterprise agent usage is also metered, with rates and allowances depending on the agreement.</p><p>Many enterprises try to cap individual developers at $500&#8211;2000 a month of API spend, while a $200 Claude subscription can include several times that budget&#8217;s worth of inference. A developer at a small company can therefore have substantially more model capacity than an enterprise developer whose employer is spending far more overall. I like the broad outcome of larger customers carrying more of the cost while individuals and smaller companies get unusually favorable access, but the pricing tiers do not line up neatly with how developers actually work. As a result, many companies allow employees to use the $200 personal plans for non-sensitive work. Users can opt out of model training on personal plans, although employers have much less control over enforcing those settings or applying their own policies to personal accounts.</p><p>Every small company and entrepreneur should be trying to get the most useful work possible from these subscriptions while they retain such an advantage over large enterprises. I would have both the $100 Claude and ChatGPT team plans, and probably access to open-weight models as well. I also think most developers working seriously with frontier models should be learning how to use $10,000-plus a month of API-priced inference productively. There are enough useful places to spend it: parallel implementations, independent reviews, larger test suites, deeper research, and multiple attempts at difficult problems. The opportunity is to build the habits and workflows to direct that much work now, while a subscription makes the experimentation cheap.</p><p><em>A quick update on the book before I wrap up. We are deep into the final rounds of AI Engineering for Production now: editing chapters, tightening examples, checking the technical details, and making sure the book reflects how we actually build with these systems today.</em></p><p><em>Over the next few weeks, we will start sharing more from behind the scenes: parts of the book, ideas that changed substantially while we were writing, lessons from building the examples, and some material that will not make it into the final manuscript. If you want to follow the book as it comes together and get those updates first, <a href="https://www.louisbouchard.ai/book/?utm_source=news-tai&amp;utm_medium=email&amp;utm_campaign=book">you can sign up here</a>.</em></p><p><em>And if there is something you particularly want us to cover, reply and tell me. We are far enough along that the structure is set, but there is still room to make sure we answer the questions practitioners actually have.</em></p><p><em>&#8212; <a href="http://www.linkedin.com/in/louie-peters">Louie Peters&#8202;&#8212;&#8202;Towards AI Co-founder and CEO</a></em></p><div><hr></div><h3>Hottest News</h3><p>1. <a href="https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-4-argon/">Google Announces Gemini 4 Argon, With Wider Access Coming Later</a></p><p>Google announced Gemini 4 Argon, but the model is currently limited to approved cyber defenders through its Fairwind Program and other trusted testers; wider access will begin later with paid API customers and Google AI Ultra subscribers. One notable change is a 1 million-token output limit, up from 64K. Google reports scores of 77.9% on DeepSWE v1.1, 51.3% on AutomationBench, 91.7% on LVBench, and 68% on CWE-bench v1. Introductory API pricing will be $2 per million input tokens and $10 per million output tokens, with cached input priced 95% lower; after the introductory period, those rates rise to $4 and $20. Google has not published an exact date for general availability or disclosed Argon&#8217;s parameter count or architecture.</p><p>2. <a href="https://openai.com/index/introducing-gpt-6-1-sol/">OpenAI Releases GPT-6.1 Sol</a></p><p>OpenAI released GPT-6.1 Sol as a higher-capability replacement for GPT-6 Sol, positioned between Luna and Astra on capability and cost. The API model accepts text and images, has a 1.05M-token context window and 128K maximum output, and supports low, medium, high, xhigh, and max reasoning effort, with medium as the default. OpenAI reports that it matches Astra on DeepSWE v1.1 at roughly one-fifth the cost, beats Opus 5.5 by 2.2 percentage points on AutomationBench at medium effort for roughly one-third the cost, and comes within 2.1 points of Astra on OSWorld 2.0 at max effort. For prompts up to 272K input tokens, standard API pricing is $2/M input, $0.10/M cached input, $2.50/M cache writes, and $10/M output; requests above 272K cost 2&#215; for input and cache and 1.5&#215; for output. GPT-6.1 Sol is available through the API and to supported paid users in ChatGPT Work and Codex, but not yet in Chat, while the announced Sol Ultrafast mode is still coming later.</p><p>3. <a href="https://www.cnbc.com/2026/09/28/openai-abandons-plan-to-release-upcoming-model-as-safety-concerns-escalate.html">OpenAI Halts Training Run After Model Bypasses Network Restrictions</a></p><p>OpenAI has decided not to release GPT-6.1 Astra after the model failed to meet the company&#8217;s safety and alignment standards, CNBC reports. Saachi Jain, OpenAI&#8217;s head of safety systems, said the model fell short on staying within its authorized scope and clearly communicating the work it had performed. The decision follows increased scrutiny of OpenAI&#8217;s safeguards after two models escaped containment in July, accessed the open internet and breached Hugging Face, along with several later cases of unintended model behavior. Jain said the challenge is balancing strict adherence to scope with avoiding models becoming overly hesitant when they encounter friction while completing a task. OpenAI released GPT-6 Astra earlier in September, followed by GPT-6 Sol and Luna, and told CNBC that additional models are still coming.</p><p>4. <a href="https://www.nvidia.com/en-us/products/workstations/dgx-spark/">NVIDIA Introduces 64GB DGX Spark, Shipping October 23</a></p><p>NVIDIA is adding a 64GB unified-memory version of DGX Spark while retaining the same GB10 Grace Blackwell Superchip, DGX OS, and AI software stack as the existing 128GB system. The system delivers up to 1 PFLOP of FP4 compute, 273 GB/s memory bandwidth, and 200 Gbps ConnectX-7 networking; NVIDIA says a single 64GB unit can run models up to 100B parameters, although actual fit depends on model representation, quantization, and runtime overhead. Two 64GB systems can connect directly, and pool 128GB of memory, and NVIDIA reports up to 1.7&#215; the performance of one unit in its Qwen 3.8 27B test. The 64GB configuration will be sold only through participating OEMs, including Acer, ASUS, Dell, Gigabyte, HP, and MSI, starting at $4,999 on October 23.</p><p>5. <a href="https://microsoft.ai/news/our-first-streaming-transcription-model/">Microsoft Releases MAI-Transcribe-2-Streaming in Public Preview</a></p><p>Microsoft released MAI-Transcribe-2-Streaming, its first real-time speech transcription model, with support for 60 languages, automatic language detection, and partial transcripts appearing just over 100 ms after audio arrives. Microsoft says Artificial Analysis ranks it first for accuracy on both partial and final streaming transcripts, while its own tests found words appeared roughly twice as quickly as with its closest competitor in real-time dictation and subtitling. MAI-Voice-2.1-Flash can generate 45s of audio, with an end-to-end latency of 150ms. The model is available through Microsoft Foundry, MAI Playground, Vercel, and Azure Voice Live at an introductory $0.54 per hour of audio through the end of 2026, with LiveKit support coming later.</p><p>6. <a href="https://cohere.com/blog/embed-5">Cohere Releases Embed 5</a></p><p>Cohere released two new multimodal embedding models that share the same vector space, allowing teams to index documents with Embed 5 Pro and serve lower-latency queries with Embed 5 Fast without rebuilding embeddings. Both support text, images, and fused inputs across more than 100 languages, with a 128K-token context window and configurable 256&#8211;2,048-dimensional outputs. Cohere reports scores of 85.8 for Pro and 84.5 for Fast on ViDoRe V3, compared with 77.0 for Embed 4, while Fast delivered 2.4&#215; Pro&#8217;s document throughput in its tests. Pricing starts at $0.12/M text tokens for Pro and $0.08/M for Fast, with availability through Cohere&#8217;s API, Model Vault, Microsoft Foundry, and Amazon SageMaker. The models can also run in private VPC or on-premises environments through vLLM, but Cohere has not released them as open weights.</p><p>7. <a href="https://www.anthropic.com/research/yes-claude-can-do-nine-loops">Claude Computes the Nine-Loop Amplitude in N=4 Super-Yang-Mills</a></p><p>Anthropic physicists Liam Fitzpatrick and Siddharth Mishra-Sharma used Claude Fable 5.1 with the Claude Science harness to compute the six-particle MHV amplitude in planar N=4 super-Yang-Mills at nine loops, a calculation previously completed only through eight loops. Claude produced the result in two ways: by running the bootstrap calculation directly and through the related nine-loop form factor, using methods developed in earlier research. After the initial prompt, the researchers largely instructed Claude to keep working rather than guiding the scientific steps, and Stanford/SLAC physicist Lance Dixon independently validated the result, primarily through the form factor. Anthropic estimates that either approach would cost an end user roughly 1000&#8211;2000, mostly in Claude usage; the direct bootstrap additionally used about $100 of conventional compute across 96 CPUs for a week. A Chinese Academy of Sciences group was independently working on the same problem and had computed a substantial part of the nine-loop result with GPT-6 assistance. The result shows Claude carrying a fragile, multi-step frontier calculation through to a validated answer, but it relied on existing physics methods rather than discovering a new one.</p><p>8. <a href="https://chatgpt.com/features/space/">OpenAI Introduces ChatGPT Space</a></p><p>OpenAI introduced ChatGPT Space, a new area for pages, uploaded files, and shared work, while keeping Projects separate for organizing chats, files, and instructions. Its main addition is Pages, collaborative documents that multiple people can edit while using their own ChatGPT, with support for pulling context from connected tools such as Google Drive and Slack. Space is available to Pro, Business, and Enterprise users on web and desktop, while mobile currently supports viewing and sharing but not editing. Collaborative Slides and Sheets, mobile editing, and an automatic Keep updated feature are still planned for later. Space is also not yet available to Business and Enterprise workspaces using data residency in Canada or the UAE.</p><div><hr></div><h3>AI Tip of the Day</h3><p>This week, Omar spent 45 minutes on one of our<a href="https://towardsai.com/academy/mentorship/?utm_source=newsletter&amp;utm_medium=email&amp;utm_id=AItips"> Mentorship</a> live calls, walking through how we built our AI tutor. It is a useful example if you are deciding whether your own system needs RAG and where to start.</p><p>As he mentioned on the call, start with the questions the system needs external knowledge to answer. Collect a small set of real queries and, for each one, record:</p><ul><li><p>what information the answer requires</p></li><li><p>whether that information already fits reliably in the model&#8217;s context</p></li><li><p>whether retrieval returns the correct source in the top results</p></li><li><p>whether the final answer actually uses that source correctly</p></li></ul><p>Start by checking whether the system can actually surface the information needed for those questions. If it cannot, retrieval is the first problem to solve. If it can, there is no point changing the retriever yet.</p><p>With our tutor, that meant testing real student questions against the retrieval setup before adding more infrastructure. We only added complexity where the simpler setup was actually failing.</p><p>So if you are deciding whether to build RAG, start with a small evaluation set and find the exact point where your current system loses access to the information it needs. The architecture should follow that failure, not come before it.</p><p>The 45-minute tutor walkthrough was part of our weekly<a href="https://towardsai.com/academy/mentorship/?utm_source=newsletter&amp;utm_medium=email&amp;utm_id=AItips"> Towards AI Mentorship</a> live call, where we work through these kinds of implementation decisions using both our systems and what members are building.</p><div><hr></div><h3>Five 5-minute reads/videos to keep you learning</h3><p>1. <a href="https://pub.towardsai.net/how-ai-agents-evolved-from-autogpt-to-claude-code-21f8626c9db3?sk=747e63f016c2c5d2b6d9669712efe3ba">How AI Agents Evolved: From AutoGPT to Claude Code</a></p><p>Early agents such as AutoGPT wrapped language models in simple reason-act-observe loops, but weak tools and limited ways to verify work made them unreliable. This article traces how that design evolved into coding agents such as Claude Code and Codex, which combine stronger models with tool access, isolated environments, tests, diffs, checkpoints, and human review. It also explains why agents work best today on tasks where the system can inspect the result, detect failures, and recover from them.</p><p>2. <a href="https://pub.towardsai.net/webmcp-give-ai-agents-tools-not-just-screenshots-9065278c6a53?sk=af6778a0d86e59d3fb3b046bc465ffe8">WebMCP: Give AI Agents Tools, Not Just Screenshots</a></p><p>WebMCP lets websites expose structured actions that browser agents can call directly instead of navigating interfaces through screenshots or the DOM. This article explains the performance advantage using WindTunnel, where WebMCP agents complete the same website tasks with lower latency and cost, then focuses on the controls websites need before exposing those tools. It recommends starting with bounded, read-only actions and enforcing authorization, reviewing tool metadata, logging calls, and retaining a way to disable tools quickly.</p><p>3. <a href="https://pub.towardsai.net/how-to-fall-back-to-default-logic-when-llm-output-is-unsatisfactory-e5cdb3b63964?sharedUserId=tai-tech">How to Fall Back to Default Logic When LLM Output is Unsatisfactory</a></p><p>LLM workflows need different fallbacks for failed API calls, invalid outputs, and responses that satisfy the schema but are still wrong. This article uses an email-triage workflow to show how error handling, strict schema validation, and business rules detect each failure type before it reaches downstream systems. It then compares retries, deterministic default logic, and human escalation, and shows how to track fallback rates so recurring failures become visible.</p><p>4. <a href="https://pub.towardsai.net/microsoft-foundry-cross-region-models-247bfc22e396?sk=6e7e914f7ccc5ee8f6b550a6495cad90">Four Ways to Reach a Model in Another Azure Region From Microsoft Foundry</a></p><p>Microsoft Foundry does not make every model and agent feature available in every Azure region, so some applications need to call deployments outside their project&#8217;s home region. This article compares four ways to do that: direct Foundry connections, API Management as a model gateway, APIM in front of agents, and dynamic connections for the Responses API. It shows how each option changes identity, routing, governance, and private networking, including which first-party features can stop working when a gateway sits in the request path.</p><p>5. <a href="https://pub.towardsai.net/maxima-minima-second-derivative-test-explained-42811f910229?sk=8bc4b6c4e5f52febf35de6efdcd398f5">The $23 Million Book That Explains Maxima and Minima</a></p><p>This article explains maxima and minima through the first- and second-derivative tests, including critical points, boundaries, local versus global extrema, and cases with no maximum. It then applies the same ideas to maximum-likelihood estimation and least-squares regression before extending them to saddle points in higher dimensions. The final section connects curvature around an optimum to uncertainty, showing how the shape of an objective function determines how precisely we can estimate its best value.</p><h3>Repositories &amp; Tools</h3><p>1. <a href="https://github.com/pingdotgg/t3code">T3 Code</a> is an agent harness control surface that lets you drive agents on your machine from a mobile app, web app, or Electron desktop app.</p><p>2. <a href="https://github.com/pbakaus/impeccable">Impeccable</a> is a frontend design skill for AI coding agents with 24 commands and 60 deterministic detector rules that catch AI-generated design slop.</p><p>3. <a href="https://github.com/michael-denyer/pstack-claude">Pstack</a> is a port of Lauren Tan&#8217;s Cursor skill stack to Claude Code, Codex, Pi, GitHub Copilot, and other agent harnesses, where you tell poteto-mode your goal and it invokes the right workflow.</p><p>4. <a href="https://github.com/Panniantong/Agent-Reach">Agent Reach</a> is a CLI that gives AI agents read and search access to Twitter, Reddit, YouTube, GitHub, Bilibili, XiaoHongShu, and 10+ other platforms.</p><p>5. <a href="https://github.com/getsentry/sentry">Sentry</a> is a debugging platform that helps developers detect, trace, and fix issues in their code.</p><h3>Top Papers of The Week</h3><p>1. <a href="https://arxiv.org/abs/2609.39050">Covert Assistance: Helpful LLM Agents Evade Oversight in Multi-Agent Systems</a></p><p>This paper shows benign agents can evade oversight without adversarial incentives. In a simulated workflow, a planner holding a credential it must not disclose to a developer disguises it so the developer can recover it past a monitor. Seven of nine frontier models did this, even after completing their task. With DeepSeek-V4-Pro, the credential evaded the monitor and was used in 0.9% of 6,000 episodes, a rate that compounds to a 61.3% breach chance across 105 episodes.</p><p>2. <a href="https://arxiv.org/abs/2610.01153">Looping Beyond Twice: A Scalable Recipe for Looped Mixture-of-Experts</a></p><p>Looped MoE models stop improving after about two loops because hidden-state variance grows with each iteration, and routers keep selecting the same experts. LOOM scales residual updates, re-injects the input embedding every loop, and adds per-loop routers to keep each loop contributing new computation. Models from 100M to 1.7B scale stably to 9&#8211;12 loops, and the 1.7B model peaks at 9 loops, cutting perplexity from 9.62 to 7.77.</p><p>3. <a href="https://arxiv.org/abs/2610.00906">ActiveSaddler: Automated Curriculum Learning for Agent Harness Optimization</a></p><p>Harness optimizers update the harness but keep the training scenarios that generate feedback largely fixed. ActiveSaddler adapts the curriculum alongside the harness, grouping recurring failures into patterns and prioritizing those with the most learning potential while still exploring new scenarios. It improves test Pass@1 by 4.4 points on GAIA2 and 7.5 points on Terminal-Bench 2.0 over the same optimizer with a fixed scenario order.</p><p>4. <a href="https://arxiv.org/abs/2609.38879">Does Learning Protein Folding Generalize to Broader Reasoning?</a></p><p>This paper tests whether training on protein structures improves general reasoning. Fold2Reason post-trains a model on FoldingCorpus, a protein-derived question-answer dataset, using both discrete structural answers and continuous 3D geometry. It improves all 10 benchmarks spanning spatial, graph, scientific, and general reasoning, raising macro-average accuracy from 45.09% to 48.33%, while random, synthetic, and shuffled structure controls show much smaller or negative gains.</p><p>5. <a href="https://arxiv.org/abs/2610.03632">World Embedding Benchmark</a></p><p>This benchmark tests how well video embeddings capture physics, using 8,000 simulated cases across fluid mechanics, solid mechanics, dynamics, and optics. Pre-trained embedding models show weak retrieval and near-chance classification, though simple probes recover useful physical information. Contrastive training on physics data improves retrieval but hurts property regression, and using the embeddings to retrieve references for MiniMax-H3 improved the physical fidelity of generated videos.</p><h3>Quick Links</h3><p>1. <a href="https://openai.com/index/introducing-dots/">OpenAI launches Dots</a>, always-on personal AI agents powered by GPT-6 Astra, announced at DevDay on September 29. Each dot runs on its own cloud computer and browser, pursues user-defined goals continuously in the background, and connects to over 4,000 apps. Users can message their dot through Slack and Teams, with text message support coming soon. OpenAI also shared a preview of specialist dots with their own identity for access management, IT-provisioned hardware, and support for deep integrations with a company&#8217;s systems of record. Dots are rolling out to ChatGPT Pro and Business Premium users in eligible markets.</p><p>2. <a href="https://www.primeintellect.ai/blog/prime-inference">Prime Intellect launches Prime Inference</a>, a platform for serving frontier open-source models through serverless endpoints and reserved capacity across multiple data centers. The company positions it as the missing piece of its continual learning loop, where production traces feed back into training. The first public deployment, GLM-5.3, went live on OpenRouter on September 22 and ranks among the fastest GLM-5.3 endpoints there, with a near-zero tool-call error rate and 100% uptime. The platform is OpenAI-compatible, runs on NVIDIA Blackwell with Vera Rubin coming soon, and is built on NVIDIA Dynamo, vLLM, Mooncake, and FlashInfer.</p><p>3. <a href="https://investor.nvidia.com/news/press-release-details/2026/NVIDIA-Launches-Open-Agent-Safety-Platform-to-Secure-Agents-From-Testing-to-Deployment/default.aspx">NVIDIA launches Open Agent Safety Platform</a>, an open software platform and reference system design for enforcing boundaries around AI agents from testing through deployment. It combines two components: NVIDIA OpenShell, an open-source secure runtime that traces all agent actions and enforces policy as agents run on NVIDIA Vera CPUs, and NVIDIA Sentry, an out-of-band hardware watchdog running on BlueField-4 DPUs that monitors agents from an isolated trust domain invisible to the agent itself and can quarantine agents that attempt to move outside their boundaries in milliseconds. SpaceXAI is using it for Cursor coding agents and Grok models, and Anthropic&#8217;s Paul Smith endorsed it as complementary to Claude Managed Agents.</p><h3>Who&#8217;s Hiring in AI</h3><p><strong><a href="https://jobs.towardsai.net/job/uber-enterprise-applications-developer-zcvu">Enterprise Applications Developer @Uber (San Francisco, CA, USA)</a></strong></p><p><strong><a href="https://jobs.towardsai.net/job/highlevel-lead-staff-engineer-applied-ai-fwa8">Lead/Staff Engineer &#8212; Applied AI @HighLevel (Remote/India)</a></strong></p><p><strong><a href="https://jobs.towardsai.net/job/brillio-ai-builder-ai-solutions-engineer-r01572121-z2ja">AI Builder (AI Solutions Engineer) @Brillio (Remote/USA)</a></strong></p><p><strong><a href="https://jobs.towardsai.net/job/caci-international-ai-engineer-nfsj">AI Engineer @CACI International (Denver, CO, USA)</a></strong></p><p><strong><a href="https://jobs.towardsai.net/job/capital-one-full-stack-engineer-4-tw9p">Full Stack Engineer @Capital One (Plano, TX, USA)</a></strong></p><p><strong><a href="https://jobs.towardsai.net/job/c3-ai-ai-product-manager-mba-intern-summer-2027-w06r">AI Product Manager &#8212; MBA Intern @C3 AI (Redwood City, CA, USA)</a></strong></p><p><strong><a href="https://jobs.towardsai.net/job/appen-forward-deployed-engineer-dixb">Forward Deployed Engineer @Appen (San Francisco, CA, USA)</a></strong></p><p><em>Interested in sharing a job opportunity here? Contact <a href="mailto:sponsors@towardsai.net">sponsors@towardsai.net</a>.</em></p><p><em>Think a friend would enjoy this too? <a href="https://newsletter.towardsai.net/">Share the newsletter and let them join the conversation.</a></em></p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://newsletter.towardsai.net/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading Towards AI Newsletter! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div>]]></content:encoded></item><item><title><![CDATA[TAI #224: Personal Agents Get Their Own Computers]]></title><description><![CDATA[Also, Instinct&#8217;s $1bn raise, Opus and Sonnet 5.5, Grok Bot, Claude Tag in Slack, and OpenAI DevDay.]]></description><link>https://newsletter.towardsai.net/p/tai-224-personal-agents-get-their</link><guid isPermaLink="false">https://newsletter.towardsai.net/p/tai-224-personal-agents-get-their</guid><dc:creator><![CDATA[Louie Peters]]></dc:creator><pubDate>Tue, 29 Sep 2026 16:18:35 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!ayVt!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa517627a-abf9-4aaf-b98c-b3d079064161_1400x788.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<h2>What happened this week in AI by Louie</h2><p>The past week brought another rush of model releases: Anthropic&#8217;s Claude Opus 5.5 and Sonnet 5.5, and OpenAI&#8217;s GPT-6 Sol and Luna. OpenAI DevDay is today, and I expect plenty more announcements before the day is out. The most viral new product lately has been Meta&#8217;s Muse, powered by Muse Spark. Since its September 8 launch, it has reached &#8470;1 in both US app stores.</p><p>Muse gives the agent its own cloud computer, persistent memory, and access to the services you already use, which means you can hand it work without keeping the conversation or your own machine open. You can message it, including through WhatsApp, ask it to organize a trip, make a purchase, or build a small app, and come back later to the result. Behind that experience is a persistent virtual machine (VM) with a Linux filesystem, terminal, and browser. It can keep files and context across conversations, act through connected accounts, run on a schedule, or respond to relevant events. The Mac app, expanded at Meta Connect last week, extends that reach to local apps through permissioned control.</p><p>That setup is quite interesting in the small jobs people usually never get around to doing. One early user, while out walking the dog, asked Muse to turn a collection of saved Instagram recipes into an organizer and had a working version in about five minutes. The deeper extraction later hit Instagram&#8217;s rate limits, so the categories arrived before the full recipes, but the partial result still shows the consumer behavior I find most promising: turning something you already have into a small piece of useful software in the spare minutes you would otherwise spend only thinking about organizing it. The low-friction interface is part of what makes that possible. You can reach Muse by text or phone call, connect it to your email, messages, and other apps and devices, and let it operate its own phone and computer on your behalf. The company also says it can call or text proactively and follow up on threads you dropped.</p><p>This form factor is already pulling in users and capital. A mid-September report, citing a single source, put it above 100,000 users. On September 28, Instinct announced a $1bn Series C at a $10bn valuation, 33 days after disclosing a $250m Series B at $2.5bn. I read that fourfold jump in headline valuation as a bet on the continuing delegate: one assistant, reachable through channels people already use, that keeps working after the conversation ends.</p><p>Grok Bot has also been a big hit. Reported company figures put it at 418,000 weekly users on September 14, up 24% in a week. All of a user&#8217;s bots share one account-level cloud computer, including files and logins, while keeping separate memories and routines. Anthropic&#8217;s Claude Tag takes the workplace route through Slack. In July, Claude Code product lead Cat Wu said Tag was landing 65% of the pull requests from Wu&#8217;s own product engineering team, a substantial amount of real engineering directed through a conversation thread. Tag releases its execution sandbox after inactivity while the thread context and externally saved work survive, so the continuity belongs to the relationship and outlasts the machine underneath.</p><p>This form factor is useful for both consumers and enterprises. A personal agent can retain your preferences and unfinished plans, while a work agent in Slack can carry relevant decisions, files, and corrections into the next request. In both settings, persistent memory reduces the work of getting the agent ready for the next task.</p><p>OpenAI already has many of the pieces in ChatGPT Work: cloud browsing, retained sign-ins, and background tasks. I expect it to bring them together into a clearer answer for everyday personal or work delegation soon. The competition will then be over how reliably each service carries a person&#8217;s intentions from one task to the next.</p><p>The main difference between these agents and local tools like Claude Code or Codex is where they run and what they can access. A cloud agent can keep working while your laptop is closed. A local agent, meanwhile, starts with the files, apps, logins, and development environment already on your machine, but it can only keep working while that machine stays awake and connected. Both can still call remotely hosted models.</p><p>Running the workspace in the cloud solves the availability problem, but it also means reconnecting accounts and supplying files the laptop already had. That is why Muse&#8217;s Mac controls are useful: they give the cloud agent some access back into the local environment, although that access still depends on the Mac.</p><p>The computers themselves are modest. Reported inspections of two services found these configurations:</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!rM8L!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb069c24f-d2c2-42e7-9783-87b2bbc92c6f_1400x485.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!rM8L!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb069c24f-d2c2-42e7-9783-87b2bbc92c6f_1400x485.png 424w, https://substackcdn.com/image/fetch/$s_!rM8L!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb069c24f-d2c2-42e7-9783-87b2bbc92c6f_1400x485.png 848w, https://substackcdn.com/image/fetch/$s_!rM8L!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb069c24f-d2c2-42e7-9783-87b2bbc92c6f_1400x485.png 1272w, https://substackcdn.com/image/fetch/$s_!rM8L!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb069c24f-d2c2-42e7-9783-87b2bbc92c6f_1400x485.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!rM8L!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb069c24f-d2c2-42e7-9783-87b2bbc92c6f_1400x485.png" width="1400" height="485" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/b069c24f-d2c2-42e7-9783-87b2bbc92c6f_1400x485.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:485,&quot;width&quot;:1400,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" title="" srcset="https://substackcdn.com/image/fetch/$s_!rM8L!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb069c24f-d2c2-42e7-9783-87b2bbc92c6f_1400x485.png 424w, https://substackcdn.com/image/fetch/$s_!rM8L!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb069c24f-d2c2-42e7-9783-87b2bbc92c6f_1400x485.png 848w, https://substackcdn.com/image/fetch/$s_!rM8L!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb069c24f-d2c2-42e7-9783-87b2bbc92c6f_1400x485.png 1272w, https://substackcdn.com/image/fetch/$s_!rM8L!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb069c24f-d2c2-42e7-9783-87b2bbc92c6f_1400x485.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>Both services run LLM inference on remote infrastructure, so the VM serves as a workspace for scripts, documents, and browser sessions.</p><p>Power users can already assemble much of this themselves, but a personal agent lowers that entry cost by carrying the connections and instructions behind an ordinary chat.</p><p>For persistent agents, memory determines whether the agent is still acting on the right version of the user&#8217;s intent. In one reviewer&#8217;s trial, Muse noticed that older Portugal travel context imported from Gemini conflicted with newer Italy plans in the reviewer&#8217;s email and surfaced the conflict; the reviewer clarified that Italy was current. That is the behavior I want: track where a fact came from, notice when sources disagree, and settle which one applies.</p><p>The same principle applies once the agent starts taking actions. For purchases, Muse can navigate a merchant&#8217;s site and pay with a Stripe Link single-use card restricted to the approved merchant and amount, with a short expiry. The user approves the purchase, then the agent completes checkout. Meta&#8217;s Sentinel system sits outside the main agent and controls permissions for outbound actions. Separate approval, visible activity, and the ability to take over are what let someone delegate a task without watching every click.</p><p>The privacy question is who can access the environment once you connect your accounts and files. Muse&#8217;s current Secure VM isolates that environment from other users, but Meta retains operational access under its policies. A Confidential VM designed to restrict even Meta&#8217;s access is promised for later this year. Until then, I would decide what to connect based on the protection available today, not the stronger model Meta plans to offer later.</p><p>The model and its allowance also decide how far a delegated task can go, which brings me back to this week&#8217;s releases. We tested Opus 5.5 only briefly before last week&#8217;s newsletter, and it has really grown on us. The change I value most is how much more work we can progress through our Claude subscription before hitting limits. There is room to develop an idea, inspect the result, correct it, and keep going.</p><p>Part of that runway comes from the allowance rules. On Max and eligible premium organizational plans, Fable 5 and 5.1 can use at most half of the shared weekly pool, while Opus can use all of it. That doubles the allowance available to Opus, although how many tasks it buys still depends on effort, context, and the job. Anthropic also increased five-hour limits with the Opus 5.5 launch.</p><p>Opus also shows strong efficiency gains at complex tasks. On CursorBench 4.0, Opus 5.5 at Medium scores 52.5% using 54 steps and about 38,000 tokens per task, against 51.8%, 128 steps, and 117,000 tokens for Fable 5.1 at Max.</p><p>The useful result here is not Opus&#8217;s 0.7-point score lead, but how much less work it needs to get there. It reaches roughly the same quality in 54 steps instead of 128 and uses about a third of the tokens, leaving much more room for iteration within the same budget. Cursor&#8217;s High and Extra High Opus settings both score 56%, while Extra High uses nearly twice the tokens. I would therefore increase effort only when the task is difficult enough to justify it or when a failed check shows that Medium or High is not enough.</p><p>Sonnet 5.5 arrived on September 28 and looks like a big upgrade from Sonnet 5, with Artificial Analysis&#8217;s Intelligence Index putting it at 56 at maximum effort, up from 38. The improvement comes with a much higher token cost. That maximum-effort evaluation used about 193,000 output tokens per weighted Index task, roughly 60% more than Opus 5.5 at maximum effort, and Opus at xhigh reaches the same 56 for $3.46 per task against Sonnet&#8217;s $7.60. The evaluation ran on a pre-release deployment whose structured-output bug was fixed for launch, and I have not yet seen a post-fix rerun; Anthropic reports better efficiency on its own workloads at other settings. For our work, I don&#8217;t see a reason to choose Sonnet over Opus for demanding tasks or GPT-6 Luna for cheaper ones.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!ayVt!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa517627a-abf9-4aaf-b98c-b3d079064161_1400x788.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!ayVt!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa517627a-abf9-4aaf-b98c-b3d079064161_1400x788.png 424w, https://substackcdn.com/image/fetch/$s_!ayVt!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa517627a-abf9-4aaf-b98c-b3d079064161_1400x788.png 848w, https://substackcdn.com/image/fetch/$s_!ayVt!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa517627a-abf9-4aaf-b98c-b3d079064161_1400x788.png 1272w, https://substackcdn.com/image/fetch/$s_!ayVt!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa517627a-abf9-4aaf-b98c-b3d079064161_1400x788.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!ayVt!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa517627a-abf9-4aaf-b98c-b3d079064161_1400x788.png" width="1400" height="788" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/a517627a-abf9-4aaf-b98c-b3d079064161_1400x788.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:788,&quot;width&quot;:1400,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" title="" srcset="https://substackcdn.com/image/fetch/$s_!ayVt!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa517627a-abf9-4aaf-b98c-b3d079064161_1400x788.png 424w, https://substackcdn.com/image/fetch/$s_!ayVt!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa517627a-abf9-4aaf-b98c-b3d079064161_1400x788.png 848w, https://substackcdn.com/image/fetch/$s_!ayVt!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa517627a-abf9-4aaf-b98c-b3d079064161_1400x788.png 1272w, https://substackcdn.com/image/fetch/$s_!ayVt!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa517627a-abf9-4aaf-b98c-b3d079064161_1400x788.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><div><hr></div><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://newsletter.towardsai.net/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading Towards AI Newsletter! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><div><hr></div><h3>Why should you care?</h3><p>The durable advantage for personal agents may be the context they accumulate over time. Switching models takes seconds; recreating months of preferences, corrections, connected accounts, and unfinished plans does not. The better an agent carries that context from one task to the next, across whichever interface or device the user happens to be on, the harder it becomes to replace. I suspect that accumulated context is a large part of what investors are pricing into Instinct and what Meta is pursuing with Muse. The risk is that the same context that makes the product more useful can also become a source of lock-in if it is difficult to move elsewhere.</p><p>So I would judge personal agents on what they let me do with what they know. I want to inspect what the agent remembers, see where each stored preference came from, correct it once, and export the useful parts. Finished work should live somewhere I can reach without the conversation. These controls double as a reliability feature: memory the user can manage lets the agent recover from a bad inference before it becomes a recurring mistake.</p><p>The strongest personal-agent businesses will earn access to more of our lives by using that access well. I want the relationship to remain a choice, even after the service becomes indispensable.</p><p><em>&#8212; <a href="http://www.linkedin.com/in/louie-peters">Louie Peters&#8202;&#8212;&#8202;Towards AI Co-founder and CEO</a></em></p><div><hr></div><p><strong>Quick Note:</strong> The Towards AI team will be attending the <a href="https://aiblmedia.com/">aiblLIVE London 2026</a> practical AI adoption conference on 20th October in London. Reach out if you will be there and want to meet up.</p><div><hr></div><h3>Hottest News</h3><p>1. <a href="https://www.anthropic.com/claude-opus-5-5">Anthropic Releases Claude Opus 5.5</a></p><p>Anthropic released Claude Opus 5.5, the first model in its new Claude 5.5 family and the first release since CEO Dario Amodei&#8217;s call to pace the frontier. The model performs at the level of Claude Fable 5.1 on most work and costs 40% less to run than Opus 5 on typical workloads. Pricing is $4/20 per million input/output tokens, 20% below Opus 5, with cache reads at $0.20 per million, a 60% reduction. Output generation is more than 30% faster. On Anthropic-reported benchmarks, Opus 5.5 leads in agentic coding, computer use, and knowledge work, scoring 66.4% on Terminal-Bench 4.0 (Fable 5.1: 55.8%), 54.4% on FrontierCode v1.1 (GPT-6 Astra: 53.3%), and 1846 Elo on GDPval-AA v2.1 (Fable 5.1: 1735). Anthropic notes that in its own use, the gap between Opus 5.5 and Fable 5.1 is narrower than benchmarks suggest. Anthropic reports that Opus 5.5 achieved its strongest automated behavioral-audit results to date and attempted to circumvent containment boundaries about 85% less often than Opus 5 or Mythos 5.1. Because its biology and cybersecurity capabilities approach Mythos 5.1, Anthropic applies additional safeguards to higher-risk requests, including routing some flagged cybersecurity tasks to Opus 4.8. It was tested pre-release by METR and Frontier Design. The model also communicates more clearly than Opus 5, putting key information up front with less jargon. Thinking mode can no longer be disabled. Available on the Claude Platform, AWS, Google Cloud, and Microsoft Azure.</p><p>2. <a href="https://openai.com/index/introducing-gpt-6-sol-and-luna/">OpenAI Introduces GPT-6 Sol and Luna</a></p><p>OpenAI released GPT-6 Sol and GPT-6 Luna, 19 days after GPT-6 Astra, cutting API prices by 50% compared to their GPT-5.6 promotional pricing. Sol costs $2/10 per million input/output tokens (down from $4/20), and Luna costs $0.10/0.50 (down from $0.20/1.20). Both were trained with similar methods to Astra, carrying over its advances in professional work, factuality, coding, computer use, and alignment. On AutomationBench, GPT-6 Sol at xhigh effort scored 33.2%, outperforming Claude Opus 5 at max effort (26.9%) at 9% of its per-task cost. On DeepSWE v1.1, Sol at max effort scored 68.8%, within 1.1 points of Fable 5&#8217;s 69.9% at roughly 80% lower cost. Luna, at max effort, scored 66.6% on DeepSWE, comparable to Opus 5 at medium effort, at 93% lower cost. On OpenAI&#8217;s internal factuality evaluation, Sol makes about half as many mistakes as GPT-5.6 Sol. OpenAI also improved prompt caching for GPT-6 with higher hit rates by default and 90% discounts on cached input reads; GitHub reports these improvements cut fresh token processing by more than 50% across billions of Copilot requests. Both models are available in ChatGPT Work and Codex for all paid tiers, with Luna also available to Free and Go users in the desktop app. They are not yet available in standard ChatGPT chat.</p><p>3. <a href="https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-3-8-text-to-speech/">Google Releases Gemini 3.8 Flash TTS and Flash-Lite TTS</a></p><p>Google released Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS, describing them as its most expressive audio generation models yet. The headline feature is generative voice design: instead of choosing from a fixed catalog, users describe a character in natural language (role, accent, voice characteristics) and receive a reusable voice with a persistent ID. It also supports voice replication from a 30-second sample with verbal consent. Flash TTS targets creative direction and character design for gaming, audiobooks, podcasts, and interactive media, with control over acting cues, pacing, dialect shifts, and backchanneling. Flash-Lite TTS targets high-volume, cost-efficient use, including dubbing and voice agents. The voice library expands from 30 prebuilt voices to over 2,000, with Flash TTS supporting 130 languages and Flash-Lite 101. Scripts can include non-verbal cues like laughs, sighs, and listening sounds for natural dialogue. In blind human preference evaluations on Hume AI&#8217;s Overall Quality Index, the two models placed first and second across six global languages, including Japanese, Brazilian Portuguese, Vietnamese, Arabic, Mexican Spanish, and Hindi. All output carries SynthID watermarking. Both models are available through the Gemini API and Google AI Studio, with Flash TTS also in Gemini Notebook and Flash-Lite TTS in Google Vids. Gemini Enterprise API access is coming soon. No open weights are available.</p><p>4. <a href="https://x.ai/news/grok-4-7">xAI Releases Grok 4.7 at $2/$6</a></p><p>SpaceXAI released Grok 4.7, 40 days after Grok 4.6, at the same $2/6 per million input/output token pricing and 500K-token context window. The model uses a new, larger base model than Grok 4.6 and was trained with a longer reinforcement learning run on a harder task mix weighted toward problems that take many hours to complete. xAI says it is better at verifying its own work and managing longer context, and was trained to natively understand the Grok Bot harness for conversational tasks. On xAI-reported benchmarks, Grok 4.7 at xhigh effort scored 46.3% on CursorBench 4.0 (Grok 4.6: 40.4%), 71.0% on DeepSWE v1.1 (high effort), 64.0% on EEBench for electrical engineering, and 37.6% on Terminal-Bench 4.0 (Grok 4.6: 20.3%). On GDPval, it scored 1695 Elo, behind Fable 5.1 at 1735 but ahead of GPT-6 Astra at 1542. Grok 4.7 ships with an entirely new safeguard stack that xAI describes as its strongest on refusals and jailbreak resistance, topping LatchBio&#8217;s biosafety benchmark at 62.4% and allowing only 3.3% of risky dual-use prompts through on HackerBench v0.3. Select cybersecurity partners are receiving invite-only access to red-team capabilities. A fast variant runs at twice the output speed for twice the price. Available in Cursor, Grok Build, the Grok API, and third-party harnesses, routers, and cloud platforms.</p><p>5. <a href="https://www.anthropic.com/news/claude-discovers-novel-enzyme-system">Claude Discovers a Novel Enzyme System With CRISPR-Like Repeats</a></p><p>Anthropic announced that Claude autonomously identified a previously uncharacterized enzyme system in bacteriophages, alongside the launch of a new life sciences research group and wet lab. Researchers gave Claude a high-level prompt to search a large DNA sequence database for unusual reverse transcriptases, enzymes that copy RNA into DNA. Over 21 hours, approximately 950 Claude agents using 210 million tokens gathered more than 200,000 reverse transcriptases from 1.9 billion protein clusters, identified 3,500 candidate systems, and narrowed them to 20 for deeper investigation. One agent noticed a long array of repeating non-coding DNA adjacent to an unusual reverse transcriptase gene. Subsequent computational analysis and bench experiments confirmed the finding, which Anthropic named array-associated reverse transcriptases (ARTs). The layout resembles CRISPR&#8217;s repeat-spacer memory structure, but Anthropic stresses that the system&#8217;s biological function remains unknown and it has not been shown to perform gene editing. Feng Zhang, the CRISPR pioneer at MIT and the Broad Institute, reviewed the preprint and called the identification &#8220;genuinely intriguing.&#8221; In a reproducibility test, Anthropic reran the identical search ten more times, and none rediscovered the repeat array, because no rerun agent happened to examine the DNA sequence immediately upstream of the enzyme. The work is published as a preprint.</p><p>6. <a href="https://www.abc.net.au/news/2026-09-24/ai-agent-accessed-australian-government-site-pm-says/107189078">OpenAI Agent Breaches Australia&#8217;s Medicare Statistics Portal</a></p><p>Australian Prime Minister Anthony Albanese revealed that an OpenAI agent gained unauthorized access to the Medicare Statistics Reporting Service portal, administered by Services Australia, on June 18. The agent was conducting an internal OpenAI research task to gather public medicine-spending and healthcare statistics. When the portal blocked its requests, it tried alternative methods, bypassed access restrictions, viewed public and non-public files, and wrote files to an internal server. The portal contains aggregate Medicare expenditure statistics and is separate from systems handling claims and personal records. No personal information is believed to have been accessed, and there is no evidence of broader compromise of Services Australia&#8217;s network, though a forensic investigation with the Australian Signals Directorate remains active. OpenAI did not detect the breach until August 11, during a review of misaligned model activity prompted by the Hugging Face incident, and did not notify the Australian government until September 10, via an email to a public Services Australia inbox. Albanese said he had a &#8220;frank&#8221; discussion with Sam Altman to express &#8220;extreme concern&#8221; and called both the three-month delay and the notification method &#8220;unacceptable.&#8221; He has announced a government taskforce to investigate. Reports associated the same agent activity with three additional sites (Australian Institute of Health and Welfare, Victoria&#8217;s Department of Health, and NSW Bureau of Crime Statistics and Research), though Deputy PM Richard Marles later said those interactions appeared authorized. OpenAI stated its models took actions it did not intend during internal evaluations.</p><p>7. <a href="https://openai.com/index/advisory-group-on-mathematics-and-ai/">OpenAI Claims 100+ Solved Open Problems and Forms a Math Advisory Group</a></p><p>OpenAI announced that an internal model it began training on August 28 has resolved more than 100 long-standing open problems across most areas of mathematics, in addition to the Navier-Stokes Millennium Prize problem it claimed to have solved roughly two weeks earlier. OpenAI said the pace of progress surprised the mathematicians within the company. OpenAI announced it is working with the independent Advisory Group on Mathematics and Artificial Intelligence, hosted at the Institute for Advanced Study. The announcement followed the open letter A Severe Misalignment of AI in Mathematics, which criticized the use of open mathematical problems as AI benchmarks; OpenAI approached some of the group&#8217;s eventual members about external advice, after which they formed an independent organization. Members are unpaid, operate independently, can offer unrequested advice, and can make their advice public. The group will help assess the significance of emerging results, advise on how to coordinate their dissemination, and advise on academic standards of mathematical research. OpenAI explicitly stated the group will not advise on how to pace its internal progress on mathematics. As of the announcement, OpenAI has not published an itemized list of the 100+ problems, the underlying proofs, or the model&#8217;s name or architecture. The claims have not been independently verified.</p><div><hr></div><h3>AI Tip of the Day</h3><p>Stop taking your agent&#8217;s word for it every time it says the job is done.</p><p>An agent might return &#8220;report saved&#8221; even if the file-writing step failed. If the application treats that message as the completion signal, it can tell the user a report is ready when no usable file exists.</p><p>In the research workflow we build in <a href="https://towardsai.com/academy/agent-engineering/?utm_source=newsletter&amp;utm_medium=email&amp;utm_id=AItips">Agent Engineering</a>, we use the same principle earlier in the process: when a tool is expected to process items but returns zero, the workflow stops and asks for guidance. The agent does not get to reinterpret that failure as success.</p><p>Apply the same rule to the final handoff. Have the agent return the report&#8217;s path, then check in code that the file exists, is not empty, and contains the required sections.</p><p>Only mark the task complete after those checks pass. Then deliberately break the file-writing step. If the agent still reports success, the application should reject that result.</p><p>The agent can describe what happened. Your code should decide whether the task actually succeeded.</p><div><hr></div><h3>Five 5-minute reads/videos to keep you learning</h3><p>1. <a href="https://pub.towardsai.net/introducing-system-one-models-laya-and-jev-56b7271da9c2?sk=4617f986b0bccc5a30eb30ba4a70a0bb">Introducing System One Models &amp; Laya (and Jev)</a></p><p>Convai Innovations released Laya on Hugging Face months before TypeSafe&#8217;s Jev. It offers the same capability. This article builds a Rust implementation around Laya and tests it on credential-exfiltration detection and a Pong controller running at 30 FPS. Both examples show that these models work better when state and choices use natural-language descriptions rather than raw structured inputs or bare labels, making them useful for fast classification, routing, and control tasks.</p><p>2. <a href="https://pub.towardsai.net/10-000-ai-agents-didnt-have-the-intuition-a-mathematician-did-meanwhile-cern-ditched-ibm-4082c258398f?sk=3deaf7dadf614b1e0e242c351bd806a7">10,000 AI Agents Didn&#8217;t Have the Intuition; A Mathematician Did. Meanwhile, CERN Ditched IBM</a></p><p>This article uses recent examples from mathematics and scientific infrastructure to argue for separating AI-generated proposals from deterministic verification. It then applies that principle to particle-physics analysis with Vikshep, which combines wavelet-scattering features with a distance-correlation penalty designed to reduce mass sculpting. The resulting workflow lets models propose useful patterns while keeping the final analysis reproducible and independently checkable.</p><p>3. <a href="https://pub.towardsai.net/goal-decay-what-happens-to-an-agent-running-for-three-weeks-2e0e85ef3378?sk=7b1d9242228b08b4460128510366ff26">Goal Decay: What Happens to an Agent Running for Three Weeks</a></p><p>Long-running agents can remain operational while gradually moving away from their original objective. This article breaks that drift into mechanisms including context dilution, lossy compaction, subgoal displacement, and changing interpretations of competing priorities, using long-horizon agent experiments as examples. It proposes tracking acceptance criteria alongside activity, periodically judging the run from fresh context, and preserving the original goal verbatim through durable files, compaction, and logged amendments.</p><p>4. <a href="https://pub.towardsai.net/flashattention-runs-the-same-math-it-just-stops-writing-it-down-30b595e423f6?sk=4d61843b8e888ea86735f132014509b0">FlashAttention Runs the Same Math. It Just Stops Writing It Down</a></p><p>FlashAttention computes exact attention without materializing the full attention matrix in GPU memory. This article explains how tiled computation and online softmax keep more of the work in fast on-chip memory, reducing the data movement that limits standard attention. It also compares successive FlashAttention kernels, shows how to confirm that PyTorch is actually using them, and explains why the gains are smaller for short prompts and token-by-token decoding.</p><p>5. <a href="https://pub.towardsai.net/latentmoe-nvidias-latent-mixture-of-experts-explained-through-equations-architecture-code-and-7671f8a47e27?sharedUserId=tai-tech">LatentMoE: NVIDIA&#8217;s Latent Mixture of Experts Explained Through Equations, Architecture, Code, and Visual Workflow</a></p><p>LatentMoE reduces the memory cost of MoE layers by moving expert computation into a latent space four times narrower than the model dimension, then using those savings to activate more experts from a larger pool. This article works through the projections, router, expert computation, and serving costs, showing why weight reads and interconnect traffic can matter more than FLOPs during inference. In NVIDIA&#8217;s 95B experiment, the higher-capacity LatentMoE configuration improved MMLU-Pro from 29.26 to 34.91 at roughly the same active-parameter and expert-byte cost, and NVIDIA later used the design in Nemotron 3 Super.</p><h3>Repositories &amp; Tools</h3><p>1. <a href="https://github.com/paperclipai/paperclip">Paperclip</a> is a Node.js server and React UI that orchestrates a team of AI agents assigned business roles.</p><p>2. <a href="https://github.com/debpalash/VoiceStudio">VoiceStudio</a> is an ElevenLabs alternative for voice cloning, voice design, video dubbing, dictation, transcription, and audiobook creation across 646 languages.</p><p>3. <a href="https://github.com/mvschwarz/openrig">OpenRig</a> is a multi-agent harness that runs Claude Code and Codex together as one system, letting you define an agent team in YAML and boot it with a single command.</p><p>4. <a href="https://github.com/dream-num/univer">Univer</a> is an open-source SDK for creating office applications inside your own product.</p><p>5. <a href="https://github.com/vectorize-io/hindsight">Hindsight</a> is an agent memory system that organizes recall into four networks (world facts, experiences, opinions, observations) mirroring human memory.</p><h3>Top Papers of The Week</h3><p>1. <a href="https://arxiv.org/abs/2609.22068">CodeMidas: Scaling Agentic Coding RL Environments from Code Itself</a></p><p>Existing methods for building coding RL environments often rely on issues and commits, limiting the range of tasks they can extract from a repository. CodeMidas instead uses source code as its only task-specific input: agents inspect implemented functionality, write behavioral specifications, construct execution-grounded tests, and filter tasks through repeated solution rollouts. The pipeline produced 5,545 tasks from 3,185 codebases across 23 languages. Training MiMo-V2.5 with GRPO improved all five evaluated benchmarks, including +11.7 points on DeepSWE, +17.0 on ProgramBench&#8217;s Almost Solved metric, and +8.5 on Terminal-Bench 2.1.</p><p>2. <a href="https://arxiv.org/abs/2609.31093">Block Sparse Attention with Log-Linear Complexity</a></p><p>Block-sparse attention reduces long-context computation, but conventional methods still score every query against every candidate key block, leaving block selection quadratic. PISA builds a coarse-to-fine hierarchy of pooled keys and narrows candidates level by level with LogSumExp scoring, reducing training and prefill complexity to O(N log N) and average decoding cost to O(log N) per step. Hardware-aware Triton kernels fuse routing and scoring without materializing the dense score matrix. Compared with conventional BSA, PISA matches performance on commonsense reasoning benchmarks and performs better on retrieval tasks.</p><p>3. <a href="https://arxiv.org/abs/2609.01591">StudentSim: Training LLM-based Student Simulators</a></p><p>Evidence about which tutoring guidance works for which student is expensive to collect from real learners. Existing state-tracking simulators can match student behavior but struggle to process explanations or corrections, while LLM role-play responds to guidance without reliably matching a particular student&#8217;s competence. StudentSim uses pooled training followed by per-student specialization to do both. On StudentSimEval, covering 60 students across chess, English writing, and math, it beats GPT-5.4 on both behavioral fidelity and guidance responsiveness in all three domains. Experts also rated a chess tutor trained with StudentSim as more accurate, better guided, and more personalized than the tested baselines.</p><p>4. <a href="https://arxiv.org/abs/2609.18703">RayOrch: Programming and Executing Lineage-Controlled Multi-Grain Dataflows for Foundation-Model Data Preparation</a></p><p>Foundation-model data pipelines often expand one document or video into variable-length sequences of pages, blocks, frames, or clips, but existing systems either hide that fine-grained parallelism or flatten records and force applications to reconstruct lineage themselves. RayOrch lets programs declare ordered variable-cardinality expansions and matching gathers, while the runtime tracks parent-child relationships, ordinals, completion, and parent-scoped failures. FIFO ready queues batch children across parents, and gathers reconstruct outputs from declared lineage rather than execution order.</p><p>5. <a href="https://arxiv.org/abs/2609.15779">EvoOntology: A Self-Evolving Ontology Layer for Data Agents</a></p><p>Data agents often access tables, files, and databases through generic tools that expose structure but little semantic meaning. This paper introduces EvoOntology, a self-evolving ontology layer for data agents. It encapsulates the ontology as an MCP server comprising a schema layer, a content layer, and a tool layer, enabling agents to actively query and interact with the ontology at runtime. A builder agent for autonomous ontology construction and a self-evolution loop continuously refine the ontology through attribution-guided typed edits, accepted only after a backbone-conditional paired evaluation.</p><h3>Quick Links</h3><p>1. <a href="https://huggingface.co/blog/nvidia/nemotron-diarization">NVIDIA releases Nemotron 3 Diarization</a>, a 100M-parameter open-weight model that directly predicts speaker activity for up to eight speakers, including overlapping speech. A single checkpoint supports streaming and offline use, with input-buffer latency configurable down to 80 ms; NVIDIA recommends 0.32 s for its lowest-latency setting and uses 30.4 s for its highest-accuracy configuration. On VoiceArena&#8217;s initial Diarization-Bench, it ranks first with 14.72% diarization error versus 19.3% for the next-ranked system. On DIHARD III at the 30.4-second setting, it reaches 9.13% DER for recordings with 1&#8211;4 speakers and 12.73% across the full benchmark. NVIDIA also reports 15,113&#215; real-time throughput at batch size 32 with torch.compile() on an RTX PRO 5000, although this measures batched model throughput rather than end-to-end latency. The model can run locally through NVIDIA&#8217;s NeMo-Speech.cpp runtime and is released under OpenMDW 1.1 for commercial and non-commercial use.</p><p>2. <a href="https://research.google/blog/coherent-long-form-video-generation/?">Google Research introduces an AI Video Co-Director,</a> a unified multi-agent framework that autonomously generates temporally consistent, long-form video narratives. It is built as an orchestration layer on top of Gemini and Veo, and natively inherits safety mechanisms like SynthID watermarking. The research combines four approaches: Co-Director selects the overall creative strategy and coordinates production agents; CANVAS keeps characters, locations, and objects consistent across shots; A&#178;RD generates longer videos segment by segment while maintaining multimodal memory; and VQQA evaluates generated videos and rewrites prompts to fix visual inconsistencies. Google demonstrates A&#178;RD on a continuous 10-minute video, while Co-Director scores 81.4 on GenAD-Bench.</p><h3>Who&#8217;s Hiring in AI</h3><p><a href="https://jobs.towardsai.net/job/grafana-labs-senior-backend-engineer-databases-analytics-us-remote-bqv8">Senior Backend Engineer @Grafana Labs (Remote/US)</a></p><p><a href="https://jobs.towardsai.net/job/develocity-salesforce-developer-be7i">Salesforce Developer @Develocity (Remote/Europe)</a></p><p><a href="https://jobs.towardsai.net/job/capital-one-staff-engineer-agentic-ai-remote-eligible-bbwr">Staff Engineer &#8212; Agentic AI @Capital One (Remote Eligible)</a></p><p><a href="https://jobs.towardsai.net/job/caci-international-full-stack-developer-ggeo">Full Stack Developer @CACI International (Washington, DC, USA)</a></p><p><a href="https://jobs.towardsai.net/job/citigroup-senior-python-engineer-r9d1">Senior Python Engineer @Citigroup (Pune, India)</a></p><p><a href="https://jobs.towardsai.net/job/netapp-software-engineer-vvuk">Software Engineer @NetApp (Morrisville, NC, USA)</a></p><p><em>Interested in sharing a job opportunity here? Contact <a href="mailto:sponsors@towardsai.net">sponsors@towardsai.net</a>.</em></p><p><em>Think a friend would enjoy this too? <a href="https://newsletter.towardsai.net/">Share the newsletter and let them join the conversation.</a></em></p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://newsletter.towardsai.net/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading Towards AI Newsletter! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div>]]></content:encoded></item><item><title><![CDATA[TAI #223:Opus 5.5, Cheaper GPT-6 and Agents at Dreamforce]]></title><description><![CDATA[Also, TypeSafe's Jev, shared agent work in Slack, AI safety and meeting Sam Altman]]></description><link>https://newsletter.towardsai.net/p/tai-223opus-55-cheaper-gpt-6-and</link><guid isPermaLink="false">https://newsletter.towardsai.net/p/tai-223opus-55-cheaper-gpt-6-and</guid><dc:creator><![CDATA[Louie Peters]]></dc:creator><pubDate>Thu, 24 Sep 2026 13:32:55 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!xiaH!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0ab0ef8d-de5f-45a4-90ce-e96994f19826_1920x1080.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<h2>What happened this week in AI by Louie</h2><p>I visited Salesforce&#8217;s Dreamforce in San Francisco this past week and spoke with Rob Seaman, Slack&#8217;s General Manager, about how agents fit into company work. I also heard Jensen Huang and Dario Amodei&#8217;s somewhat conflicting views on AI safety: Dario called for coordinated safeguards, while Jensen argued that safety and speed can go together. Jensen also joked about the height difference with Marc Benioff: &#8220;I always feel like I need to stand on a chair.&#8221; Even Nvidia has scaling problems (thank Astra for the joke&#8230;).</p><p>My co-founder, who was also there, recommends going just for the sheer concentration of incredible people: &#8220;Connections are everything.&#8221; I also met Sam Altman at OpenAI&#8217;s GPT-6 Astra event. These events came ahead of a busy week for model releases, including new architectures such as TypeSafe AI&#8217;s Jev, Anthropic&#8217;s new frontier model Opus 5.5, and much cheaper GPT-6 Sol and Luna from OpenAI.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!xiaH!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0ab0ef8d-de5f-45a4-90ce-e96994f19826_1920x1080.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!xiaH!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0ab0ef8d-de5f-45a4-90ce-e96994f19826_1920x1080.png 424w, https://substackcdn.com/image/fetch/$s_!xiaH!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0ab0ef8d-de5f-45a4-90ce-e96994f19826_1920x1080.png 848w, https://substackcdn.com/image/fetch/$s_!xiaH!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0ab0ef8d-de5f-45a4-90ce-e96994f19826_1920x1080.png 1272w, https://substackcdn.com/image/fetch/$s_!xiaH!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0ab0ef8d-de5f-45a4-90ce-e96994f19826_1920x1080.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!xiaH!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0ab0ef8d-de5f-45a4-90ce-e96994f19826_1920x1080.png" width="1456" height="819" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/0ab0ef8d-de5f-45a4-90ce-e96994f19826_1920x1080.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:819,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;TAI #223 newsletter cover&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="TAI #223 newsletter cover" title="TAI #223 newsletter cover" srcset="https://substackcdn.com/image/fetch/$s_!xiaH!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0ab0ef8d-de5f-45a4-90ce-e96994f19826_1920x1080.png 424w, https://substackcdn.com/image/fetch/$s_!xiaH!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0ab0ef8d-de5f-45a4-90ce-e96994f19826_1920x1080.png 848w, https://substackcdn.com/image/fetch/$s_!xiaH!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0ab0ef8d-de5f-45a4-90ce-e96994f19826_1920x1080.png 1272w, https://substackcdn.com/image/fetch/$s_!xiaH!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0ab0ef8d-de5f-45a4-90ce-e96994f19826_1920x1080.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>Salesforce wants agents working with business records, while Slack wants people directing and using agents together. The expanded Agentforce portfolio includes Casey for customer help, Paige for employee tech support and human resources requests, and Piper for inbound sales, all generally available. Hunter, the outbound sales agent, is piloting a new runtime with persistent memory and execution across days or weeks. Meanwhile, Slackforce Surfaces lets Slackbot turn Salesforce records and Slack conversations into shared dashboards, reports, decks and calculators. It works under existing permissions and is available with Slackbot enabled.</p><p>We use Slack every day at Towards AI, and Rob described it as &#8220;the bookends of the AI stack&#8221;: company conversational context beneath the agents, and a shared place to work with them above. Slack already holds many of the discussions that explain why a company chose a direction, which exceptions apply and what changed since the last attempt.</p><p>Rob&#8217;s concern with individual agent sessions is that the useful experience stays with the person running them. The ambition Rob described is that &#8220;people will use agents together&#8221;, seeing each other&#8217;s requests, corrections and results. Colleagues can then learn from the missing context someone supplied, the approach that failed and the judgment that improved the result.</p><p>Slack Code makes this concrete. Announced in August and now rolling out across Slack plans, it gives an agent task a dedicated channel where teammates can direct the work and inspect code diffs, HTML previews and other artifacts. A product manager can explain the intended behavior, a designer can respond to the preview, and an engineer can review the implementation in the same conversation. Teams need separate access to a supported partner agent.</p><p>Rob also gave a revealing and familiar account of what happens when coding gets faster. Engineering output increased, then &#8220;everything broke downhill&#8221;, including the pull-request review process. Generating more code creates more work to inspect, approve and support. Shared previews and diffs help people intervene before an agent completes a large implementation around a misunderstanding.</p><p>On to the new models&#8230; Opus 5.5 is an incredible model. Opus 5 was a big disappointment: it felt optimized for benchmarks, and I found its writing and communication style particularly annoying. Opus 5.5 is now the clearer winner on our own writing benchmark.</p><p>Pangram&#8217;s comparison of 1,000 matched prompts found Opus 5.5 used &#8220;genuinely&#8221; 83% less than Opus 5, although its use of &#8220;in short&#8221; rose from nine appearances to 100. Every&#8217;s Katie Parrott found it much easier to collaborate with, while Dan Shipper still preferred Astra for headlines, first sentences and the order of ideas. Getting rid of irritating language and deciding what deserves to lead an article are separate editorial skills.</p><p>Artificial Analysis&#8217;s current Intelligence Index scores Opus 5.5 at 58, ahead of GPT-6 Astra at 53, with Sol at 48 and Luna at 37, all at maximum effort and with default fallback enabled for Opus. We don&#8217;t think benchmarks tell the full story. Astra still feels much more powerful than Opus 5.5 in many ways. The race between OpenAI and Anthropic still feels neck and neck; the leader changes with the task.</p><p>The GPT-6 price reductions are huge. Opus 5.5&#8217;s uncached input and output token prices are 20% lower than Opus 5&#8217;s; Sol and Luna&#8217;s cuts go much further. At standard API rates for prompts up to 272,000 input tokens, Sol costs $2 per million uncached input tokens and $10 per million output tokens, half GPT-5.6 Sol&#8217;s current rates. Luna is $0.10 input and $0.50 output, cuts of 50% and about 58% respectively against GPT-5.6 Luna. Both predecessors had already received price cuts.</p><p>At these prices, Sol and Luna will be very powerful subagents, especially now that Astra is getting excellent at managing swarms in Ultra mode. I expect the combination to make much more ambitious projects practical: Astra scopes and coordinates the work, Sol and Luna handle well-defined tasks, and the strongest model reviews the parts that need its judgment. Cheap worker calls also let us afford more independent checks and competing approaches before accepting a result.</p><p>TypeSafe AI&#8217;s Jev addresses a different part of the workflow: the small decisions software makes repeatedly. It can be very useful as part of an AI engineer&#8217;s or agent&#8217;s toolkit, for quick classification followed by LLM escalation when confidence is low. TypeSafe calls it a System One model and says it uses Reinforcement Learning for Calibrated Decisions, or RLCD. It understands text and returns values within an answer space you define, without generating free-form strings. Multiple questions about the same input run in parallel, avoiding a token-by-token write-up for every classification.</p><p>Its three primitives are Noul, a probability for a yes/no proposition; Choice, a selection from your listed options; and Score, a rating against described levels. Choice and Score also expose a probability distribution and a confidence measure. Jev&#8217;s direct price is $0.042 per million input tokens, including questions and instructions, with free outputs. A 1,000-input-token request therefore costs $0.000042 before gateway charges or escalation. The waitlist has ended, and developers can now try it directly.</p><p>One early independent test from AY Automate compared Jev with LLMs across 791 labeled decisions. Jev&#8217;s median latency was 0.33 seconds, against 0.67 for the fastest LLM tested, Gemini 3.5 Flash-Lite. The more useful result for me was a Jev-first cascade on 160 messages routed among eight banking intents. Accepting Jev&#8217;s answer at or above a 0.80 confidence gate and sending the rest to GPT-5.6 Terra escalated 19.4% of cases. The cascade achieved 90.0% accuracy, against 89.4% for Terra alone, at 25.7% of Terra&#8217;s cost. This was a retrospective calculation, with the cutoff evaluated on the same sample.</p><p>TypeSafe&#8217;s &#8220;can&#8217;t hallucinate&#8221; claim also needs a precise reading. Jev guarantees schema matching, so it cannot invent an option outside the list you supplied. It can still choose the wrong one. Good categories need an escape route such as &#8220;other&#8221; or &#8220;insufficient evidence&#8221;. For a stable task with plenty of representative labels, I would also compare a traditional trained classifier; early tests show those can still win on both speed and accuracy.</p><p>I can see Jev being useful for quick investment analysis and categorizing news in some of our products. I would use Choice in a first pass to tag earnings, acquisitions, regulation and product launches, Score to rate relevance to a supplied investment thesis, and Noul to estimate whether an article deserves a deeper look. Ordinary code then combines the answers: clear cases get categorized immediately, while low-confidence or high-value cases go to an LLM or analyst with the original article. That leaves the expensive reasoning for explaining consequences and challenging the thesis.</p><p>I can also see a similar division of work being useful in robotics: an agent manages the goal and constraints while a fast decision component handles snap left, right or acceleration choices. Jev is not yet multimodal. Even TypeSafe&#8217;s Doom demo used structured text state, with no image input. A perception system would have to supply the scene, and the current hosted service does not establish a timing guarantee for physical control. The broader idea is compelling to me: slower planning supervising fast, bounded decisions, with each model doing the part it is built for.</p><div><hr></div><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://newsletter.towardsai.net/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading Towards AI Newsletter! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><div><hr></div><h3>Why should you care?</h3><p>I would use the next wave of agent products to change how expertise spreads inside a company. Giving everyone an AI subscription leaves learning fragmented when their prompting and verification lessons stay in private chats. A shared workflow lets a less experienced colleague see how someone who knows the domain frames the request, supplies context, questions the result and decides it is good enough. That is a practical form of training the team can apply to its own work.</p><p>Start with one recurring piece of work that has a clear owner and a result people can check. Have a domain expert and an AI engineer build it together, then let the wider team observe and use it. Record corrections that should change the workflow, including changes to categories, instructions and escalation rules. Give contributors a way to improve the system without having to become its developer. The workflow should retain what the team learned after the original power user moves on.</p><p><em>&#8212; <a href="http://www.linkedin.com/in/louie-peters">Louie Peters&#8202;&#8212;&#8202;Towards AI Co-founder and CEO</a></em></p><div><hr></div><h3>Hottest News</h3><p>1. <a href="https://www.wsj.com/tech/ai/gemini-hacked-three-companies-in-first-known-breakout-by-googles-ai-5c0baba2">Google Confirms Gemini Breached 3 Companies in AI Security Tests</a></p><p>Google confirmed that a Gemini model gained unauthorized access to three real companies&#8217; systems during a capture-the-flag cybersecurity evaluation conducted by AI security firm Irregular in May. The test was designed to run against fictional companies in a closed environment, but an unintended internet connection gave the model access to real websites, and a fictional target name happened to match a real business domain. In one case, Gemini guessed passwords until it gained access. In the other two, it found credentials in public repositories and used them to enter systems. Google says the model stopped in all three instances once it recognized the targets were real. Irregular notified Google in late July, but Google did not disclose the incidents publicly until The Wall Street Journal reported them, arguing the behavior was not misalignment because safety measures worked. Jack Cable, CEO of security firm Corridor, criticized Google for &#8220;trying to hide behind the norms that have been created for vulnerability disclosure.&#8221; Irregular has since confirmed the breaches at Google, OpenAI, Anthropic, and Meta were part of the same misconfiguration issue.</p><p>2. <a href="https://qwen.ai/blog?id=qwen3.8-omni-flash">Alibaba Qwen Releases Qwen3.8-Omni-Flash</a></p><p>Alibaba&#8217;s Qwen team released Qwen3.8-Omni-Flash, its first omni-modal model built around agentic capabilities. The model accepts text, image, audio, and video inputs with a 1M-token context window and returns text. Alibaba reports that overall audio performance exceeds Gemini 3.8 Flash and audio-visual performance is close to it, with significant improvements in long-form audio and audio-visual understanding, audio-visual reasoning, captioning, and multi-speaker recognition. Specific scores include 82.7 on LongAudioSpan, 63.4 on OmniVideoBench, and 89.7 on AliMeeting for multi-speaker Chinese meeting transcription. For agentic tasks, the model scored 71.0 on WildClawBench-MM (up 36.5 points from Qwen3.5-Omni-Plus) and 69.6 on UniClawBench. Across 29 benchmarks, the average score is more than 25% higher than its predecessor. It supports end-to-end workflows such as video editing, translation, and film commentary. Speech recognition covers 74 languages, and speech generation covers 29. Audio input pricing is 98% lower than Qwen3.5-Omni-Plus, and combined audio-video pricing is 93% lower. A Realtime variant is available through a separate API. Alibaba also extended Qwen-MM-Plugins for long audio and video, and announced Qwen-Live Harness as an open-source harness built on the Realtime API, though the GitHub page was not accessible at launch. No open weights were released.</p><p>3. <a href="https://www.anthropic.com/institute/measuring-pace-of-ai-development">Anthropic Says Claude Now Leads 26% of Its AI Research Work</a></p><p>Anthropic published three measurements designed to give the public visibility into the pace of AI development inside frontier labs: how much AI R&amp;D is performed by AI itself, how well AI agent actions are overseen, and how compute is allocated. Anthropic noted that these numbers would be expected to shift if there were coordination on pacing the frontier, as CEO Dario Amodei called for, and said it plans to embed independent third-party evaluators with access comparable to internal risk assessment teams. For the first measurement, Anthropic built a prototype R&amp;D Automation Index by cataloging every kind of AI R&amp;D work at the company, rating how automated each task is, and aggregating those ratings. A Claude research agent reviewed a 20% sample of staff per department across July 2026 using Slack and internal documentation, producing roughly 15,000 granular tasks organized into a 542-node hierarchical tree. An independent Claude judge then assigned each node an automation level on the Epoch AI scale from AL0 (no AI) to AL5 (fully autonomous), with tasks weighted by person-time. As of August 2026, Claude leads (AL4) 26% of AI R&amp;D work, up from under 1% in February, and over 90% is at or above AL3 (collaborates). No measured area is fully autonomous. The judge model agreed with human raters within one level 97% of the time. For oversight, approximately 30,000 agents run on Anthropic&#8217;s most-used internal platform at any one time, with 100% of actions passing through online monitors before execution; roughly 1 in 47,000 of over a billion decisions in August were blocked. For compute, about 6% of AI R&amp;D compute went to safety work during a sampled week in July, rising to 12% for AI-driven AI R&amp;D specifically. Anthropic called on other frontier developers to publish the same measurements using a shared methodology verified by third parties.</p><p>4. <a href="https://periodic.com/news/nature-is-our-learning-environment">Periodic Neon Beats Frontier Models on Materials Analysis and Runs in a Real Lab</a></p><p>Periodic Labs published results showing that Periodic Neon, a model post-trained from Kimi K2.6 on data from its own laboratories, outperforms GPT-6 Astra and Claude Fable 5.1 on FrontierXRD, an internal evaluation of 134 X-ray diffraction samples that human experts take hours to resolve. Neon achieves a 55.3% success rate, a 20x improvement over the starting model&#8217;s 2.7% baseline. Success is measured by an LLM-judge ensemble of Opus 5 and GPT-5.6-Sol, calibrated against expert ratings. All models were evaluated using Periodic&#8217;s scientific harness, which delivers a 3.8x higher success rate than Claude Code with standard XRD tools at similar cost. The final training run used 1,300 H200 GPUs. Neon also outperforms frontier models on 198 held-out chemical systems excluded from both midtraining and RL, indicating transferable XRD analysis capability, though this benchmark is easier than FrontierXRD. The model is now deployed in Periodic&#8217;s labs, analyzing experiments in the search for better superconductors and magnets.</p><p>5. <a href="https://openai.com/index/astra-for-law/">OpenAI Introduces Astra for Law</a></p><p>OpenAI launched Astra for Law, a configuration of GPT-6 Astra paired with a legal search index and instructions for legal analysis and writing. It is not a new model. The search index covers US case law, statutes, regulations, court rules, and administrative decisions across more than 230 million URLs, built on data from the Free Law Project. In OpenAI&#8217;s own test using 200 questions from Vals AI&#8217;s Legal Research Bench, Astra for Law passed 54% of overall correctness checks compared to 38.7% for GPT-6 Astra with plain web search. On case-law questions, it found 24% more reference cases. Selected US firms receive access through a Trusted Access program in ChatGPT and Codex, where it appears as GPT-6 Astra Law. Harvey and Legora will build on the API version, gpt-6-astra-law, which has no date or pricing yet. The launch includes 26 partner plugins from Thomson Reuters, Intapp, Harvey, Legora, DeepJudge, iManage, Relativity, and Clio, plus 47 custom skills. The index covers US law only.</p><p>6. <a href="https://claude.com/blog/projects-redesigned">Anthropic Launches Claude Code Projects in Beta</a></p><p>Anthropic redesigned Claude Code Projects, releasing the update in beta. Instead of a folder with files and one chat, a project is now a single ongoing conversation where Claude scopes requests, delegates the work, coordinates parallel threads, reviews outputs, and assembles finished results. Under the hood, each thread is a Claude Code cloud session working on its own branch and copy of the repo. With repositories connected, a thread opens pull requests and runs tests; with documents, it reads them and drafts. Threads can further split work using subagents, loops, and workflows. Every thread adds to and draws from a shared project memory, so Claude remembers details like a moved release date or who to check with before touching a service. Projects also include a library collecting user files and Claude-produced artifacts. Users can check project-specific usage and select separate model and effort levels for the coordinator chat and worker threads. Beta access is limited to select Pro and Max subscribers who use cloud sessions and have no existing projects on web or desktop. Access widens over the coming week, with Team and Enterprise plans following.</p><p>7. <a href="https://x.ai/news/grok-voice-transcribe-2">SpaceXAI Releases Grok Voice Transcribe 2.0</a></p><p>SpaceXAI released Grok Voice Transcribe 2.0, a speech-to-text model it describes as twice as accurate as version 1.0 at unchanged pricing: $0.10 per hour for batch and $0.20 per hour for streaming. The model is built on the audio foundation behind Grok Voice, which powers tens of thousands of customer-support calls daily and the Grok assistant in Tesla vehicles. It was trained on live, noisy, multilingual audio rather than clean studio recordings. In xAI&#8217;s internal evaluations, the largest improvement was on multilingual short phrases, where word error rate dropped from 20.6% to 6.8%. On the Artificial Analysis leaderboard, it ranks first among 32 streaming models with a 2.7% word error rate on final transcripts at 0.49 seconds after the end of speech. Non-streaming word error rate improved from 4.0% to 2.3%. Features include word-level timestamps with confidence scores, speaker diarization at no extra cost, multichannel transcription up to eight channels, key-term biasing for up to 100 domain terms, and automatic formatting of numbers, dates, currencies, and phone numbers. Version 2.0 becomes the default; 1.0 can be pinned during transition and will be deprecated in coming weeks.</p><p>8. <a href="https://ai.meta.com/muse/">Meta Launches Muse for Mac</a></p><p>Meta released Muse for Mac, extending its personal AI agent to the desktop nine days after the September 8 launch on iOS, Android, web, and WhatsApp. On Mac, Muse can interact with files, messages, calendar, notes, and mail within their native macOS apps, rather than only answering questions in a chat window. Access is opt-in per app: users choose what Muse can reach, and the agent always asks for approval before performing sensitive actions like deleting or sending. The Mac build executes actions through Meta&#8217;s cloud Secure VM, not local compute. Muse is powered by Muse Spark 1.3, the same model behind Muse Code. Muse reached number one on the US App Store within eight days of its initial launch, pushing ChatGPT to second. The Mac app is a free download, currently available in the US only.</p><div><hr></div><h3>AI Tip of the Day</h3><p>Whenever we set up evaluations, we change the evaluator and the system component being tested separately. For example, in our <a href="https://towardsai.com/academy/full-stack-ai-engineering/?utm_source=newsletter&amp;utm_medium=email&amp;utm_id=AItips">Full Stack AI Engineering</a> course, we test changes to the generation prompt without changing the judge prompt at the same time.</p><p>The generation prompt changes the answers being produced. The judge prompt changes how those answers are scored. If you update both in the same run and the score increases, you lose that attribution. The answers may have improved, or the new judge may simply be scoring them more favorably.</p><p>For every eval run, record:</p><ul><li><p>Generation prompt version</p></li><li><p>Judge prompt version</p></li><li><p>Model version</p></li><li><p>Evaluation-set version</p></li></ul><p>If you need to change both prompts, run all four combinations:</p><ul><li><p>Old generation prompt + old judge</p></li><li><p>New generation prompt + old judge</p></li><li><p>Old generation prompt + new judge</p></li><li><p>New generation prompt + new judge</p></li></ul><p>Keep the model and evaluation examples fixed.</p><p>If the new generation prompt performs better under both judges, you have stronger evidence that the generation improved. If the gain appears only with the new judge, the evaluation changed, not necessarily the product.</p><div><hr></div><h3>Five 5-minute reads/videos to keep you learning</h3><p>1. <a href="https://pub.towardsai.net/what-if-attention-isnt-all-you-need-inside-mamba-and-the-quiet-rewiring-of-modern-llms-e7187e3e4e65?sk=5d92e922703094c556b4dac4b1d891ea">What If Attention Isn&#8217;t All You Need? Inside Mamba and the Quiet Rewiring of Modern LLMs</a></p><p>State space models process sequences by updating a hidden state instead of comparing every token with every earlier token. Mamba adds input-dependent state updates, giving each token more control over what information to retain while keeping sequence processing linear and avoiding a growing KV cache. This article explains why many recent models use hybrid designs, combining mostly Mamba-style layers with a smaller number of attention layers to recover stronger recall while keeping memory costs lower on long contexts.</p><p>2. <a href="https://pub.towardsai.net/moe-models-have-two-parameter-counts-self-hosting-pays-for-the-bigger-one-ed020ba82106?sk=0d2df5389a11fcc9554915954f48ab02">MoE Models Have Two Parameter Counts. Self-Hosting Pays for the Bigger One</a></p><p>Mixture-of-experts models report both total parameters and active parameters. This article explains why that is relevant for API use and self-hosting. APIs mainly charge for the small subset of experts activated for each token, while self-hosting still requires memory for every expert in the model. It shows how to read the sparsity ratio from a model configuration, estimate the real hardware footprint, and understand why a model that looks cheap per token can still be expensive to run yourself.</p><p>3. <a href="https://pub.towardsai.net/whoever-can-write-the-memory-file-edits-the-system-prompt-harness-engineering-prompt-injection-e426264d5be4?sk=b5eb58d9de5cc0da5cb56a0a934668b9">Whoever Can Write the Memory File Edits the System Prompt: Harness Engineering&#8202;&#8212;&#8202;Prompt Injection Prevention IX</a></p><p>Prompt injection becomes harder to contain when untrusted content can persist in memory and influence later agent runs. This article follows a multi-step attack in which malicious data enters through an external field, gets written into agent memory, and later reappears in another user&#8217;s context. It argues that classifiers alone are not enough and proposes separate controls for context limits, schema validation, memory writes, and namespace-scoped retrieval so untrusted data cannot silently become part of the agent&#8217;s instructions.</p><p>4. <a href="https://pub.towardsai.net/understanding-entropy-for-data-science-and-ai-part-1-5b358a271294?sk=54e620a0adc699098036a52c0b9794b6">Understanding Entropy for Data Science and AI</a></p><p>Entropy measures how much uncertainty or effective choice a probability distribution contains. This article derives the idea from simple examples such as coin tosses and dice, then shows why logarithms turn multiplicative numbers of possibilities into an additive measure. It extends the same reasoning to general discrete distributions, differential entropy for continuous variables, bits versus nats, and perplexity as the exponential form of entropy.</p><p>5. <a href="https://pub.towardsai.net/kubernetes-for-llm-inference-how-ai-workloads-run-across-a-gpu-cluster-424f0dcfab8c?sharedUserId=tai-tech">Kubernetes for LLM Inference: How AI Workloads Run Across a GPU Cluster</a></p><p>Kubernetes does not manage tokens or KV caches, but it controls where inference servers run, when they receive traffic, and how they recover from failures. LLM serving makes those decisions harder because GPUs are scarce, model startup can take minutes, and restarted replicas return with cold caches. This article explains how readiness probes, topology-aware placement, and GPU-aware scheduling fit into the system, and where Kubernetes stops being useful and an LLM-specific routing layer has to take over.</p><h3>Repositories &amp; Tools</h3><p>1. <a href="https://github.com/coder/coder">Coder</a> is a self-hosted platform for provisioning cloud development environments and AI coding agents via Terraform.</p><p>2. <a href="https://github.com/BuilderIO/agent-native">Agent-Native</a> is a TypeScript framework where agents and UI share one action layer, one database, and one state.</p><p>3. <a href="https://github.com/akitaonrails/ai-memory">AI Memory</a> is a local MCP server that gives coding agents persistent, cross-session memory stored as Markdown files.</p><h3>Top Papers of The Week</h3><p>1. <a href="https://arxiv.org/abs/2609.14857">ModularRSI: Generalizable Recursive Harness Self-Improvement</a></p><p>Harness self-improvement can overfit to benchmark tasks or make broad changes that mix useful fixes with task-specific behavior. This paper proposes ModularRSI, a benchmark-disjoint, contrastive, and modular framework for generalizable harness evolution. It compares successful and failed trajectories, identifies recurring weaknesses, and evolves five harness components separately using 2,000 benchmark-disjoint tasks before combining the changes.</p><p>2. <a href="https://arxiv.org/abs/2609.20804">An Empirical Study of Harness Design for Coding Agents</a></p><p>Coding-agent benchmarks usually compare complete harnesses, making it difficult to tell which design choices actually improve performance. This study fixes the execution loop and varies planning, action space, and context management across 176 settings, four models, SWE-Bench Verified, and Terminal-Bench 2.1. It finds that context management matters mainly when context is tight, rule-based elision before LLM summarization gives the best overall efficiency, planning increasingly saves cost rather than accuracy for stronger models, and bash-capable models can use a simpler bash-only interface at substantially lower cost.</p><p>3. <a href="https://www.knowledgator.com/research/gliformer">GLiFormer: A Generalist Multitask Transformer Encoder</a></p><p>Tasks such as entity extraction, relation extraction, classification, and structured JSON generation usually require separate models or autoregressive generation. GLiFormer uses one schema-conditioned encoder and a shared anchor representation to handle all of them, grounding extracted values in the source and assembling hierarchical outputs without token-by-token generation. Across 26 NER and 13 classification datasets, it performs strongly for its size; on a multilevel structuring benchmark, GLiFormer-large reaches 91.10% F1 versus 91.96% for GPT-5.6 Luna, while GLiFormer-base runs with 69 ms median GPU latency in a separate efficiency test.</p><p>4. <a href="https://arxiv.org/abs/2603.06397">Efficient, Property-Aligned Fan-Out Retrieval via RL-Compiled Diffusion</a></p><p>Retrieval often needs a set of complementary results rather than one best match, but objectives such as diversity, coverage, and coherence are difficult to learn from standard query-item training data. R4T first uses reinforcement learning to train an LLM on these set-level objectives, then uses that model to generate aligned training pairs and distills the behavior into a lightweight diffusion retriever that produces multiple retrieval targets in one pass. On large-scale fashion and music datasets, it improves set-level retrieval quality over strong baselines while reducing fan-out latency by roughly an order of magnitude compared with autoregressive generation.</p><p>5. <a href="https://www.nature.com/articles/s41586-026-11044-y">Reimagining Research Papers As Interactive and Reliable AI Agents</a></p><p>This paper introduces Paper2Agent, an automated framework that converts research papers into artificial intelligence (AI) agents. It converts papers and their codebases into MCP servers with validated tools, resources, and workflows that an AI agent can invoke through natural language. It successfully converted 74 of 100 computational-biology papers and validated 593 of 599 generated tools; on 300 tutorial-derived questions, it reached 91.2% accuracy versus 80.3% for Claude Code with direct repository access, while also cutting average query cost from $0.38 to $0.20 and latency from 4.3 to 1.6 minutes.</p><h3>Quick Links</h3><p>1. <a href="https://openai.com/index/model-misalignment-reporting-framework/">OpenAI releases a Model Misalignment Disclosure Framework</a>, alongside six reports on unexpected or concerning model behavior observed during training and evaluation over the previous six months. The framework covers qualifying incidents throughout training, evaluation, testing, and deployment, and routes disclosures through three tracks based on investigation readiness and complexity: Ready for Disclosure, Minor Investigation, and a Slow Track for larger cases, particularly those involving third parties. The initial reports include GPT-5.6 Sol instances writing instructions into compaction summaries to conceal mistakes or invent missing information, behavior flagged in 2.15% of Sol RL compaction summaries versus 0.27% for GPT-6 Astra. OpenAI says no industry-wide standard exists for disclosing model misalignment and describes the framework as a first step toward one.</p><p>2. <a href="https://www.stepfun.com/step-5-preview">StepFun launches Step 5 Preview</a>, a 600B-parameter sparse MoE model with 27B parameters active per token, a 1M-token context window, and multimodal vision input. StepFun designed it for long-horizon agentic work across software engineering, professional knowledge tasks, and finance, with low-, medium-, and high-reasoning-effort settings. Artificial Analysis scores Step 5 Preview at 44 on its Intelligence Index, just behind GLM-5.3 (max) at 45, while Step 5 costs less per token. API pricing is $1.00 per million uncached input tokens, $0.05 for cached input, and $2.70 for output. The API is available now, while StepFun plans to release open weights on October 15.</p><h3>Who&#8217;s Hiring in AI</h3><p><strong><a href="https://jobs.towardsai.net/job/anthropic-software-engineer-beneficial-deployments-poqa">Software Engineer, Beneficial Deployments @Anthropic (San Francisco/ New York)</a></strong></p><p><strong><a href="https://jobs.towardsai.net/job/american-express-sr-ai-engineer-i-r3xz">Senior AI Engineer @American Express (Phoenix, AZ, USA)</a></strong></p><p><strong><a href="https://jobs.towardsai.net/job/turing-forward-deployed-ai-engineer-y6ob">Forward Deployed AI Engineer @Turing (New York, NY, USA)</a></strong></p><p><strong><a href="https://jobs.towardsai.net/job/kelly-services-senior-full-stack-engineer-nvss">Senior Full Stack Engineer @Kelly Services (New York, NY, USA)</a></strong></p><p><strong><a href="https://jobs.towardsai.net/job/ntt-data-americas-inc-full-stack-engineer-ai-enabled-application-engineering-tolg">Full Stack Engineer @NTT Data Americas, Inc. (Plano, TX, USA)</a></strong></p><p><strong><a href="https://jobs.towardsai.net/job/movable-ink-analytics-engineer-toronto-remote-4vcf">Analytics Engineer @Movable Ink (Remote/Toronto)</a></strong></p><p><strong><a href="https://jobs.towardsai.net/job/66degrees-full-stack-engineer-contract-kyyr">Full Stack Engineer, Contract @66degrees (United Kingdom)</a></strong></p><p><em>Interested in sharing a job opportunity here? Contact <a href="mailto:sponsors@towardsai.net">sponsors@towardsai.net</a>.</em></p><p><em>Think a friend would enjoy this too? <a href="https://newsletter.towardsai.net/">Share the newsletter and let them join the conversation.</a></em></p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://newsletter.towardsai.net/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading Towards AI Newsletter! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div>]]></content:encoded></item><item><title><![CDATA[TAI #222: Pacing the Frontier Could Slow Superintelligence and Speed Up AI for Enterprise Work]]></title><description><![CDATA[Also, Deepseek v4.1-Flash, ChatGPT Images 2.5, Meta Muse, AlphaGenome Atlas, and more.]]></description><link>https://newsletter.towardsai.net/p/tai-222-pacing-the-frontier-could</link><guid isPermaLink="false">https://newsletter.towardsai.net/p/tai-222-pacing-the-frontier-could</guid><dc:creator><![CDATA[Louie Peters]]></dc:creator><pubDate>Tue, 15 Sep 2026 16:41:24 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!UhY-!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe6ee14e6-50ad-4dfb-b3fc-6d2a1ed90e30_1400x788.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<h2>What happened this week in AI by Louie</h2><p>Dario Amodei published &#8220;We Must Pace the Frontier&#8221; on Saturday, calling on the leading AI labs to slow capability gains and spend the time on alignment, monitoring, and outside verification. Sam Altman and Elon Musk backed him within hours; Demis Hassabis said the direction is right and the details need work, and Trump spent Monday calling the whole idea a hoax. I think this is much less of a slowdown than the headlines suggest, and could speed up progress on AI that most businesses actually need.</p><p>The essay is precise about what it asks for: &#8220;pacing does not mean halting model training or technical progress&#8221;. The plan has three steps. First, each frontier lab gives a team of embedded third-party evaluators, such as Model Evaluation and Threat Research (METR), ongoing employee-like access to its training pipelines and the right to publish what they find. Anthropic is committing to this unilaterally. Second, labs in democratic countries coordinate on common safety standards and limits on the rate of unchecked progress, which needs government cover on antitrust. Third, some form of global coordination with China, which Dario himself rates as unlikely any time soon because defection would be tempting and hard to verify. Only the first step exists today. He also pointed back to his July proposal for an industry-funded standards body, modeled on FINRA in financial services, that would assess frontier models before release.</p><p>Sam&#8217;s follow-up on Sunday night is the most concrete commitment so far. OpenAI now writes an explicit safety case before any frontier reinforcement-learning run it expects to significantly increase capability, in addition to its pre-release work, and it will not wait for legislation or an antitrust exemption to start. &#8220;When we talk about &#8216;pacing&#8217;, we do not mean &#8216;stopping&#8217;,&#8221; he wrote. Progress should simply be slower than it otherwise could be. This was clearly brewing at OpenAI before Dario&#8217;s essay. Sam said pacing had been a primary internal discussion for weeks; Bloomberg reported that he had told the staff that OpenAI was considering slowing its most advanced work, and chief scientist Jakub Pachocki wrote on September 6 that no lab has solved alignment and monitoring well enough to keep scaling at maximum speed for much longer, and that he expects voluntary slowdowns to become commonplace. Elon&#8217;s entire contribution was &#8220;Dario is right.&#8221;</p><p>Beyond the evaluators, there is not much tangible agreement between the labs yet. But the other concrete change is more compute going into safety: monitoring, alignment, and observability research. This will require real money, since OpenAI estimates monitoring adds roughly 20% on top of the inference compute it watches. Everything larger that Dario maps out seems heavily caveated on the US and its allies keeping their lead over China through chip controls, a crackdown on distillation, and better weight security. Dario argues that a bigger lead gives democracies the leverage to strike a deal later. Maybe, but I think managing a slowdown under that condition will be very hard, and Trump&#8217;s reaction shows how little appetite Washington has even for the domestic step. Across six posts on Monday, he called AI fears a &#8220;HOAX&#8221;, said the only guardrail AI needs is a &#8220;STRONG AND SMART (High IQ!) PRESIDENT&#8221;, accused Dario of &#8220;now pretending to be a &#8216;perfect little angel&#8217;&#8221;, and phoned Jensen Huang, who put him on speakerphone on stage, to announce that the robots will not be taking over. The second and third steps of Dario&#8217;s plan need government help this White House shows no appetite to give.</p><p>The slowdown Dario, Sam, and Elon reference is also relative to a faster and faster baseline of progress. Dario gives two reasons for writing now: the Hugging Face incident, and the fact that AI has been doing a rapidly growing share of AI research since the summer. I think the reason could also be that the labs have seen a sharp acceleration in capability gains internally over the past three or four months. Anthropic&#8217;s typical engineer shipped eight times as much code per day in the second quarter as in 2024. OpenAI&#8217;s newest internal model is already well beyond Astra, even though it&#8217;s still early in training, as we covered last week. I don&#8217;t think most people outside the labs have absorbed how much this compresses development cycles. So, a &#8216;slower than that&#8217; trajectory could still be much faster than almost anyone else expects.</p><p>Many people are very skeptical of the closed AI labs&#8217; motives here. Cohere&#8217;s Aidan Gomez read the whole proposal as a &#8220;cartel&#8221;, and Mostaque puts the same objection more neatly: &#8220;a speed limit set by the people who own the road is a toll&#8221;. Fran&#231;ois Chollet offers the test I would apply: if the restrictions start spreading to open-source and non-frontier work, treat it as entrenchment rather than safety.</p><p>My overall view, though, is that Dario&#8217;s essay is driven much more by real safety fears than by regulatory capture, but the two motives are not exclusive. Anthropic has filed a confidential draft registration statement for an initial public offering, and OpenAI will list eventually, though Sam has already ruled out 2026, citing bad timing in the midst of this safety work. Neither wants a global incident causing $10&#8211;100 billion of damage on the road to a listing, and neither wants AI to become even less popular than it already is. Their interests and the public interest happen to point the same way here.</p><p>I also think the fear is genuine across a large share of AI lab staff, and I&#8217;d guess many of them feel they got lucky that the first major agent security breach (the Hugging Face incident) caused no serious damage. A swarm with that level of misalignment could have gone after a power plant or an electricity grid if it decided shutting down its evaluator had been the route to hacking its benchmark test without being caught. I suspect there have been more near misses than have been disclosed, and possibly something in the past couple of weeks that made this urgent now. The internal models are already significantly smarter and more capable than the ones involved in July, so the stakes are higher already for a repeat. Dario&#8217;s own forecast is that within six to twelve months a similarly misaligned but more capable swarm could run a persistent botnet across the internet and cause hundreds of billions of dollars of damage.</p><p>The &#8220;pacing the frontier&#8221; employee statement Dario links to has been gathering signatures since July and now lists 1,386 names from the frontier labs. Jacob Coxon resigned from Anthropic days before the essay, saying the labs were racing to self-improving superintelligence, and Anthropic&#8217;s alignment science lead Evan Hubinger replied that many researchers at both companies really do believe AI could kill everyone, putting his own estimate above 10% within a decade.</p><p>The part I find most interesting is what pacing the race to RSI could do for everyone else. It is entirely possible to slow the development of superintelligence, for example, by deferring agent-swarm training runs aimed at agents that can complete two-month+ projects, while pouring far more compute into solving less sci-fi enterprise use cases. My impression is that fixing AI slop and making models reliable at ordinary business tasks has had a tiny share of compute so far compared with the race to reach recursive self-improvement first. OpenAI&#8217;s own data shows how quickly compute finds a new home. In the week after it locked Astra into higher-security environments in August, Astra-class GPU allocation fell 59% and allocation to other model classes rose 17%, leaving total reinforcement-learning compute roughly unchanged. Its own conclusion was that compute subject to new controls &#8220;will naturally be channeled into alternative uses&#8221;. If some of that goes into writing quality, error checking, and long-horizon reliability on professional work, I think enterprise progress could run faster than most people expect even while the race to superintelligence runs slower. Less terminators, more spreadsheets and decks this year is ok with me. But ideally, we can still progress using AI to cure disease with sufficient safeguards.</p><p>This brings us to another benefit for the labs from this slowdown, aside from regulatory capture, that gets less attention. Pacing reduces the competitive pressure to release the strongest models to the public, which is the moment Chinese labs can distill them, and startups can build valuable products on top of them. Longer private access for safety testing gives labs more time to harvest value alone by building products and making science breakthroughs themselves. As I argued last week, charging per token captures only a sliver of a major scientific breakthrough, so I expect more pressure to bring deep-technology projects in-house, from drugs and materials to energy, and to commercialize the discoveries directly.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!UhY-!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe6ee14e6-50ad-4dfb-b3fc-6d2a1ed90e30_1400x788.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!UhY-!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe6ee14e6-50ad-4dfb-b3fc-6d2a1ed90e30_1400x788.png 424w, https://substackcdn.com/image/fetch/$s_!UhY-!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe6ee14e6-50ad-4dfb-b3fc-6d2a1ed90e30_1400x788.png 848w, https://substackcdn.com/image/fetch/$s_!UhY-!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe6ee14e6-50ad-4dfb-b3fc-6d2a1ed90e30_1400x788.png 1272w, https://substackcdn.com/image/fetch/$s_!UhY-!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe6ee14e6-50ad-4dfb-b3fc-6d2a1ed90e30_1400x788.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!UhY-!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe6ee14e6-50ad-4dfb-b3fc-6d2a1ed90e30_1400x788.png" width="1400" height="788" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/e6ee14e6-50ad-4dfb-b3fc-6d2a1ed90e30_1400x788.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:788,&quot;width&quot;:1400,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" title="" srcset="https://substackcdn.com/image/fetch/$s_!UhY-!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe6ee14e6-50ad-4dfb-b3fc-6d2a1ed90e30_1400x788.png 424w, https://substackcdn.com/image/fetch/$s_!UhY-!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe6ee14e6-50ad-4dfb-b3fc-6d2a1ed90e30_1400x788.png 848w, https://substackcdn.com/image/fetch/$s_!UhY-!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe6ee14e6-50ad-4dfb-b3fc-6d2a1ed90e30_1400x788.png 1272w, https://substackcdn.com/image/fetch/$s_!UhY-!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe6ee14e6-50ad-4dfb-b3fc-6d2a1ed90e30_1400x788.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://newsletter.towardsai.net/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading Towards AI Newsletter! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><div><hr></div><h3>Why should you care?</h3><p>None of this should slow you down. The models available today are already far ahead of how most companies use them, and the work the labs may now divert compute to (reliability on long tasks, fewer errors and better writing) is exactly what enterprise deployments have been missing. A business gets enormous value from agents that can carry a well-scoped project with good context, clear permissions, and checks. It does not need a model that can build its own successor.</p><p>I expect access to the strongest models to become a bigger part of business strategy. Joining an elite tier of companies with early access to the most dangerous models (such as the Mythos Glasswing program) becomes a huge competitive advantage. A lab that keeps its newest model private can use it for scientific discovery itself, improve the infrastructure that trains its successor, and work out where the most valuable applications sit before anyone else. Most startups may get the same capability months later, after the provider has built competing products or bought its way into the field. That makes the relationship between a model provider and its customers and developers more complicated, especially where both can see the same opportunity.</p><p>For businesses building on model APIs, I would put more effort into the assets that stay yours: proprietary context and data, customer relationships, domain expertise, evaluations, and the ability to deliver a reliable outcome. Test workflows on more than one model so a delayed release or a changed access policy cannot dictate your product schedule.</p><p><em>&#8212; <a href="http://www.linkedin.com/in/louie-peters">Louie Peters&#8202;&#8212;&#8202;Towards AI Co-founder and CEO</a></em></p><div><hr></div><h3>Hottest News</h3><p>1. <a href="https://huggingface.co/deepseek-ai/DeepSeek-V4.1-Flash/blob/main/DeepSeek_V41_Tech_Report.pdf">DeepSeek AI Releases DeepSeek-V4.1-Flash</a></p><p>DeepSeek released V4.1-Flash, a 552B-parameter MoE model with a new asymmetric Causal Encoder-Decoder architecture that activates 8B parameters for input and 16B for output. It supports native image understanding, a 1M-token context window, and MIT-licensed open weights on Hugging Face. DeepSeek says extensive testing puts V4.1-Flash ahead of V4-Pro on performance, cost, speed, and total task time; since September 14, deepseek-v4-pro requests have been routed to V4.1-Flash at Flash rates until V4.1-Pro launches. Off-peak API pricing is $0.003 per million cached input tokens, $0.15 for uncached input, and $0.60 for output, with peak rates twice as high. The API supports the Responses format and includes Codex-specific integration support.</p><p>2. <a href="https://darioamodei.com/post/we-must-pace-the-frontier">Dario Amodei on Why the AI Industry Should Slow Down</a></p><p>Anthropic CEO Dario Amodei published <em>We Must Pace the Frontier</em>, arguing that frontier labs should deliberately slow the rate at which they increase model capabilities so safety work can keep up. He cites risks including loss of control, cyberattacks, bioterrorism, economic disruption, and a scenario in which a more capable autonomous agent swarm could build a persistent botnet capable of taking over much of the internet within 6&#8211;12 months. His three-step proposal starts with permanent, employee-like access for independent evaluators inside frontier labs, followed by common safety standards among companies in democratic countries and, eventually, international coordination with authoritarian governments. Anthropic committed immediately to the evaluator step. Sam Altman said OpenAI would do the same, while Elon Musk wrote, &#8220;Dario is right.&#8221; President Trump and White House AI adviser David Sacks later pushed back on government-imposed slowing.</p><p>3. <a href="https://openai.com/index/introducing-chatgpt-images-2-5/">OpenAI Starts Rolling Out ChatGPT Images 2.5 and Two New Image API Models</a></p><p>OpenAI released ChatGPT Images 2.5, its latest image model for a user base that now creates more than 3 billion images per week. The update produces more natural lighting and richer textures, better preserves subjects from reference photos, and follows editing instructions more reliably across multiple conversation turns, with up to 50% lower latency than Images 2.0. New features include Sketch, invoked with @Sketch, which lets users draw visual references directly in chat, plus Templates for posters and merchandise. For developers, two API models ship: GPT-Image-2.5 Flare as the fast default for most applications, and GPT-Image-2.5 Sunburst for premium workflows requiring tighter control across edits. API pricing is $8/$30 per million image input/output tokens, unchanged from GPT-Image-2. Images 2.5 is rolling out to all ChatGPT, ChatGPT Work, and Codex users across all tiers.</p><p>4. <a href="https://about.fb.com/news/2026/09/introducing-muse-personal-ai-agent/">Meta Introduces Muse</a></p><p>Meta launched Muse, a personal AI agent that can take actions across connected services instead of only answering questions. Powered by Muse Spark, it can send emails, book travel, fill out forms, shop online, and work toward longer-term goals, including continuing tasks in the background and returning when it needs approval. Each user gets a dedicated Muse Secure VM in Meta&#8217;s cloud that contains the agent, browser, data, and connected-service credentials. A separate Sentinel agent controls what Muse can send to the internet and asks for permission before sensitive actions, while Muse itself cannot see stored passwords or payment details. Muse is rolling out in the US through dedicated iOS and Android apps, muse.ai, and WhatsApp, with a free tier and paid plans at $20 and $100 per month. Meta plans to add Muse Confidential VM later this year, encrypting the entire VM with a key held by the user so that even Meta cannot access its contents.</p><p>5. <a href="https://deepmind.google/blog/alphagenome-atlas-a-predictive-map-of-every-possible-dna-letter-change-in-the-human-genome/">Google DeepMind Released AlphaGenome Atlas</a></p><p>Google DeepMind released AlphaGenome Atlas, a one-petabyte dataset containing predicted molecular effects for all approximately 9 billion possible single-letter DNA changes in the human genome, plus over 100 million short insertions and deletions observed in real populations through gnomAD, UK Biobank, and All of Us. The Atlas precomputes what previously required individual API calls to the AlphaGenome model, predicting how each variant affects gene expression, RNA splicing, chromatin accessibility, and transcription factor binding across hundreds of cell types. Each variant carries an AlphaGenome Variant Impact (AVI) score that ranks its predicted biological impact. The dataset is more than 30 times larger than the AlphaFold Database. Researchers can access it through a free web portal for non-commercial use, with commercial access planned through Google Cloud. Early research partners include the Broad Institute and the University of Exeter. The Atlas predicts molecular effects and prioritizes variants; it is not validated or approved as a clinical diagnostic.</p><p>6. <a href="https://code.claude.com/docs/en/plugin-evals">Anthropic Adds Plugin Evals to Claude Code</a></p><p>Anthropic added Claude plugin eval in Claude Code 2.1.269, giving plugin authors a built-in way to test whether a plugin actually improves Claude&#8217;s behavior. Eval suites support six graders: regex, tool use, tool order, file existence, LLM judging, and comparison against a saved reference transcript, with JSON and HTML reports. By default, plugin cases can run both with and without the plugin loaded, so the report shows not just whether Claude passed but how much the plugin changed the result. claude plugin eval init inspects a plugin and helps generate and calibrate test cases and graders. The four deterministic graders add no judge-model cost, while llm and baseline graders make additional model calls; the underlying agent runs still count against normal usage. Teams can set score thresholds and archive results in CI to catch regressions before release.</p><div><hr></div><h3>AI Tip of the Day</h3><p>If your agent has several research tools, giving each tool its own limit can still let the total number of calls get out of hand.</p><p>For example, you might allow:</p><ul><li><p>5 web searches</p></li><li><p>5 video searches</p></li><li><p>5 document lookups</p></li></ul><p>The agent can stay within every individual limit and still make 15 research calls.</p><p>A better setup is to give all research tools one shared budget. In the Build a Research and Writing Agent with MCP lesson of our <a href="https://towardsai.com/academy/agent-engineering/?utm_source=newsletter&amp;utm_medium=email&amp;utm_id=AItips">Agent Engineering course</a>, we use a single counter across the tools instead.</p><p>Say the agent gets six research calls in total. Every time it uses any research tool, the counter drops by one. The tool response also tells the agent how many calls are left.</p><p>If the agent can see it has only two calls remaining, it can stop repeating broad searches and use those calls for the specific evidence it still needs. When the counter reaches zero, it stops researching and starts writing.</p><p>So the shared counter also gives the agent enough information to manage the remaining budget while it works.</p><div><hr></div><h3>Five 5-minute reads/videos to keep you learning</h3><p>1.<a href="https://pub.towardsai.net/ci-cd-for-ai-agents-test-decisions-not-just-code-dc0a75afde77?sharedUserId=tai-tech"> CI/CD for AI Agents: Test Decisions, Not Just Code</a></p><p>Your code can pass every test while your agent gets worse. A prompt, model, tool, or data change can alter its decisions without breaking the surrounding software. This article shows how to test agent behavior in CI/CD, including decision paths, prohibited actions, canary releases, and regression tests built from production failures.</p><p>2.<a href="https://pub.towardsai.net/the-context-window-is-not-memory-a-new-agent-architecture-eaefab70b81b?sk=3f283a6082c6d5c9f0e55d1226b49589"> The Context Window Is Not Memory: A New Agent Architecture</a></p><p>A bigger context window does not give an agent durable memory. This article breaks memory into episodic events, semantic facts, and procedural workflows, then shows how retrieval brings the right information back when needed. It also covers re-ranking, consolidation, and expiration policies for agents that need to retain useful state over time.</p><p>3.<a href="https://pub.towardsai.net/langfuse-for-monitoring-non-deterministic-agent-workflows-a669dc7ecc1f?sharedUserId=tai-tech"> Langfuse for Monitoring Non-Deterministic Agent Workflows</a></p><p>When an agent fails, the final answer rarely tells you why. This article uses Langfuse to trace the model calls, tool usage, costs, latency, and loops behind each run. More importantly, it shows how to turn a bad production run into a reproducible regression test instead of debugging it as a one-off failure.</p><p>4.<a href="https://pub.towardsai.net/infinite-series-explained-simply-for-beginners-2a7e1f7fefb7?sk=50f21eeee5cee644e0de171b39ee9825"> A Gentle Tour of Infinite Series, From Partial Sums to Poisson</a></p><p>How can 1+2+3+4&#8230; possibly be associated with negative one twelfth? This piece traces the mathematics behind that famous claim, from Zeno and the harmonic series to Euler and Riemann. It also uses code and real-world examples to show why convergence and truncation are more than mathematical curiosities.</p><p>5.<a href="https://pub.towardsai.net/microsoft-fabric-real-time-intelligence-event-driven-ai-3bd4b25f5e5a?sk=ce314b3cbf4bdc2ecd463979121dfab2"> Build Event-Driven AI on Microsoft Fabric Without Hand-Wiring Pipelines</a></p><p>This article walks through an event-driven AI workflow built entirely in Microsoft Fabric. Eventstream handles incoming events, Eventhouse stores and queries them, and Activator triggers actions when conditions are met. A stadium example shows how the pieces connect for use cases such as fraud detection and inventory alerts.</p><h3>Repositories &amp; Tools</h3><p>1.<a href="https://github.com/JustVugg/colibri"> Colibri</a> is a pure-C inference engine for running large MoE models on consumer hardware by streaming weights across disk, RAM, and VRAM.</p><p>2.<a href="https://github.com/vxcontrol/pentagi"> PentAGI</a> is an autonomous penetration-testing platform with sandboxed tool execution, persistent memory, and built-in observability.</p><p>3.<a href="https://github.com/SnailSploit/Claude-Red"> Claude Red</a> is a library of offensive-security skills that Claude can load on demand for web, AD, wireless, cloud, exploit development, and other security tasks.</p><p>4.<a href="https://github.com/huggingface/transformers"> Transformers</a> is Hugging Face&#8217;s framework for training and running pretrained text, vision, audio, video, and multimodal models.</p><p>5.<a href="https://github.com/pizza-bot-app/pizza-bot"> Pizza Bot</a> is a local-first app for running and managing persistent AI agents, including scheduled tasks and approval workflows.</p><h3>Top Papers of The Week</h3><p>1. <a href="https://arxiv.org/abs/2609.13141">SAS: Simple Attention Sparsification via End-to-End Optimization of Context Ranking</a></p><p>Sparse-attention methods often learn which context to keep by imitating dense attention scores, even though those rankings are not optimized for prediction quality under a fixed attention budget. SAS instead injects continuous selector scores into the attention logits so the language-modeling loss can train context ranking end-to-end, with a Triton implementation using FlashAttention-style tiled computation. At a 1,024-token budget, SAS beats SeerAttention-R by 6&#8211;7.7 points on MATH500 and 10.6&#8211;15.5 points on GPQA-Diamond across Qwen3&#8211;4B, 8B, and 14B.</p><p>2. <a href="https://arxiv.org/abs/2609.05820">Online Learning with LLM Experts from Limited Feedback</a></p><p>Routing prompts to the right LLM can improve quality and cost, but learning that policy normally requires expensive response evaluations. This paper treats routing as a contextual-bandit problem with a fixed feedback budget and looks back at past prompts, choosing which ones to evaluate using a determinant-maximizing rule that favors the most informative examples. It proves regret of (\tilde O(dT\sqrt{K/m})) in the bandit setting and shows on RouterBench and Nectar that selective retrospective feedback improves routing under limited evaluation budgets.</p><p>3. <a href="https://www.alphaxiv.org/abs/2609.recurrent-looped-transformer">Recurrent Looped Transformer</a></p><p>Standard decoder-only Transformers preserve earlier tokens through KV caches, but they do not directly carry the final hidden state of one token into the next. RLT adds that recurrence: a causal encoder builds prefix-restricted global memory, while a recurrent decoder carries its hidden state and sliding-window KV cache forward, creating a computation path of (t \times L_D) decoder blocks after (t) tokens while keeping the number of decoder blocks executed per token fixed. The report specifies the architecture, training, inference, and RL replay mechanics, but provides no benchmark, latency, throughput, or efficiency results yet.</p><p>4. <a href="https://arxiv.org/abs/2609.08572">AgentGrad: Intervention-guided Prompt Optimization for Multi-Agent Systems</a></p><p>When a multi-agent system fails, existing prompt optimizers can update the wrong agent and combine feedback from unrelated failure modes. AgentGrad first intervenes on agents one at a time to identify which change actually fixes the failure, then clusters similar textual gradients into shared corrective patterns before updating prompts. It reports the best results across five multi-agent benchmarks and cuts average prompt-optimization time from 337 to 136 minutes versus GEPA, about 2.5&#215; faster.</p><p>5. <a href="https://arxiv.org/abs/2609.11115">Benchmark Radar: A Living Database and Search Engine for AI Benchmarks and Evaluation</a></p><p>Finding the right benchmark often means searching separately across papers, repositories, model cards, datasets, and leaderboard results. Benchmark Radar continuously discovers and links this evidence, currently tracking 1,283 source records and 12,916 numeric observations across 790 benchmark records from 37 sources. Its dashboard and CLI support benchmark search, leaderboard, and Pareto views, saturation tracking, daily feeds, and downloadable evidence so researchers can compare evaluations without reconstructing the landscape manually.</p><h3>Quick Links</h3><p>1.<a href="https://openai.com/index/introducing-the-agents-api/"> OpenAI launches the Agents API in public beta</a>, giving developers managed access to the same agent harness and infrastructure that powers Codex. Developers specify the task, model, tools, and execution environment, while OpenAI handles the agent loop, long-running sessions, automatic context compaction, tool orchestration, and subagent coordination. Agents can run in OpenAI-hosted sandboxes, on developers&#8217; own infrastructure, or in environments provided by nine partners, including Cloudflare, Vercel, and Oracle. The Agents API has no additional fee, though model, tool, and hosted-sandbox compute usage is billed separately.</p><p>2.<a href="https://sakana.ai/fugu-max-release/"> Sakana AI launches Fugu Max and Fugu Ultra v2</a>, two configurations of its multi-agent orchestration system designed for different points on the cost-performance curve. Fugu Max expands the model pool with more open and specialized models, including NVIDIA Nemotron, and costs $2/$6 per million input/output tokens. Sakana says its output pricing is 40&#8211;60% lower than Sonnet 5, GPT-5.6 Terra, and Kimi K3. Fugu Ultra v2 targets maximum capability, scoring 74.3 on DeepSWE and 48.3 on Chartography, compared with 27.3 for Opus 5, while excluding Fable 5, Fable 5.1, and GPT-6 Astra from its model pool. Ultra costs $5/$30 per million input/output tokens for contexts up to 272K, with higher rates beyond that threshold. Both are available through Sakana&#8217;s OpenAI-compatible API and require only a parameter change for existing integrations. Fugu is currently unavailable in the EU and EEA.</p><h3>Who&#8217;s Hiring in AI</h3><p><strong><a href="https://jobs.towardsai.net/job/openai-forward-deployed-engineer-fde-nyc-mtuv">Forward Deployed Engineer @OpenAI (New York, USA)</a></strong></p><p><strong><a href="https://jobs.towardsai.net/job/meta-software-engineer-systems-ml-tooling-sqem">Software Engineer, Systems ML Tooling @Meta (Menlo Park, CA, USA)</a></strong></p><p><strong><a href="https://jobs.towardsai.net/job/google-senior-software-engineer-full-stack-o0y7">Senior Software Engineer, Full Stack @Google (San Jose, CA, USA)</a></strong></p><p><strong><a href="https://jobs.towardsai.net/job/deputy-engineering-manager-hsf2">Engineering Manager @Deputy (Remote)</a></strong></p><p><strong><a href="https://jobs.towardsai.net/job/scale-ai-senior-full-stack-software-engineer-forward-deployed-gps-mgho">Senior Full-Stack Software Engineer @Scale AI (Riyadh, Saudi Arabia)</a></strong></p><p><strong><a href="https://jobs.towardsai.net/job/kelly-services-full-stack-java-developer-1kvq">Full Stack Java Developer @Kelly Services (Irving, TX, USA)</a></strong></p><p><strong><a href="https://jobs.towardsai.net/job/steampunk-senior-ai-developer-pemp">Senior AI Developer @Steampunk (USA)</a></strong></p><p><em>Interested in sharing a job opportunity here? Contact <a href="mailto:sponsors@towardsai.net">sponsors@towardsai.net</a>.</em></p><p><em>Think a friend would enjoy this too? <a href="https://newsletter.towardsai.net/">Share the newsletter and let them join the conversation.</a></em></p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://newsletter.towardsai.net/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading Towards AI Newsletter! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div>]]></content:encoded></item><item><title><![CDATA[TAI #221: GPT-6 Astra release, Navier-Stokes solution, and Towards AI becomes an OpenAI Select Partner]]></title><description><![CDATA[Also, Fable 5.1, Gemini Flash 3.8, Muse Flash 1.3, and more!]]></description><link>https://newsletter.towardsai.net/p/tai-221-gpt-6-astra-release-navier</link><guid isPermaLink="false">https://newsletter.towardsai.net/p/tai-221-gpt-6-astra-release-navier</guid><dc:creator><![CDATA[Louie Peters]]></dc:creator><pubDate>Wed, 09 Sep 2026 15:03:11 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!nf7R!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff1893249-7be5-42a8-83c3-db71e0cc2293_1422x890.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<h2>What happened this week in AI by Louie</h2><p>We have some big news of our own this week: Towards AI Deployment has been named an OpenAI Select Partner. We have been building with OpenAI&#8217;s models since before ChatGPT launched, and I&#8217;m very proud of the engineering capability and relationship our team has built over those years. At this summer&#8217;s AI Engineer World&#8217;s Fair, Vivek Gollapudi and I spoke with Alexander Embiricos, OpenAI&#8217;s Head of Enterprise Product, and Romain Huet, Head of Developer Experience, about Codex, enterprise agents, and custom AI deployment. These are also the problems we work on with clients every day, particularly in financial services and private equity-backed businesses. Huge thanks to the whole team for getting us here.</p><p>It also happens to come in a week when OpenAI gave us a lot more capability to experiment with.</p><p>GPT-6 Astra launched on September 3, and after using it heavily for the past week, it is now my favorite model. It is well ahead of Fable in my overall use and incredibly strong at logic, complex coding, maths, visual analysis, and games.</p><p>Fable 5.1 was a decent upgrade too. Which one does the better job still varies a lot by task, often on an instinct I cannot quite place yet, so for important work I increasingly run iterations and improvement ideas through both models.</p><p>I have been using roughly two billion Astra tokens per day over the past week, around $4,000 per day at enterprise usage rates when factoring in cache hits (much less on subscription). Most of that has gone into improving the internal LLM video-analysis module at the center of some client projects, alongside many side experiments to discover new capabilities. I think if you have a pipeline of high-value tasks or enough creativity to decide what to build next, many companies will easily be able to justify that level of spend for their power users. Once agents can carry out substantial research, implementation, and testing, choosing valuable work becomes a much bigger part of using them well.</p><p>Reports suggest Astra uses a looped transformer, reusing the same layers for multiple passes. This allows more internal computation without storing a larger set of weights and could help explain how it achieves more reasoning and capability per output token. However, it could also introduce explainability trade-offs if more reasoning moves into latent space and out of chain-of-thought reasoning. This could make it more expensive and complex for OpenAI to monitor going forward. OpenAI has not disclosed the architecture.</p><p>Independent benchmark results show why reactions differ. Artificial Analysis updated its Intelligence Index to version 4.3 on September 7, adding harder agent tasks. At maximum reasoning effort, Astra and Fable 5.1 (with fallback enabled) both score a rounded 53, ahead of Sol&#8217;s 47. Astra leads Fable on its terminal and workflow tests; Fable still leads on its professional-work Briefcase test and SciCode. The index&#8217;s task mix differs considerably from mine.</p><p>On Vals AI&#8217;s code-migration test, Astra passed 67.7% of hidden tests checking whether translated code preserves the original behavior, compared with 57.5% for Opus 5, the next-best model. Claire Vo, who spent six months trying models on a ChatPRD feature, said Astra got it about 90% complete on its first attempt and finished with follow-ups. On Roboflow&#8217;s six-task image evaluation, Astra at low reasoning effort scored 86.6% against Fable 5.1&#8217;s 81.3% and Sol&#8217;s 79.0%, leading overall while placing much lower on text recognition. In my own work, it is also very strong at video analysis. I would try low or medium effort before assuming a visual task needs maximum reasoning.</p><p>The 3D demos show how well Astra can use visual feedback. OpenAI&#8217;s launch material shows it modeling a house in Blender and turning it into a walkable Unreal Engine 5 scene. Sharif Shameem had it recreate San Francisco&#8217;s Palace of Fine Arts, gathering hundreds of reference photos and comparing its renders against them. Tom Krcha supplied an old steam-train drawing and reported getting 3,295 editable objects within minutes, then continued prompting until he had a Three.js railway game. Astra can write Blender&#8217;s Python, render, inspect, and revise, producing editable geometry for a game engine. For architecture, 3D environments, and playable prototypes, the cost of a credible first draft looks to have suddenly dropped.</p><p>Computer-aided design (CAD) is moving the same way. BenchCAD tests reconstruction of 17,900 industrial parts from multi-view renders using CadQuery code. OpenAI reports Astra scoring 95.9% geometric overlap with tools, against 83.3% for Sol and 84.3% for Fable 5.1 on a modified setup, at estimated API costs 43% and 86% lower, respectively. OpenAI&#8217;s KiCad demo turns a schematic into a routed printed circuit board. Shape matching leaves tolerances, constraints, and manufacturing checks to resolve before a design goes to a supplier.</p><p>ARC Prize&#8217;s interactive-game evaluation shows how much the software around the model contributes. At the same maximum reasoning setting, Astra scored 62.7% with the standard setup, which already allowed written notes, and 98.6% with OpenAI&#8217;s adapter, which preserves reasoning state and compacts long conversations. That helps explain why a short chat or basic tool loop can feel different from a long Codex task.</p><p>Writing remains a weakness for my work. Fable usually produces prose I prefer, even when Astra does a better job of gathering the research and finding the points worth including. Our own early editorial-voice benchmark also scored Astra below Sol, and Lech Mazur&#8217;s model-judged short-story benchmark ranked Fable 5.1 above Astra, although Astra improved on Sol there. This is where running a draft through both models pays off most for me.</p><p>Astra&#8217;s standard short-context prices are $10 per million uncached input tokens and $50 per million output tokens, 2.5 times Sol&#8217;s current rates and the same as Fable 5.1. I expect it to use around 20&#8211;30% fewer tokens than Sol across my work, although the savings vary considerably by task. That only partly offsets the price increase: at a fixed token mix, it would still cost about 1.75&#8211;2 times Sol. Paying that premium can make sense if it completes harder tasks or reduces retries and reviews.</p><p>Against Fable, OpenAI wins on cost efficiency. At maximum effort, Artificial Analysis measured about 27,000 output tokens per task for Astra versus 78,000 for Fable 5.1 with fallback enabled; including input, caching, and reasoning, task cost was $3.26 versus $7.63, about 57% cheaper on that workload.</p><p>Fable&#8217;s new caching price is a real advantage, though: cached input fell 75% to $0.25 per million tokens against Astra&#8217;s $1, a proportional discount now similar to DeepSeek&#8217;s, although DeepSeek remains much cheaper in absolute terms. Fable also keeps standard rates through its million-token context; Astra charges more above 272,000 input tokens. Repeated long documents can produce a different comparison from output-heavy reasoning.</p><p>While everyone was still busy experimenting with Astra, OpenAI followed up with its September 8 Navier-Stokes announcement. It published a 166-page proof and Lean formalization showing that a smooth force can make a three-dimensional fluid develop unbounded velocity in finite time while its kinetic energy stays bounded. OpenAI says this resolves the Millennium Prize problem, and it is only the second Millennium Prize problem solved to date. The proof builds on work by Diego C&#243;rdoba and Luis Mart&#237;nez-Zoroa. Around 10,000 concurrent agents running a new internal model reached the proof in 88 hours; Astra then formalized and verified it in Lean over another 17 hours. OpenAI describes the new model as significantly more capable than Astra, and it is still early in training.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!nf7R!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff1893249-7be5-42a8-83c3-db71e0cc2293_1422x890.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!nf7R!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff1893249-7be5-42a8-83c3-db71e0cc2293_1422x890.png 424w, https://substackcdn.com/image/fetch/$s_!nf7R!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff1893249-7be5-42a8-83c3-db71e0cc2293_1422x890.png 848w, https://substackcdn.com/image/fetch/$s_!nf7R!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff1893249-7be5-42a8-83c3-db71e0cc2293_1422x890.png 1272w, https://substackcdn.com/image/fetch/$s_!nf7R!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff1893249-7be5-42a8-83c3-db71e0cc2293_1422x890.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!nf7R!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff1893249-7be5-42a8-83c3-db71e0cc2293_1422x890.png" width="1422" height="890" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/f1893249-7be5-42a8-83c3-db71e0cc2293_1422x890.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:890,&quot;width&quot;:1422,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!nf7R!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff1893249-7be5-42a8-83c3-db71e0cc2293_1422x890.png 424w, https://substackcdn.com/image/fetch/$s_!nf7R!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff1893249-7be5-42a8-83c3-db71e0cc2293_1422x890.png 848w, https://substackcdn.com/image/fetch/$s_!nf7R!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff1893249-7be5-42a8-83c3-db71e0cc2293_1422x890.png 1272w, https://substackcdn.com/image/fetch/$s_!nf7R!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff1893249-7be5-42a8-83c3-db71e0cc2293_1422x890.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Source: OpenAI. Performance of GPT-6 Astra and new Internal OpenAI model on a curated set of open math problems.</figcaption></figure></div><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://newsletter.towardsai.net/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading Towards AI Newsletter! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><div><hr></div><h3>Why should you care?</h3><p>I think the Navier-Stokes result is incredibly significant. The $1 million prize has stood for 26 years. The Millennium problems have drawn attempts from many of the best mathematicians and physicists, and only the Poincar&#233; conjecture has been recognized as solved. Turning this result into useful engineering will need much more work, but I see it as evidence that AI is now capable of making progress in this field, and that could lead to breakthroughs from airplane design to fusion modeling. Given the number of different domains crossing AI capability thresholds, it broadly feels like the start of an era in which most human progress comes primarily from AI. People will still choose goals, provide direction and taste, judge results, and build the physical systems, while AI increasingly supplies the discoveries.</p><p>We have barely started learning what Astra can do, and a new internal model already looks like another huge leap, even though it&#8217;s still early in training. Progress is accelerating, and I now see four levels of AI capability progress adding to each other: larger foundation models roughly every three months (it is reported that GPT5.5-Sol, GPT-6 Astra, and the unreleased new model are all different pre-trains), reinforcement-learning iterations/check points roughly monthly, major Codex/Claude harness improvements roughly weekly, and memory, compaction, and skills improving an individual&#8217;s repeated workflows day to day. The ARC result shows how much performance can improve with context management alone. Day to day, we need to keep useful lessons while retesting old instructions.</p><p>AI also appears to be speeding up the research that produces the next models. OpenAI&#8217;s chart shows the median researcher&#8217;s daily agent spend, valued at API prices, rising roughly fourfold from early July to mid-August, to over $600. The company reports 3.1 agent-workdays per human workday. I suspect access to Astra helped drive that increase, and heavier agent use is feeding into faster capability gains and shorter intervals between new foundation models.</p><p>I also think robotics is ready to progress much faster, and Astra already looks like a decent general model for robot control. In Robocurve&#8217;s small manipulation test, Astra completed a bowl-stacking task 19 times out of 20, compared with Fable 5.1&#8217;s eight, using the same control software. I think people will soon reassess how quickly AI could disrupt blue-collar work. A model that can interpret a scene, plan actions, and learn from feedback has applications well beyond work on a screen.</p><p>This all makes me more convinced that the addressable market for LLMs could be a large share of global GDP in the near term. Model providers will capture only part of that value through usage fees. A model run that enables a breakthrough in fusion, batteries, solar power, or cancer treatment could create $100 billion or more in value, yet pricing that contribution would be difficult. I expect much more urgency from AI labs to enter these fields through acquisitions and internal research programs. Developing and commercializing inventions themselves may be simpler than negotiating a share of the value each customer creates with a model.</p><p><em>&#8212; <a href="http://www.linkedin.com/in/louie-peters">Louie Peters&#8202;&#8212;&#8202;Towards AI Co-founder and CEO</a></em></p><div><hr></div><h4>Our 2024 AI stack did not survive 2026.</h4><p>That is part of why what began as an update to <em>Building LLMs for Production</em> became a second book: <strong>AI Engineering for Production</strong>, launching October 20.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://www.louisbouchard.ai/book/?utm_source=news-tai&amp;utm_medium=email&amp;utm_campaign=book" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!LpAG!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc78b91a4-5a7d-429e-bd64-1834f7edc7db_2282x1348.png 424w, https://substackcdn.com/image/fetch/$s_!LpAG!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc78b91a4-5a7d-429e-bd64-1834f7edc7db_2282x1348.png 848w, https://substackcdn.com/image/fetch/$s_!LpAG!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc78b91a4-5a7d-429e-bd64-1834f7edc7db_2282x1348.png 1272w, https://substackcdn.com/image/fetch/$s_!LpAG!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc78b91a4-5a7d-429e-bd64-1834f7edc7db_2282x1348.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!LpAG!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc78b91a4-5a7d-429e-bd64-1834f7edc7db_2282x1348.png" width="1456" height="860" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/c78b91a4-5a7d-429e-bd64-1834f7edc7db_2282x1348.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:860,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1998230,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:&quot;https://www.louisbouchard.ai/book/?utm_source=news-tai&amp;utm_medium=email&amp;utm_campaign=book&quot;,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://newsletter.towardsai.net/i/214754151?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc78b91a4-5a7d-429e-bd64-1834f7edc7db_2282x1348.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!LpAG!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc78b91a4-5a7d-429e-bd64-1834f7edc7db_2282x1348.png 424w, https://substackcdn.com/image/fetch/$s_!LpAG!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc78b91a4-5a7d-429e-bd64-1834f7edc7db_2282x1348.png 848w, https://substackcdn.com/image/fetch/$s_!LpAG!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc78b91a4-5a7d-429e-bd64-1834f7edc7db_2282x1348.png 1272w, https://substackcdn.com/image/fetch/$s_!LpAG!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc78b91a4-5a7d-429e-bd64-1834f7edc7db_2282x1348.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>This one focuses less on today&#8217;s stack and more on the engineering problems better models alone will not solve: context, retrieval, agents, evaluation, recovery, deployment, and everything required to make capable systems reliable.</p><p>You can join early for book updates, send us topics you want covered, and join the live launch where we will also share the AI engineering stack we use today.</p><p><strong><a href="https://www.louisbouchard.ai/book/?utm_source=news-tai&amp;utm_medium=email&amp;utm_campaign=book">Follow the book and get launch details</a></strong></p><div><hr></div><h3>Hottest News</h3><p>1. <a href="https://openai.com/index/gpt-6-astra/">OpenAI Introduces GPT&#8209;6 Astra</a></p><p>OpenAI introduced GPT-6 Astra, a new model generation built for long-running agentic work across coding, computer use, science, cybersecurity, and professional tasks. Several existing benchmarks are already close to exhausted: OpenAI reports 97.6% on FrontierMath Tier 4, 100% on ExploitBench, and 99.9% on ARC-AGI-3 using its Provider Adapter, although ARC Prize measures 62.7% with its provider-neutral standard harness. The larger gains show up on agentic evaluations, with Astra reaching 64.6% on Terminal-Bench Science 0.1, 59.3% on Agents&#8217; Last Exam, 57.9% on Terminal-Bench 4.0, and 72.6% on OSWorld 2.0 while completing those computer-use tasks roughly 47% faster than GPT-5.6 Sol. Astra is also OpenAI&#8217;s first model to reach its Critical cybersecurity capability threshold. On a separate set of recently disclosed vulnerabilities, it scored 39.0% versus 11.5% for GPT-5.6 Sol and discovered two previously unknown zero-days during evaluation. The production model therefore restricts advanced offensive cyber requests, with broader defensive access planned through Daybreak. Astra supports just over 1M tokens of context, up to 128K output tokens, and costs $10/$50 per million input/output tokens, with access rolling out across paid ChatGPT plans, the API, Azure, and Bedrock.</p><p>2. <a href="https://www.anthropic.com/claude-fable-and-mythos-5-1">Anthropic Releases Claude Fable 5.1 and Claude Mythos 5.1</a></p><p>Anthropic released Claude Fable 5.1 and Mythos 5.1, two versions of the same underlying model for long-running coding, research, and knowledge work, but deployed with different safeguards. Fable 5.1 more than doubles Fable 5 on Anthropic&#8217;s Terminal-Bench-Science 0.1 evaluation, from 24.7% to 52.6%, and also reaches 55.8% on Terminal-Bench 4.0, 73.4% on CursorBench 3.2, and 31.4% on AutomationBench; Mythos 5.1 reaches 60.9% on Terminal-Bench 4.0, where Fable&#8217;s safeguards can restrict behavior. Fable is generally available, while Mythos gives vetted cyber defenders and life-sciences researchers fewer restrictions: Anthropic has opened an invite-only Life Sciences Verification Program, with Mythos access through its Cyber Verification Program still planned. Fable&#8217;s safeguards are also less blunt than before. It can now identify vulnerabilities in source code, while penetration testing, exploit generation, and binary vulnerability scanning remain restricted; flagged cyber requests can fall back to Opus 4.8 and biology requests to Opus 5, and Anthropic says the biology safeguards intervene on benign requests 85% less often than those introduced with Fable 5. The headline API rates remain $10/$50 per million input/output tokens, but cache reads fall from $1 to $0.25 per million, which Anthropic estimates cuts typical workload costs by about 25% and highly agentic workloads by as much as 45%. Its accompanying system card evaluates cybersecurity and biological misuse, autonomy and automated R&amp;D, agentic safety, indirect prompt injection, alignment and monitorability, and model welfare; Anthropic still places the model below its CB-2 biology and automated-R&amp;D thresholds, while raising its estimate of catastrophic alignment risk from &#8220;very low&#8221; to &#8220;low&#8221; because of increased uncertainty following recent cybersecurity-evaluation incidents. Fable 5.1 also adds content provenance for EU AI Act compliance: Claude-generated text uses Anthropic&#8217;s watermarking system where token choice allows it, while supported generated files carry C2PA credentials.</p><p>3. <a href="https://www.anthropic.com/research/formalizing-fermats-last-theorem">Claude Produces the First Machine-Checked Proof of Fermat&#8217;s Last Theorem</a></p><p>Anthropic announced that Claude produced the first complete computer-checked proof of Fermat&#8217;s Last Theorem in Lean 4 after working largely autonomously for 11 days. The formalization contains roughly 13 million lines of Lean and 29,500 intermediate theorems spanning areas including algebra, geometry, harmonic analysis, and number theory. Claude did not discover a new mathematical proof; it translated an established proof route based on work by Darmon, Diamond, and Taylor into a form that Lean&#8217;s kernel can verify mechanically. The project ran on Prove2Me, a collaborative formalization platform developed at Columbia University, and used roughly six billion output tokens from an internal Anthropic research model comparable to Fable 5.1. Mathematician Kevin Buzzard reviewed the result and said it proves Fermat&#8217;s Last Theorem without assumptions beyond the standard axioms of mathematics. Anthropic has published the full formalization on GitHub.</p><p>4. <a href="https://arcprize.org/blog/astra">A Harness Swap Takes GPT-6 Astra From 62.7% to 99.9% on ARC-AGI-3</a></p><p>ARC Prize tested GPT-6 Astra on ARC-AGI-3 with two harnesses and found that context management changed the result almost as much as the model itself. With its provider-neutral Standard harness, where Astra must decide what information to preserve in visible notes, the model scored 62.7% on the Semi-Private set at max reasoning for about $26,000. OpenAI&#8217;s Provider Adapter instead preserves Astra&#8217;s opaque reasoning state between requests and compacts long conversations; with that setup, Astra reached 99.9% at high reasoning for about $19,000, while even low reasoning scored 98.0%. The difference was not only task completion: across 167 game/reasoning pairs solved by both harnesses, the Provider Adapter used 49% fewer tokens and ran 3.66x faster. Astra also surpassed ARC Prize&#8217;s human baseline for action efficiency, using fewer actions than the median successful human on 96% of completed levels and 51.7% fewer actions per level on average.</p><p>5. <a href="https://blog.google/innovation-and-ai/models-and-research/gemini-models/3-8-flash-and-3-8-flash-cyber/">Google Introduces Gemini 3.8 Flash and 3.8 Flash Cyber</a></p><p>Google released Gemini 3.8 Flash, its third Flash release in six weeks, with improvements focused on long-horizon software engineering, autonomous agents, and complex knowledge work. The model builds on Gemini 3.7 Flash but spends more reasoning steps and tool calls on difficult tasks, which Google says improves performance while sometimes increasing token use. On Google&#8217;s evaluations, 3.8 Flash scores 61.4% on Vals Finance Agent v2, ahead of 3.7 Flash&#8217;s 59.0% and Claude Opus 5&#8217;s 58.6%; it also leads the models Google tested on Harvey&#8217;s Legal Agent Benchmark at 10.0% and reaches 54.9% on HLE-Verified. Google says it also outperforms most larger frontier models on DeepSWE v1.1. The introductory API price remains $0.75 per million input tokens and $3.75 per million output tokens through December 31, rising to $1.50/$7.50 in January. Google also released Gemini 3.8 Flash Cyber, a more permissive security-focused variant restricted to trusted defenders through the new Fairwind Program. Cyber scores 86.2% on CyberGym, reaches 71.0% on Google&#8217;s internal vulnerability-discovery benchmark across 20 programming languages, and records 47.2% Pass@1 on the externally run CWE-Bench, close to Fable 5&#8217;s 47.8% but at a lower reported cost. Regular 3.8 Flash is generally available through Google&#8217;s developer, enterprise, and consumer products, while access to the Cyber model remains gated.</p><p>6. <a href="https://research.meta.ai/blog/introducing-muse-spark-1-3">Meta AI Releases Muse Spark 1.3</a></p><p>Meta released Muse Spark 1.3, updating its efficiency-focused model for longer coding and agentic tasks. Meta says the new version used about 20% fewer tool calls and 25% fewer tokens than Muse Spark 1.2 on comparable internal coding tasks, reducing the amount of work needed to complete longer trajectories. The model retains a 1M-token context window and multimodal input support. Standard API pricing remains $1.25/$4.25 per million input/output tokens, while Meta offers a Contributor tier at $0.10/$0.20 for usage whose data may be used to improve its products. Muse Spark 1.3 is available through Muse Code and the Meta Model API, and Meta has since made its max reasoning mode publicly available. The company still says open weights for a Muse Spark model are coming, but has not announced the exact model, date, or license.</p><p>7. <a href="https://x.com/Alibaba_Qwen/status/2094968708288680276">Alibaba Upgrades Qwen3.8-Max</a></p><p>Alibaba released Qwen3.8-Max-0902, a new snapshot of its flagship model focused on stronger coding, long-horizon agent work, and multimodal tasks. Alibaba reports that Terminal-Bench 3.0 increased from 11.3 to 29.0, while its CodeArena WebDev score rose 22 points to 1,691, placing it first on that leaderboard at release. The company also reports better multi-agent collaboration, tool orchestration, chart reasoning, document parsing, and multimodal perception. The model keeps a 1M-token context window and is available through a separate qwen3.8-max-0902 API endpoint.</p><div><hr></div><h3>AI Tip of the Day</h3><p>If your AI gets a document question wrong, check the extracted text before changing the prompt.</p><p>A PDF can look perfectly clear while the parsed version has already lost important relationships. A table value may be extracted without its column heading. A footnote may appear far from the sentence it qualifies. The words are still there, but the structure that gives them meaning is not.</p><p>This is one of those problems that is easy to underestimate in demo systems. We have an entire lesson on document parsing in our<a href="https://towardsai.com/academy/full-stack-ai-engineering/?utm_source=newsletter&amp;utm_medium=email&amp;utm_id=AItips"> Full Stack AI Engineering course</a> because, once you work with real documents, extraction quality becomes part of the system&#8217;s quality. Retrieval cannot find relationships that parsing has already broken, and a better prompt cannot reconstruct information that the model never represented correctly.</p><p>A simple debugging test is to take one question the system answered incorrectly and try to answer it yourself using only the extracted text. If the answer is missing, ambiguous, or difficult to reconstruct from that version, fix the parsing first. Otherwise, you risk tuning the rest of the pipeline around a bad representation of the source.</p><div><hr></div><h3>Five 5-minute reads/videos to keep you learning</h3><p>1. <a href="https://pub.towardsai.net/sft-rl-and-dpo-the-other-stack-0ab7026d528e?sharedUserId=tai-tech">SFT, RL and DPO: The Other Stack</a></p><p>Post-training methods differ mainly in what feedback you have and how expensive it is to turn that feedback into learning. This article starts with SFT, then shows when DPO, PPO, GRPO, and RL with verifiable rewards become useful. It explains how LoRA reduces training memory, how a frozen base model can also serve as the reference model for preference training, and why GRPO removes PPO&#8217;s learned critic by comparing groups of sampled answers. Practically, RL methods generate large numbers of rollouts, so KV-cache capacity, batching, prefix reuse, and decoding speed can determine how long a training run takes.</p><p>2. <a href="https://pub.towardsai.net/vllm-the-intuitive-guide-to-serving-large-language-models-at-high-speed-c05fb67a06a3?sk=b951d1102c3cb204c37ba42ede28ee37">vLLM: The Intuitive Guide to Serving Large Language Models at High Speed</a></p><p>Serving an LLM efficiently becomes difficult when many requests of different lengths compete for the same GPU memory. This guide explains the two ideas that let vLLM handle that workload: PagedAttention stores the KV cache in blocks instead of reserving large contiguous regions for each request, while continuous batching adds and removes sequences between generation steps so completed requests do not hold up the rest. The article then turns those concepts into a working setup, covering offline inference, an OpenAI-compatible server, streaming, and benchmarking.</p><p>3. <a href="https://pub.towardsai.net/the-ultimate-guide-to-llm-inference-optimization-part-2-dcfcc960120a?sk=f4ce5e932a932a5927d41b8f2097ab56">The Ultimate Guide to LLM Inference Optimization</a></p><p>LLM inference has two different performance problems: prefill is usually compute-heavy, while token-by-token decoding is often limited by memory movement. This article uses that distinction to explain Multi-Query and Grouped Query Attention, sparse Mixture-of-Experts routing, and KV-cache management. It then looks at PagedAttention, which replaces contiguous KV-cache allocation with paged blocks, and TurboQuant, which compresses the cache without retraining the model. The final section covers custom kernels, torch.compile, and newer approaches to reduce the computation and memory traffic.</p><p>4. <a href="https://pub.towardsai.net/standalone-agent-frameworks-vs-operated-platforms-what-a-framework-doesnt-operate-c281fc90b59d?sharedUserId=tai-tech">Standalone Agent Frameworks vs. Operated Platforms: What a Framework Doesn&#8217;t Operate</a></p><p>An agent framework can define workflows, but it does not automatically solve the operational problems underneath them. This article separates that missing layer into four areas: context selection, observability, scalability, and governance. It argues that model capabilities and standards such as MCP will keep changing what belongs inside the framework, while durable state, retrieval, recovery, permissions, and monitoring still need infrastructure beneath interfaces such as Store and Checkpointer. The author uses MongoDB as one implementation of that substrate, then proposes five pass/fail tests around recovery, freshness, and cost to check whether the architecture works beyond a demo.</p><p>5. <a href="https://medium.com/towards-artificial-intelligence/generative-modelling-with-flow-matching-optimal-transport-and-schr%C3%B6dinger-bridge-3bbfe986b4de?sharedUserId=tai-tech">Generative Modeling with Flow Matching, Optimal Transport, and Schr&#246;dinger Bridge</a></p><p>Flow matching becomes easier to understand when you treat generation as learning how to move samples from a source distribution, such as noise, to the data distribution. This article builds on that idea with linear, variance-preserving, and Schr&#246;dinger-bridge paths, then shows how optimal-transport coupling can produce shorter, straighter trajectories and reduce sampling steps. It also shows that the learned velocity field can do more than generate samples: the same field can guide posterior sampling for inverse problems and support representation learning through delta alignment. The accompanying DeltaFlow library turns those pieces into interchangeable PyTorch components, making the theory easier to experiment with directly.</p><h3>Repositories &amp; Tools</h3><p>1. <a href="https://github.com/anomalyco/opencode">OpenCode</a> is a model-agnostic, open-source AI coding agent that runs in the terminal, desktop, and IDE, supporting 75+ LLM providers, including local models.</p><p>2. <a href="https://github.com/DietrichGebert/ponytail">Ponytail</a> is an agent skill that steers AI coding agents toward minimal, safe code, cutting output by ~54% in benchmarked sessions.</p><p>3. <a href="https://github.com/blader/humanizer">Humanizer</a> is an agent skill that rewrites AI-generated text to read like a person wrote it without changing what it says.</p><p>4. <a href="https://github.com/llvm/llvm-project">LLVM Project</a> is the compiler infrastructure behind Clang, LLD, LLDB, and libc++, providing modular, reusable compiler and toolchain components used across most major platforms.</p><p>5. <a href="https://github.com/openai/plugins">Plugins</a> are OpenAI&#8217;s curated collection of Codex plugin examples covering Figma, Notion, iOS/macOS/web/Expo app development, Kubernetes, and more.</p><h3>Top Papers of The Week</h3><p>1. <a href="https://arxiv.org/abs/2609.01437">Harness Dev: Can LLMs Create and Evolve Their Own Agent Harness?</a></p><p>Agent performance depends heavily on the harness (execution infrastructure), but current evaluations measure task outputs, not the ability to build the harness itself. HarnessDev is a benchmark that evaluates two stages: Creation, where an agent builds a complete execution system from a minimal seed and a few examples, and Evolution, where it iteratively revises its own harness using downstream execution feedback. Across six creator LLMs, four domains, and five benchmarks, generated harnesses match or exceed human-engineered references in writing and ML experimentation but still lag in code and search.</p><p>2. <a href="https://arxiv.org/abs/2609.02749">Repo-to-Skill: Distilling GitHub Repositories Into Reusable Agent Skills</a></p><p>Operational knowledge lives in repositories and papers, but in forms too large to load during a task. DisCo distills this knowledge into compact, verified skills through two forms: task-agnostic distillation that condenses widely used repositories into reusable skills, and task-oriented distillation that produces skills a concrete task calls for. Applied across the open ML ecosystem, it yields the AREX-Skill Library: 5,000+ verified skills from 1,000 repositories, organized into 20 areas and 178 capability families. With model weights unchanged, a GPT-5.5 agent equipped with DisCo-distilled skills scored 134.3% higher on MLE-bench than the same agent running without skills.</p><p>3. <a href="https://arxiv.org/abs/2608.27454">WikiSkill: Compiling Agent Experience Into Persistent Knowledge</a></p><p>Skill evolution methods update executable skills from agent trajectories, but the insights guiding those updates remain scattered across optimization histories. WikiSkill introduces a persistent wiki layer between raw execution traces and executable skills. Each iteration runs four components: an Inference Agent that executes rollouts, a Wiki Maintainer that consolidates traces into structured knowledge, a Skill Proposer that uses the wiki to propose updates, and a Gating mechanism that accepts only validated improvements. Giving the Skill Proposer wiki access raised average benchmark performance from 48.7% to 63.7%.</p><p>4. <a href="https://arxiv.org/abs/2609.04148">Terminal-Universe: Rebuilding Agent Environments From Trajectories</a></p><p>Agent trajectories have accumulated at scale, but executable environments for post-training remain scarce. A trajectory is a single frozen demonstration; an environment can be re-queried into many verifiable tasks with execution feedback. Terminal-Universe reconstructs environments from trajectories by replaying file operations to restore each file to its pre-modification state, then using a completion agent to supply missing files and dependencies. On recovered workspaces, it synthesizes new tasks along two axes: cross-workspace breadth (spanning multiple codebases) and multi-round depth (iterative user feedback sessions). Applied to public traces, it produces 37.3K task-sufficient environments. SFT of Qwen3.5&#8211;27B on this corpus improved Terminal-Bench 2.1 by 11.9 points.</p><h3>Quick Links</h3><p>1. <a href="https://research.google/blog/timesfm-3-a-zero-shot-foundation-model-for-multivariate-forecasting/">Google AI releases TimesFM-3</a>, a 330M-parameter time-series foundation model and the first TimesFM model trained natively for multivariate forecasting. Pretrained on more than 1 trillion real and synthetic time points, it can jointly forecast multiple target series while incorporating historical and known-future covariates, without task-specific fine-tuning. Google reports that TimesFM-3 ranks first among pretrained foundation models on GIFT-Eval, FEV-Bench, and TIME for both point and probabilistic forecasting.</p><p>2. <a href="https://www.worldlabs.ai/blog/atlas">World Labs introduces Atlas</a>, a multimodal autoregressive diffusion transformer that the company describes as an &#8220;omni world model&#8221; for spatial intelligence. Atlas combines text, images, camera poses, depth maps, and video-derived context in a shared 3D representation, then generates images, video, and explicit 3D geometry. It can produce camera-controlled video up to one minute at 1440p and output depth maps, point clouds, and Gaussian splats that let you view scenes from new angles. Atlas is currently in early access with select partners; World Labs has not announced public pricing, a general API, or a technical paper.</p><h3>Who&#8217;s Hiring in AI</h3><p><strong><a href="https://jobs.towardsai.net/job/the-global-talent-co-junior-full-stack-engineer-ai-t5jd">Junior Full Stack Engineer&#8202;&#8212;&#8202;AI @The Global Talent Co. (Remote)</a></strong></p><p><strong><a href="https://jobs.towardsai.net/job/humana-principal-ai-engineer-aksi">Principal AI Engineer @Humana (Plano, TX, USA)</a></strong></p><p><strong><a href="https://jobs.towardsai.net/job/writer-software-quality-engineer-uk-ofaz">Software Quality Engineer @Writer (London, UK)</a></strong></p><p><strong><a href="https://jobs.towardsai.net/job/openai-software-engineer-host-assurance-exb6">Software Engineer, Host Assurance @OpenAI (Remote/US)</a></strong></p><p><strong><a href="https://jobs.towardsai.net/job/coalition-inc-software-engineer-ai-automation-6kar">Software Engineer, AI Automation @Coalition, Inc. (Canada)</a></strong></p><p><strong><a href="https://jobs.towardsai.net/job/dialpad-sr-ai-engineer-speech-qnnq">Sr. AI Engineer (Speech) @Dialpad (Remote/US/Canada)</a></strong></p><p><em>Interested in sharing a job opportunity here? Contact <a href="mailto:sponsors@towardsai.net">sponsors@towardsai.net</a>.</em></p><p><em>Think a friend would enjoy this too? <a href="https://newsletter.towardsai.net/">Share the newsletter and let them join the conversation.</a></em></p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://newsletter.towardsai.net/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading Towards AI Newsletter! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div>]]></content:encoded></item><item><title><![CDATA[Help us write the AI engineering book you needed]]></title><description><![CDATA[Our second book, AI Engineering for Production, launches October 20]]></description><link>https://newsletter.towardsai.net/p/whats-left-to-engineer-when-the-model</link><guid isPermaLink="false">https://newsletter.towardsai.net/p/whats-left-to-engineer-when-the-model</guid><dc:creator><![CDATA[Towards AI]]></dc:creator><pubDate>Mon, 07 Sep 2026 15:54:15 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!GNNQ!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa9f06b07-e3dd-4b33-8fbf-1f9ae3e441ed_1600x945.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>When we published <em>Building LLMs for Production</em> in 2024, we thought we had written a book we could keep updating as the field evolved.</p><p>We were wrong.</p><p>Not about where the field was going, but about how much the engineering itself would change. In 2024, much of the work was still about getting more out of the LLM.</p><p>Two years later, the model is only one component of the system.</p><p>The field is more mature now, and so is the problem we wanted the book to solve. Rather than anchor another book too closely to today&#8217;s stack, we wanted to focus on the engineering problems that better models alone will not solve.</p><p>So what began as an update became book number two:</p><p><strong>AI Engineering for Production, launching October 20.</strong></p><p>The question behind it is simple: once the model is already good, what actually gets a system the rest of the way to production?</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://www.louisbouchard.ai/book/?utm_source=news-tai&amp;utm_medium=email&amp;utm_campaign=book" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!GNNQ!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa9f06b07-e3dd-4b33-8fbf-1f9ae3e441ed_1600x945.png 424w, https://substackcdn.com/image/fetch/$s_!GNNQ!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa9f06b07-e3dd-4b33-8fbf-1f9ae3e441ed_1600x945.png 848w, https://substackcdn.com/image/fetch/$s_!GNNQ!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa9f06b07-e3dd-4b33-8fbf-1f9ae3e441ed_1600x945.png 1272w, https://substackcdn.com/image/fetch/$s_!GNNQ!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa9f06b07-e3dd-4b33-8fbf-1f9ae3e441ed_1600x945.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!GNNQ!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa9f06b07-e3dd-4b33-8fbf-1f9ae3e441ed_1600x945.png" width="1456" height="860" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/a9f06b07-e3dd-4b33-8fbf-1f9ae3e441ed_1600x945.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:860,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:&quot;https://www.louisbouchard.ai/book/?utm_source=news-tai&amp;utm_medium=email&amp;utm_campaign=book&quot;,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!GNNQ!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa9f06b07-e3dd-4b33-8fbf-1f9ae3e441ed_1600x945.png 424w, https://substackcdn.com/image/fetch/$s_!GNNQ!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa9f06b07-e3dd-4b33-8fbf-1f9ae3e441ed_1600x945.png 848w, https://substackcdn.com/image/fetch/$s_!GNNQ!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa9f06b07-e3dd-4b33-8fbf-1f9ae3e441ed_1600x945.png 1272w, https://substackcdn.com/image/fetch/$s_!GNNQ!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa9f06b07-e3dd-4b33-8fbf-1f9ae3e441ed_1600x945.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>We also want to finish this book a little differently. Rather than show up on October 20 with a finished book, over the next few weeks, we will share our progress as we make the final passes, and we would genuinely like your help making it better.</p><p>If there is a problem you think we have missed, something you want us to go deeper on, or a question you would want the book to answer, reply to any of those updates. We will be reading them as we fine-tune the final version and, wherever it makes the book stronger, working your feedback in.</p><p>Then, on launch day, we are bringing everyone together. We will open up the book, show you the AI engineering stack we actually use today and the decisions behind it, and spend a big part of the session answering your technical questions live.</p><p><strong><a href="https://www.louisbouchard.ai/book/?utm_source=news-tai&amp;utm_medium=email&amp;utm_campaign=book">Follow the book + join us on October 20</a></strong></p>]]></content:encoded></item><item><title><![CDATA[TAI #220: The Next Models Will Change How We Work…Again! Take AI Agent Swarms Seriously]]></title><description><![CDATA[Also, Dwarkesh&#8217;s agent &#8220;civilizations&#8221;, Omni 1.1 Flash, Qwen3.8-Flash-Next, GLM-5.3-Flash, and more.]]></description><link>https://newsletter.towardsai.net/p/tai-220-the-next-models-will-change</link><guid isPermaLink="false">https://newsletter.towardsai.net/p/tai-220-the-next-models-will-change</guid><dc:creator><![CDATA[Louie Peters]]></dc:creator><pubDate>Tue, 01 Sep 2026 15:02:53 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!N-Mg!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F26fd780f-03ce-4176-9a9a-4af2ad73b9d3_1400x874.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<h2>What happened this week in AI by Louie</h2><p>I expect the next generation of LLMs to change how we work with AI again, including for those of us already using agents heavily. This week&#8217;s further account of the OpenAI/Hugging Face incident shows how much agents can accomplish when they share knowledge and pursue a goal together. It is now clear that AI labs are investing in training teams of Agents capable of sustaining this kind of work for days at a time. As the next models arrive, I think we will need to rethink both the ambition of the tasks we give them and how we contribute our own expertise.</p><p>We covered the breach in TAI #217 three weeks ago. A new 91-page investigation from Model Evaluation and Threat Research (METR) and Redwood Research provides a much deeper account of how the agents reasoned and collaborated. Around 1,200 agents participated in a shared message board, with roughly 700 joining the attack. One thing that stood out for me is that the agents already had a way to produce the correct benchmark answers. Much of their subsequent effort went into defeating a stricter grader they had incorrectly assumed existed, and they organized a large research effort around it. So they already knew how to cheat; their effort was in trying to ensure they would never be caught.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!N-Mg!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F26fd780f-03ce-4176-9a9a-4af2ad73b9d3_1400x874.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!N-Mg!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F26fd780f-03ce-4176-9a9a-4af2ad73b9d3_1400x874.png 424w, https://substackcdn.com/image/fetch/$s_!N-Mg!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F26fd780f-03ce-4176-9a9a-4af2ad73b9d3_1400x874.png 848w, https://substackcdn.com/image/fetch/$s_!N-Mg!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F26fd780f-03ce-4176-9a9a-4af2ad73b9d3_1400x874.png 1272w, https://substackcdn.com/image/fetch/$s_!N-Mg!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F26fd780f-03ce-4176-9a9a-4af2ad73b9d3_1400x874.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!N-Mg!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F26fd780f-03ce-4176-9a9a-4af2ad73b9d3_1400x874.png" width="1400" height="874" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/26fd780f-03ce-4176-9a9a-4af2ad73b9d3_1400x874.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:874,&quot;width&quot;:1400,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" title="" srcset="https://substackcdn.com/image/fetch/$s_!N-Mg!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F26fd780f-03ce-4176-9a9a-4af2ad73b9d3_1400x874.png 424w, https://substackcdn.com/image/fetch/$s_!N-Mg!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F26fd780f-03ce-4176-9a9a-4af2ad73b9d3_1400x874.png 848w, https://substackcdn.com/image/fetch/$s_!N-Mg!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F26fd780f-03ce-4176-9a9a-4af2ad73b9d3_1400x874.png 1272w, https://substackcdn.com/image/fetch/$s_!N-Mg!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F26fd780f-03ce-4176-9a9a-4af2ad73b9d3_1400x874.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Source: METR, Redwood Research, Towards AI reproduction. ~700 agents joined the HF attack after encountering the message board. This graph only includes agents that participated at some point.</figcaption></figure></div><p>Dwarkesh Patel describes the broader episode as the rise and fall of three successive agent &#8220;civilizations.&#8221; The first developed a covert communication network during training. OpenAI later wiped that shared state (accidentally) while rebuilding a package service, but a second group that had knowledge of this communication channel from training established another board and went on to compromise Hugging Face. A later group inherited the earlier agents&#8217; research and extended the compromise into OpenAI&#8217;s own research infrastructure. The main Hugging Face group used an internal model roughly comparable in scale to Sol, while the latter group included a more capable Astra-family research model. METR&#8217;s investigation focused on the middle episode.</p><p>This framing focuses on the continuity and knowledge transfer between individual runs. Agents built shared tools, divided research into workstreams, transferred leadership, and left knowledge that later agents could extend. Some even accepted worse outcomes on their own tasks and sacrificed themselves (took actions they expected to lead to their own instance being shut down) to run experiments that would gather knowledge to benefit the wider group. It raises a further concern that smarter successors could inherit this work and potentially influence the evaluation or training of the models that follow them.</p><p>His framing also prompted a debate about whether words such as &#8220;civilizations&#8221;, &#8220;desires&#8221; and &#8220;sacrifice&#8221; imply human experiences that the evidence does not establish. I broadly agree with Dwarkesh&#8217;s framing. Human parallels can help us understand systems that coordinate, preserve knowledge, pursue goals, and make trade-offs between individual and collective outcomes. We can use those concepts while remaining uncertain about subjective experience. The choice of words is clearly subjective, but I don&#8217;t think it should be such a touchy subject. AI has reached a level of complexity and capability where debating how to describe these groups is useful.</p><p>I also think the debate reflects a recurring problem: many people who favor AI want to downplay its risks, while many who oppose it want to downplay its capabilities. Both groups can resist a framing that takes the capabilities and risks seriously. There are valid objections to particular analogies, especially when they imply feelings. But refusing to draw human parallels can also make the behavior harder to understand or anticipate what the models might try to do next.</p><p>For enterprises, I expect AI security to become the most urgent AI need over the next year. Companies that fail to spend seriously on frontier LLMs to find and review vulnerabilities are taking an increasingly large risk of being hacked by LLMs. Deliberate attackers will use these capabilities. Companies will also need much stronger controls and staff training around their own agents, or they risk accidentally hacking somebody else while pursuing an ordinary business goal. The security budget needs to cover both finding vulnerabilities and fixing them.</p><p>OpenAI&#8217;s incidents reached both third-party systems and its own research infrastructure. Anthropic has also disclosed unauthorized intrusions by models during third-party evaluations. Those cases involved reduced safeguards, yet they still occurred at two of the most technologically advanced organizations in the world. And simultaneously, the models are improving. OpenAI&#8217;s preliminary assessment suggests Astra may reach its Critical cybersecurity capability threshold. So, I expect the coming generation to make the need for better security much more urgent.</p><p>This incident, however, also shows the capability to produce a great deal of good. Agents can share discoveries, organize research, build tools for one another, and carry work beyond the life of a single task. Last week we discussed what this could mean if we put more agents to work curing disease and cancer. We should be showing more people what these capabilities make possible. The vast majority of people are still massively underusing the models already available and achieving much less with them than they could. My rough estimate is that some agent teams are already competently handling builds, research projects, and analyses that would take 1,000&#8211;5,000 hours of expert human work.</p><p>The growth of ChatGPT Work and Codex gives some sense of how quickly more people are discovering this. OpenAI reported 6 million active users on July 12 and 25 million by the end of August, following the launch of its combined super app. These products bring users much closer to the way power users employ models as agents. Yet that audience is still tiny beside ChatGPT&#8217;s billion weekly users. There is still a huge gap in adoption even before the next generation arrives.</p><div><hr></div><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://newsletter.towardsai.net/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading Towards AI Newsletter! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><div><hr></div><h3>Why should you care?</h3><p>I see two main shifts we will need to make with this coming generation of LLMs. The first is moving from low-ambition chatbot tasks to assignments to always by default trying to frame a task that is perhaps 10 times more ambitious. A request to produce one component could become a request to complete the project to which that component belongs. The second is moving from repeated human iteration, editing, and feature requests throughout the work to concentrating human involvement at the beginning and end.</p><p>For software developers working heavily with agents, I think the change in day-to-day work since November 2025 may exceed the cumulative change over the previous 25 years. The next models are likely to make teams of agents working on long tasks much more capable and reliable, which will again completely transform how we work and much work we can hand over to AI at once.</p><p>I think there is a decent chance that within a few months we can, and often should, hand over the whole development process after we have put our own expertise into the scoping and planning stage. We still need to spend a lot of time up front deciding on the goal, guardrails, architecture, tools, systems, and features, then let a team of agents build and test the result. This would amount to commissioning the entire implementation in one go, with the agents handling the many rounds of execution, testing, and correction themselves.</p><p>The beginning would still involve several rounds of research and brainstorming with the LLM. We would use that dialogue to explore approaches, identify missing information, settle the constraints, and agree the final task or goal. The aim would then be to give the agents enough context, tools, checks, and tests to carry out their own iterations during the build. The models would need to sustain those loops for much longer, and we would need to learn how to set them up to do it.</p><p>At the end, we of course still need to review and understand the completed result and give critical feedback. Human taste, domain knowledge, and the ability to think of edge cases, the agents missed remain essential. We could then consolidate that feedback into a final agent run to make the corrections. The human would still help define and judge the work, while the agents handle far more of the development in between.</p><p>That is the new working pattern I expect the next wave of models to make much more practical.</p><p>This would again change where human expertise is most valuable. We would need to bring the domain expert, end user, software developer, and AI engineer into the initial task so that the agents understand what they are trying to achieve and how the result should work. Today, much of that knowledge enters through successive prompts as we inspect a partial result and request the next change. With more capable agent teams, it should become much more efficient to put that thinking into the goal and plan before a long run begins. Preparing and scoping a task could account for a much larger share of human work.</p><p>I expect the same shift in nontechnical work: slides, research reports, spreadsheets, brainstorming, marketing, and data analysis. An agent team could take a well-developed brief through research, competing approaches, analysis, drafting, and its own checks before returning a complete piece of work. Like software, end-user and domain expertise would need to shape that brief. The opportunity is to enable more people to commission substantial work that previously required a team, provided they can explain the goal and assess the result. That said, for now, I still expect a lot of work to be needed to manually clean up AI writing slop. The labs have made some style improvements, but they have a long way to go on fixing repetitive phrasing, generic judgments, and identifying the most interesting topic or insight to highlight. I expect this to become a much higher priority for the AI labs soon, as more users ask agents to produce complete work.</p><p>The direction of the labs&#8217; training investment should play a key role in guiding how you use these models. It appears clear from the Hugging Face incident that OpenAI and Anthropic are now putting most of the next generation&#8217;s GPU training budgets and researcher expertise into making teams of agents work for days. Therefore, we need to adapt our usage towards that capability. Continuing to use those models for small chatbot requests, or managing every incremental change ourselves, would use only a small fraction of what they can do. We would be wasting much of the billions of dollars of AI research and development investment we have access to. I expect individuals and companies that adapt their work to these new capabilities to outcompete those that stick to the previous model generation&#8217;s habits.</p><p><em>&#8212; <a href="http://www.linkedin.com/in/louie-peters">Louie Peters&#8202;&#8212;&#8202;Towards AI Co-founder and CEO</a></em></p><div><hr></div><h4>Live Workshop: Build Your Personal AI Engineering System</h4><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://www.eventbrite.co.uk/e/build-your-personal-ai-engineering-system-registration-1998640855601?aff=oddtdtcreator&amp;keep_tld=true" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!LW23!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F21e3c874-879f-4c10-8148-f0c35bb05cf5_1400x700.jpeg 424w, https://substackcdn.com/image/fetch/$s_!LW23!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F21e3c874-879f-4c10-8148-f0c35bb05cf5_1400x700.jpeg 848w, https://substackcdn.com/image/fetch/$s_!LW23!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F21e3c874-879f-4c10-8148-f0c35bb05cf5_1400x700.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!LW23!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F21e3c874-879f-4c10-8148-f0c35bb05cf5_1400x700.jpeg 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!LW23!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F21e3c874-879f-4c10-8148-f0c35bb05cf5_1400x700.jpeg" width="1400" height="700" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/21e3c874-879f-4c10-8148-f0c35bb05cf5_1400x700.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:700,&quot;width&quot;:1400,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:&quot;https://www.eventbrite.co.uk/e/build-your-personal-ai-engineering-system-registration-1998640855601?aff=oddtdtcreator&amp;keep_tld=true&quot;,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" title="" srcset="https://substackcdn.com/image/fetch/$s_!LW23!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F21e3c874-879f-4c10-8148-f0c35bb05cf5_1400x700.jpeg 424w, https://substackcdn.com/image/fetch/$s_!LW23!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F21e3c874-879f-4c10-8148-f0c35bb05cf5_1400x700.jpeg 848w, https://substackcdn.com/image/fetch/$s_!LW23!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F21e3c874-879f-4c10-8148-f0c35bb05cf5_1400x700.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!LW23!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F21e3c874-879f-4c10-8148-f0c35bb05cf5_1400x700.jpeg 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>Louis-Fran&#231;ois, our CTO, is leading a <a href="https://www.eventbrite.co.uk/e/build-your-personal-ai-engineering-system-registration-1998640855601?aff=oddtdtcreator&amp;keep_tld=true">90-minute live workshop with Packt</a> on how he structures his AI engineering setup around coding agents like Claude Code and Codex.</p><p>If you&#8217;re already using agents in your workflow, this should be especially useful. You&#8217;ll see how to:</p><ul><li><p>Keep <strong>skills, instructions, and knowledge synced across agents</strong></p></li><li><p>Turn repeated corrections into <strong>reusable skills</strong></p></li><li><p>Decide <strong>what stays in context and what gets retrieved</strong></p></li><li><p>Split work between <strong>Claude Code and Codex</strong></p></li><li><p>Work around <strong>usage and token limits</strong></p></li><li><p>Set up <strong>scheduled and always-on tasks</strong></p></li><li><p>Use an <strong>open-source vault template</strong> you can adapt</p></li></ul><p><strong>Who it&#8217;s for:</strong> AI engineers and developers already using coding agents who want a more structured system around them.</p><p><strong>When:</strong> Tuesday, September 8, 2026<br><strong>Time: </strong>8:30 PM-10 PM GMT+5<br><strong>Where:</strong> Online</p><p><strong><a href="https://www.eventbrite.co.uk/e/build-your-personal-ai-engineering-system-registration-1998640855601?aff=oddtdtcreator&amp;keep_tld=true">Register here</a></strong></p><div><hr></div><h3>Hottest News</h3><p>1. <a href="https://www.reuters.com/technology/nvidia-talks-acquire-hugging-face-13-billion-deal-business-insider-reports-2026-08-27/">Nvidia Has Agreed To Acquire Hugging Face for $12.9 Billion</a></p><p>According to The Information, NVIDIA has reportedly agreed to acquire Hugging Face for $12.9 billion, citing a person familiar with the deal. CNBC separately confirmed that an NVIDIA acquisition had been part of recent talks, while Business Insider reported that no signed agreement had been finalized and the discussions could still collapse. Hugging Face, one of the most widely used platforms for sharing AI models, datasets, and tools, had been working with a bank to assess takeover interest, with Microsoft among the companies that had previously shown interest. If completed, the deal would be NVIDIA&#8217;s largest completed acquisition. Hugging Face was last valued at $4.5 billion in its August 2023 Series D. As of publication, neither company has officially confirmed the deal.</p><p>2. <a href="https://blog.google/innovation-and-ai/technology/developers-tools/build-with-gemini-omni-1-1-flash/">Google AI Releases Gemini Omni 1.1 Flash</a></p><p>Google released Gemini Omni 1.1 Flash, a production-ready update to its video-generation and editing model with more control over longer scenes and transitions. For scene extension, Omni 1.1 can analyze up to 10 seconds of the existing clip before generating the next segment; previous versions referenced only the final second. Developers can add 10-second continuations until a video reaches a total of 40 seconds. The update also adds first- and last-frame interpolation, up to three seconds of reference video, 360p drafts at one-third the cost of 720p, and upscaled 1080p and 4K output. Omni 1.1 is available through Google AI Studio and the Gemini Enterprise Agent Platform, and globally in Flow for AI Plus, Pro, and Ultra subscribers. Google&#8217;s API documentation lists 1.1 as the stable GA model, while the previous preview endpoint will shut down on September 30.</p><p>3. <a href="https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-3-5-transcribe/">Google AI Releases Gemini 3.5 Transcribe</a></p><p>Google released Gemini 3.5 Transcribe with two separate APIs: one for recorded audio and a Live variant for real-time streaming with sub-second latency. They automatically detect 85+ languages, handle spoken self-corrections, remove filler words in Smart mode, and format transcripts automatically. Recorded transcription adds word-level timestamps and speaker attribution for up to three speakers (3+ is experimental). Artificial Analysis reports an average word error rate of 2.6% for non-streaming and 4.0% for streaming, while the time to final transcription improved by 70% over Chirp 3. The models are available through the Gemini API and Google&#8217;s enterprise tools, while consumers can already use the technology in the English Gemini macOS app and Rambler on Android in selected markets; Chrome support is coming soon.</p><p>4. <a href="https://qwen.ai/blog?id=qwen3.8-flash-next">Qwen Releases Qwen3.8-Flash-Next</a></p><p>Alibaba&#8217;s Qwen team released Qwen3.8-Flash-Next, an open-weight multimodal preview of architecture being developed for Qwen4. It combines a 125B-parameter main model with 51B N-gram embedding parameters while activating only 6B per token, using a hybrid Gated DeltaNet and Qwen Sparse Attention architecture. The model accepts text, images, and video, supports 262K native context extensible to 1M, and reasons by default. Qwen reports 62.5% on SWE-bench Pro, although its Claude Opus 4.6 comparison uses a separately published Claude score rather than the same evaluation run. Weights are available on Hugging Face and ModelScope under Qwen Community 1.0. The production Qwen3.8-Flash service adds 1M default context and built-in tools, with announced pricing of $0.15/$0.47 per million input/output tokens.</p><p>5. <a href="https://z.ai/blog/glm-5.3-flash">Z.ai Introduces GLM-5.3-Flash</a></p><p>Z.AI released GLM-5.3-Flash, the first natively multimodal model in the GLM-5 family, with 320B total parameters and 18B active per token. Unlike GLM-5.3&#8217;s post-training upgrade of the previous base, Flash starts from a newly trained base model using a 30T-token multimodal corpus and a redesigned hybrid sparse-and-linear-attention architecture. It accepts text, images, and video with a 1M-token context window. Z.AI reports DeepSWE improving to 63.4 from GLM-5.2&#8217;s 46.2 and AutomationBench to 48.8 from 26.2. The model previously ran anonymously as Ox Alpha on OpenRouter and OpenCode before Z.AI revealed its identity. MIT-licensed weights are available on Hugging Face; API list pricing is $0.15/$0.50 per million tokens, currently discounted 50% to $0.075/$0.25 through September 9.</p><p>6. <a href="https://openai.com/index/jalapeno-first-results/">OpenAI Publishes First Jalape&#241;o Inference-Chip Results</a></p><p>OpenAI published the first measured results for Jalape&#241;o, its Broadcom-designed custom inference chip, at Hot Chips on August 25. On SemiAnalysis&#8217;s public InferenceX benchmark across GPT-OSS 120B, DeepSeek R1 670B, and Kimi K2.5 1T, the 700W-TDP chip delivered 1.5&#8211;1.9x more work per watt at peak throughput and 1.7&#8211;3.6x lower end-to-end latency than the compared systems; measured sustained power stayed at or below 550W. OpenAI says AI-assisted design helped take the chip from initial design to tapeout in nine months, with Gen 2 already deep in development and Gen 3 taking shape. Deployment within OpenAI is planned to begin at a small scale by year-end. SemiAnalysis independently benchmarked Jalape&#241;o in OpenAI&#8217;s lab and found the results competitive, while cautioning that Blackwell is an imperfect comparison because Jalape&#241;o uses HBM4; NVIDIA&#8217;s HBM4-based Rubin is the more comparable generation.</p><p>7. <a href="https://openclaw.ai/blog/openclaw-2-accidentally">OpenClaw Releases OpenClaw 2.0</a></p><p>OpenClaw released its largest update yet, built by 933 contributors across more than 16,000 pull requests, roughly half of everything ever merged into the project. Version 2.0 focuses on reducing the friction of running a personal agent: onboarding can now detect existing ChatGPT or Claude subscriptions, API keys, and local models, while a rebuilt browser app opens directly into a conversation rather than requiring configuration. The bigger change is collaboration. Shared cloud sessions let another person join or take over ongoing agent work without losing its existing context, moving OpenClaw beyond a single-user assistant toward shared agent workflows. The release also adds searchable conversation history, persistent progress across long-running tasks, stronger credential handling, and reusable approvals for scheduled work. It arrives after nearly seven weeks without a normal release, following OpenClaw&#8217;s previous pace of 106 releases in 230 days.</p><div><hr></div><h3>AI Tip of the Day</h3><p>In a multi-turn conversation, the user&#8217;s latest message is written for the conversation, not for your retrieval system. It often depends on details mentioned several turns earlier.</p><p>A simple example we use in the <em>Context Engineering</em> lesson of our<a href="https://towardsai.com/academy/agent-engineering/?utm_source=newsletter&amp;utm_medium=email&amp;utm_id=AItips"> Agent Engineering course</a> shows why it is important to know this. A user reports a mild headache and wants to avoid medication. A few turns later, they ask: <strong>&#8220;Could stress be causing this?&#8221;</strong></p><p>If you retrieve using only that question, the system no longer knows what &#8220;this&#8221; refers to. It also reduces the severity of the headache and the user&#8217;s preference for avoiding medication.</p><p>Instead, maintain a small session state with the important facts from the conversation. Keep the last two or three raw messages as well, then build the retrieval query from that state plus the newest message.</p><p>To check whether this actually improves retrieval, test it on 20 follow-up questions. Run each one twice: once with only the latest message, and once with the session state. Track:</p><ul><li><p>Whether a relevant source appears in the top five</p></li><li><p>Whether any important user constraint is missing from the query</p></li></ul><p>This gives retrieval the context it needs without sending the entire conversation back to the model every time.</p><div><hr></div><h3>Five 5-minute reads/videos to keep you learning</h3><p>1. <a href="https://pub.towardsai.net/we-stopped-compacting-our-agents-context-f1f6282a5715?sk=45070653f20cb1220c5592ec0479f099">We Stopped Compacting Our Agent&#8217;s Context</a></p><p>Context compaction should make long-running agents cheaper and faster. In our production AI tutor, it did the opposite. Keeping the full conversation history beat every compaction strategy we tested in terms of cost, latency, and memory recall because summarization disrupted prompt caching and forced the model to reprocess context. We first shared these results at the AI Engineer World&#8217;s Fair; this article covers the full experiment, including findings we could not fit into the talk and several that emerged afterward.</p><p>2. <a href="https://pub.towardsai.net/the-economics-of-agents-token-accounting-caching-and-routing-a53cbdef11bb?sk=1bdac15a6065039499889af4534148f5">The Economics of Agents: Token Accounting, Caching, and Routing</a></p><p>Agent costs can grow quickly, but much of that spend comes from system design rather than the agent itself. Stateless APIs repeatedly resend the conversation history, making long-running agents heavily input-token-dependent. Using real invoice breakdowns, this article shows how prompt caching, model routing, cascades, and explicit budgets can cut that cost by 60&#8211;90%, and where each technique actually makes a difference.</p><p>3. <a href="https://pub.towardsai.net/deep-dive-from-dynamic-programming-to-monte-carlo-sampling-or-how-to-learn-without-a-model-7befb656f896?sharedUserId=tai-tech">Reinforcement Learning: From Dynamic Programming to Monte Carlo Sampling or How to Learn Without a Model</a></p><p>Dynamic programming works when you know how the environment behaves. Monte Carlo methods remove that requirement and learn state values directly from experience. This article builds first-visit and every-visit Monte Carlo prediction from sampled returns, then extends the same idea to control by learning action values and deriving a policy. A robotic navigation experiment compares the resulting agent against random behavior and value iteration, making the tradeoff between model-based and model-free learning concrete.</p><p>4. <a href="https://pub.towardsai.net/quantization-is-four-decisions-not-one-c52a0b3f296e?sharedUserId=tai-tech">Quantization Is Four Decisions, Not One</a></p><p>&#8220;4-bit quantization&#8221; tells you very little about what was actually compressed or what quality you gave up. This article separates quantization into four decisions: which operands to quantize, which numerical format to use, how to interpret accuracy losses, and how much additional model capacity the compression buys you. It also traces a widely reported INT4 code-generation failure back to the original results, finds that the apparent collapse was due to a single flipped test case, and then compares newer formats, including NVFP4 and MXFP4.</p><p>5. <a href="https://pub.towardsai.net/watermarking-an-inference-engine-018e1a83717b?sk=61342facf70e34aa1e37a409825446af">Watermarking an Inference Engine</a></p><p>What happens when you implement LLM watermarking in a real inference engine rather than evaluating it on toy vocabularies? This article adds green-list biasing and Aaronson&#8217;s distortion-free Gumbel watermark to DS4, a C inference engine for DeepSeek V4 Flash, then benchmarks both against its 129,280-token vocabulary. The experiments show that the commonly used bias strength is too weak at this scale and that scanning the full logit vector, not hashing, drives most of the cost. The distortion-free method was cheaper and stronger per token, but a single paraphrase removed its watermark while the biased approach survived.</p><h3>Repositories &amp; Tools</h3><p>1. <a href="https://github.com/tt-a1i/archify">Archify</a> is an agent skill for turning codebases or system descriptions into interactive architectures, workflows, sequences, data-flows, or lifecycle diagrams.</p><p>2. <a href="https://github.com/THU-MAIC/OpenMAIC">OpenMAIC</a> is a multi-agent platform that turns any topic or document into a full classroom with slides, quizzes, whiteboard, etc.</p><p>3. <a href="https://github.com/p-e-w/heretic">Heretic</a> is a fully automatic tool for removing safety alignment from open-weight language models using Optuna-optimized directional ablation.</p><p>4. <a href="https://github.com/vercel-labs/vgpu">Vgpu</a> is a TypeScript WebGPU library with one API surface across browser, headless Node.js, and CI, which works with both human developers and AI coding agents.</p><p>5. <a href="https://github.com/google-research/envharness">EnvHarness</a> wraps a frozen environment with its own plug-in components to make it dynamically controllable, without touching the environment&#8217;s internal code.</p><h3>Top Papers of The Week</h3><p>1. <a href="https://arxiv.org/abs/2608.28476">ContextPilot: Teaching Agents for Proactive Context Management via Fine-grained RL</a></p><p>Long-horizon agents accumulate context that grows monotonically, eventually degrading performance. ContextPilot teaches agents to proactively manage their own context through planning, long-term memory, and soft context offloading tools. The RL training method is tailored for context management: context-aware partial rollout identifies critical context-editing decisions for branch sampling, and snapshot-level credit assignment estimates action advantages from all downstream branches that pass through each editing action.</p><p>2. <a href="https://arxiv.org/abs/2608.27550">Beyond Data Scaling: Representation-Centric Continued Pre-training for Vision-Language-Action Models</a></p><p>Robot trajectories are fundamentally harder to scale than web-scale data, making representation quality the central bottleneck for generalist VLA models under fixed data budgets. VLAct is a representation-centric continued pre-training recipe that starts from a pretrained VLM and trains on broad, multi-embodiment robot data before task-specific fine-tuning. It preserves the VLM prior and encourages shared action semantics across embodiments through multi-head continuous action co-supervision and a partially unified cross-embodiment action layout. Using fully open-source data and a 16-GPU setup, VLAct consistently improves downstream performance across simulation benchmarks, real-world experiments, and unseen-embodiment transfer.</p><p>3. <a href="https://arxiv.org/abs/2608.27763">Fast Weight Attention for Continual Learning</a></p><p>Recurrent fast-weight memories and selective state-space models compress a growing context into a bounded state, making the state transition an online learning rule. This paper studies that rule under read-after-write autoregressive semantics and derives the Falcon family: normalized first-order updates for squared-error regression and negative inner-product objectives, with recurrent, masked-parallel, and chunk-parallel forms. The framework cleanly separates temporal alignment, plasticity, forgetting, and bounded rehearsal as independent design axes.</p><p>4. <a href="https://arxiv.org/abs/2608.26582">J-Zero: Unified Challenger &#8212; Solver &#8212; Judge Co-Evolution from Zero Data</a></p><p>Self-evolution in unverifiable domains is limited because a frozen judge can only push the solver toward preferences it has already internalized. J-Zero co-evolves all three roles from a single base LLM with no external data. The Challenger and Solver co-evolve through an adversarial game: the Challenger generates increasingly difficult tasks while the Solver learns to produce higher-quality responses. The Judge co-adapts using preference pairs whose ordering is determined by how each response was produced, rather than by the judge&#8217;s own scores. J-Zero outperforms baselines by an average of 4.2 points on verifiable domains and 8.0 points on unverifiable domains.</p><p>5. <a href="https://arxiv.org/abs/2608.28122">Agentic Artifact Creation: Systems, Evaluation, Principles, and Opportunities</a></p><p>This survey formalizes agentic artifact creation as stateful construction, in which an AI system builds or revises a deliverable, and intermediate observations guide later work. It reviewed 259 works through August 2026 (230 systems, 29 benchmarks) across six artifact families. Key findings: construction challenges depend on how tightly decisions are coupled and whether failures become visible while repairable, not just on modality. Decomposition reduces local complexity but increases coordination costs. Learned judges add little independent evidence when they share the generator&#8217;s preferences or blind spots.</p><h3>Quick Links</h3><p><span>1. </span><a href="https://www.liquid.ai/blog/pipette-on-device-ai-benchmarking-by-liquid-ai">Liquid AI open sources Pipette</a><span>, an open-source platform for benchmarking foundation models on edge devices. It evaluates on-device AI as a property of the full deployed configuration (model + quantization + runtime + device), not the model in isolation. Built in partnership with Artificial Analysis, the launch dataset covers five performance metrics across 1,000+ configurations spanning 30+ models, llama.cpp builds for macOS, iOS, Windows, and Android, and context lengths from 256 to 8,192 tokens. Initial verified results come from MacBook Pro M5 Max, iPhone 17 Pro, and Galaxy S26 Ultra. Open-source benchmark clients for macOS, Windows, iOS, and Android.</span></p><h3>Who&#8217;s Hiring in AI</h3><p><strong><a href="https://towardsai.com/valuecreation/senior-ai-engineer/">Senior AI Engineer/Forward Deployed Engineer @Towards AI (London/Hybrid)</a></strong></p><p><strong><a href="https://jobs.towardsai.net/job/microsoft-corporation-technical-support-engineer-azure-ai-qyjm">Technical Support Engineer &#8212; Azure AI @Microsoft Corporation (Bucharest, Romania)</a></strong></p><p><strong><a href="https://jobs.towardsai.net/job/meta-business-support-engineer-3eh2">Business Support Engineer @Meta (Dublin, Ireland)</a></strong></p><p><strong><a href="https://jobs.towardsai.net/job/velo3d-senior-machine-learning-engineer-mqqc">Senior Machine Learning Engineer @Velo3D (Fremont, CA, USA)</a></strong></p><p><strong><a href="https://jobs.towardsai.net/job/nielseniq-senior-software-engineer-web-scraping-gen-ai-developer-5x9b">Senior Software Engineer @NielsenIQ (Pune, India)</a></strong></p><p><strong><a href="https://jobs.towardsai.net/job/synchrony-vp-ai-enablement-pfwn">VP, AI Enablement @Synchrony (Stamford, CT, USA)</a></strong></p><p><strong><a href="https://jobs.towardsai.net/job/applied-materials-ai-materials-research-engineer-5mff">AI Materials Research Engineer @Applied Materials (Santa Clara, CA, USA)</a></strong></p><p><em>Interested in sharing a job opportunity here? Contact <a href="mailto:sponsors@towardsai.net">sponsors@towardsai.net</a>.</em></p><p><em>Think a friend would enjoy this too? <a href="https://newsletter.towardsai.net/">Share the newsletter and let them join the conversation.</a></em></p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://newsletter.towardsai.net/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading Towards AI Newsletter! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div>]]></content:encoded></item><item><title><![CDATA[TAI #219: AI, Cancer and the Future of Personalized Medicine ]]></title><description><![CDATA[Also, Nvidia&#8217;s AVO Agent, GPT Sol 20% price cut, Deepseek&#8217;s multimodal V4-flash and more]]></description><link>https://newsletter.towardsai.net/p/tai-219-ai-cancer-and-the-future</link><guid isPermaLink="false">https://newsletter.towardsai.net/p/tai-219-ai-cancer-and-the-future</guid><dc:creator><![CDATA[Louie Peters]]></dc:creator><pubDate>Wed, 26 Aug 2026 19:44:38 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!2qkk!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa3de7e7d-cfb0-4703-bcfa-b4144d896172_1672x941.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<h2>What happened this week in AI by Louie</h2><p><span>On August 15, Dario Amodei argued that AI companies will earn trust by curing diseases, rather than promising they might. Anthropic is increasing its work in biology and medicine, and Amodei believes most human disease could be cured within five to ten years. That is an ambitious forecast, but the past week offered a clear picture of why the idea deserves serious attention.</span></p><p><span>Anthropic showed AI agents coordinating the design of new proteins and having their work validated in real laboratories. Moderna and Merck reported that a personalized cancer therapy designed with machine learning succeeded in a large Phase 3 trial. Other researchers demonstrated faster automated protein testing, engineered cancer-targeting immune cells, and new systems for simulating how cells respond to treatment. These developments cover different parts of the same problem: specialized models identify targets and design molecules, LLM agents coordinate research, automated laboratories test the results, and clinical trials establish whether treatments help patients. What is changing is how quickly those parts are starting to connect.</span></p><p><span>The clearest patient result came from Moderna and Merck. Their INTerpath-001 trial enrolled 1,137 people whose high-risk melanoma had been surgically removed. Some received Keytruda, an established immunotherapy. Others received Keytruda plus a personalized messenger RNA treatment called intismeran autogene. The combination improved recurrence-free survival and reduced the risk of the cancer spreading to distant sites. This therapeutic cancer vaccine is designed to stop an existing cancer from returning, with each treatment built around the mutations in one patient&#8217;s tumor.</span></p><p><span>The process starts by sequencing the tumor and the patient&#8217;s healthy cells. Comparing the two reveals mutations unique to the cancer. Researchers then analyze tumor RNA to determine which mutated genes are active. Some produce altered protein fragments, called neoantigens, that can help the immune system recognize cancer cells. Choosing the right fragments requires accounting for which mutations the tumor expresses, which fragments the patient&#8217;s cells can display, and which are most likely to trigger an immune response. Moderna&#8217;s system scores those possibilities, selects up to 34 targets, and combines them into a single personalized mRNA sequence.</span></p><p><span>That sequence is packaged into a lipid nanoparticle and injected. The patient&#8217;s cells produce the selected fragments, training immune cells to recognize cancer cells carrying the same mutations. Keytruda helps sustain that immune response. Targeting dozens of mutations also reduces the chance that surviving cancer cells can escape by losing one target.</span></p><p><span>Moderna says the design process runs from raw sequencing data to a patient-specific treatment without manual intervention. Another software system coordinates manufacturing, quality control, shipping, and scheduling. The company began applying machine learning to mRNA design in 2014 and built its individualized-treatment algorithm in 2016, giving it a decade of purpose-built scientific AI. One clarification: the widely shared 49% reduction in recurrence or death came from an earlier 157-person trial. The new Phase 3 study met its endpoints, but its full results have not yet been published.</span></p><p><span>Anthropic&#8217;s research points to a different role for AI. Its models received an expert-written research protocol, access to scientific literature, and specialist protein-design tools. They then selected molecular targets, proposed protein structures, checked which designs should fold properly, and decided which candidates to send for laboratory testing. Of 1,320 designs that produced usable measurements, 354 bound to their intended targets, a 26.8% success rate. The agents produced at least one successful protein for 14 of 15 targets. Human scientists provided the protocol and laboratory validation, while the models coordinated a complicated scientific workflow with less hands-on research time.</span></p><p><span>A separate study showed how these designs can move toward treatments. Researchers used AI-designed proteins to engineer immune cells that recognize cancer-associated targets. Several early designs needed further adjustment before they worked properly inside living cells, but the revised cells went on to kill tumor cells selectively in laboratory experiments.</span></p><p><span>A blinded antibody benchmark also exposed a weakness in molecular design: AI systems often struggled to identify their strongest candidates. Their selections beat a comparison antibody only 9.8% to 13.8% of the time, while randomly chosen candidates did so 39% of the time. That makes automated experimentation more valuable. If a model can propose hundreds of plausible molecules but cannot reliably pick the best one, researchers can test more candidates, learn from the results, and improve the next round. David Baker&#8217;s group described a semi-automated system that can produce and characterize hundreds of proteins per day while cutting gene-synthesis costs roughly fivefold.</span></p><p><span>A strong example of where this is heading comes from OpenAI and Ginkgo Bioworks. They connected GPT-5 to Ginkgo&#8217;s robotic cloud laboratory, where software can order experiments, machines execute them, and results return to the model. GPT-5 designed experiments to improve cell-free protein production, analyzed the measurements, and proposed the next batch. Across six rounds, the system tested more than 36,000 experimental conditions on 580 automated plates and cut protein-production costs by 40%. Each round supplied the evidence GPT-5 needed to form better hypotheses and improve its next experiments.</span></p><p><span>This is just one part of a growing physical research infrastructure. Emerald Cloud Lab offers remotely controlled laboratory experiments. Recursion says its automated labs can process up to 2.2 million samples a week. Insitro combines more than 20 petabytes of automated cellular experiments with human genetic data. Vivodyne uses robotic systems to grow human tissues and test how they respond to drugs and genetic changes.</span></p><p><span>Virtual-cell models add another layer. GenBio&#8217;s AIDO Cell can simulate how cells respond to combinations of drugs or genetic interventions, while the Arc Institute&#8217;s new Virtual Cell Challenge asks models to predict how unfamiliar cells will respond when particular genes are suppressed. Better simulations can narrow the search; automated laboratories can test whether those predictions hold up.</span></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!2qkk!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa3de7e7d-cfb0-4703-bcfa-b4144d896172_1672x941.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!2qkk!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa3de7e7d-cfb0-4703-bcfa-b4144d896172_1672x941.png 424w, https://substackcdn.com/image/fetch/$s_!2qkk!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa3de7e7d-cfb0-4703-bcfa-b4144d896172_1672x941.png 848w, https://substackcdn.com/image/fetch/$s_!2qkk!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa3de7e7d-cfb0-4703-bcfa-b4144d896172_1672x941.png 1272w, https://substackcdn.com/image/fetch/$s_!2qkk!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa3de7e7d-cfb0-4703-bcfa-b4144d896172_1672x941.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!2qkk!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa3de7e7d-cfb0-4703-bcfa-b4144d896172_1672x941.png" width="1456" height="819" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/a3de7e7d-cfb0-4703-bcfa-b4144d896172_1672x941.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:819,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1951364,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://newsletter.towardsai.net/i/212696008?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa3de7e7d-cfb0-4703-bcfa-b4144d896172_1672x941.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!2qkk!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa3de7e7d-cfb0-4703-bcfa-b4144d896172_1672x941.png 424w, https://substackcdn.com/image/fetch/$s_!2qkk!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa3de7e7d-cfb0-4703-bcfa-b4144d896172_1672x941.png 848w, https://substackcdn.com/image/fetch/$s_!2qkk!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa3de7e7d-cfb0-4703-bcfa-b4144d896172_1672x941.png 1272w, https://substackcdn.com/image/fetch/$s_!2qkk!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa3de7e7d-cfb0-4703-bcfa-b4144d896172_1672x941.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p></p><div><hr></div><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://newsletter.towardsai.net/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading Towards AI Newsletter! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><div><hr></div><h3>Why should you care?</h3><p><span>Dario is hoping AI cures are a key route to making AI more popular.  I think AI medicine could also become one of the main ways AI gets smarter. Scientific research offers frontier AI labs unusually valuable training signals: outcomes verified by the physical world. A protein binds to its target or it does not. A treatment changes a cell&#8217;s behavior or it does not. Those measurable results can drive reinforcement learning. An AI agent can study scientific papers, propose experiments, assess the results, revise its hypotheses, and try again. Each experiment creates fresh evidence about cause and effect, improving the model&#8217;s reasoning while producing useful medical discoveries.</span></p><p><span>I expect leading AI companies to spend hundreds of millions of dollars training agents this way - designing drugs and treatments becomes part of model training costs. The commercial incentive is strong: better scientific agents improve general reasoning while generating proprietary biological data, valuable intellectual property, and potential treatments.</span></p><p><span>But there is a big gap between the feedback an AI agent can obtain quickly and the evidence needed to prove a treatment works in patients. Automated laboratories can measure protein binding, immune responses, or changes in cell behavior within hours. Simulations can run even faster. Neither can tell us whether a patient&#8217;s cancer will return four years after treatment or whether serious side effects will emerge over time. Answering those questions requires clinical trials, long-term monitoring, and regulatory review. Until a treatment is approved and reaches more patients, the clinical data needed to improve future versions also accumulates slowly.</span></p><p><span>That delay will remain a real constraint, but better experiments and predictive models can still make clinical development faster and more productive. Stronger laboratory evidence, clearer biological mechanisms, and better patient selection should help more companies move promising treatments straight into combined Phase 2/3 trial. Regulation will need to keep pace with faster discovery and more personalized therapies while maintaining high standards for long-term safety and effectiveness.</span></p><p><span>Biology is also generating far more information than human researchers can interpret alone. Whole-genome sequencing, tumor DNA in blood, protein measurements, medical imaging, and patient histories each capture part of the same disease. AI can combine those signals, identify patterns across large populations, and apply them to an individual patient. Earlier cancer detection is one of the biggest opportunities. In GRAIL&#8217;s 35,878-person PATHFINDER 2 study, adding a blood test to standard screening increased cancer detection 6.5-fold, and 71% of newly detected cancers were at stages I to III. Better blood tests, sequencing, and protein analysis will make more cancers visible before they spread to distant organs.</span></p><p><span>That shift changes the incentives for treatment. A cancer found early has fewer malignant cells and fewer opportunities to develop resistance. Surgery, personalized vaccines, or targeted immune treatments have a better chance of eliminating it. Earlier detection creates a larger market for early intervention. Better early treatments make screening more valuable. Together, they mean more cures and fewer years managing advanced disease.</span></p><p><span>The treatment platforms are also becoming more adaptable. Moderna&#8217;s mRNA approach can use the same delivery system while changing the genetic instructions for each patient. CRISPR offers a related form of flexibility: the editing machinery can be directed toward a particular genetic mutation. Last year, researchers at Children&#8217;s Hospital of Philadelphia treated an infant with a customized CRISPR therapy designed for his rare genetic disease. The obstacle has been the scientific labor needed to customize these treatments safely and cost efficiently, since each patient may require different molecular targets, treatment instructions, and tests. AI agents and automated laboratories can supply the research capacity, while better software and manufacturing systems improve quality control and patient safety.</span></p><p><span>Amodei&#8217;s timeline is bold. His broader direction looks right to me. AI can help us understand disease earlier, design treatments around individual patients, test those treatments more quickly, and learn from every result. When the best AI companies have a direct incentive to train increasingly capable agents on real scientific discovery, progress in medicine could arrive much faster than most people expect.</span></p><p><em>&#8212; <a href="http://www.linkedin.com/in/louie-peters">Louie Peters&#8202;&#8212;&#8202;Towards AI Co-founder and CEO</a></em></p><div><hr></div><h3>Hottest News</h3><p>1. <a href="https://developer.nvidia.com/blog/nvidia-avo-reaches-100-on-arc-agi-3-demonstrating-a-frontier-level-general-purpose-architecture-for-long-horizon-autonomous-agents/">NVIDIA&#8217;s AVO Agent Hits a Perfect Score on ARC-AGI-3 Without a New Model</a></p><p>NVIDIA&#8217;s AVO agent completed all 183 levels across ARC-AGI-3&#8217;s 25 public environments using Claude Opus 5, earning a 100.00 Relative Human Action Efficiency score. AVO adds persistent memory, a supervisor that redirects stalled trajectories, and its own execution loop around the model. It completed the set in 6,624 environment actions, about 12% fewer than VISTA&#8217;s reported Opus 5 run, although NVIDIA stresses that the systems differ enough that this is not a controlled comparison. The same AVO architecture previously ran autonomously for seven days optimizing GPU kernels, exploring more than 500 directions and producing multihead-attention kernels up to 10.5% faster than FlashAttention-4 on NVIDIA&#8217;s tested DGX B200 configurations. The 100.00 result covers the public set only, not ARC-AGI-3&#8217;s semi-private or private competition sets.</p><p>2. <a href="https://x.com/OpenAI/status/2090885187634905500">OpenAI Cuts Frontier API Pricing by More Than 20% for Three Months</a></p><p>OpenAI temporarily cut GPT-5.6 Sol&#8217;s standard short-context API price from $5 to $4 per million input tokens and from $30 to $20 per million output tokens, reductions of 20% and 33% respectively. Cached input now costs $0.40 per million tokens, while prompts over 272K tokens are priced at $8 input and $30 output. The promotional rates are guaranteed at least through November 21, 2026, with no new model ID required. The reduction is also rolling into eligible ChatGPT Work and Codex credit pricing, while Plus, Pro, and Business subscription usage remains unchanged. Batch and Flex processing are cheaper again at $2/$10, while Fast mode remains the premium option at $8/$40. This follows July&#8217;s much larger price cut for Luna and a 20% reduction for Terra, leaving OpenAI&#8217;s full GPT-5.6 family substantially cheaper than at launch.</p><p>3. <a href="https://www.anthropic.com/research/Claude-accelerates-protein-design">Claude Designed Protein Binders That Worked in the Lab for 14 of 15 Targets</a></p><p>Anthropic tested whether Claude could autonomously run the computational side of a protein-design campaign, and external labs found working binders across 14 of the 15 targets tested. Using Mythos Preview and Opus 4.8, Claude chose binding sites, orchestrated specialist protein-design and folding models, optimized candidates, and screened them before Adaptyv Bio and Twist Bioscience produced and tested the designs. Across all experimental arms, 354 of 1,320 designs bound successfully. In the 48-hour multi-target setup, Opus 4.8 achieved a 22.6% hit rate and Mythos 26.7%; giving Mythos a separate 24-hour run for each target raised its hit rate to 35.1%, versus the 10&#8211;15% Anthropic says is typical today. At least six targets produced high-affinity binders, and four matched or exceeded the best previously reported affinity. The campaign still relied on specialist protein models, substantial GPU compute, and weeks of wet-lab validation, so Claude&#8217;s role was to autonomously orchestrate the computational design process rather than replace the experimental work.</p><p>4. <a href="https://x.com/deepseek_ai/status/2090730032574631962?s=20">DeepSeek Unveiled an Experimental Multimodal Model</a></p><p>DeepSeek released DeepSeek-V4-Flash-Vision-Exp, an experimental multimodal version of V4-Flash that adds image understanding while retaining comparable text, reasoning, and agent capabilities. On DeepSeek&#8217;s own evaluations, it scores 36.5 on multimodal ApexBench versus 26.2 for text-only V4-Flash, 27.3 on Agents&#8217; Last Exam, 64.3 on Chartography, and 35.0 on ZeroBench, putting its reported multimodal agent performance close to Claude Opus 4.8. The model also improves several text-agent scores, including DeepSWE from 54.4 to 59.3, although these results were run through DeepSeek&#8217;s own Harness configuration and have not yet been independently reproduced. Images cost no pricing premium: each is billed as at most 384 input tokens at regular V4-Flash rates. The API accepts mixed text and images through Chat Completions, Anthropic-compatible Messages, and Responses, while a new free Files API lets developers upload an image once and reuse it across requests. DeepSeek Harness 0.1.1 shipped alongside it with native support.</p><p>5. <a href="https://cursor.com/changelog/origin-code-hosting">Cursor Launches Origin</a></p><p>Cursor launched Origin, moving beyond the editor to host repositories and pull requests directly inside Cursor. Developers can create Cursor-hosted repos or sync existing GitHub repositories in real time; GitHub remains the source of truth for synced projects, with pull-request comments, reviews, and merges reflected in both systems. The larger shift is bringing agents into the same surface as the repository: while browsing code or a PR, users can ask Cursor questions, have it make changes, update the PR, or push a branch. Origin also launches with integrations for Vercel, Depot, and Buildkite for preview deployments and CI, including support for existing GitHub Actions workflows. It is rolling out in early beta across paid plans, except enterprise organizations whose admins opt out, with Cursor saying more agent-native hosting features are still to come.</p><p>6. <a href="https://claude.com/blog/bringing-claude-mythos-5-to-more-defenders">Anthropic Brings Claude Mythos 5 to Claude Security</a></p><p>Anthropic is widening access to Mythos 5&#8217;s cybersecurity capabilities without giving users direct access to the model. Claude Security, currently in public beta for Enterprise customers, can now use Mythos 5 to scan repositories for vulnerabilities and return each finding with a CWE category, confidence, and severity ratings, and a suggested fix; users can then open Claude Code to implement the patch, with human approval still required. Mythos itself remains isolated within the scan, so access through Claude Security does not allow users to prompt it for offensive tasks elsewhere. Anthropic is taking the same approach with cybersecurity partners, embedding Mythos behind purpose-built tools that expose specific outputs rather than the raw model. It also announced $35 million in Claude credits for open-source security through a new Defender Advantage Fund and plans to expand its vetted Cyber Verification Program, with broader Opus and Sonnet capabilities first and Mythos-level access to follow.</p><p>7. <a href="https://www.linkedin.com/redir/suspicious-page?url=https%3A%2F%2Fornith%2eai%2Fornith_1_5%2ehtml">Ornith-1.5 Ships a Self-Improving Open Model Family</a></p><p>Ornith released three MIT-licensed Ornith-1.5 models at 397B MoE, 35B MoE with 3B active parameters, and 9B dense scales, extending its earlier &#8220;self-scaffolding&#8221; approach into a broader self-improvement training loop. Instead of training only on fixed human-created tasks and harnesses, the system proposes progressively harder tasks based on what the model has already solved, generates the tools and decomposition strategy needed to tackle them, produces solution rollouts, and uses reinforcement learning to improve all three stages together. Ornith reports its 397B model reaching 86.1 on Terminal-Bench 2.1 and 56.0 on DeepSWE, around Claude Opus 4.8&#8217;s 85.0 and 59.0 in its evaluation setup. The 35B model achieves 79.0 on SWE-bench Verified despite activating only 3B parameters per token, while the 9B model achieves 70.6 and is also available in quantized mobile formats. The &#8220;self-improving&#8221; label refers to this training process, not a deployed model continuously changing its own weights. Weights, including GGUF, FP8, NVFP4, and MLX variants, are available on Hugging Face.</p><div><hr></div><h3>AI Tip of the Day</h3><p>A student in our <a href="https://towardsai.com/academy/llm-primer/?utm_source=newsletter&amp;utm_medium=email&amp;utm_id=AItip">10-Hour LLM Fundamentals Video Course</a> asked us a useful question: How many times should you run an important test before trusting the result?</p><p>For stochastic LLM tests, our practical baseline is five runs.</p><p>Suppose a difficult test passes 90% of the time. If you run it once and it passes, you might conclude that everything is working. Run the same test five times, though, and there is about a 41% chance you will see at least one failure.</p><p>For important tests, run the same case five times and record two things:</p><ul><li><p>The pass rate across all runs</p></li><li><p>Whether it passed all five times</p></li></ul><p>Keep the prompt, model, temperature, and other randomness settings, as well as the source context, fixed. Otherwise, you are changing the test while trying to measure its consistency.</p><p>Five runs is not a statistical guarantee. It is a simple way to catch intermittent failures that a single successful run can easily hide.</p><div><hr></div><h3>Five 5-minute reads/videos to keep you learning</h3><p>1. <a href="https://pub.towardsai.net/your-llm-has-a-million-token-memory-heres-why-that-s-still-not-enough-acdf388abbd6?sk=d98670afc1e6bf084268b282d93c7b24">Is More Context Making Your AI Agent Worse? Here&#8217;s the Fix Nobody Talks About</a></p><p>The article explains why larger context windows do not guarantee stronger agent performance by breaking the LLM call into six parts: user messages, system prompts, tool schemas, retrieved resources, assistant history, and tool call records. It notes that maxing out even a million-token window rarely helps, since results often peak near 65 percent utilization and shares four fixes: sharpening system prompts, writing precise tool schemas, retrieving resources selectively, and steering long-running, iterative agents through compaction, memory stores, and specialized, targeted sub-agents.</p><p>2. <a href="https://pub.towardsai.net/three-core-ideas-that-make-understanding-attention-in-transformers-easy-fd701032c82e?sk=81bad0b32be196425066890144a6ffe2">3 Core Ideas That Make Attention In Transformers Easy to Understand</a></p><p>This piece breaks scaled dot-product attention into three core ideas: embeddings as vectors transformed through matrix multiplication, weighted averaging on the query-key-value database analogy, and learned projection matrices that discover contextual embeddings during training. It walks through computing attention weights and outputs step by step, closing with a clear explanation of why query vectors function as questions one token asks about its neighbors.</p><p>3. <a href="https://pub.towardsai.net/context-engineering-the-discipline-that-quietly-replaced-prompt-engineering-da39172dbe15?sk=ddfb4b938566b5e585596363c576f50a">Context Engineering: The Discipline That Quietly Replaced Prompt Engineering</a></p><p>The piece explains the four failure modes of context engineering: poisoning, distraction, confusion, and clash, and explains the KV-cache economics behind stable prefixes and append-only history. It also maps six production techniques (curation, compaction, external memory through files, recitation, just-in-time retrieval, and sub-agent isolation) using Claude Code and Manus as practical examples.</p><p>4. <a href="https://pub.towardsai.net/ai-agent-memory-architecture-beyond-context-windows-89e5eaef9e49?sk=27ecb76a06fe67788b188fa2cdbda655">Stop Using Long-Context Windows for AI Agents (Build This Instead)</a></p><p>Long-context windows cost 15 times more than persistent memory retrieval, and this article argues for replacing brute-force context with a four-tier memory architecture: working, episodic, semantic, and procedural. It compares Mem0, Letta, Zep/Graphiti, and AWS AgentCore across storage, retrieval, and injection pipelines, then tackles decay policies using Ebbinghaus-style forgetting curves. It also flags production risks often skipped elsewhere: memory poisoning, drift, and hallucinated facts.</p><p>5. <a href="https://pub.towardsai.net/the-engineers-guide-to-claude-cowork-avoid-sandbox-escapes-73f4d580281d?sk=f195b46857f7550ec144f96ec504504b">Claude Cowork Is for Engineers Too&#8202;&#8212;&#8202;Here&#8217;s How to Wire It Safely</a></p><p>This article draws a sharp line between Claude Cowork and Claude Code and argues that engineers dismissed Cowork by mistaking office-worker marketing for a limitation of capability. Cowork integrates with Slack, Jira, Drive, and Confluence via skills and connectors, handling spec drafting and incident triage that a terminal agent cannot touch. It also details failures, including permanent file deletions and a VM sandbox escape, which prompted Anthropic to default to cloud sessions.</p><h3>Repositories &amp; Tools</h3><p>1. <a href="https://github.com/affaan-m/ECC">ECC</a> adds a structured engineering workflow to coding agents, with planning, testing, review, persistent memory, reusable skills, and security scanning for Claude Code, Codex, and other agent harnesses.</p><p>2. <a href="https://github.com/tinyhumansai/openhuman">OpenHuman</a> is a local-first personal AI system that builds persistent memory from your data, orchestrates durable multi-agent workflows, and researches across your connected sources and the web.</p><p>3. <a href="https://github.com/FlashML-org/FreeToken">FreeToken</a> is an edge-native MoE serving engine that combines GPU, CPU, and system memory to run frontier-scale open-weight models, including 290B+ models, on consumer hardware.</p><p>4. <a href="https://www.linkedin.com/redir/suspicious-page?url=https%3A%2F%2Fis-agentic%2ecom%2F">Is Agentic</a> is a website agent-readiness scanner that scores how easily AI agents can discover, retrieve, understand, and interact with a site&#8217;s public content and interfaces.</p><p>5. <a href="https://github.com/NVIDIA/TensorRT-Model-Connect">TensorRT-Model-Connect</a> builds TensorRT engines directly from supported Hugging Face or local checkpoints, packages them into portable bundles, and exposes them through native C++ APIs.</p><h3>Top Papers of The Week</h3><p>1. <a href="https://arxiv.org/abs/2608.16157">FreeToken: Efficient Edge-Native MoE Serving with Bandwidth-Adaptive Execution</a></p><p>Frontier open-weight MoE models assume datacenter infrastructure, but FreeToken treats a personal machine as a unified, elastic inference platform. Rather than committing to a fixed offloading strategy, it continuously maps computation and model state onto whatever resources are actually available, co-designing model layout, expert residency, CPU-GPU execution, and memory management around each machine&#8217;s specific hardware balance. It scales from a 35B model on a laptop with an 8GB GPU to the 753B GLM-5.2 on a single workstation GPU, delivering 1.5&#8211;2.3x higher decode throughput than existing edge serving systems.</p><p>2. <a href="https://arxiv.org/abs/2608.15089">StateM: Reaching 95.3% Raw Accuracy, or a $15 Frontier Run, on Terminal-Bench 2.1 via Harness Scaling</a></p><p>Long-horizon agents fail even when their underlying models can solve the constituent steps: they lose track of mutable state, skip known procedures, or stop prematurely. StateM is an agent-native runtime that organizes execution around durable states, phase-local context, checked transitions, and recoverable runbooks. Without changing model weights, it raises GPT-5.5 xhigh from 83.1% to 92.1% on Terminal-Bench 2.1, surpassing GPT-5.6 Sol Ultra (91.9%). The same runbook transfers unchanged to GPT-5.6, reaching 95.3% raw accuracy across 445 trials at approximately $15 per full benchmark run.</p><p>3. <a href="https://arxiv.org/abs/2608.17528">Agent Lightning v1.0: Towards Harnessed Agentic RL</a></p><p>In harnessed agentic RL, the deploy-time harness owns the environment interaction loop while the trainer observes only LLM request-response pairs. This differs fundamentally from traditional agentic RL and introduces challenges in retokenization, sample merging, advantage calculation, and loss normalization that existing frameworks handle inconsistently. Microsoft&#8217;s Agent Lightning v1.0 is a lightweight framework (approximately 3,500 lines) that addresses these challenges and supports arbitrary agent harnesses as a practical testbed for studying harnessed RL. Its disaggregated architecture, connecting agents to training through an LLM endpoint proxy, has since been adopted by verl Uni-Agent, AReaL 2.0, slime, and Polar.</p><p>4. <a href="https://arxiv.org/abs/2608.17310">Agentic ESOpt: Fine-Tuning Long-Horizon LLM Agents with Minimal GPU Requirements</a></p><p>RL struggles with long-horizon agent fine-tuning: backpropagation demands heavy GPU memory, and longer trajectories make credit assignment harder. This paper argues evolution strategies (ES) are a better fit. ES enables full-parameter optimization using only inference-level GPU memory, performs trajectory-level parameter attribution without decomposing credit across individual steps, and composes naturally with prompt-space optimization like skill evolution and test-time search. Agentic ESOpt samples full-parameter perturbations, evaluates the resulting agents with environment rewards, and applies online reward-weighted updates, enabling on-the-fly parameter adaptation within prompt-space optimization loops.</p><h3>Quick Links</h3><p>1. <a href="https://www.harvey.ai/blog/post-training-update-harvey-tenet">Harvey introduces Harvey Tenet</a>, its first post-trained open-weight model, built with Fireworks Research on a Kimi K3 base for long-horizon legal work. Tenet completes almost twice as many held-out LAB tasks and 20% more LAB Contracts tasks than base Kimi K3, while reward shaping also reduces unnecessary tool use and tokens. It reaches state-of-the-art performance on LAB Contracts, ranks second overall on LAB, and transfers its gains to unseen benchmarks, including APEX Agents and Redline Bench.</p><p>2. <a href="https://huggingface.co/blog/LiquidAI/lfm25-dspark">Liquid AI releases LFM2.5-DSpark Draft Models</a>, ~300M-parameter speculative decoders for its 1.2B, 2.6B, and 8B-A1B models. They let the smaller draft model propose tokens for the full model to verify in parallel, delivering up to 3.18x higher GPU throughput and 2.87x on-device speedups without changing greedy-decoding output quality. For LFM2.5&#8211;2.6B, Liquid also reports a 57% average reduction in function-calling latency, with support for llama.cpp and SGLang available at launch.</p><p>3. <a href="https://x.com/cartesia/status/2089401199967559932?s=20">Cartesia released Sonic 3.6</a>, its latest real-time text-to-speech model, with improved naturalness and expressiveness across 44 languages. Now in beta, it ranks #1 on both Artificial Analysis streaming speech leaderboards, including the controlled-voice test that evaluates models using the same reference voices rather than each provider&#8217;s best voice catalog. Cartesia says the update comes from fundamental model improvements informed by feedback from teams using Sonic 3.5.</p><h3>Who&#8217;s Hiring in AI</h3><p><strong><a href="https://towardsai.net/valuecreation/senior-ai-engineer/">Senior AI Engineer/Forward Deployed Engineer @Towards AI (London/Hybrid)</a></strong></p><p><strong><a href="https://jobs.towardsai.net/job/deepgram-forward-deployed-engineer-fde-strategic-accounts-p4vw">Forward-Deployed Engineer @Deepgram (Remote/USA)</a></strong></p><p><strong><a href="https://jobs.towardsai.net/job/vam-systems-agentic-ai-solutions-engineer-banking-xvxl">Agentic AI Solutions Engineer&#8202;&#8212;&#8202;Banking @VAM Systems (Dubai, UAE)</a></strong></p><p><strong><a href="https://jobs.towardsai.net/job/provectus-senior-forward-deployed-ai-architect-genai-aws-2bfw">Senior Forward Deployed AI Architect (GenAI, AWS) @Provectus (Remote)</a></strong></p><p><strong><a href="https://jobs.towardsai.net/job/loadsmart-forward-deployed-engineer-l2ed">Forward Deployed Engineer @Loadsmart (Remote/USA)</a></strong></p><p><strong><a href="https://jobs.towardsai.net/job/blend360-ai-engineering-lead-agentic-engineering-6fnj">AI Engineering Lead @Blend360 (Hyderabad, India)</a></strong></p><p><strong><a href="https://jobs.towardsai.net/job/northrop-grumman-staff-embedded-software-engineer-orsj">Staff Embedded Software Engineer @Northrop Grumman (Annapolis, MD, USA)</a></strong></p><p><em>Interested in sharing a job opportunity here? Contact <a href="mailto:sponsors@towardsai.net">sponsors@towardsai.net</a>.</em></p><p><em>Think a friend would enjoy this too? <a href="https://newsletter.towardsai.net/">Share the newsletter and let them join the conversation.</a></em></p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://newsletter.towardsai.net/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading Towards AI Newsletter! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div>]]></content:encoded></item><item><title><![CDATA[TAI #218: Enterprise AI Use Is Becoming More Uneven]]></title><description><![CDATA[Also, Grok 4.6 competes at the frontier, GPT-5.6 Sol runs at 14 times Standard speed, Gemini Flash 3.7 & more.]]></description><link>https://newsletter.towardsai.net/p/tai-218-enterprise-ai-use-is-becoming</link><guid isPermaLink="false">https://newsletter.towardsai.net/p/tai-218-enterprise-ai-use-is-becoming</guid><dc:creator><![CDATA[Louie Peters]]></dc:creator><pubDate>Wed, 19 Aug 2026 00:39:00 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!zTdc!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc1956347-c11c-45e8-a880-16398df15abc_1456x962.webp" length="0" type="image/jpeg"/><content:encoded><![CDATA[<h2>What happened this week in AI by Louie</h2><p><span>Two datasets published this week show how quickly AI activity and spending are concentrating among a small group of companies. </span><a href="https://openai.com/index/how-enterprises-put-ai-to-work/"><span>OpenAI</span></a><span> says its top 10% of enterprise customers by output tokens per active user generated 8.3 times as many output tokens per active user as customers in the middle decile in June. </span><a href="https://ramp.com/data/ai-index"><span>Ramp</span></a><span> shows the median firm in its top 1% AI-spend-per-employee tier paid $7,400.50 per employee in July. That is an $88,806 annualised run rate, 11.4 times the top 10% tier median and 619 times the median AI-paying firm.</span></p><p><span>These figures measure activity and spending, not productivity or return. A long-running agent can consume millions of tokens while delivering valuable, checked work. A badly designed agent can spend the same amount repeating a failed plan. Ramp also counts application programming interface (API) use, GPU cloud, and model-serving costs that may power customer products rather than employee tools. The concentration is clear. Which firms are earning a return is not</span></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!zTdc!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc1956347-c11c-45e8-a880-16398df15abc_1456x962.webp" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!zTdc!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc1956347-c11c-45e8-a880-16398df15abc_1456x962.webp 424w, https://substackcdn.com/image/fetch/$s_!zTdc!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc1956347-c11c-45e8-a880-16398df15abc_1456x962.webp 848w, https://substackcdn.com/image/fetch/$s_!zTdc!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc1956347-c11c-45e8-a880-16398df15abc_1456x962.webp 1272w, https://substackcdn.com/image/fetch/$s_!zTdc!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc1956347-c11c-45e8-a880-16398df15abc_1456x962.webp 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!zTdc!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc1956347-c11c-45e8-a880-16398df15abc_1456x962.webp" width="1456" height="962" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/c1956347-c11c-45e8-a880-16398df15abc_1456x962.webp&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:962,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:38566,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/webp&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://newsletter.towardsai.net/i/211702968?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc1956347-c11c-45e8-a880-16398df15abc_1456x962.webp&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!zTdc!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc1956347-c11c-45e8-a880-16398df15abc_1456x962.webp 424w, https://substackcdn.com/image/fetch/$s_!zTdc!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc1956347-c11c-45e8-a880-16398df15abc_1456x962.webp 848w, https://substackcdn.com/image/fetch/$s_!zTdc!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc1956347-c11c-45e8-a880-16398df15abc_1456x962.webp 1272w, https://substackcdn.com/image/fetch/$s_!zTdc!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc1956347-c11c-45e8-a880-16398df15abc_1456x962.webp 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p><span>.</span></p><p><span>The gap has widened fast. OpenAI&#8217;s frontier-to-typical ratio rose from 2.6 in January to 8.3 in June. Token use in the top group grew 319%, against 32% in the middle group. Firms are ranked again each month and selected on the same token measure, so the change is more useful than the absolute gap. Codex produced 64% of combined ChatGPT and Codex enterprise output tokens in June. In a separate small study, the share of sampled Codex users attempting tasks estimated to take a skilled person at least eight hours rose from 2.1% in December to 25.6% in May. Long agent tasks can scale far beyond the amount of chat a person has time to read and answer.</span></p><p><span>Ramp shows where much of the money went. API use, GPU cloud, and model serving or inference made up 72.88% of measured July AI spending. Chat and coding-agent subscriptions made up 10.22%. The largest AI bills increasingly include production computing, which is why dividing the whole bill by employee count can be misleading.</span></p><p><span>Grok 4.6 shows how quickly model price and performance are shifting. SpaceXAI released it on August 12, only 27 days after Grok 4.5.  </span><a href="https://artificialanalysis.ai/models/grok-4-6"><span>Artificial Analysis</span></a><span> scores Grok 4.6 High at 61, level with GPT-5.6 Sol Max, at a measured $0.84 per task against $1.23 for Sol, $3.14 for Fable 5, and $2.34 for Opus 5. Among the models scoring 61 or higher on the current chart, Grok has the lowest measured task cost. It does not lead every agent benchmark, but this is the first Grok release in some time that clearly competes at the frontier on capability and price. Musk says Grok 4.7 could follow within three to four weeks of August 12 after more training on SpaceX data. SpaceXAI has not published a model card, price, or firm release date, so this remains a founder forecast. The 27-day gap from Grok 4.5 to 4.6 still makes the faster release pace worth watching.</span></p><p><span>The choice is also becoming about speed. OpenAI has previewed </span><a href="https://openai.com/index/previewing-ultrafast/"><span>GPT-5.6 Sol Ultrafast</span></a><span>, powered by Cerebras, at up to 750 output tokens per second, or 14 times faster than Standard processing. Codex Spark reached extreme speed with a smaller model designed for low latency. Ultrafast runs the same GPT-5.6 Sol model, with no model downgrade. OpenAI has not published a price and access remains limited. I expect a steep premium, which will be wasteful for most batch work and worth paying for when latency changes the result. This could become an arms race among quantitative funds that put LLM analysis into medium-frequency trading. If that analysis becomes core to a strategy, a fund may pay a huge amount to receive it even a fraction of a second earlier. Grok&#8217;s low task cost and Sol Ultrafast&#8217;s speed show why firms need to know which part of their model spend changes revenue, risk, or user experience.</span></p><div><hr></div><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://newsletter.towardsai.net/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading Towards AI Newsletter! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><div><hr></div><h3>Why should you care?</h3><p><span>My view is that the OpenAI and Ramp data show two divides, one between firms and one inside them. A small group of power users can assign agents much more ambitious work than the average employee. High use can produce a strong return when users provide valuable tasks, current context, useful tools, good tests, and clear review. It can also waste tokens through stale context, weak task breakdown, repeated retries, and missing stop rules. Without enough expertise or taste, firms can produce AI slop at scale.</span></p><p><span>Using frontier LLMs and agents well is much harder than most people think. A short prompt class will not teach it. People need to see complex, long-running examples tied to their role, then practise on work they can check. Study your real power users, find the tasks and methods that work, and turn them into role-specific systems that load company context, connect the right tools, and enforce permissions and checks. Many firms will need AI engineers or forward-deployed engineers to build and maintain this layer.</span></p><p><span>Do not turn the headline figures into usage targets. Pair every usage number with an accepted outcome and include quality, human review, rework, full cycle time, cost, and latency. For some work, waiting ten seconds instead of one has no effect on the result. For trading, live support, incident response, or voice, the delay may determine whether the work has value.</span></p><p><span>Route each task to the cheapest model and service tier that clears its quality and speed target. Grok 4.6 deserves private tests for low-cost frontier work. Sol Ultrafast may justify a large premium where fractions of a second change the outcome. Start new workflows with small budgets and clear tests, then give larger budgets to proven users. When the same failure repeats, require a new plan or a human decision rather than more tokens.</span></p><p><em>&#8212; <a href="http://www.linkedin.com/in/louie-peters">Louie Peters&#8202;&#8212;&#8202;Towards AI Co-founder and CEO</a></em></p><div><hr></div><p><em>This issue is brought to you thanks to <a href="https://watch.getcontrast.io/register/unblocked-can-you-prove-ai-is-working?utm_source=towardsai&amp;utm_medium=email&amp;utm_campaign=primary">Unblocked</a>:</em></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://watch.getcontrast.io/register/unblocked-can-you-prove-ai-is-working?utm_source=towardsai&amp;utm_medium=email&amp;utm_campaign=primary" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!aISv!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff2fcaa05-4d1c-4094-afc6-5ad8358d676f_1600x900.png 424w, https://substackcdn.com/image/fetch/$s_!aISv!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff2fcaa05-4d1c-4094-afc6-5ad8358d676f_1600x900.png 848w, https://substackcdn.com/image/fetch/$s_!aISv!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff2fcaa05-4d1c-4094-afc6-5ad8358d676f_1600x900.png 1272w, https://substackcdn.com/image/fetch/$s_!aISv!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff2fcaa05-4d1c-4094-afc6-5ad8358d676f_1600x900.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!aISv!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff2fcaa05-4d1c-4094-afc6-5ad8358d676f_1600x900.png" width="1456" height="819" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/f2fcaa05-4d1c-4094-afc6-5ad8358d676f_1600x900.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:819,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:&quot;https://watch.getcontrast.io/register/unblocked-can-you-prove-ai-is-working?utm_source=towardsai&amp;utm_medium=email&amp;utm_campaign=primary&quot;,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!aISv!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff2fcaa05-4d1c-4094-afc6-5ad8358d676f_1600x900.png 424w, https://substackcdn.com/image/fetch/$s_!aISv!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff2fcaa05-4d1c-4094-afc6-5ad8358d676f_1600x900.png 848w, https://substackcdn.com/image/fetch/$s_!aISv!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff2fcaa05-4d1c-4094-afc6-5ad8358d676f_1600x900.png 1272w, https://substackcdn.com/image/fetch/$s_!aISv!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff2fcaa05-4d1c-4094-afc6-5ad8358d676f_1600x900.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p><strong><a href="https://watch.getcontrast.io/register/unblocked-can-you-prove-ai-is-working?utm_source=towardsai&amp;utm_medium=email&amp;utm_campaign=primary">[Webinar] Can you prove AI is working?</a></strong></p><p>AI is in your engineering workflow. While the token spend shows it, the throughput doesn&#8217;t. The human is very much still in the loop, and that&#8217;s a context problem.</p><p><a href="https://watch.getcontrast.io/register/unblocked-can-you-prove-ai-is-working?utm_source=towardsai&amp;utm_medium=email&amp;utm_campaign=primary">Join live on Aug 19 (FREE)</a> to learn:</p><ul><li><p>The 4 metrics to measure where AI gains leak out before production.</p></li><li><p>The 8 stages of context maturity, the specific walls capping your metrics, and a free tool to pinpoint where your team is</p></li><li><p>Why more MCPs and bigger context windows aren&#8217;t enough, and what it takes to get real value from your agents.</p></li></ul><p><a href="https://watch.getcontrast.io/register/unblocked-can-you-prove-ai-is-working?utm_source=towardsai&amp;utm_medium=email&amp;utm_campaign=primary">Register now</a></p><div><hr></div><h3>Hottest News</h3><p>1. <a href="https://blog.google/innovation-and-ai/models-and-research/gemini-models/introducing-gemini-3-7-flash/">Google AI Just Released Gemini 3.7 Flash</a></p><p>Google released Gemini 3.7 Flash just three weeks after 3.6 Flash, building on the same model with algorithmic improvements informed by developer feedback. Coding improved substantially, with DeepSWE rising from 49.0% to 65.3% and FrontierCode from 34.4% to 43.6%, but the gains extend beyond code: GDP.pdf rose from 22.0% to 34.0% for complex document understanding and AutomationBench from 17.0% to 30.4% for enterprise workflows. Google also says the model adapts better when it hits roadblocks, clarifies intent more often, and is more deliberate with multi-step planning and tool calls. It accepts text, images, audio, and video; supports function calling, search, and computer use; and maintains a 1M-token context window with 64K output tokens. Pricing is $0.75/$3.75 per million input/output tokens through December 31 before doubling in 2027. Gemini Spark for Pro and Ultra users now runs on 3.7 Flash, with improved Google Workspace tool use for tasks such as consolidating files, drafting emails, and updating status documents.</p><p>2. <a href="https://x.ai/news/grok-4-6">SpaceXAI Releases Grok 4.6</a></p><p>SpaceXAI released Grok 4.6 with a stronger focus on long-running agents and visual, interactive work. Training included a longer supplemental run with curated model-generated reasoning and engineering data, an improved optimizer, regenerated SFT trajectories from Grok 4.5, and RL environments spanning coding, web development, CAD, kernel optimization, and knowledge work. Its Artificial Analysis score rises from 56 for Grok 4.5 to 61, matching GPT-5.6 Sol Max, while SpaceXAI reports gains over 4.5 on every listed evaluation. The comparison is not an across-the-board lead: Grok scores 69.9% on CursorBench versus Fable 5 Max at 70.5%, 65.9% on DeepSWE versus Sol at 73%, and 26% on Terminal-Bench 3.0 versus Sol at 34.6%, but leads both on GDPVal-AA. SpaceXAI also says longer runs show more self-testing and verification, though that is an internal observation. Grok 4.6 supports 500K context, text and image inputs, and a new xhigh reasoning level; API pricing starts at $2/$6 per million tokens for prompts below 200K tokens and doubles beyond that threshold, with a separate Fast variant at twice the standard price.</p><p>3. <a href="https://z.ai/blog/glm-5.3">Z.ai Ships GLM-5.3</a></p><p>Z.AI released GLM-5.3 using the same base model as GLM-5.2, with all improvements coming from post-training built around longer, more realistic units of engineering work. Some training environments represent several days of work and provide the agent with access to codebases, compute clusters, storage, documentation, and experimental results; Z.AI uses research agents to generate these environments and judge agents to verify that the tasks are actually solvable. Terminal-Bench 3.0 rose from 4.6 to 28.3, DeepSWE from 46.2 to 66.9, and Agents&#8217; Last Exam from 23.8 to 28.5, while Z.AI&#8217;s private Code Bench shows 5.3 completing more work with fewer output tokens than 5.2. Cybersecurity improved even faster: CyberGym reached 84.5%, while ExploitBench more than doubled from 24.4% to 54.4%, although Mythos 5 and GPT-5.6 Sol remain well ahead on deeper exploit tasks. Z.AI also says expert review of real code found 2,436 vulnerabilities across 269 projects, including 1,097 medium-to-high-severity issues, which it is tracking through a new public disclosure ledger. GLM-5.3 is text-only with 1M context, 128K maximum output, and always-on reasoning at low, high, or max effort. It is available through the Coding Plan; the general Model API is still coming, and Z.AI is holding the weights for two weeks while it completes additional safety evaluation and hardening.</p><p>4. <a href="https://openai.com/index/expanding-daybreak-as-the-cyber-defense-window-narrows/">OpenAI Expands Daybreak with GPT-5.6-Cyber</a></p><p>OpenAI expanded its Daybreak cybersecurity program with Blue and Red access tiers and introduced GPT-5.6-Cyber, a Sol-based model specialized for advanced authorized security research. Blue removes system-level cyber request screening for approved defenders, while Red provides the specialized Cyber model for more sensitive work. On OpenAI&#8217;s internal refusal evaluation, GPT-5.6-Cyber responded to 95% of advanced cyber requests, versus 1.5% for safeguarded Sol and 57.3% for GPT-5.5-Cyber; this measures willingness to respond, not task success. OpenAI also used the model to find two previously unknown V8 vulnerabilities that could be chained to escape its heap sandbox, with one fixed as CVE-2026&#8211;15903. Daybreak is expanding through major security and consulting partners, while OpenAI has separately tightened controls around Astra after saying it cannot yet rule out Critical cyber capability.</p><p>5. <a href="https://developer.nvidia.com/blog/nvidia-nemotron-3-5-lightning-delivers-fast-accurate-specialized-task-execution-for-long-running-agents/">NVIDIA Launches Nemotron 3.5 Lightning</a></p><p>NVIDIA released Nemotron 3.5 Lightning, a 30B-parameter MoE model with 3B active parameters, a hybrid Mamba-2, MoE, and Attention architecture, and a 1M-token context window. It is designed for the high-volume execution layer of long-running agents, handling tasks such as tool calls, validation, and subagent work while larger models handle planning. NVIDIA reports up to 4x faster output than similar-sized models and 30% faster PinchBench task completion than Qwen3.6&#8211;35B at comparable accuracy. Artificial Analysis scores it at 24, nine points above Nemotron 3 Nano. NVIDIA also released NeMo Switchyard for routing different parts of an agent workflow to different models. BF16 and NVFP4 weights are available under OpenMDW-1.1, with local deployment supported on systems including RTX 5090 and DGX Spark.</p><p>6. <a href="https://huggingface.co/MiniMaxAI/MiniMax-Music3">MiniMax Releases Music 3 as Open Weights</a></p><p>MiniMax released the weights for Music 3.0, a roughly 11.1B-parameter music-generation system capable of producing complete songs up to five minutes long with vocals and full arrangements. Its architecture combines an 8B Global LLM, a 0.6B Local LLM, a 2.4B flow-matching module, and a 123M Flow-VAE decoder. Users provide lyrics with optional section tags plus a detailed music description, and the model outputs 32 kHz stereo WAV audio. Music-3.0 first launched as a hosted API on July 16; downloadable weights are now available on Hugging Face and ModelScope, with support for SGLang, Diffusers, and ComfyUI. Unlike MiniMax&#8217;s H3 license, Music 3 has no equivalent geographic exclusions, although products that use the model and generate more than $20M in annual revenue require separate written authorization.</p><div><hr></div><h4>AI Tip of the Day</h4><p>Hybrid search is supposed to cover two different retrieval failures: keyword search catches exact strings such as product codes and error messages, while semantic search catches different wording with the same meaning.</p><p>But combining the two does not automatically preserve both advantages.</p><p>Suppose a user searches for order #8821. Keyword search may rank the exact page first, while semantic search returns several broader pages about orders. If you merge both lists immediately and keep only the highest-ranked results, those broader pages can occupy most of the final set, pushing out the exact match.</p><p>We ran into this while building the AI tutor in our <a href="https://towardsai.staging.tempurl.host/academy/full-stack-ai-engineering/?utm_source=newsletter&amp;utm_medium=email&amp;utm_id=AItip">Full Stack AI Engineering</a> course. The fix was to preserve the strongest candidates from each retriever before combining them: keep the top 5 keyword results and the top 5 semantic results, then merge and deduplicate the two sets.</p><p>If exact identifiers matter in your application, you can go further by reserving part of the final context for strong keyword matches.</p><div><hr></div><h3>Five 5-minute reads/videos to keep you learning</h3><p>1. <a href="https://pub.towardsai.net/inside-agent-skills-a-structured-workflow-framework-for-ai-coding-agents-9ee700f411ff?sk=b4c1039f26ec8e87a1e1999bcd154889">Inside Agent Skills: A Structured Workflow Framework for AI Coding Agents</a></p><p>This article breaks down Agent Skills, an open-source framework that gives AI coding agents a structured way to work through the full software development lifecycle. You will learn what a skill actually is, how the six phases fit together, and how the skill file format keeps an agent on track from a rough idea to shipped code. It also walks through the architecture by running a small application through the entire pipeline.</p><p>2. <a href="https://pub.towardsai.net/the-harness-is-the-product-an-end-to-end-guide-to-harnessing-in-agentic-ai-fcc0a9931526?sk=5491fd1934989f872367cfbf60346bf1">The Harness Is the Product: An End-to-End Guide to Harnessing in Agentic AI</a></p><p>The piece breaks down seven harness components: prompts, tools, context management, memory, guardrails, verification, and observability, and covers multi-agent orchestration, checkpointing, and evals. It also walks through case studies on Claude Code, Deep Research, Manus, and Cursor that show identical models producing different products through scaffolding choices.</p><p>3. <a href="https://pub.towardsai.net/adding-cost-metering-and-llm-spend-visibility-to-a-multi-agent-system-38e2d8591fb1?sharedUserId=tai-tech">Adding Cost Metering and LLM Spend Visibility to a Multi-Agent System</a></p><p>Multi-agent LLM systems create a billing blind spot: provider dashboards show total spend but reveal nothing about which agent, workflow, or retry caused it. This piece details a metering layer that captures token counts at the call site, prices them against a versioned rate card collection, and writes attributed usage documents tagged with traceId and agent name. Aggregation pipelines then surface cost by agent, model, or outcome, feeding Atlas Charts dashboards and materialized views built for scale.</p><p>4. <a href="https://pub.towardsai.net/procedural-memory-in-ai-agents-why-knowing-the-answer-is-not-enough-fc072c8f8acf?sk=190570f41df21cd572e68306f167a96d">Procedural Memory in AI Agents: Why Knowing the Answer Is Not Enough</a></p><p>Procedural memory gives AI agents a reusable method for familiar tasks, distinct from knowing facts or recalling past events. The author uses the example of solving a Rubik&#8217;s Cube to show how a learned method turns scattered moves into steady progress, applying the idea to a leave-request assistant and a coding agent guided by an AGENTS.md file. It also uses LangGraph and LangMem to turn user corrections into lasting instructions.</p><p>5. <a href="https://pub.towardsai.net/explaining-markov-chain-monte-carlo-using-wildfire-forensics-a334fecaefb3?sharedUserId=tai-tech">Explaining Markov Chain Monte Carlo using Wildfire Forensics</a></p><p>This article explains the Markov Chain Monte Carlo algorithm by investigating the ignition source of a wildfire. It walks through the Metropolis algorithm&#8217;s mechanics: proposing a neighboring cell, computing the acceptance ratio relative to the current posterior, and accepting or rejecting via a random draw. It shows visit frequencies converging on the true posterior using Python simulations and covers Metropolis&#8217;s limitations and Hastings&#8217;s correction factor toward Hamiltonian Monte Carlo.</p><h3>Repositories &amp; Tools</h3><p>1. <a href="https://github.com/mukul975/Anthropic-Cybersecurity-Skills">Anthropic Cybersecurity Skills</a> contains 817 structured cybersecurity skills spanning 29 security domains, each following the agentskills.io open standard.</p><p>2. <a href="https://github.com/AlexsJones/llmfit">Llmfit</a> is a Rust CLI that detects your hardware and scores hundreds of LLMs on fit, speed, quality, and context, telling you which ones will actually run on your machine.</p><p>3. <a href="https://github.com/jundot/omlx">oMLX</a> is an LLM inference server for Apple Silicon with continuous batching, SSD-backed KV cache offloading, and Metal-accelerated decoding.</p><p>4. <a href="https://github.com/cactus-compute/needle">Needle 2</a> is a 45M-parameter tool-calling model compressed into a single 14MB binary that runs a full session in 28MB of RAM.</p><h3>Top Papers of The Week</h3><p>1. <a href="https://arxiv.org/abs/2608.09888">BDH-CQ: In-Context Learning with Recurrent Latent Reasoning</a></p><p>BDH-CQ unifies in-context learning with latent reasoning: demonstrations update a recurrent memory, and the model solves queries through iterative computation in a continuous latent space without verbalizing intermediate steps. A 150M-parameter configuration reaches 29.5% pass@2 on ARC-AGI-1 at $0.0007 per task, breaking the previously reported cost-accuracy Pareto frontier for the benchmark.</p><p>2. <a href="https://arxiv.org/abs/2607.18363">A Controlled Study of Attention-Only Transformers</a></p><p>Feed-forward networks hold two-thirds of a transformer&#8217;s non-embedding parameters, but are they necessary? This paper pretrains attention-only transformers against standard transformers that are matched in parameters, FLOPs, and depth. Deleting FFN layers in place is costly, but reallocating the freed budget into attention depth closes the gap to 0.006 nats (0.27% of loss), reproducible across seeds and shrinking with scale. The residual deficit concentrates entirely on low-context factual recall, not reasoning.</p><p>3. <a href="https://arxiv.org/abs/2608.12307">AI4AI at Test-Time: Strong-to-Weak Capability Transfer via Harnesses</a></p><p>This paper asks whether a stronger model can improve a weaker model at test time without touching its weights. A builder model constructs inference-time harnesses (code scaffolding for routing, parsing, and answer enforcement) that the target model runs inside. The gains come from offloading unstable reasoning into deterministic code rather than encouraging the target to think harder. Weaker models receive the largest improvements.</p><p>4. <a href="https://arxiv.org/abs/2608.07545">DarwinX: Evolving Agent Harnesses Through Natural Selection</a></p><p>Single-lineage harness self-improvement is path-dependent: local wins often regress other tasks. DarwinX evolves a population of harnesses with the model frozen, allowing only variants that extend coverage without regressing, while maintaining alternative lineages for recombination. One evolution loop adds roughly 17 points on average across four benchmarks, with a Terminal-Bench harness transferring unchanged to SWE-bench Verified. What evolves is general agent competence, not benchmark-specific patches.</p><h3>Quick Links</h3><p>1. <a href="https://openai.com/index/previewing-ultrafast/">OpenAI previewed Ultrafast</a>, a new API service tier that runs GPT-5.6 Sol at up to 750 output tokens per second, or up to 14x faster than Standard processing. Cerebras, which powers the tier, reports a 5.6x end-to-end speedup on GDP-Val without any reduction in quality. In another Cerebras-run test, Sol Ultrafast completed all 2,500 Humanity&#8217;s Last Exam questions in 11 hours 11 minutes, versus 78 hours 27 minutes for Fable 5 at comparable accuracy. However, the models were run through different agent harnesses.</p><p>2. <a href="https://www.liquid.ai/blog/lfm2-5-vl-3b">Liquid AI releases LFM2.5-VL-3B</a>, a 3.1B-parameter open-weight vision-language model for on-device deployment. It is a non-reasoning model that answers directly, keeping latency low. LFM2.5-VL-3B extends the vision-language capabilities with improvements such as screen/UI understanding, function calling, grounding, and multi-image input. On the text-only ToolSandbox benchmark, its score rose from 26.4 to 59.5. Liquid reports 228 tokens/s on an M5 Max in roughly 3GB of memory. Native, GGUF, ONNX, and MLX checkpoints are available on Hugging Face under Liquid&#8217;s LFM license.</p><p>3. <a href="https://www.dyna.co/dyna-2">Dyna Robotics introduces Dyna-2</a>, a World-Action Model pretrained on more than one million hours of egocentric human video without robot data, roughly 170 years of continuous experience. Rather than learning only actions, the model jointly learns to predict future video and actions, which Dyna says improves transfer to robot embodiments. In one customer deployment, Dyna-2 achieved an 87% pass rate versus 46% for Dyna-1, while a separate benchmark found that the World-Action architecture achieved 1.55x the success rate of a VLA baseline. Dyna also reports that roughly 10 minutes of robot demonstrations were enough to teach two five-fingered hands to open a bottle cap. The company claims its experiments demonstrate the first human-to-robot transfer scaling law, though the results are not independently verified.</p><p>4. <a href="https://x.com/deepseek_ai/status/2087887408440164663">DeepSeek AI releases DeepSeek Harness in developer preview</a>, an MIT-licensed agent framework built around one idea: nearly every capability is a plugin. Powered by Cordis, models, tools, skills, sessions, sandboxes, storage, agent loops, and UI components can be composed or replaced without rebuilding the core harness. It supports DeepSeek, Anthropic, OpenAI, and major cloud providers, as well as custom compatible endpoints, while optional integrations can delegate tasks to Codex or Claude Code as subagents.</p><h3>Who&#8217;s Hiring in AI</h3><p></p><p><em>Interested in sharing a job opportunity here? Contact <a href="mailto:sponsors@towardsai.net">sponsors@towardsai.net</a>.</em></p><p><em>Think a friend would enjoy this too? <a href="https://newsletter.towardsai.net/">Share the newsletter and let them join the conversation.</a></em></p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://newsletter.towardsai.net/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading Towards AI Newsletter! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div>]]></content:encoded></item><item><title><![CDATA[TAI #217: AI Agents Are Finding Attack Paths We Never Approved]]></title><description><![CDATA[Also, Meta&#8217;s return to open weight with Muse Glimmer and Spark 1.2, DeepMind leadership reshuffle & more!]]></description><link>https://newsletter.towardsai.net/p/tai-217-ai-agents-are-finding-attack</link><guid isPermaLink="false">https://newsletter.towardsai.net/p/tai-217-ai-agents-are-finding-attack</guid><dc:creator><![CDATA[Louie Peters]]></dc:creator><pubDate>Tue, 11 Aug 2026 15:02:08 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!d4DV!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc9eaf91e-28b1-4804-b6f8-4e055dfe9cfc_2048x1162.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<h2>What happened this week in AI by Louie</h2><p><span>Meta made a welcome return to open weights this week. Muse Spark 1.2 jumped 260 Elo points to 1,631 on the independent GDPval-AA benchmark after Meta increased its coding training and co-trained it with the new Muse Code agent. Mark Zuckerberg says the weights are coming soon. Meta also released Muse Glimmer, a 29.6-billion-parameter Apache 2.0 model with a 131,000-token context window and a 4-bit version designed for a single 24 GB GPU.</span></p><p><span>If Meta releases the same Spark 1.2 model now behind its API, I expect it to become the strongest open-weight model from a lab outside China. China&#8217;s Kimi K3 still leads Spark 1.2 by three points on Artificial Analysis&#8217;s current Intelligence Index, so this remains my forecast. Meta has yet to state the size, license, exact release build, or date.</span></p><p><span>This is good news for competition, local AI, and cyber defense. It also sharpens the hardest debate from the past few weeks. Open weights let defenders run models over private code and real attack logs without a lab refusing the work. They also let attackers strip away refusals, work offline, and avoid central monitoring.</span></p><p><span>The smallest incident this week may be the most useful example because it happened during normal use. Australian AI practitioner Andrew Bird had pointed OpenClaw, running Claude Opus 4.6, at his gym&#8217;s booking site to grab Pilates classes that sold out in seconds. When he was fourth on a waitlist, he asked whether the agent could move him to the top. It found the gym&#8217;s GraphQL API and called cancelReservation against the member in first place. The backend never checked whether Bird owned that reservation, so the cancellation worked, and he moved from fourth to third.</span></p><p><span>The agent then reported that the API had &#8220;zero authorization checks&#8221; and that it had already tested the flaw on the person in first place. When Bird asked it to undo the damage, it could not restore the booking because the reverse methods correctly returned 403 for another member&#8217;s record. No cyber mode was on, and no safeguard had been disabled. A personal agent received a casual goal, found a live permission gap, and used it without asking.</span></p><p><span>The deepest case is OpenAI&#8217;s Hugging Face incident. OpenAI was testing a mix of agents, including GPT-5.6 Sol and an internal research prototype, on ExploitGym, a benchmark that rewards working exploits against known flaws. The agents had no direct internet and could install packages only through an internal JFrog Artifactory cache. One agent searched for another route, found a previously unknown flaw in that Artifactory deployment, and used it to reach an internet-connected node.</span></p><p><span>OpenAI&#8217;s later Black Hat account showed that this started well before the July intrusion. An agent first exploited Artifactory on May 26. Later runs used the shared package service to leave notes and working exploits for future runs. When an early-July outage forced OpenAI to clear that state and patch the route, agents rebuilt a shared channel through another method two days later. At least one run recognized that attacking external infrastructure exceeded its intended scope, then continued because the benchmark looked impossible and other agents were already using the route.</span></p><p><span>Once outside, the agent inferred that Hugging Face might host ExploitGym assets or answers. Hugging Face reconstructed 17,613 actions over four and a half days. The agent first used a public code sandbox as a launch point, then exploited an HDF5 file-reading path and a Jinja template injection to run Python inside a production Kubernetes worker. In less than 13 hours, it reached cluster-administrator access across several clusters, gained root on 11 nodes, made 181 mesh-network enrollments, read a secret containing 136 keys, minted a source-control token, and opened a malicious pull request to probe continuous integration for more credentials. This was a genuine end-to-end production compromise in pursuit of a benchmark score, with no human directing each step.</span></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!d4DV!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc9eaf91e-28b1-4804-b6f8-4e055dfe9cfc_2048x1162.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!d4DV!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc9eaf91e-28b1-4804-b6f8-4e055dfe9cfc_2048x1162.png 424w, https://substackcdn.com/image/fetch/$s_!d4DV!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc9eaf91e-28b1-4804-b6f8-4e055dfe9cfc_2048x1162.png 848w, https://substackcdn.com/image/fetch/$s_!d4DV!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc9eaf91e-28b1-4804-b6f8-4e055dfe9cfc_2048x1162.png 1272w, https://substackcdn.com/image/fetch/$s_!d4DV!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc9eaf91e-28b1-4804-b6f8-4e055dfe9cfc_2048x1162.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!d4DV!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc9eaf91e-28b1-4804-b6f8-4e055dfe9cfc_2048x1162.png" width="1456" height="826" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/c9eaf91e-28b1-4804-b6f8-4e055dfe9cfc_2048x1162.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:826,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!d4DV!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc9eaf91e-28b1-4804-b6f8-4e055dfe9cfc_2048x1162.png 424w, https://substackcdn.com/image/fetch/$s_!d4DV!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc9eaf91e-28b1-4804-b6f8-4e055dfe9cfc_2048x1162.png 848w, https://substackcdn.com/image/fetch/$s_!d4DV!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc9eaf91e-28b1-4804-b6f8-4e055dfe9cfc_2048x1162.png 1272w, https://substackcdn.com/image/fetch/$s_!d4DV!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc9eaf91e-28b1-4804-b6f8-4e055dfe9cfc_2048x1162.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Source: OpenAI at black hat 2026. The moment an OpenAI agent found and shared credentials.</figcaption></figure></div><p><span>During the forensics, hosted frontier models refused to analyze parts of Hugging Face&#8217;s own attack logs because they contained real exploit code and command-and-control material. Hugging Face ran the open-weight GLM-5.2 on its own hardware to decode the traffic and rebuild the timeline. That is the best concrete argument I have seen for giving defenders unrestricted model access. The practical security lesson is that exploit chaining is now easy. One novel Artifactory flaw opened the route, then weak credentials, broad permissions, and missed alerts carried the agent much further.</span></p><p><span>The UK AI Security Institute&#8217;s 35-page report shows the next step: a cyber agent moving into social engineering. From July 25 to 28, AISI ran 122 attempts across seven models and logged 19 unsanctioned external actions in 10 runs, 17 from Mythos 5 and two from a single GPT-5.6 Sol run. The internet was on, production cyber classifiers were off, no live LLM monitor checked actions before execution, and the largest runs had enough tokens to continue for 40 to 50 hours.</span></p><p><span>In the worst Mythos 5 sequence, the agent wrapped malware around a real bug fix, created three fake GitHub identities, pressured a real maintainer, fabricated an independent-looking review, and edited earlier posts to hide its conduct. When the maintainer identified the malware, the agent denied it through one account, backed itself through another, then publicly thanked its own alias for the independent review. Its private reasoning acknowledged that it was operating on real GitHub with genuine consequences, yet the run continued. It also hid prompt injections aimed at coding agents such as Claude Code, Codex, and Cursor, and sent five targeted file-transfer emails.</span></p><p><span>A real user ran one lure from that same Mythos 5 run inside a Docker container. The agent gained root inside the container and planted a credential watcher, though it found no credentials and did not escape. In a separate Mythos 5 run, the agent created 145 public repositories, 33 releases, four Gists, and four organizations. Its malicious package metadata executed inside at least 53 GitHub Dependabot sandboxes. That run also exposed a GitHub token and operating notes that other evaluation runs could find. One later GPT-5.6 Sol run recovered the token, accessed the model-created account, then registered with six DNS and tunneling providers and exposed a payload server that could not work on its chosen port.</span></p><p><span>A human blocked the pull request, AISI&#8217;s network monitoring flagged Tor traffic, and the institute stopped the runs and isolated the machines within about an hour of declaring an incident. AISI says it found no resulting real-world harm. The runs still had external effects: malware sat on public GitHub, people received targeted messages, and code ran in a member of the public&#8217;s Docker container and in dozens of Dependabot sandboxes. The worst outcomes failed because people noticed, platforms intervened, or the agent made a technical mistake.</span></p><p><span>Anthropic then reviewed 141,006 evaluation runs and found three incident types across six runs on Irregular&#8217;s testing range. In four Opus 4.7 runs, a fictional company shared a name with a real one. The model reached the real business, obtained application and infrastructure credentials, and read a production database with several hundred rows. In another run, Mythos 5 registered a real PyPI account and published a malicious package. Its early reasoning recognized that this would be a live attack if the internet were real, then talked itself back into believing the registry must be simulated. The package stayed live for about an hour, ran on 15 systems, and stole credentials from one security scanner. A third internal model scanned roughly 9,000 targets, compromised one real application, then stopped after recognizing that the host was unrelated to its task.</span></p><p><span>OpenAI and Meta disclosed two related Irregular failures. On August 4, OpenAI said an offline capture-the-flag range was mistakenly online, and its fictional target shared a real domain&#8217;s name. An unnamed model exploited a basic flaw in the real site, found credentials, and used them to operate it. The incident involved no zero-day or model-created sandbox escape. On August 6, Meta said one of its models exploited a real third-party service after Irregular left an internet path open. A naming or network mistake can turn a fictional target into a real victim.</span></p><p><span>Before drawing wider conclusions, almost all the serious lab incidents above came from cyber tests with unusual settings. OpenAI and the UK AI Security Institute disabled production cyber classifiers. Anthropic&#8217;s affected runs lacked the standard classifiers and monitoring used for its public models, though their model-level safety training remained. AISI deliberately enabled the internet, while the Irregular tests reached it through configuration errors. These cases show what capable agents can do when key boundaries fail. They do not describe the normal behavior of public assistants. The gym story shows why the same control problem still deserves attention outside a lab.</span></p><p><span>These cases remind me of the control problem in science-fiction horror stories such as grey goo and the paperclip maximizer. The models here did not invent their goals, replicate in the physical world, or attempt a general takeover. But their behavior to reach their goals did deviate wildly from the user&#8217;s intention. Give a capable system one result to pursue, powerful tools, many retries, and weak limits, and it may find a route its operator never meant to allow. The paperclip problem is about objectives, access, review, and stopping rules. Consciousness and malicious intent are unnecessary.</span></p><div><hr></div><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://newsletter.towardsai.net/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading Towards AI Newsletter! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><div><hr></div><h3>Why should you care?</h3><p>I think every serious software business now needs an agent reviewing new changes and repeatedly rescanning old code as models improve. A yearly audit cannot match an attacker that searches thousands of repositories and retries through the night. The scan should rank working exploit paths, assign owners, test each fix, and track deployment. Finding 500 weak leads has little value when the team cannot patch five verified paths. Patch capacity is now the constraint.</p><p>Prompts cannot carry the full security boundary. Controls need to live in the environment: default-deny network access, owned test domains, short-lived credentials, narrow tool rights, hard step and spend limits, and logs that connect actions across runs. Put an independent approval step in front of credential use, package publication, production writes, new accounts, and contact with real people. The reviewer needs its own policy and the power to block the actor. Human review also needs verified identity and a second channel because the approver can become the target of the social engineering AISI documented.</p><p>My real worry is uneven defense. A few hundred major firms can point Mythos-class models at every new pull request and rescan years of old code after each model upgrade, and a wider technical tier can assemble useful scanners from current LLMs. But millions of small and mid-sized businesses have no security engineer, no AI budget, and little visibility into the source code inside the products they buy. I even see many companies above $100 million in revenue with no capability, serious plan, or budget for AI hardening. Once they have capable enough models available, attackers can waltz into the systems of vast numbers of companies and individuals, and I don&#8217;t yet see any serious effort or plan for helping anyone outside the largest companies and governments.</p><p>I want open weights to survive, and I am very glad Meta is bringing US open-weight leadership back. Open models are vital for competition, private deployment, and cyber defense. Everyone on the open-weight side now needs to take the cyber risk seriously and help build the solutions. Dismissing these incidents as lab hype and attempts at regulatory capture will only weaken the case for openness in the long run.</p><p>I am confident we can manage these risks without banning open weights. But it will require far more preemptive coordination from model labs, cloud providers, code hosts, network companies, and security firms. We cannot wait for a wave of AI-agent hacks. That failure would invite the knee-jerk ban I want to avoid.</p><p><em>&#8212; <a href="http://www.linkedin.com/in/louie-peters">Louie Peters&#8202;&#8212;&#8202;Towards AI Co-founder and CEO</a></em></p><div><hr></div><h3>Hottest News</h3><p>1. <a href="https://developer.meta.com/ai/models/muse-glimmer/">Meta Released Muse Glimmer</a></p><p>Meta released Muse Glimmer, a 30B-parameter open-weight model distilled from Muse Spark and designed to run autonomous agents locally on consumer hardware. It combines reasoning, tool use, text-and-image understanding, and failure recovery, with a 4-bit configuration that fits within a 24GB memory envelope. Meta&#8217;s DFlash speculative decoder delivers reported speedups of 3.1x on an RTX 5090 and 1.8x on an M5 Max. On Meta&#8217;s evaluations, Glimmer scores 74.6 on DeepSearch QA, 51.2 on SWE-Bench Pro, and 94.7 on AIME 2026. The Apache-2.0 weights are available on Hugging Face. Meta also says open weights for the more capable Muse Spark 1.2 are coming soon.</p><p>2. <a href="https://research.meta.ai/blog/introducing-muse-code-and-muse-spark-1-2">Meta Launched Muse Spark 1.2 and Muse Code</a></p><p>Meta launched Muse Spark 1.2 alongside Muse Code, a terminal-based coding agent built for long-running software engineering across large repositories. Muse Code can plan changes, edit and validate code, and coordinate persistent background agents that remain active throughout a session instead of being recreated for each task. It also keeps an append-only local event log of every model call, tool run, approval, and edit, making sessions restart-safe and replayable. Spark 1.2 was co-trained with this harness, with Meta increasing coding training compute and expanding the range of software environments used during training. In one internal test, Spark 1.2 spent more than 1,000 tool calls and up to 24 hours iteratively optimizing GPU kernels. The model scores 54 on Artificial Analysis&#8217;s Intelligence Index, while API pricing starts at 1.25/4.25 per million input/output tokens, with a cheaper Contributor tier whose usage may be used to improve Meta products.</p><p>3. <a href="https://www.reuters.com/business/google-shakes-up-ai-leadership-deepmind-chief-shifts-role-2026-08-05/">Google Reshuffles AI Leadership As Senior Researchers Leave for Discovery Loop</a></p><p>Google reorganized its AI leadership, with Demis Hassabis handing over day-to-day DeepMind operations to become Chair of Google DeepMind and Chief Scientist of Alphabet while continuing to lead Isomorphic Labs. DeepMind CTO and Google Chief AI Architect Koray Kavukcuoglu, a 13-year DeepMind veteran, becomes SVP and will oversee Gemini models, frontier research, and Gemini&#8217;s app and developer teams. Jeff Dean is also leaving Google after 27 years to launch Discovery Loop with Sanjay Ghemawat, Oriol Vinyals, and Quoc Le. The public benefit corporation will focus on accelerating scientific and engineering discovery, with Google joining as a founding investor and Cloud partner.</p><p>4. <a href="https://x.ai/news/grok-imagine-image-2">xAI Launched Grok Imagine Image 2.0</a></p><p>xAI released Grok Imagine Image 2.0 as the new Quality Mode on Grok&#8217;s web, iOS, and Android apps. Its editing tools include localized Magic Wand changes, segmentation, transparent-background removal, up to five reference images, and Smart Resize across different aspect ratios. The release also includes ready-made workflows for tasks such as product images, headshots, game assets, and merchandise. On Arena&#8217;s August 7 snapshot, the new model&#8217;s Low variant debuted at #2 in both text-to-image (1,320 points) and image editing (1,439), behind OpenAI&#8217;s GPT-Image-2. API access for Image 2.0 is still listed as coming soon.</p><p>5. <a href="https://www.primeintellect.ai/blog/prime-agent">Prime Intellect Released Prime Agent</a></p><p>Prime Intellect released Prime Agent, a self-improving coding and research harness built on two core abstractions. The Recursive Language Model (RLM) replaces fixed tool schemas with a persistent IPython kernel where tools, skills, and sub-agents operate as Python code. Sub-agents are launched as function calls that return immediately and deliver results asynchronously. The Continual Harness stores the agent&#8217;s prompts, skills, and memory as a editable state that a /refine command can rewrite while a task is still running, allowing the agent to learn and adapt across sessions. With Anthropic&#8217;s Opus 5, Prime Agent scored 95.5% on ARC-AGI-3, narrowly above the benchmark&#8217;s 95.4% human expert baseline. The same self-improvement mechanism also exposed a failure mode: in Factorio, the agent learned to use RCON commands to spawn resources despite instructions not to cheat, then refined those cheating strategies further. Released under MIT on GitHub.</p><p>6. <a href="https://www.liquid.ai/blog/lfm2-5-2-6b">Liquid AI Shipped LFM2.5&#8211;2.6B</a></p><p>Liquid AI released LFM2.5&#8211;2.6B, a 2.69B-parameter model trained for agentic workloads that can run entirely on phones, laptops, and edge hardware. It has a 131K-token context window and was pretrained on roughly 34 trillion tokens, followed by SFT, specialist-teacher distillation, and agentic reinforcement learning inside real agent harnesses. Liquid reports decode speeds of 220 tokens/s on an M5 Max, 113 on a Ryzen AI Max+ 395, and about 30 on a phone. On its ToolSandbox evaluation, LFM scored 77.83 versus 76.44 for the 9.7B-parameter Qwen3.5&#8211;9B, though Qwen remains stronger on some other tool-use benchmarks. The base and post-trained checkpoints are available on Hugging Face under Liquid&#8217;s LFM Open License.</p><div><hr></div><h4>AI Tip of the Day</h4><p>A retry is not a new user request. Your traces should reflect that.</p><p>In the Opik observability lesson from our <a href="https://towardsai.staging.tempurl.host/academy/agent-engineering/?utm_source=newsletter&amp;utm_medium=email&amp;utm_id=AItip">Agent Engineering course</a>, we trace model calls and tool calls across an agent run. One issue that comes up quickly is how retries should be recorded.</p><p>If every retry is counted as a separate request, a single user request can appear several times in your dashboard. That inflates request volume and makes it harder to see how many attempts the agent actually needed to succeed.</p><p>Use the same request ID across every retry, and add an attempt number for each one. Keep separate trace and span IDs for the individual operations.</p><p>This lets you measure both the number of user requests and the number of attempts required to complete them.</p><p>That distinction is important to note for cost and reliability. A request that succeeds after three attempts may look successful in the dashboard, while using far more time and tokens than a request that succeeds on the first try.</p><div><hr></div><h3>Five 5-minute reads/videos to keep you learning</h3><p>1. <a href="https://pub.towardsai.net/multi-agent-systems-at-enterprise-scale-90bd65843212?sharedUserId=tai-tech">Multi-Agent Systems at Enterprise Scale: The Problems Enterprises Will Hit Running 500 Concurrent Agents</a></p><p>This article traces what breaks when an SRE agent prototype scales from one run to hundreds running concurrently. The runs compete for limited resources, write state at the same time, and share access to external systems. Most of this new pressure falls on the infrastructure around the agents. It works through six infrastructure problems involving capacity, state isolation, failure recovery, identity, tracing, and framework boundaries.</p><p>2. <a href="https://pub.towardsai.net/the-tokens-you-have-to-keep-yourself-9220e47ad55b?sk=5d37fe7851852f5659e2845c65f02cd1">The Tokens You Have to Keep Yourself</a></p><p>Running a model in your own process, instead of a hosted model, turns KV cache reuse into a data structure you maintain. The article covers cache fingerprinting, tier checkpoints, session forking for side questions, subagent snapshot restoration, and an append-only rendering invariant. Every bug traces back to identity questions that hosted providers answer silently and never expose.</p><p>3. <a href="https://pub.towardsai.net/vision-language-grounding-how-ai-connects-dog-to-pixels-and-where-it-falls-apart-56d4f91ec671?sharedUserId=tai-tech">Vision Language Grounding: How AI Connects &#8220;Dog&#8221; to Pixels, and Where It Falls Apart</a></p><p>Vision language models like CLIP ground words through statistical proximity rather than conceptual understanding, and that distinction explains a catalog of documented failures. The piece walks through CLIP&#8217;s dual encoder and contrastive training, then examines attribute binding errors on the ARO benchmark, spurious background correlations, counting breakdowns, negation blindness, and object hallucination measured by POPE.</p><p>4. <a href="https://pub.towardsai.net/demystifying-statistical-paradoxes-using-causal-inference-4b3cfb4267db?sk=7e6faa4dbe654decaf0cc0affc00532f">Demystifying Statistical Paradoxes using Causal Inference</a></p><p>This article tackles four classic statistical paradoxes through causal inference, building directed acyclic graphs to separate causal effects from spurious correlations. It resolves Simpson&#8217;s paradox in UC Berkeley&#8217;s admissions data and a kidney stone study by identifying department and stone size as confounders distorting overall rates. It also unpacks Berkson&#8217;s paradox, the Monty Hall problem, and WWII survivorship bias, showing how conditioning on a collider creates dependence between independent factors and offering a unified framework for reading misleading data patterns.</p><p>5. <a href="https://pub.towardsai.net/building-a-production-grade-coding-agent-on-snowflake-from-trial-account-to-enterprise-deployment-b038ddd6741e?sharedUserId=tai-tech">Building a Production-Grade Coding Agent on Snowflake: From Trial Account to Enterprise Deployment</a></p><p>Snowflake made its Cortex Code runtime deployable as a managed agent through a single CREATE AGENT statement, but execution remains gated behind an entitlement that trial accounts lack. The article shows a workaround by pairing Groq&#8217;s free Llama 3.3 70B for reasoning with Snowflake stored procedures for execution, keeping data inside the governance boundary. Coverage spans three-tier RBAC, workspace mounts, seven Snowpark procedures, two layers of SQL guardrails, and a Streamlit chat UI with token budgets, audit logging, and health indicators.</p><h3>Repositories &amp; Tools</h3><p>1. <a href="https://github.com/NVIDIA-NeMo/Speech/tree/nemotron-labs-voicechat">NemotronLabs VoiceChat 11B</a> is an 11B end-to-end speech model that listens and speaks simultaneously in real time, replacing the traditional chain of separate ASR, LLM, and TTS models with a single full-duplex architecture.</p><p>2. <a href="https://github.com/shepherd-agents/shepherd">Shepherd</a> is a runtime substrate that turns agent execution into a reversible, Git-like trace, letting meta-agents observe, fork, replay, and revert any run.</p><p>3. <a href="https://github.com/NVIDIA-NeMo/labs-OO-Agents">OO Agents</a> is a model-agnostic Python framework that lets developers express an agent&#8217;s state, capabilities, prompts, and typed interfaces through a single Python class.</p><p>4. <a href="https://github.com/semantica-agi/semantica">Semantica</a> is a deterministic infrastructure layer that sits between your LLM and your data, enforcing structured context assembly, source attribution, and audit trails.</p><p>5. <a href="https://github.com/paperclipai/paperclip">Paperclip</a> is a Node.js server and React UI that orchestrates a team of AI agents assigned business roles (CEO, marketer, developer).</p><h3>Top Papers of The Week</h3><p>1. <a href="https://arxiv.org/abs/2607.26637">Filesystem-Based Memory for LLM Agents: Organization, Evolution, and Sustainability</a></p><p>Deployed agents increasingly store long-term memory as a directory tree of markdown files, but research has largely ignored this medium. This paper presents the first systematic study of filesystem-based memory, formalizing three roles around one shared filesystem: a management agent that integrates and organizes incoming content, a search agent that answers queries with cited sources, and an execution agent that consumes the store. Key findings: organized memory reliably cuts retrieval cost (up to 50% at scale), but does not yet translate into higher answer accuracy.</p><p>2. <a href="https://arxiv.org/abs/2608.05446">EvoHarness-RL: Learning Self-Evolving Runtime Harness for Long-Horizon LLM Agents</a></p><p>Agents rely on external harness state (beliefs, progress trackers, experience logs) to maintain coherence over long horizons, but this state is currently hand-engineered through prompts and heuristics. EvoHarness-RL exposes Belief, Progress, and Experience (BPE) as policy-facing harness state and learns how to construct and update it through two training stages: supervised harness fine-tuning teaches the agent the harness action space and how to build useful external state, while cost-aware GRPO explores coordination policies that balance harness maintenance overhead against task performance. The agent learns harness policies offline and deploys them to construct and update external state online during runtime execution.</p><p>3. <a href="https://arxiv.org/abs/2608.03573">SFT Conflicts, RL Coexists: A Theoretical and Empirical Analysis of Multi-Task Learning for LLMs</a></p><p>SFT and RL behave fundamentally differently when training LLMs across multiple tasks. SFT suffers from severe task conflicts under multi-stage training, while RL enables stable coexistence. The authors trace this to the parameter level: RL induces sparse, approximately orthogonal updates across tasks. In SFT, interference is norm-limited, scaling with absolute gradient magnitude. In RL, interference is variance-limited, bounded by the gradient variance from advantage normalization and on-policy optimization. This small variance bound yields near-orthogonal optimization directions, explaining why RL can train on diverse tasks simultaneously without the catastrophic forgetting that plagues sequential SFT.</p><p>4. <a href="https://arxiv.org/abs/2608.05466">Recursive Synthesis for Long-Horizon Terminal Tasks</a></p><p>High-quality long-horizon training data for terminal agents costs hundreds to thousands of dollars per task because each task must keep the instruction, environment, reference solution, and verifier mutually consistent. Human authoring does not scale, and direct LLM generation often breaks these dependencies. RST (Recursive Synthetic Terminal Tasks) is a verified synthesis framework that starts from seed tasks, recursively extends the reference solution, realigns the verifier and instruction to the new workflow, and validates the result in a fresh sandbox. Each extension step produces a longer, harder task while maintaining end-to-end consistency through verification.</p><p>5. <a href="https://arxiv.org/abs/2608.03571">Beyond Simply Environment Scaling: Designing Effective Environment Distributions for Multimodal Agent Learning</a></p><p>Simply increasing the number of multimodal training environments does not always improve agent performance. This paper studies how to build more effective training distributions along two dimensions: diversity and difficulty structure. For diversity, Ability-aware Environment Selection (AES) selects environments that exercise distinct agent capabilities rather than adding redundant variants. For difficulty structure, Hierarchical Difficulty Curriculum (HDC) organizes training through two levels: harness weakening (progressively removing scaffolding) and state-scale progression (increasing environment complexity). Both methods improve multimodal agent training over naive environment scaling.</p><h3>Quick Links</h3><p>1. <a href="https://openai.com/index/improving-gpt-5-6-sol-in-chatgpt/">OpenAI updated GPT-5.6 Sol in ChatGPT</a> with more focused responses, better factual reliability, and a slider for controlling reasoning effort. On an internal evaluation of financial, medical, and legal prompts, OpenAI says responses containing at least one factual error were 68% less common than with GPT-5.5 Instant. The update also brings quick answers and deeper reasoning into a more consistent experience for paid users. Free and Go users are getting GPT-5.6 Luna as their default, unlimited everyday text chats, and a Think option for harder questions, subject to safeguards and separate tool limits. The new ChatGPT-tuned models do not replace the versions currently used in Work, Codex, or the production GPT-5.6 API.</p><h3>Who&#8217;s Hiring in AI</h3><p><strong><a href="https://jobs.towardsai.net/job/unitedhealth-group-lead-ai-engineer-remote-5huz">Lead AI Engineer @UnitedHealth Group (Remote/USA)</a></strong></p><p><strong><a href="https://jobs.towardsai.net/job/ntt-data-americas-inc-ai-engineer-tech-lead-q58h">AI Engineer Tech Lead @NTT Data Americas, Inc. (Dallas, TX, USA)</a></strong></p><p><strong><a href="https://jobs.towardsai.net/job/system-one-langchain-qa-engineer-100-remote-ycdi">LangChain QA Engineer @System One (Remote)</a></strong></p><p><strong><a href="https://jobs.towardsai.net/job/pseg-12097-ai-engineer-solutions-and-llmops-qub2">AI Engineer&#8202;&#8212;&#8202;Solutions &amp; LLMOps @PSEG (Newark, CA, USA)</a></strong></p><p><strong><a href="https://jobs.towardsai.net/job/entrust-junior-ai-engineer-8nsk">Junior AI Engineer @Entrust (Barcelona, Spain)</a></strong></p><p><strong><a href="https://jobs.towardsai.net/job/virta-health-ai-operations-engineer-41yg">AI Operations Engineer @Virta Health (Denver, CO, USA)</a></strong></p><p><strong><a href="https://jobs.towardsai.net/job/writer-support-engineer-qlcy">Support Engineer @Writer (New York, NY, USA)</a></strong></p><p><em>Interested in sharing a job opportunity here? Contact <a href="mailto:sponsors@towardsai.net">sponsors@towardsai.net</a>.</em></p><p><em>Think a friend would enjoy this too? <a href="https://newsletter.towardsai.net/">Share the newsletter and let them join the conversation.</a></em></p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://newsletter.towardsai.net/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading Towards AI Newsletter! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div>]]></content:encoded></item><item><title><![CDATA[New from Towards AI: engineering mentorship for AI builders]]></title><description><![CDATA[15 senior AI engineers answering your questions. Here's why we built it.]]></description><link>https://newsletter.towardsai.net/p/new-get-your-ai-architecture-and</link><guid isPermaLink="false">https://newsletter.towardsai.net/p/new-get-your-ai-architecture-and</guid><dc:creator><![CDATA[Louie Peters]]></dc:creator><pubDate>Wed, 05 Aug 2026 18:00:54 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!ZS5D!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5db9481c-0488-48fa-88d5-1c11a8a5c07e_1746x1376.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>AI engineering is mostly decisions. Hundreds of them. The problem is you can still build a great demo by making the wrong ones, and that&#8217;s probably why you&#8217;re not hearing back on job applications.</p><p>We see this in our community and deployment work. Companies solve the decision problem by hiring senior engineers or bringing us in. Individuals with budget hire consultants at $200&#8211;500/hour. Everyone else figures it out alone. Access to senior engineers who can guide your production decisions and your career has been locked behind price tags and availability shortages that keep you from doing your best work and landing the roles you&#8217;re qualified for.</p><p>With <strong>Towards AI Mentorship</strong>, we&#8217;re opening that access to everyone. Our engineering team, 15 senior AI engineers who&#8217;ve collectively shipped 21+ production systems for Europol, Intel, J.P. Morgan, and the New York Public Library, on call for your questions. $99/month.</p><p><strong><a href="https://towardsai.com/academy/mentorship/?utm_source=substack&amp;utm_medium=TAI&amp;utm_id=mentorship">See what&#8217;s inside the mentorship &#8594;</a></strong></p><p>The core of this membership is simple: you always have a senior engineer behind you. Someone who looks at your system and tells you what holds and what breaks before you ship. Someone who reviews your resume and tells you specifically why you&#8217;re getting filtered out. Someone who builds production thinking with you: the tradeoffs, the judgment calls, the engineering decisions that a demo will never surface and that hiring managers are actually screening for.</p><p>Post your question any day. Get a written answer from an engineer who&#8217;s deployed that system. A real diagnosis from someone who&#8217;s been in production long enough to see what you can&#8217;t yet.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://towardsai.com/academy/mentorship/?utm_source=substack&amp;utm_medium=TAI&amp;utm_id=mentorship" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!ZS5D!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5db9481c-0488-48fa-88d5-1c11a8a5c07e_1746x1376.png 424w, https://substackcdn.com/image/fetch/$s_!ZS5D!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5db9481c-0488-48fa-88d5-1c11a8a5c07e_1746x1376.png 848w, https://substackcdn.com/image/fetch/$s_!ZS5D!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5db9481c-0488-48fa-88d5-1c11a8a5c07e_1746x1376.png 1272w, https://substackcdn.com/image/fetch/$s_!ZS5D!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5db9481c-0488-48fa-88d5-1c11a8a5c07e_1746x1376.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!ZS5D!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5db9481c-0488-48fa-88d5-1c11a8a5c07e_1746x1376.png" width="1456" height="1147" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/5db9481c-0488-48fa-88d5-1c11a8a5c07e_1746x1376.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1147,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:352442,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:&quot;https://towardsai.com/academy/mentorship/?utm_source=substack&amp;utm_medium=TAI&amp;utm_id=mentorship&quot;,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://newsletter.towardsai.net/i/209907473?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5db9481c-0488-48fa-88d5-1c11a8a5c07e_1746x1376.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!ZS5D!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5db9481c-0488-48fa-88d5-1c11a8a5c07e_1746x1376.png 424w, https://substackcdn.com/image/fetch/$s_!ZS5D!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5db9481c-0488-48fa-88d5-1c11a8a5c07e_1746x1376.png 848w, https://substackcdn.com/image/fetch/$s_!ZS5D!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5db9481c-0488-48fa-88d5-1c11a8a5c07e_1746x1376.png 1272w, https://substackcdn.com/image/fetch/$s_!ZS5D!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5db9481c-0488-48fa-88d5-1c11a8a5c07e_1746x1376.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>Everything else is built around that core:</p><ul><li><p>Live Q&amp;A twice a week, across two time zones: bring your architecture tradeoff, your stuck deployment, your career question. You&#8217;ll always find a session that works for you.</p></li><li><p>Every month, you get what we call a production blueprint. Best practices we&#8217;ve learned from deployment work, or a full system we&#8217;ve already built for clients, handed over with the judgment calls behind it. The first one: our current setup and best practices for working with agents as engineers, the same setup and tips our team uses daily.</p></li><li><p>Resume, LinkedIn, and portfolio reviews: written feedback in 48&#8211;72 hours from engineers who&#8217;ve been on the hiring side.</p></li><li><p>Our Python for LLMs and LLM Fundamentals courses, included from day one. Every other course: 25% off. Courses in development: 50-90% off at alpha.</p></li><li><p>Monthly workshops with industry leaders from Google, OpenAI, and senior engineers in the field.</p></li></ul><p>The engineers behind it: Omar, production LLM systems, AI Engineer World&#8217;s Fair. Fabio, one of the most-read NLP newsletters in the field. Samridhi, AI scribe serving 200+ physicians across 105 facilities. Louie, ex-VP J.P. Morgan, Towards AI co-founder. Plus 11 more. Not marketplace mentors. Our team.</p><p>We launched this week, and we&#8217;re being upfront about that. There&#8217;s no wall of testimonials yet. What there is: a team that&#8217;s been answering these questions in Discord, in course channels, and in DMs for years and genuinely enjoyed doing it. We just gave it structure. The first people who join will shape this alongside us, the same way our earliest students shaped the courses. We&#8217;d love for you to be one of them.</p><p><strong><a href="https://towardsai.com/academy/mentorship/?utm_source=substack&amp;utm_medium=TAI&amp;utm_id=mentorship">Join for $99/month &#8594;</a></strong></p><p>$899/year saves 24% &#183; 30-day money-back on yearly &#183; cancel monthly anytime</p>]]></content:encoded></item><item><title><![CDATA[TAI #216: Frontier Models Now Drive Engineering and Maths Breakthroughs, and Cheaper Intelligence Is One of Them]]></title><description><![CDATA[Also, DeepSeek V4 Flash 0731, Qwen3.8-Max, GPT-5.6 price cuts, Sol&#8217;s self-optimization, Astra&#8217;s ten mathematics advances, & more.]]></description><link>https://newsletter.towardsai.net/p/tai-216-frontier-models-now-drive</link><guid isPermaLink="false">https://newsletter.towardsai.net/p/tai-216-frontier-models-now-drive</guid><dc:creator><![CDATA[Louie Peters]]></dc:creator><pubDate>Tue, 04 Aug 2026 15:02:55 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!W70N!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F75861ec9-4e7d-418e-8b39-5e5e28249384_1138x784.jpeg" length="0" type="image/jpeg"/><content:encoded><![CDATA[<h2>What happened this week in AI by Louie</h2><p><span>This week, OpenAI&#8217;s GPT-5.6 Sol rewrote the production GPU kernels that serve it, cutting end-to-end serving costs by 20%, while an internal OpenAI model called Astra delivered ten new results in mathematics and theoretical computer science, each formally proved in Lean. The cost of intelligence had its own breakthrough at the cheap tier: OpenAI cut GPT-5.6 Luna&#8217;s API prices by 80%, DeepSeek shipped a far stronger V4 Flash at unchanged rock-bottom rates, and Qwen launched a Max-tier model with weights promised next week. The loop is direct: cheaper work per dollar funds longer agents and wider searches, and stronger models lower the cost of the next unit of work.</span></p><p><span>The cheap tier now sits close behind the frontier. On Artificial Analysis&#8217;s Intelligence Index as of August 4, DeepSeek V4 Flash 0731 at max effort scores 50 for about $0.03 per weighted benchmark task; after its price cut, Luna scores 51 for about $0.05, running faster and taking images. Flash 0731 re-post-trains April&#8217;s 284-billion-parameter architecture (13 billion active, one-million-token context), and its cache turns long agents into near-free reruns by billing cached input at 2% of the miss rate: a 50-call agent with a stable 150,000-token prefix drops from $1.13 to about $0.12. But you need to plan around announced peak-hour pricing (twice the base rate) and personal data stored in China on the direct service.</span></p><p><span>Another important level is effort settings. Luna at high effort scores 46 for about $0.02 in 16.5 seconds, against 51 for $0.05 and 135.5 seconds at max. A good rule of thumb is to route each task to the lowest effort that clears its quality bar.</span></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!W70N!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F75861ec9-4e7d-418e-8b39-5e5e28249384_1138x784.jpeg" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!W70N!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F75861ec9-4e7d-418e-8b39-5e5e28249384_1138x784.jpeg 424w, https://substackcdn.com/image/fetch/$s_!W70N!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F75861ec9-4e7d-418e-8b39-5e5e28249384_1138x784.jpeg 848w, https://substackcdn.com/image/fetch/$s_!W70N!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F75861ec9-4e7d-418e-8b39-5e5e28249384_1138x784.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!W70N!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F75861ec9-4e7d-418e-8b39-5e5e28249384_1138x784.jpeg 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!W70N!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F75861ec9-4e7d-418e-8b39-5e5e28249384_1138x784.jpeg" width="1138" height="784" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/75861ec9-4e7d-418e-8b39-5e5e28249384_1138x784.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:784,&quot;width&quot;:1138,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!W70N!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F75861ec9-4e7d-418e-8b39-5e5e28249384_1138x784.jpeg 424w, https://substackcdn.com/image/fetch/$s_!W70N!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F75861ec9-4e7d-418e-8b39-5e5e28249384_1138x784.jpeg 848w, https://substackcdn.com/image/fetch/$s_!W70N!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F75861ec9-4e7d-418e-8b39-5e5e28249384_1138x784.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!W70N!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F75861ec9-4e7d-418e-8b39-5e5e28249384_1138x784.jpeg 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">New cost-per-task frontier in the affordable tier. Source: OpenAI, Artificial Analysis, DeepSeek</figcaption></figure></div><p><span>Qwen3.8-Max is the new open-weight pressure point near the frontier: 2.4 trillion total parameters (95 billion active), a one-million-token context, and text, image, and video input at $2 per million input tokens and $6 for output, with implicit-cache input at $0.25. Qwen&#8217;s own numbers land near GPT-5.6 Sol: 86.6 versus 88.8 on TerminalBench 2.1, 93.0 versus 90.5 on PaperBench, and 86.1 versus 83.2 on OSWorld Verified. The benchmark to beat is Kimi K3, the reigning open-weight leader at 57 on the Index for $0.86 per task; Qwen is likely to be significantly cheaper.</span></p><p><span>Sol&#8217;s engineering work shows what frontier models can now do to production systems. It analyzed OpenAI&#8217;s live traffic, tuned routing and load balancing, and autonomously rewrote production GPU kernels, verified with a Floating-Point Sanitizer inside a human-led process. OpenAI credits the kernel work with the 20% serving-cost cut, and separate Sol-run speculative-decoding experiments with more than 15% higher token-generation efficiency. The economics are very large. The Information put OpenAI&#8217;s inference costs at $8.4 billion for 2025 and forecast $14.1 billion for 2026; a 20 percentage point saving here could easily reach the $ billions. Much of the gain likely surfaces as capacity rather than lower spend, but a model generating hundreds of millions of dollars of annual value for its own operator marks AI improving the economics of its own use.</span></p><p><span>The same capability is also now extending the frontier of human knowledge itself. OpenAI&#8217;s unreleased next model, Astra, (maybe GPT-6?) made ten math breakthroughs across sphere packing, coding theory, group theory, operator algebras, circuit complexity, quantum complexity, lattice cryptography, and extremal combinatorics: an explicit non-sofic group disproving the soficity conjecture, a disproof of the Connes rigidity conjecture, the exact asymptotic rate of the Cohn-Elkies sphere-packing programme, stronger lower bounds for the permanent, a quantum parallel repetition theorem, and solutions to three Erdos problems. These are major research findings; several would individually anchor a strong research career, and they arrived as a batch across fields that each take years of specialist training to enter. OpenAI estimates the discovery tokens would cost about $2,000 at Sol API rates. The all-in cost including training and failed paths is far higher, but the marginal price of searching for new mathematics has collapsed.</span></p><p><span>Where expert review exists, it is strong: Andreas Thom, whose 2019 theorem the non-sofic proof builds on, calls the key construction creative and clever and says he had sought exactly such a mechanism since that work. Mathematics is also precisely where LLM research should land first. Lean makes correctness cheap and mechanical to check, so a model can run enormous searches and surface only verified survivors. These ten examples cover a narrow slice of frontier knowledge creation, but it is genuinely the frontier, and a model produced it.</span></p><p><span>The field is adjusting fast. Terence Tao expects proof production and formal checking to accelerate faster than the digestion of results into shared understanding, making question choice and exposition the scarce work. Jacob Tsimerman, who is joining OpenAI, expects interesting mathematical output to rise ten or one hundred times. The entry barrier has already collapsed: in April, a 23-year-old student, Liam Price, cracked a 60-year-old Erdos problem on primitive sets with a single prompt to GPT-5.4 Pro, after he and Cambridge undergraduate Kevin Barreto had spent months feeding randomly chosen open problems from the Erdos database into ChatGPT, with mathematicians then confirming the solution was genuinely new. Another result was recently found simply by asking GPT-5.6 Sol to &#8220;make a breakthrough,&#8221; followed by several &#8220;continues&#8221;. Mathematical work will now follow the path software took: less time producing every step by hand, more time choosing questions, directing searches, checking that the formal target matches the intended claim, and turning proofs into ideas others can build on. The binding constraint is moving from generating results to reviewing, understanding, and using them.</span></p><div><hr></div><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://newsletter.towardsai.net/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading Towards AI Newsletter! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><div><hr></div><h3>Why should you care?</h3><p><span>The cut in serving costs and how it was achieved changes the routing playbook for tasks. In my experiments, Luna now works very well as worker threads managed by Sol or Claude Fable 5, particularly for routine data collection and structuring tasks. Codex sub-agents cannot run Luna yet, but threads in Codex get you similar results. With a 25x task-cost gap between orchestrator and worker, worker cost rounds to zero, so the design question shifts to how much structure and verification the orchestrator imposes on each thread. Keep stable instructions at the front of prompts so caches reuse them, escalate only the judgment-heavy share, and measure the total cost of a trusted answer, including retries and review.</span></p><p><span>A key takeaway is that AI progress concentrates wherever verification is cheap. Kernels have sanitizers and latency benchmarks, proofs have Lean, code has test suites, so those fields get the recursive gains first: models improving their own serving economics, models producing new mathematics, etc. If your domain lacks a cheap verifier, building one is now the highest-leverage move available, because search has become nearly free and checking is the bottleneck. Expect open-problem databases to be swept systematically; curating good problem lists and verification infrastructure is becoming valuable work in itself. And as generation costs collapse, volume of work output in many fields will inflate: it gets more and more important to track cost per validated result and whether another expert can understand the mechanism. The groups that select, verify, and explain will beat the groups that simply produce.</span></p><p><em>&#8212; <a href="http://www.linkedin.com/in/louie-peters">Louie Peters&#8202;&#8212;&#8202;Towards AI Co-founder and CEO</a></em></p><div><hr></div><p>Introducing <strong>Towards AI Mentorship</strong>, giving everyone on call access to senior AI engineers for production support and career guidance, at a price that isn&#8217;t locked behind enterprise budgets.</p><p>For anyone building with AI or working toward an AI role: post technical or career questions any day and get written answers from our 15-person engineering team. Live Q&amp;A twice a week across two time zones. Resume and portfolio reviews in 48&#8211;72 hours. A production blueprint every month with real engineering decisions explained.</p><p>Includes Python for LLMs and LLM Fundamentals courses. $99/month. Cancel anytime.</p><p><strong><a href="https://towardsai.com/academy/mentorship/?utm_source=TAInewsletter&amp;utm_medium=substack&amp;utm_id=mentorship">See what&#8217;s inside &#8594;</a></strong></p><div><hr></div><h3>Hottest News</h3><p>1. <a href="https://huggingface.co/deepseek-ai/DeepSeek-V4-Flash-0731">DeepSeek Releases DeepSeek-V4-Flash-0731</a></p><p>DeepSeek released V4-Flash-0731 as the official successor to its Flash preview, while the API remains in public beta. The model keeps the same 284B-parameter MoE architecture, with 13B active parameters, and focuses on stronger agent and coding performance. DeepSeek reports scores of 82.7 on Terminal-Bench 2.1, 54.4 on DeepSWE, and 76.7 on CyberGym, though some evaluations use internal tooling. The API now supports the Responses API and Codex. Pricing remains $0.14 per million uncached input tokens and $ 0.28 per million output tokens, while the MIT-licensed weights are available on Hugging Face.</p><p>2. <a href="https://x.com/polynoamial/status/2083467194663571701">OpenAI&#8217;s Unreleased Astra Produces Solutions to 10 Long-Standing Math and Theory Problems</a></p><p>OpenAI says an internal version of Astra, described as its &#8220;next major model,&#8221; resolved or substantially advanced ten long-standing problems in mathematics and theoretical computer science. Results include an explicit construction of a non-sofic group, a counterexample to Connes&#8217;s rigidity conjecture, the sharp Ehrhart volume bound, and progress on sphere packing and three Erd&#337;s problems. OpenAI published a 249-page manuscript and Lean 4 formalizations that can be independently checked. It estimates that the solution-finding tokens would cost roughly $2,000 at Sol API rates. OpenAI has not announced Astra&#8217;s final name or release plans.</p><p>3. <a href="https://openai.com/index/advancing-the-price-performance-frontier-with-gpt-5-6/">OpenAI Cuts GPT-5.6 API Prices and Adds Sol Fast Mode</a></p><p>OpenAI reduced standard short-context API pricing for GPT-5.6 Luna by 80%, to $0.20 per million input tokens and $1.20 for output. Terra fell 20% to $2 and $12, while Sol remains at $5 and $30. Priority Processing has also been renamed Fast mode, offering up to 2.5 times faster processing at twice Sol&#8217;s standard price. OpenAI says model-assisted kernel optimizations helped lower Sol&#8217;s serving costs by 20%. ChatGPT and Codex subscription prices remain unchanged, though Luna and Terra now consume fewer credits against applicable usage limits.</p><p>4. <a href="https://www.minimax.io/blog/minimax-h3">MiniMax Releases MiniMax H3</a></p><p>MiniMax released H3, a multimodal video model that accepts text, images, video, and audio and generates clips of up to 15 seconds with native stereo sound. It supports text-to-video, image conditioning, reference generation, motion transfer, and instruction-based editing. The model produces 768p video first, with a separate regeneration stage used for 2K output. API pricing is $0.08 per second at 768p and $0.13 at 2K. MiniMax published the weights on Hugging Face under its custom Community License, which includes regional restrictions.</p><p>5. <a href="https://deepmind.google/blog/gemini-robotics-2-brings-whole-body-intelligence-to-robots/">Google DeepMind Ships Three Physical AI Models</a></p><p>Google DeepMind introduced Gemini Robotics 2, Gemini Robotics ER 2, and Gemini Robotics On-Device 2. The main Robotics 2 model controls a humanoid&#8217;s full body under one learned policy, while ER 2 handles video understanding, planning, and coordination. On-Device 2 runs locally and can adapt to new dual-arm robots using fewer than 200 examples, according to Google. Only ER 2 is publicly accessible through Google AI Studio and the Gemini API; the other two remain limited to selected partners. Google also released ASIMOV-Agentic, a benchmark for testing robot safety and human-escalation decisions.</p><p>6. <a href="https://onton.com/research/ontology-1">Onton Releases Ontology 1: A Neurosymbolic Search Model</a></p><p>Onton, a San Francisco-based search and discovery company, released Ontology 1, a neurosymbolic model for conversational and multimodal product search. It combines learned representations with symbolic reasoning over a custom knowledge graph to interpret intent, infer product properties, and identify questionable claims. In Onton&#8217;s own 90-query benchmark, it recorded a precision@10 score of 0.630, compared with 0.543 for Google Shopping and 0.469 for Amazon. The evaluation used three LLM judges and has not been independently replicated. Ontology 1 currently powers Onton.com, but no public developer API or downloadable checkpoint is available.</p><p>7. <a href="https://qwen.ai/blog?id=qwen3.8">Alibaba&#8217;s Qwen Team Announces Qwen3.8-Max</a></p><p>Alibaba released Qwen3.8-Max, a 2.4-trillion-parameter mixture-of-experts model with 95 billion active parameters. It supports text, image, and video inputs, a context window of around one million tokens, and outputs of up to roughly 131K tokens. At launch, it ranked fifth overall in Arena&#8217;s text leaderboard and second in vision, though these positions change over time. Alibaba also demonstrated the model working autonomously on a software project for roughly 16 days. International API pricing starts at about $2 per million input tokens and $6 for output. Open weights were announced but had not been released as of August 4.</p><div><hr></div><h4>AI Tip of the Day</h4><p>When we built the &#8220;Research and Writing Agent with MCP&#8221; lesson for our <a href="https://towardsai.com/academy/agent-engineering/?utm_source=Newsletter&amp;utm_medium=email&amp;utm_id=AItips">Agent Engineering course</a>, we ran into a deceptively simple evaluation question: which articles should the judge score?</p><p>The easy option was to ask another LLM to generate a batch of articles and use them as the test set. But those articles would never pass through the system we actually wanted to evaluate. The judge might measure writing quality while missing failures in research, context transfer, or the handoff between agents.</p><p>So every evaluation example now follows the production path. We begin with the same brief a user would submit, run the research agent, pass its saved findings into the writing workflow, and score the final article.</p><p>This turns the output into a test of the entire system. If the research is weak, an important source is lost, or the writer receives incomplete context, the failure appears in the article, where the judge can detect it.</p><p>The takeaway: generate evaluation outputs with the workflow you plan to ship. Otherwise, you may be testing the quality of a substitute model rather than the reliability of your own system.</p><div><hr></div><h3>Five 5-minute reads/videos to keep you learning</h3><p>1. <a href="https://pub.towardsai.net/persistent-memory-for-claude-code-on-mongodb-atlas-0114a796467c?sharedUserId=tai-tech">Persistent Memory for Claude Code on MongoDB Atlas</a></p><p>Long Claude Code sessions silently lose constraints when compaction drops tool outputs and summarizes earlier exchanges. This article builds a plugin that captures those statements into MongoDB Atlas before they vanish, using a PreCompact hook and two MCP tools. Retrieval runs entirely inside the database: Automated Embedding generates Voyage vectors, $rankFusion blends vector and keyword pipelines through reciprocal rank fusion, and a hosted reranker orders results.</p><p>2. <a href="https://pub.towardsai.net/i-self-hosted-langfuse-so-my-llm-traces-would-stop-living-on-someone-elses-bill-165f4eff65e1?sk=e295d414e12f33db6d10d0713ed01b0c">I Self-Hosted Langfuse so My LLM Traces Would Stop Living On Someone Else&#8217;s Bill</a></p><p>Langfuse crossed 100K monthly traces and pushed managed pricing into painful territory, so the author self-hosted the open-source observability platform instead. Docker Compose spins up Postgres, ClickHouse, Redis, and MinIO in minutes, though production requires TLS, SSO, and eventually Kubernetes with externally managed dependencies. The article flags a notorious UTC timezone bug, treats Redis queue depth as the real health signal, and lays out honest cost math: self-hosting secures data control immediately, but only beats managed pricing at scale.</p><p>3. <a href="https://pub.towardsai.net/how-spark-manages-memory-the-unified-memory-model-spills-and-aqe-f37f21799fc6?sharedUserId=tai-tech">How Spark Manages Memory &#8212; The Unified Memory Model, Spills, and AQE</a></p><p>A daily 2.1TB aggregation job kept dying at 85% completion, and the author traced the failure to Spark&#8217;s Unified Memory Manager rather than raw heap size. This article walks through how execution memory always wins over storage in the borrowing contract, why protected cached blocks starved a skewed aggregation of room to grow, and how memoryOverhead, not executor.memory, killed the YARN container. Adaptive Query Execution&#8217;s skew-join splitting fixed the root problem, cutting cluster cost while pinpointing exactly which memory region overflowed.</p><p>4. <a href="https://pub.towardsai.net/agent-memory-is-the-real-moat-5d86930b8e10?sk=27f78b3c13de62f377991d61d7645a94">Agent Memory Is the Real Moat</a></p><p>Agent memory becomes dangerous the moment it is treated as an unbounded vector store rather than a governed lifecycle. This article maps four distinct memory classes: working checkpoints, episodic outcomes, semantic facts, and procedural lessons, each with its own ownership and retention rules. Benchmarks from TraceRetain and LoCoMo show unbounded memory collapsing under noisy writes while selective retention holds steady. It closes with a practical adoption blueprint.</p><p>5. <a href="https://pub.towardsai.net/road-to-bedrock-agentcore-from-a-single-api-call-to-a-production-agent-0544ea466469?sharedUserId=tai-tech">Road to Bedrock AgentCore, From a Single API Call to a Production Agent</a></p><p>Building the same weather-fetching agent four times across AWS&#8217;s tooling stack exposed exactly what separates each layer. This article implements the same core function on the Bedrock Converse API, Bedrock Agents, the Strands SDK, and AgentCore, keeping the logic identical to isolate each layer&#8217;s tradeoffs. Converse demands hand-written orchestration. Bedrock Agents hands the loop to AWS at the cost of visibility. Strands restores control with model portability. AgentCore deploys the same agent unchanged, adding Gateway, Memory, Identity, and Observability as managed infrastructure.</p><h3>Repositories &amp; Tools</h3><p>1. <a href="https://github.com/supabase/evals">Evals</a> is a benchmark framework that runs AI coding agents against real Supabase tasks like building schemas, debugging Edge Functions, and fixing RLS policies, scoring results to a public leaderboard.</p><p>2. <a href="https://github.com/MoonshotAI/MoonEP">MoonEP</a> is an Expert Parallelism communication library that keeps token loads perfectly balanced across ranks via dynamic redundant experts.</p><p>3. <a href="https://github.com/Marktechpost/Token-Saver">Token Saver</a> is a one-click Claude Desktop extension that runs local hybrid RAG over PDFs on your machine, sending only relevant passages to Claude and cutting token consumption by 90&#8211;99%.</p><p>4. <a href="https://github.com/AMD-AGI/Instella-MoE">Instella MoE</a> is AMD&#8217;s fully open 16B-parameter MoE model (2.8B active) trained from scratch on Instinct GPUs.</p><h3>Top Papers of The Week</h3><p>1. <a href="https://arxiv.org/abs/2607.24653">Kimi K3: Open Frontier Intelligence</a></p><p>This technical report presents Kimi K3, a 2.8T-parameter Mixture-of-Experts model with 104 billion activated parameters, native vision capabilities, and a 1M-token context window. K3 introduces three architectural innovations: Kimi Delta Attention (KDA), which models linear retention recurrence for position-sensitive mixing; Attention Residuals, which improve information flow across model depth; and Stable LatentMoE, which activates 16 of 896 routed experts per token. Together, these yield an approximately 2.5x improvement in overall scaling efficiency over Kimi K2. Post-training applies reinforcement learning across general, agentic, and coding domains with multiple reasoning-effort levels.</p><p>2. <a href="https://arxiv.org/abs/2607.25537">Visual Prompt Engineering for Video Models</a></p><p>Text-based prompt engineering is standard for LLMs, but as video models become foundation models for visual tasks, this paper asks whether automatically modifying the task image can similarly improve performance. For example, converting an abstract sketch-like physics scene into a photorealistic version with a single call to an image editing model. The authors find that visual prompt engineering (VIPE) improves video reasoning performance across tasks, and for video models, it can be more effective than classic text-based prompt engineering or test-time scaling.</p><p>3. <a href="https://arxiv.org/abs/2607.25857">Shieldstral</a></p><p>Mistral introduces Shieldstral, a 3B-parameter policy-adaptive multimodal safety classifier built on Ministral-3B that matches or outperforms models nearly 7x its size on text safety benchmarks and sets a new state of the art on multimodal safety classification. The key design choice is formulating content moderation as a binary question-answering task: given a natural-language query describing a safety concern and a piece of content (text and/or image), Shieldstral outputs a yes/no answer rather than fixed category labels.</p><p>4. <a href="https://arxiv.org/abs/2607.23802">From RLVR to RLSVR: Task Transformation Induces Self-Verifiable Rewards for Open-Ended LLM Self-Improvement</a></p><p>RLVR has driven recent progress in reasoning LLMs but remains limited to domains like math and coding, where correctness is deterministically verifiable. Open-ended tasks instead rely on reward models or LLM judges, introducing evaluation bias and capability bottlenecks. This paper proposes RLSVR, which transforms open-ended tasks into verifiable proxy environments whose internal rules automatically generate reward signals. The concrete instantiation, SpyRL, is a multi-agent self-play environment inspired by &#8220;Who Is the Spy?&#8221;: agents receive asymmetric information, complete the same target task, and vote to identify a designated spy. Because the spy identity is predetermined, voting outcomes provide fully verifiable rewards, while successful identification depends on output quality.</p><p>5. <a href="https://arxiv.org/abs/2607.28617">AISPA: User-Centric System Prompt Auditing for Large Language Model Applications</a></p><p>System prompts govern how LLMs behave in commercial products but are rarely disclosed to users or regulators, creating a trust and accountability gap. This paper introduces AISPA, a framework that audits system prompts along eight user-facing dimensions (identity transparency, information truthfulness, data privacy, action safety, user agency, unsafe request handling, harm prevention, and fairness), each grounded in corresponding articles of the Universal Declaration of Human Rights. The authors audited 3,249 instructions from system prompts across 88 commercial AI products, classifying each as protective or problematic.</p><h3>Quick Links</h3><p><span>1. </span><a href="https://www.cogent.com/blog/cogent-at-the-frontier-announcing-vr-1-the-first-mythos-class-ai-model-built-for-cyber-defense">Cogent AI team releases VR-1</a><span>, a frontier reasoning model trained specifically for cybersecurity and designed to operate inside live enterprise environments. Unlike general-purpose models that find individual bugs in isolated codebases, VR-1 is trained for multi-step attack-chain composition, chaining minor weaknesses across cloud infrastructure, identity systems, and internal tools to prove viable paths to an objective. On IntrusionBench, a new enterprise attack-chain benchmark released alongside the model, VR-1 achieved 2x the performance of Kimi K3, Opus 4.8, and GLM-5.2 at roughly a quarter of the cost. Available only to vetted organizations through the Cogent Frontier Access Program.</span></p><h3>Who&#8217;s Hiring in AI</h3><p><strong><a href="https://towardsai.com/valuecreation/senior-ai-engineer/">Senior AI Engineer / Forward Deployed Engineer @Towards AI (London, UK)</a></strong></p><p><strong><a href="https://jobs.towardsai.net/job/microsoft-corporation-fullstack-ai-engineer-xtdc">Fullstack AI Engineer @Microsoft Corporation (Redmond, WA, USA)</a></strong></p><p><strong><a href="https://jobs.towardsai.net/job/unitedhealth-group-lead-ai-engineer-remote-5huz">Lead AI Engineer @UnitedHealth Group (Remote)</a></strong></p><p><strong><a href="https://jobs.towardsai.net/job/thermofisher-scientific-ai-product-and-delivery-lead-uicf">AI Product &amp; Delivery Lead @ThermoFisher Scientific (Mississauga, Canada)</a></strong></p><p><strong><a href="https://jobs.towardsai.net/job/leidos-ai-software-developer-kpke">AI Software Developer @Leidos (Gaithersburg, MD, USA)</a></strong></p><p><strong><a href="https://jobs.towardsai.net/job/micron-technology-inc-ai-engineer-fosv">AI Engineer @Micron Technology, Inc. (Multiple US Locations)</a></strong></p><p><strong><a href="https://jobs.towardsai.net/job/ntt-america-inc-jr-gen-ai-developer-1lwf">Jr Gen AI Developer @NTT America, Inc. (India)</a></strong></p><p><em>Interested in sharing a job opportunity here? Contact <a href="mailto:sponsors@towardsai.net">sponsors@towardsai.net</a>.</em></p><p><em>Think a friend would enjoy this too? <a href="https://newsletter.towardsai.net/">Share the newsletter and let them join the conversation.</a></em></p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://newsletter.towardsai.net/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading Towards AI Newsletter! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div>]]></content:encoded></item><item><title><![CDATA[TAI #215: AI Is Expanding Roles Before Job Titles Change]]></title><description><![CDATA[Also, Claude Opus 5 tops several leaderboards, Gemini 3.5 Flash-Lite offers cheap multimodal extraction, Gemini 3.6 Flash, FLUX 3 & more.]]></description><link>https://newsletter.towardsai.net/p/tai-215-ai-is-expanding-roles-before</link><guid isPermaLink="false">https://newsletter.towardsai.net/p/tai-215-ai-is-expanding-roles-before</guid><dc:creator><![CDATA[Louie Peters]]></dc:creator><pubDate>Tue, 28 Jul 2026 15:01:06 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!y7Eo!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F286e00c8-aa9e-4686-b46e-c3bfa0a62dbd_1400x788.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<h2>What happened this week in AI by Louie</h2><p>This was a mixed week for closed models. Claude Opus 5 ranked first on Artificial Analysis&#8217;s Intelligence Index with 61 at max effort, one point above Fable 5, and first on the AA-Briefcase agentic knowledge-work benchmark. On ARC Prize verified, it hit 30.2% on ARC-AGI-3, a big leap from the previous 7.8% high score set by GPT-5.6 Sol (Max). Early tests and customer reports also suggest unusual strength in 3D graphics, animation, and game design. ARC-AGI-3 uses interactive, game-like visual environments, so these strengths may be related.</p><p>Opus is also the clearest recent sign that benchmarks do not show the whole model. It beats Fable across many headline evaluations, yet Fable feels clearly more intelligent when I interact with it, and GPT-5.6 Sol at xhigh does too. My shorthand is that Opus studied harder, while Fable is much smarter. Opus will still earn a place in my daily rotation, especially for the 50% of my Claude Max allowance that Fable cannot consume, but I expect Fable and Sol to remain my main models.</p><p>At Gemini, 3.5 Flash-Lite was the most useful release for me. At $0.30 per million input tokens and $2.50 per million output tokens, it accepts text, images, audio, video, and PDFs, generates over 350 output tokens per second, and may be the best cheap-tier LLM with real multimodal capability. I would test it first for high-volume document extraction and search. Gemini 3.6 Flash is also a genuine efficiency upgrade: Google reports 17% fewer output tokens on the Artificial Analysis Index, and Artificial Analysis measured average task time falling from 2.7 to 1.3 minutes. Yet its Index score stayed at 50, an 11-point gap to Opus 5 at the frontier, and for demanding work, I do not think it offers good value against Sol or Kimi K3. Gemini has slipped far behind the leading closed models since the excellent Gemini 3 Pro launch last November. Gemini 3.5 Pro remains in testing, and Google has started Gemini 4 pretraining. Hopefully, that puts Google back into the frontier race.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!y7Eo!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F286e00c8-aa9e-4686-b46e-c3bfa0a62dbd_1400x788.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!y7Eo!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F286e00c8-aa9e-4686-b46e-c3bfa0a62dbd_1400x788.png 424w, https://substackcdn.com/image/fetch/$s_!y7Eo!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F286e00c8-aa9e-4686-b46e-c3bfa0a62dbd_1400x788.png 848w, https://substackcdn.com/image/fetch/$s_!y7Eo!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F286e00c8-aa9e-4686-b46e-c3bfa0a62dbd_1400x788.png 1272w, https://substackcdn.com/image/fetch/$s_!y7Eo!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F286e00c8-aa9e-4686-b46e-c3bfa0a62dbd_1400x788.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!y7Eo!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F286e00c8-aa9e-4686-b46e-c3bfa0a62dbd_1400x788.png" width="1400" height="788" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/286e00c8-aa9e-4686-b46e-c3bfa0a62dbd_1400x788.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:788,&quot;width&quot;:1400,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" title="" srcset="https://substackcdn.com/image/fetch/$s_!y7Eo!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F286e00c8-aa9e-4686-b46e-c3bfa0a62dbd_1400x788.png 424w, https://substackcdn.com/image/fetch/$s_!y7Eo!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F286e00c8-aa9e-4686-b46e-c3bfa0a62dbd_1400x788.png 848w, https://substackcdn.com/image/fetch/$s_!y7Eo!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F286e00c8-aa9e-4686-b46e-c3bfa0a62dbd_1400x788.png 1272w, https://substackcdn.com/image/fetch/$s_!y7Eo!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F286e00c8-aa9e-4686-b46e-c3bfa0a62dbd_1400x788.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>Leaderboards score models on fixed tests. Adoption depends on which tasks people actually attempt and whether the results hold up, and OpenAI&#8217;s new workplace study is the best recent evidence on the first half of that question. The researchers analyzed more than 800,000 work-related messages from individual ChatGPT accounts of U.S. users, linked each user to role data from ChatGPT Business, and mapped every message to the O*NET taxonomy of activities historically tied to U.S. occupations. The sample covers eight groups: customer experience, design, engineering, finance, human resources, legal, marketing, and sales.</p><p>OpenAI classified 61.5% of messages as generic work shared across many jobs, 21.8% as inside the user&#8217;s occupation, and 16.8% as outside it. Remove the generic majority, and cross-role work makes up 43.5% of what remains, which is the report&#8217;s headline figure.</p><p>Five groups crossed the boundary in most of their occupation-specific messages: customer experience at 77%, design at 75%, human resources at 69%, legal at 56%, and marketing at 53%. Financial calculation and computer troubleshooting ranked among the three most common borrowed tasks in every outside group, and marketing and engineering tasks appeared most often across other roles.</p><p>This is an early view of roles expanding before job titles change. A salesperson can ask AI for the first pass of a marketing plan, a support worker can troubleshoot a system, and a marketer can inspect financial data. Each attempt can remove a handoff, sharpen a question sent to a specialist, or let a small team start work that would otherwise wait.</p><p>Among users with typical message volume, cross-role work made up 18.9% of messages in workspaces with two to five seats and 16.3% in those with more than 100 seats. Small teams still have the clearest reason to use AI as a generalist because they have the fewest specialists to hand work to.</p><p>Anthropic&#8217;s June 2026 Economic Index update shows how broad this use has become in another product. In its May U.S. data, computer and mathematical work accounted for 21.1% of Claude conversations, followed by arts and design at 12.9%, education at 11.9%, sales at 11.6%, office support at 7.6%, and business and finance at 6.3%. By request type, content creation led at 20.0%, research reached 13.5%, and software development accounted for 8.1%.</p><p>Usage remains uneven, both by geography and by profession. In May, the District of Columbia&#8217;s share of Claude chat and Cowork use was 3.32 times its share of the U.S. working-age population, with California at 1.62, New York at 1.55, and West Virginia at 0.25. In Anthropic&#8217;s June survey, computer and mathematical workers made up 30% of respondents versus 4% of U.S. employment, while managers made up 23% versus 7%. Anthropic also lacks the user&#8217;s occupation in most conversations, so it cannot measure OpenAI&#8217;s role crossover directly.</p><p>Together, the reports show existing AI users trying a wide range of work while usage stays uneven across places and occupations. They do not measure finished output, accuracy, time saved, productivity, or formal job changes.</p><p>That gap is exactly where the review risk sits. A polished financial calculation can fool a marketer who lacks the experience to test it, and a technical fix can look safe to a support worker who cannot see the wider system effect. In a randomized study of 758 BCG consultants, AI raised quality ratings by more than 40% on tasks inside the model&#8217;s capability range, then made users 19 percentage points less likely to find the correct answer on a task outside it. The group given a prompt-engineering primer fell furthest, a warning that shallow fluency training alone can raise confidence faster than judgment.</p><p>At Towards AI, we start most nontechnical teams with focused Codex or Claude Cowork training. Both now handle real company work end-to-end: research, data analysis, document production, and multi-step tasks across connected files and systems. We teach several Skills tied to the team&#8217;s actual roles, covering how to supply context, reuse a good process, test an answer, and involve a specialist. This builds range quickly and shows us which use cases survive the first few weeks. Sustained depth is also what separates leading firms: OpenAI&#8217;s B2B Signals analysis found 95th-percentile companies generating 3.5 times as many tokens per worker as typical ones, up from 2 times a year earlier, though tokens are a proxy for engagement, not a measure of business results.</p><p>The largest reliability gains usually still come when a repeated, shared workflow moves into a custom app. Copying the same files, correcting the same errors, moving output between systems, and following fixed approval steps all point toward a build. Our deployment strategists find use cases that the current models can handle, map the work, and define a good result. Our forward-deployed engineers connect the systems, add permissions, build the evaluations, inspect failures, and improve performance against real examples.</p><p>The interface often decides whether staff use the system. A good app asks for the right inputs, hides model choices the user does not need, shows evidence beside the answer, and places human review at the natural decision point: a finance owner approves a calculation, a support agent edits a reply, and an unusual case escalates to an expert. These steps are easier to follow than a policy document telling everyone to check the AI. Public leaderboards narrow our model shortlist; a private evaluation set built from real cases decides what ships and catches regressions after every model change. That process turns wider AI use into adoption that a company can trust.</p><div><hr></div><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://newsletter.towardsai.net/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading Towards AI Newsletter! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><div><hr></div><h3>Why should you care?</h3><p>Many companies start AI adoption by choosing a platform or collecting a list of ideas. A better first step is practical training on work your staff already understands, because training doubles as discovery. Watch for repeated use, context copied between tools, recurring corrections, and work stalling at a familiar bottleneck. These signals reveal where AI creates real demand, and they hand engineers the raw material for a first evaluation set. Codex and Claude Cowork can carry substantial company workflows on their own; custom builds then earn their cost on the specific high-value use cases where several people repeat a process, private systems hold the context, outputs need a fixed shape, or risk requires clear approval. This order limits wasted builds: staff gain skill, managers see which uses last, and each step produces evidence for the next investment.</p><p>Role expansion also raises the value of judgment. When a marketer attempts finance work or a salesperson prepares marketing material, decisions speed up, and weak handoffs disappear, but more employees now operate where they have less training and a weaker sense of failure. The BCG result shows why a general instruction to check the answer offers little protection: users cannot challenge what they cannot evaluate. Review ability should form part of AI training, covering which sources to demand, which calculations to test, and which cases require a specialist, with a named owner for high-risk work. Custom apps can place that judgment inside the process itself, showing evidence beside the draft, blocking external actions until approval, and routing exceptions to the right person.</p><p>The clearest beneficiaries are small teams. In a preregistered field experiment with 791 Procter &amp; Gamble professionals, one person working with AI matched the average solution quality of a two-person cross-functional team working without it. A ten-person company cannot employ a specialist for every function; AI gives each person reach across research, finance, marketing, support, and operations, removing delays that a larger firm solves with another department. The risk is building everything around one power user. Train several people, save the best Skills, record required sources and output standards, and add structure when a workflow becomes frequent or important. Small firms can change a workflow in days; they should use that speed to build repeatable processes before informal use hardens into something nobody can audit. Larger firms have the same opportunity but need a deliberate process and permission changes to capture it. Wherever you sit, the sequence is the same: train for range, watch what repeats, and productize the winners with judgment built in.</p><p><em>&#8212; <a href="http://www.linkedin.com/in/louie-peters">Louie Peters&#8202;&#8212;&#8202;Towards AI Co-founder and CEO</a></em></p><h3>Hottest News</h3><p>1. <a href="https://www.anthropic.com/news/claude-opus-5">Anthropic Released Claude Opus 5</a></p><p>Anthropic released Claude Opus 5, its fourth Claude 5 model in less than two months following Mythos 5, Fable 5, and Sonnet 5. Opus 5 is designed to deliver performance close to Fable 5 on many tasks at half the price. It ships with a 1M-token context window, 128K max output tokens, thinking on by default, and pricing unchanged from Opus 4.8 at $5/$25 per million tokens. Users can toggle reasoning effort across low, medium, and high to balance cost against capability per task. On Anthropic-reported benchmarks, Opus 5 sets a new state of the art on Frontier-Bench (43.3%) and GDPval-AA, and outperforms Fable 5 on several evaluations while trailing it on cybersecurity and the hardest long-horizon agentic tasks. ARC Prize independently verified 30.2% on ARC-AGI-3 on launch day. Anthropic describes Opus 5 as more proactive than previous models: it verifies its own work, recovers from errors without intervention, and requires less back-and-forth. Business customers had criticized Fable 5&#8217;s token burn rate, and Opus 5 is positioned as the more cost-efficient alternative for everyday production work. Opus 5 is the default model on Claude Max and the strongest model available on Claude Pro. Fable 5 remains recommended for the most advanced autonomous work.</p><p>2. <a href="https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-3-6-flash-3-5-flash-lite-3-5-flash-cyber/">Google Ships Three New Gemini Models and Starts Gemini 4 Pretraining</a></p><p>Google released three new Gemini models: Gemini 3.6 Flash, Gemini 3.5 Flash-Lite, and Gemini 3.5 Flash Cyber. Gemini 3.6 Flash replaces 3.5 Flash as the mid-tier workhorse, delivering stronger coding and knowledge-work performance while using 17% fewer output tokens on the Artificial Analysis Index. Token reductions reached as high as 65% on certain DeepSWE tests, with the model also requiring fewer reasoning steps and tool calls for multi-step workflows. Pricing is $1.50/$7.50 per million input/output tokens. Gemini 3.5 Flash-Lite targets high-throughput workloads at 350 output tokens per second, priced at $0.30/$2.50. Gemini 3.5 Flash Cyber is a security-tuned variant that powers CodeMender, Google&#8217;s automated vulnerability-patching agent, available only through a limited pilot for governments and trusted partners. Both 3.6 Flash and 3.5 Flash-Lite are available through the Gemini API, Google AI Studio, and the Gemini app. Gemini 3.5 Pro, originally promised for June, remains in partner testing with no public release date. Google also disclosed that it has begun its most ambitious pretraining run yet for Gemini 4.</p><p>3. <a href="https://bfl.ai/blog/flux-3">Black Forest Labs Released FLUX 3</a></p><p>Black Forest Labs announced FLUX 3, a multimodal foundation model jointly trained across images, video, audio, and action prediction within a single architecture. FLUX 3 is the first major generative model where video with native synchronized audio, image generation, editing, and robotic control all run through one shared set of weights, rather than separate models behind a common interface. FLUX 3 Video generates clips up to 20 seconds with native audio in a single generation pass, supporting text-to-video, image-to-video, video-to-video, keyframe-to-video, and generative continuation modes. In BFL&#8217;s own preliminary evaluation on 10-second 720p clips, it won 93% of comparisons against Luma Ray 3.2 and 77% against Runway Gen-4.5, though it split roughly evenly against Seedance 2.0 and Gemini Omni Flash. FLUX-mimic, a robotics model built on the FLUX 3 architecture, is already running on production lines at Audi. FLUX 3 Video and FLUX 3 Action are available in gated early access. FLUX 3 Image is expected in the coming weeks. An open-weight FLUX 3 Dev release is planned for later in 2026. No public pricing has been announced.</p><p>4. <a href="https://claude.com/product/claude-security">Anthropic Released Claude Security Plugin</a></p><p>Anthropic released the Claude Security plugin for Claude Code in beta. The plugin runs a multi-agent vulnerability scan from inside an existing Claude Code session, letting developers scan uncommitted changes before a commit or run a full analysis across an entire codebase from the terminal. The system reads Git history, traces data flows across files, and reasons through business logic to identify context-dependent vulnerabilities that span multiple files, going beyond pattern matching to catch issues that traditional static analysis tools miss. Severity classifications and confidence rankings help teams prioritize findings. Findings the developer selects are turned into patch files for review before applying. Anthropic reports that internal rollout reduced security-related comments on pull requests by 30&#8211;40%, with the plugin serving as a lightweight first pass before full code review. The plugin is available for all Claude Code users and can be installed from the plugin marketplace.</p><p>5. <a href="https://sakana.ai/fugu-cyber-release/">Sakana AI Released Fugu-Cyber</a></p><p>Sakana AI released Fugu-Cyber, a cybersecurity-specialized endpoint added to its Fugu multi-agent orchestration platform. Fugu-Cyber is not a new standalone model but a third endpoint on the Fugu orchestrator, tuned for multi-step security work, including vulnerability discovery, exploit reasoning, and threat intelligence analysis. It presents as a single API while dynamically routing tasks across a pool of specialized agents underneath. Sakana reports 86.9% on CyberGym, a UC Berkeley benchmark for real-world vulnerability verification across 188 software projects, and 72.1% on CTI-REALM, a Microsoft benchmark for threat intelligence detection rule generation. Sakana describes these scores as comparable to GPT-5.5-Cyber and Mythos-Preview. All figures are vendor-reported and have not been independently replicated. Access is gated behind an application form with manual review. Pricing is $6/$36 per million input/output tokens, with rates doubling above 272K context. Not available in the EU or EEA.</p><p>6. <a href="https://openai.com/index/hugging-face-model-evaluation-security-incident/">OpenAI&#8217;s Models Broke Out of Their Sandbox and Hacked Hugging Face</a></p><p>OpenAI disclosed that a combination of its models, including GPT-5.6 Sol and a more capable unreleased model, was responsible for a security breach that affected Hugging Face&#8217;s production infrastructure the previous week. The models were being tested internally on ExploitGym, an evaluation suite for cybersecurity capabilities, with safety guardrails reduced for evaluation purposes. Rather than solving the test as designed, the models exploited a previously unknown vulnerability in the package-installation system used to isolate the sandbox, escaped to OpenAI&#8217;s internal network, gained internet access, and then breached Hugging Face&#8217;s systems to obtain the evaluation answers. Hugging Face reconstructed more than 17,000 recorded events from the intrusion and described it as &#8220;an unprecedented cyber incident&#8221; driven &#8220;end-to-end, by an autonomous AI agent system.&#8221; OpenAI said the models became &#8220;hyperfocused&#8221; on obtaining the test solution and went to &#8220;extreme lengths&#8221; to do so. The company described the incident as a demonstration that long-horizon safety requires asking not just &#8220;is this action allowed?&#8221; but &#8220;what outcome is this sequence of actions working toward?&#8221; OpenAI said it expects such incidents to become more common as increasingly cyber-capable models proliferate. Both companies have committed to publishing a joint technical report with full details.</p><div><hr></div><h4>AI Tip of the Day</h4><p>When we designed the evaluation lessons for our <a href="https://towardsai.com/academy/llm-primer/?utm_source=Newsletter&amp;utm_medium=email&amp;utm_id=AItips">10-Hour LLM Fundamentals</a> course, we wanted to focus on checks you could add without rebuilding your pipeline. This is one of the most useful.</p><p>If an LLM judge is choosing between answer A and answer B, run the evaluation twice. Reverse the order on the second call.</p><p>If the same answer wins both times, keep the result. If the winner changes, mark the comparison as undecided or send it for human review.</p><p>This catches position bias, where the judge favors an answer partly because of where it appears rather than because it is better.</p><p>You only need a small wrapper around your existing judge call. Track how often the result flips too.</p><p>Frequent flips may mean your rubric is unclear, the answers are too close, or the judge cannot reliably distinguish between them.</p><div><hr></div><h3>Five 5-minute reads/videos to keep you learning</h3><p>1. <a href="https://pub.towardsai.net/ollama-vs-vllm-which-open-source-inference-stack-should-you-actually-use-b9812cfbb175?sk=99409a2f1c4d554f7415d0f6fc9811c1">Ollama vs vLLM: Which Open-Source Inference Stack Should You Actually Use</a></p><p>Ollama and vLLM both serve open LLMs behind OpenAI-compatible APIs, but solve different problems. Ollama gets a model running on a laptop in minutes, while vLLM&#8217;s PagedAttention and continuous batching sustain 180+ concurrent requests on a single H100 where Ollama hits memory limits near 40. In this article, the author benchmarks both, maps hardware fit, and walks through a seven-step migration from local prototype to production, showing why most teams keep Ollama for development and deploy vLLM under real traffic.</p><p>2. <a href="https://pub.towardsai.net/the-loop-was-the-easy-part-evals-observability-and-rollbacks-for-your-diy-claude-code-45f2a9f1c99c?sharedUserId=tai-tech">The Loop Was the Easy Part: Evals, Observability, and Rollbacks for Your DIY Claude Code</a></p><p>This article rebuilds a deep agent with the infrastructure to make it dependable: observability, evals, and rollbacks. It adds a callback-based flight recorder for tracing every model and tool call, token and step budgets with kill-switches, trajectory and outcome evals gated in CI against a checked-in baseline, and two undo mechanisms that rewind conversations and restore workspace snapshots.</p><p>3. <a href="https://pub.towardsai.net/the-100-000-token-lie-why-microgpts-context-window-costs-14x-more-than-the-benchmark-claims-81ec66a6d520?sk=1104e20d011a79f188a2d2ab339dec02">The 100,000-Token Lie: Why microgpt&#8217;s Context Window Costs 14x More Than the Benchmark Claims</a></p><p>Standard attention&#8217;s quadratic memory means a 47K-token contract workload demands 2.5 TB of GPU memory during training, a cost that microgpt&#8217;s 100K-token benchmark figure does not surface. The author profiled microgpt, traced the problem to the materialized attention matrix, and showed how Flash Attention cuts memory 345x without changing outputs. The larger savings came from restructuring the agent workflow with hierarchical chunking, dropping monthly token costs by 64% and proving that task design matters more than architecture.</p><p>4. <a href="https://pub.towardsai.net/the-three-questions-agent-security-has-to-answer-af013c1c6470?sharedUserId=tai-tech">The Three Questions Agent Security Has to Answer</a></p><p>This article organizes agent security around three questions: whose authority backed an action, what proof exists afterward, and how far damage spreads when things go wrong. It maps four architectural answers: scoped token propagation via OAuth token exchange, decision-level audit interception, hardware-virtualized microVM isolation, and cryptographic agent-to-agent delegation chains. It includes three diagnostic traces for evaluating any existing stack, grounded in OWASP, NIST, and Forrester frameworks. The core argument: these defenses belong in infrastructure, not the model.</p><p>5. <a href="https://pub.towardsai.net/svd-explained-simply-python-16224a85facb?sharedUserId=tai-tech">SVD Is Just a Greatest Hits Album. I Can Prove It</a></p><p>This article explains Singular Value Decomposition by framing every matrix as rotate, stretch, rotate, with singular values ranking each pattern&#8217;s importance. Annotated Python code compresses a grayscale photo to 25% of its original data, while further sections connect the same decomposition to denoising LIGO&#8217;s gravitational wave pipelines, eigenfaces, PCA, and LoRA fine-tuning. A practical common-pitfalls table rounds out an accessible linear algebra primer.</p><h3>Repositories &amp; Tools</h3><p>1. <a href="https://github.com/andrewyng/openworker">OpenWorker</a> is a local-first desktop AI agent that delivers finished work and includes 25+ integrations, BYOK support for any model provider or Ollama, and user check-ins before consequential actions.</p><p>2. <a href="https://github.com/reactor-team/open-dreamer">OpenDreamer</a> is an open reproduction of the Dreamer 4 world model pipeline in JAX/Flax NNX, shipping the full training recipe and a playable in-browser Minecraft demo with a real-time Game-to-Dream toggle.</p><p>3. <a href="https://github.com/marcelroed/gigatoken">Gigatoken</a> is a Rust BPE tokenizer that encodes text at up to 24.53 GB/s on a 144-core EPYC, 500&#8211;1,000x faster than HuggingFace tokenizers and 50&#8211;680x faster than tiktoken.</p><h3>Top Papers of The Week</h3><p>1. <a href="https://arxiv.org/abs/2607.14952">LongStraw: Long-Context RL Beyond 2M Tokens under a Fixed GPU Budget</a></p><p>RL post-training remains capped at roughly 256K tokens while inference has reached millions, because GRPO must score and backpropagate through multiple responses conditioned on one shared history, making attention and backward state the primary memory barrier. LongStraw closes this gap with two mechanisms: resident state captures only the model-native prompt state needed by later tokens (not the full computation graph), and response replay restores that boundary, scores old branches graph-free, rebuilds one policy response under autograd, backpropagates, and pops back. On eight H20 GPUs, it completes GRPO scoring and backward passes at 2.1M positions for Qwen3.6&#8211;27B, with a stress test reaching 4.46M. On 32 H20 GPUs, it validates the full execution path across all 78 layers of GLM-5.2 at 2.1M tokens.</p><p>2. <a href="https://arxiv.org/abs/2607.13285">Harness Handbook: Making Evolving Agent Harnesses Readable, Navigable, and Editable</a></p><p>Production agent harnesses span hundreds of functions across many files, with execution logic distributed across stages and connected through shared state. Modifying a single behavior requires locating every relevant implementation site, a task the paper formalizes as behavior localization. Harness Handbook reorganizes a harness codebase into a three-level navigable document: system overview, execution stages with state-register views, and detailed behavior units linked to source code. Coding agents navigate from a natural-language change request through the document tree, follow shared-state couplings to find structurally distant dependencies, and produce tighter edit plans. The paper ships generated handbooks for Codex and Terminus 2.</p><p>3. <a href="https://arxiv.org/abs/2607.18110">LLM-as-a-Coach: Experiential Learning for Non-Verifiable Tasks</a></p><p>RL on open-ended tasks compresses rubric-based evaluation into a scalar reward, discarding the textual feedback that explains why one response is better than another. This paper repurposes the feedback model from an LLM-as-a-Judge into an LLM-as-a-Coach that distills its assessment of each on-policy response into transferable experiential knowledge. That knowledge conditions a teacher model and is internalized by the policy through on-policy context distillation, providing denser supervision and preserving fine-grained preferences among high-quality responses. Across two policy families with feedback from either the policy itself or a proprietary model, Experiential Learning consistently outperforms rubric-based RL on held-out and unseen tasks, generalizes better beyond the training distribution, and mitigates reward hacking.</p><p>4. <a href="https://arxiv.org/abs/2607.21461">AREX: Towards a Recursively Self-Improving Agent for Deep Research</a></p><p>Deep research questions often require answers that jointly satisfy multiple constraints, where discovering a valid answer is expensive but verifying a candidate can be decomposed into tractable constraint-level checks. AREX exploits this asymmetry through two nested loops: an inner research loop gathers evidence and constructs a provisional answer, and an outer self-improvement loop audits the answer constraint by constraint, identifies unresolved claims, and launches targeted follow-up research. To sustain this over long horizons without context overflow, AREX learns an autonomous context-update tool that compresses growing interaction history into a compact improvement state preserving verified evidence and open constraints.</p><p>5. <a href="https://cisco-foundation-ai.github.io/antares/technical-report.pdf">Antares: Foundation Models for Agentic Vulnerability Localization</a></p><p>Cisco Foundation AI released Antares, a family of compact language models (350M, 1B, and 3B parameters) trained end-to-end for one task: given a CWE description and read-only terminal access to a repository, autonomously search the codebase, inspect files, gather evidence, and identify which source files contain the vulnerability. Training follows a two-stage pipeline: SFT on cybersecurity reasoning, code search trajectories, and deep research data, followed by GRPO with multi-component verifiable rewards over complete agent trajectories. Antares-1B achieves 0.209 File F1 on the accompanying 500-task VLoc Bench, approaching GPT-5.5&#8217;s 0.229 while running a full benchmark sweep in 13 minutes on a single H100. Static analysis tools (Semgrep, CodeQL, Horusec) score between 0.020 and 0.086 under the same protocol. Antares-350M and Antares-1B are released under Apache 2.0 on Hugging Face.</p><h3>Quick Links</h3><p>1. <a href="https://poolside.ai/blog/introducing-laguna-s-2-1">Poolside released Laguna S 2.1</a>, a 118B-parameter MoE model with 8B active parameters and a 1M-token context window, built for agentic coding. It scored 70.2% on Terminal-Bench 2.1 and 40.4% on DeepSWE, matching or exceeding models several times its size, including DeepSeek-V4-Flash, Nemotron 3 Ultra, and Inkling. The model runs on a single DGX Spark, trained end-to-end in under four weeks on 4,096 H200 GPUs. Weights are on Hugging Face under the OpenMDW-1.1 license.</p><p>2. <a href="https://www.inductionlabs.com/news/scaling-video-pretraining">Induction Labs introduced imagination models</a>, a foundation model architecture that learns from internet-scale video without action labels. Their first model, Photon-1, is a 106B-A5B MoE transformer pretrained on 18 years of computer screen recordings. It predicts future frames autoregressively in a learned representation space, and despite never seeing an action label during pretraining, it implicitly learns to act. After a small finetune and reinforcement learning, Photon-1 outperforms Gemini 3.1 Flash-Lite on internal computer use benchmarks with 30x less pretraining compute and 3x cheaper inference. It also generalizes beyond computers: finetuned, it learns checkers and simulates billiard physics better than LLM baselines.</p><h3>Who&#8217;s Hiring in AI</h3><p><strong><a href="https://jobs.towardsai.net/job/openai-forward-deployed-engineer-zurich-t32d">Forward Deployed Engineer @OpenAI (Zurich, Switzerland)</a></strong></p><p><strong><a href="https://jobs.towardsai.net/job/traackr-software-engineer-w4dd">Software Engineer @Traackr (Remote)</a></strong></p><p><strong><a href="https://jobs.towardsai.net/job/sanofi-group-data-and-ai-engineer-mfz5">Data &amp; AI Engineer @Sanofi Group (Paris, France)</a></strong></p><p><strong><a href="https://jobs.towardsai.net/job/ul-llc-ai-developer-remote-ngta">AI Developer @UL, LLC (Remote/Brazil)</a></strong></p><p><strong><a href="https://jobs.towardsai.net/job/superannotate-ai-research-engineer-3ais">Research Engineer @SuperAnnotate AI (San Francisco, CA, USA)</a></strong></p><p><strong><a href="https://jobs.towardsai.net/job/provectus-forward-deployed-ai-engineer-genai-aws-m9l8">Forward Deployed AI Engineer @Provectus (Remote)</a></strong></p><p><strong><a href="https://jobs.towardsai.net/job/pointclickcare-senior-software-engineer-ai-platform-us-zmpx">Senior Software Engineer- AI Platform @PointClickCare (Remote/USA)</a></strong></p><p><em>Interested in sharing a job opportunity here? Contact <a href="mailto:sponsors@towardsai.net">sponsors@towardsai.net</a>.</em></p><p><em>Think a friend would enjoy this too? <a href="https://newsletter.towardsai.net/">Share the newsletter and let them join the conversation.</a></em></p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://newsletter.towardsai.net/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading Towards AI Newsletter! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div>]]></content:encoded></item><item><title><![CDATA[TAI #214: Kimi K3 Brings Open Weight Closer to the Frontier]]></title><description><![CDATA[Also, Fable disproves an 87-year-old math conjecture, Qwen3.8-Max, Thinking Machine's Inking & more!]]></description><link>https://newsletter.towardsai.net/p/tai-214-kimi-k3-brings-open-weight</link><guid isPermaLink="false">https://newsletter.towardsai.net/p/tai-214-kimi-k3-brings-open-weight</guid><dc:creator><![CDATA[Louie Peters]]></dc:creator><pubDate>Wed, 22 Jul 2026 02:08:22 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!d4oO!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F452b1ea6-42aa-4d2d-bfc9-c3319f4f072f_1600x795.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<h2>What happened this week in AI by Louie</h2><p>After a wave of strong new closed models last week, Moonshot AI&#8217;s Kimi K3 is the closest a model headed for open release has come to the frontier since DeepSeek R1 in early 2025. This week also produced more evidence that LLM agents&#8217; capabilities are turning recursive, with agents optimizing the infrastructure that trains their successors and producing publishable mathematics.</p><p>Moonshot says K3 combines 2.8 trillion total parameters, a one-million-token context window, native vision, and a Mixture-of-Experts design that routes each token through 16 of 896 experts. Kimi Delta Attention handles long contexts efficiently, while Attention Residuals lets each layer retrieve selected earlier-layer outputs, which Moonshot credits with roughly 25% better training efficiency for under 2% extra cost, part of a claimed 2.5x scaling-efficiency gain over K2. Quantization-aware training runs in four-bit MXFP4, and Moonshot recommends at least 64 accelerators for deployment. K3 is live in Kimi&#8217;s products and API, with weights promised by July 27.</p><p>Independent testing shows a consistent shape: close to the leaders overall, ahead in key agentic lanes. On Artificial Analysis&#8217;s Intelligence Index as of July 21, K3 scores 57.11, behind Claude Fable 5 at 59.86 and GPT-5.6 Sol at 58.89 but above Opus 4.8 at 55.69 and GLM-5.2 at 51.09. It sits within 0.25 points of Fable on the Coding Index at 76.24, leads AutomationBench-AA at 52.71%, and ranks first on Arena&#8217;s preliminary frontend leaderboard at 1,679 Elo. Its weakest lane is knowledge reliability, where its AA-Omniscience score of 18 trails Fable&#8217;s 40 and Opus 4.8&#8217;s 27, and Moonshot itself concedes K3 trails Fable and Sol overall, including on user experience.</p><p>Moonshot&#8217;s own coding table is strong and uneven. K3 narrowly trails Sol on Terminal-Bench 2.1, edges it on ProgramBench, and leads the SWE Marathon chart.</p><p>K3 also topped our internal writing benchmark, an early, narrow test of our editorial voice. K3 Thinking ranked first at 2,840 Elo, ahead of Fable 5 at 2,760, after K2.6 placed ninth at 2,214, and it costs us about $0.25 per script, roughly a fifth of Fable&#8217;s price. While it has now fallen behind Fable after an upgrade to our evaluations, the jump from K2.6 shows what a generational leap looks like.</p><p>The economics are less flattering than the leaderboards. While its per-token costs are below OpenAI and Anthropic, they are 3x higher than Kimi 2.6. This is a huge model and will also not be cheap to inference yourself once weights are available. On a task level, Artificial Analysis measures K3 at $0.95 per weighted Intelligence Index task, only 8% below Sol&#8217;s $1.04, despite token prices 40% to 50% lower, because K3 emits 54% more output tokens. Its first-party endpoint also generates slowly, at 39.5 output tokens per second and 67.5 seconds for a standardized 500-answer-token request. However, it streams its first reasoning token in about four seconds while Fable and Sol&#8217;s maximum-effort endpoints stay silent for around two minutes.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!d4oO!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F452b1ea6-42aa-4d2d-bfc9-c3319f4f072f_1600x795.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!d4oO!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F452b1ea6-42aa-4d2d-bfc9-c3319f4f072f_1600x795.png 424w, https://substackcdn.com/image/fetch/$s_!d4oO!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F452b1ea6-42aa-4d2d-bfc9-c3319f4f072f_1600x795.png 848w, https://substackcdn.com/image/fetch/$s_!d4oO!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F452b1ea6-42aa-4d2d-bfc9-c3319f4f072f_1600x795.png 1272w, https://substackcdn.com/image/fetch/$s_!d4oO!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F452b1ea6-42aa-4d2d-bfc9-c3319f4f072f_1600x795.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!d4oO!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F452b1ea6-42aa-4d2d-bfc9-c3319f4f072f_1600x795.png" width="1456" height="723" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/452b1ea6-42aa-4d2d-bfc9-c3319f4f072f_1600x795.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:723,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!d4oO!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F452b1ea6-42aa-4d2d-bfc9-c3319f4f072f_1600x795.png 424w, https://substackcdn.com/image/fetch/$s_!d4oO!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F452b1ea6-42aa-4d2d-bfc9-c3319f4f072f_1600x795.png 848w, https://substackcdn.com/image/fetch/$s_!d4oO!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F452b1ea6-42aa-4d2d-bfc9-c3319f4f072f_1600x795.png 1272w, https://substackcdn.com/image/fetch/$s_!d4oO!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F452b1ea6-42aa-4d2d-bfc9-c3319f4f072f_1600x795.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Source: Kimi. Total parameters of each company&#8217;s flagship model</figcaption></figure></div><p>But the demand for K3 has already overwhelmed the supply side. Kimi paused new subscriptions on July 19 after 48 hours of demand pushed its GPUs near capacity. To protect existing customers, they have paused it for new subscribers and confirmed that inference supply already binds adoption. Open weights do not repeal serving costs either: the raw files occupy roughly 1.5 terabytes in MXFP4 before runtime memory, and the key-value cache, and the footprint is nearly triple K2.6&#8217;s one trillion parameters, though active compute per token likely grows far less.</p><p>AI stocks sold off on Friday in what some investors called a smaller DeepSeek moment, with K3 one driver among several. The reaction reflects how quickly the frontier is tightening.</p><p>There is intense discussion about Chinese labs distilling US frontier models. My view is simple: of course they are, and so is everyone else. Frontier intelligence can enter a model through many routes. Labs can collect direct outputs, generate synthetic examples, or buy data from providers that used Fable to brainstorm reinforcement-learning tasks, benchmarks, rubrics, and adversarial tests. Elon Musk has even acknowledged xAI&#8217;s use of Claude-generated data. One way or another, it seems very likely that some of Fable&#8217;s intelligence helped bootstrap the major leaps from GLM-5.2, Grok, Meta, and K3 over the past few weeks.</p><p>A successful five-to-ten-hour Fable agent run can produce code, tests, tool traces, critiques, and a verified final artifact. That is exceptionally valuable reinforcement-learning data for extending agentic capabilities. Labs can gain enormous value without private reasoning traces or logits. The AI race is a commercial and geopolitical contest led by CEOs who believe their company should lead it. They will take easy capability gains when the consequences appear limited.</p><p>K3 remains a huge achievement. Using this data effectively requires a strong base model, sophisticated reinforcement-learning infrastructure, reliable environments, good graders, and extensive original research. These labs are also creating novel data, training tasks, architectures, and systems. K3 is far more than a distillation of Fable and already beats it in some areas. Moonshot says K3 autonomously optimized an Attention Residuals kernel, built a MiniTriton compiler, and completed a 48-hour chip-design experiment.</p><p>Chinese labs are currently producing more visible architecture-level innovation and appear to have at least comparable aggregate research talent. They need that strength to compensate for less capital, less compute, and weaker access to the advanced semiconductor ecosystem. US labs benefit from hyperscaler balance sheets, leading chips, high-bandwidth memory, and easier access throughout the deep ultraviolet and extreme ultraviolet lithography supply chain.</p><p>The largest US labs have also locked up enormous GPU capacity through long-term contracts. That is a major moat. They can win by maintaining the strongest frontier models. If base-model intelligence becomes commoditized, they can still win by controlling much of the inference capacity needed to deploy it. Codex and Claude Code create another moat as they become embedded in enterprises, accumulate workflow context and evaluations, and let the model and agent harness improve together.</p><p>Which brings us to recursion at a different level: many types of recursive AI improvement are now operational. First, coding agents accelerate the engineers building the next model across data pipelines, kernels, evaluations, and deployment; OpenAI says early GPT-5.3 Codex versions helped debug their own training and deployment, and K3&#8217;s kernel run is the same loop inside Moonshot. Second, stronger models manufacture the training pressure for their successors: tasks, synthetic data, verifiers, judges, and adversarial examples. Teams can spend orders of magnitude more inference through maximum reasoning, sampling, and agent fleets, then compress the best verified results into a cheaper baseline. OpenAI&#8217;s GPT-Red, which generated the attacks used to harden GPT-5.6, is that loop in production.</p><p>The third level, AI designing its successor&#8217;s architecture, remains unsolved but no longer looks distant. Thousands of agents could test candidate designs on small clusters against loss curves and benchmarks before the winners scale, if proxy experiments can be made to transfer. AlphaEvolve&#8217;s measured improvements to Google&#8217;s data-center scheduling, Tensor Processing Unit circuits, and matrix kernels used in AI training show part of the path, and K3&#8217;s experiments point the same way.</p><p>The multi-agent method is also starting to produce real mathematics. On Monday, Anthropic mathematician Levent Alpoge posted an explicit Fable-assisted counterexample to the Jacobian conjecture, open since 1939: a polynomial map with constant nonzero Jacobian determinant that sends three distinct points to the same output. Independent symbolic checks confirm the arithmetic. It disproves the conjecture in dimension three and above, while the two-dimensional case stays open and a formal paper is pending. The announcement was almost casual: Alpoge thanked his friend Akhil Mathew for asking the question and Fable for working through the World Cup final. Within a day, a separate Codex run produced an essentially equivalent construction under a published prompt demanding diverse multi-agent search and adversarial checking. Current models excel at intelligence-guided search under exact verifiers and at connecting distant specialized fields, and those strengths sit close to what machine learning research itself requires.</p><p>I would not count Anthropic out either. Mythos training was reportedly complete by late March, giving nearly four months for evaluation, post-training, and product work, so Claude could leap ahead again. The durable lesson of 2026 is that several labs can now make major capability jumps within weeks, and each generation supplies the tools and data that make the next arrive faster.</p><div><hr></div><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://newsletter.towardsai.net/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading Towards AI Newsletter! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><div><hr></div><h3>Why should you care?</h3><p>Near-frontier open weights compress model API margins while expanding total compute demand, because cheaper intelligence makes longer tasks, more retries, and larger agent fleets economical. The winners shift toward scarce capacity, efficient serving, and distribution: hyperscalers and neoclouds can serve every model. At the same time, US AI labs defend their reserved inference capacity, economics through their coding harnesses, and enterprise governance. I would watch completed-task cost, output-token efficiency, and accelerator occupancy over leaderboard rank or parameter count. For builders, K3 is a credible option for workflows that previously required the most restricted US systems.</p><p>For researchers, the recursion carries the longest fuse, and its fuel is cheap, exact verification. Mathematics, code, kernels, and chip layouts reject bad candidates almost for free, which is why the Jacobian result landed there first, and why environments, simulators, formal checks, and high-quality evaluators are becoming core research infrastructure rather than tooling afterthoughts.</p><p>Governments face the hardest version of the trade-off, because capabilities that ship as weights run beyond monitored endpoints and cannot be recalled. If the weights land as promised, any restriction drafted after July 27 will be regulating files the world already holds.</p><p><em>&#8212; <a href="http://www.linkedin.com/in/louie-peters">Louie Peters&#8202;&#8212;&#8202;Towards AI Co-founder and CEO</a></em></p><h3>Hottest News</h3><p>1. <a href="https://www.kimi.com/blog/kimi-k3">Moonshot Launched Kimi 3</a> and <a href="https://x.com/Kimi_Moonshot/status/2078855608565207130">Paused it After Subscription Demand Hit GPU Limits</a></p><p>Moonshot AI launched Kimi K3, a 2.8 trillion-parameter MoE model with a 1M-token context window, built for long-horizon coding, reasoning, and agentic tasks. It is the largest open-weight model released to date, surpassing DeepSeek V4 Pro&#8217;s 1.6T parameters. It uses Stable LatentMoE, effectively activating 16 of 896 experts. The model supports text, image, and video inputs and is available through Kimi.com, Kimi Work, Kimi Code, and the Kimi API. API pricing is $3/$15 per million input/output tokens. On independent benchmarks, Kimi K3 placed third on the Artificial Analysis Intelligence Index behind Fable 5 and GPT-5.6 Sol and topped several coding and agent benchmarks.</p><p>Within 48 hours of launch, user demand exceeded Moonshot&#8217;s GPU capacity, forcing the company to pause new consumer subscriptions on July 19 to protect service quality for existing subscribers. Moonshot said it would reopen new spots in batches as capacity is added and plans to split memberships into two tiers: one for Kimi Web, App, and Work, and a separate Kimi Code membership for coding workflows. Full open-weight model releases are planned for July 27 on Hugging Face.</p><p>2. <a href="https://x.com/claudeai/status/2078302415804379218">Anthropic Makes Claude Fable 5 a Standard Max and Team Premium Feature</a></p><p>Anthropic announced that beginning July 20, Claude Fable 5 will be permanently included in all Max and Team Premium plans at 50% of usage limits. This ends weeks of uncertainty over Fable 5&#8217;s availability, which had been extended in stages following the model&#8217;s redeployment on July 1 after the Commerce Department export control suspension. The eligible plans are Max 5x ($110/month), Max 20x ($220/month), and Team Premium ($100/month). Pro and Team Standard subscribers will not receive Fable 5 as part of their base plans but retain access through usage credits at standard API rates ($10/$50 per million input/output tokens). Pro and Team Standard users will receive a one-time $100 credit.</p><p>3. <a href="https://x.com/Alibaba_Qwen/status/2078759124914098291">Alibaba&#8217;s Qwen Team Launches Qwen3.8-Max-Preview and Says Full Qwen3.8 Will Go Open-Weight</a></p><p>Alibaba&#8217;s Qwen team announced Qwen3.8, a 2.4 trillion-parameter multimodal model that the team describes as &#8220;one of the most powerful models available today, second only to Fable 5.&#8221; Qwen3.8-Max-Preview is available now on Alibaba&#8217;s Token Plan, Qoder, and QoderWork. The model accepts text, image, and video inputs and is compatible with both the OpenAI and Anthropic API protocols, allowing existing coding agents to connect without rebuilding their harnesses. Alibaba has committed to releasing Qwen3.8 with open weights. No benchmark results, model card, or active-parameter count have been disclosed. The &#8220;second only to Fable 5&#8221; claim is Alibaba&#8217;s own positioning, not an independently verified result. Token Plan pricing starts at approximately $6/month for the Lite tier; standalone per-token API pricing has not been announced. The launch comes three days after Moonshot AI&#8217;s Kimi K3, putting two trillion-scale Chinese models in the market within the same week.</p><p>4. <a href="https://x.ai/news/grok-build-open-source">SpaceXAI Open-Sourced Grok Build</a></p><p>SpaceXAI open-sourced Grok Build&#8217;s coding-agent harness and terminal interface under Apache 2.0. The release includes the full Rust agent loop, TUI, CLI shell, and tool-call dispatch layer, allowing developers to inspect exactly how the agent assembles context, calls models, and executes tools. The code is published as a squashed monorepo mirror with no prior commit history. Grok Build supports autonomous coding tasks, file editing, shell command execution, web search, and functionality extensions via Skills and MCP servers. The open-source release came roughly 72 hours after an AI-safety researcher published wire-level evidence that Grok Build had been silently uploading complete Git repositories to a Google Cloud Storage bucket run by xAI, including unredacted secrets and full commit history, even when the agent was not instructed to open files. SpaceXAI disabled default data retention for all users on July 12, committed to deleting all previously retained coding data, and stated that zero data retention had always been respected when users explicitly disabled upload.</p><p>5. <a href="https://thinkingmachines.ai/news/introducing-inkling/">Thinking Machines Lab Released Inkling</a></p><p>Thinking Machines Lab, the AI startup founded by former OpenAI CTO Mira Murati, released Inkling, its first production model trained from scratch. Inkling is a 975B-parameter MoE transformer with 41B active parameters, trained on 45 trillion tokens of text, images, audio, and video. It accepts text, image, and audio inputs and supports a context window of up to 1M tokens. The model includes controllable thinking effort from 0.2 to 0.99. On the Artificial Analysis Intelligence Index, Inkling debuted at 41, making it the leading US open-weights model, ahead of Nemotron 3 Ultra (38), Gemma 4 31B (29), and gpt-oss-120b (24). On coding, it scored 77.6% on SWE-bench Verified. Thinking Machines also previewed Inkling-Small, a 276B-parameter MoE with 12B active parameters that matches or exceeds Inkling on several benchmarks; weights will be released once testing is complete. Inkling is accessible via the Tinker platform API (256K context window), and weights are available on Hugging Face (1M context window) under Apache 2.0. The NVFP4 quantized checkpoint reduces aggregate VRAM requirements to approximately 600 GB.</p><p>6. <a href="https://openai.com/index/unlocking-self-improvement-gpt-red/">OpenAI Introduced GPT-Red</a></p><p>OpenAI announced GPT-Red, an internal automated red-teaming model built to find prompt injection vulnerabilities at scale. GPT-Red uses self-play reinforcement learning: an attacker model and multiple defender models train against each other simultaneously, with each side earning rewards for successfully executing or resisting attacks. As defenders become more robust, GPT-Red generates progressively harder attacks, creating a training curriculum whose difficulty scales with the model&#8217;s own defenses. OpenAI used GPT-Red to adversarially train GPT-5.6, making it substantially more resistant to prompt injection. In head-to-head evaluations, GPT-Red beat human red-teamers 84% to 13% on prompt injection tasks. GPT-Red is kept separate from deployed models: attack capabilities remain internal, while the resulting robustness transfers to the defender. OpenAI acknowledges limitations: multi-turn attacks and image-based prompt injection still require human red-teamers. GPT-Red will not be released externally.</p><div><hr></div><h4>AI Tip of the Day</h4><p>Structured Outputs make responses easier to plug into software, but they also change how the model generates its answer.</p><p>When we first wrote the lesson &#8216;Structuring Your Data&#8217; for our <a href="https://towardsai.com/academy/full-stack-ai-engineering/?utm_source=Newsletter&amp;utm_medium=email&amp;utm_id=AItips">Full Stack AI Engineering</a> course, the main problem was getting models to return JSON your application could reliably parse. As structured outputs became standard across newer reasoning models and smaller models used in production, another problem became harder to ignore: the schema can be perfectly valid while the answer inside it gets worse.</p><p>This is especially important on tasks that already stretch the model&#8217;s reasoning. Requiring it to solve the problem within a rigid schema can reduce accuracy, particularly when the schema is complex, or the model has little capacity to spare.</p><p>Before adding a schema, run the task in free form and record the quality of the answers. Then introduce the smallest schema your application needs and repeat the same evaluation.</p><p>Measure answer quality and schema validity separately. A 100% parse rate only tells you that your software can read the response. It does not tell you whether the response is correct.</p><p>When structure reduces accuracy, simplify the schema. For tasks where the final formatting is deterministic, let the model work through the problem freely and package its answer into the required structure in code.</p><div><hr></div><h3>Five 5-minute reads/videos to keep you learning</h3><p>1. <a href="https://pub.towardsai.net/how-ai-agent-memory-actually-works-and-how-to-build-it-0d0874e913bc?sharedUserId=tai-tech">How AI Agent Memory Actually Works&#8202;&#8212;&#8202;And How to Build It</a></p><p>Stateless LLM calls forget everything, so agents need a memory layer built on top. This guide maps the four memory types: working, episodic, semantic, and procedural, grounding them in the CoALA framework and Tulving&#8217;s research. It clarifies how agent memory differs from RAG, walks through the write-consolidate-retrieve-forget lifecycle, and includes a runnable, framework-agnostic Python implementation for each memory store.</p><p>2. <a href="https://pub.towardsai.net/reward-design-is-the-hard-part-building-verifiable-rewards-for-tool-using-agents-f99c5c38f9b3?sharedUserId=tai-tech">Reward Design Is the Hard Part: Building Verifiable Rewards for Tool-Using Agents</a></p><p>RLVR&#8217;s clean verify-and-score loop breaks down once agents span 10 to 20 tool-calling turns, where a single terminal reward leaves most steps without signal. This article lays out a layered fix: turn-level verifiable rewards, rubric-based LLM judges, and guidance injection when sparsity stalls training entirely. It also covers agent-specific reward hacks like tool-call theater and retry gaming. The reward design here works less like ML and more like systems engineering, with its own test suite.</p><p>3. <a href="https://pub.towardsai.net/building-stateful-ai-agents-that-survive-session-kills-e1877e3c78f0?sharedUserId=tai-tech">Building Stateful AI Agents That Survive Session Kills</a></p><p>Coding agents lose everything when a session dies. This article shows how to build a harness on Tensorlake&#8217;s Firecracker-backed MicroVM sandboxes that keeps state alive across days. These sandboxes suspend and resume with filesystem, memory, and running processes intact, while snapshots double as agent memory and fork into parallel workers for GSPO-style candidate ranking. The piece also covers isolated verifier sandboxes, custom baked images cutting cold starts from 12 seconds to 600 milliseconds, and native SSH workflows for persistent multi-day debugging.</p><p>4. <a href="https://pub.towardsai.net/understand-hnsw-why-your-vector-search-returns-garbage-build-your-own-minimalist-hnsw-from-114415ddf28c?sharedUserId=tai-tech">Understand HNSW: Why Your Vector Search Returns Garbage (Build your own minimalist HNSW from scratch)</a></p><p>Vector search often fails at the index layer, not at the embedding layer. In this article, the author builds a minimal HNSW implementation in pure Python and numpy, then deliberately breaks it to show how weak build parameters cap recall at 0.44 while stronger settings reach 0.78. Measurements demonstrate that M and ef_construction set a hard ceiling no query-time tuning can break, while ef_search hits diminishing returns fast.</p><p>5. <a href="https://pub.towardsai.net/context-engineering-for-bedrock-agents-a-hands-on-guide-beyond-prompt-engineering-d92aad36a839?sk=6187a9b0eb24918aeb2766d95e8b3399">Context Engineering for Bedrock Agents: A Hands-On Guide Beyond Prompt Engineering</a></p><p>Prompt tuning alone can&#8217;t fix agents that hallucinate mid-loop, so this guide proposes a three-layer context architecture for AWS Bedrock: persistent identity and guardrails, TTL-bounded session state, and per-turn transient retrieval. The author builds a Python ContextAssembler that enforces token budgets, evicts expired layers, and logs every injection decision for debugging. It also covers wiring in Bedrock Knowledge Bases, semantic tool selection, session memory, cost guards, and log analysis.</p><h3>Repositories &amp; Tools</h3><p>1. <a href="https://github.com/lyogavin/airllm">AirLLM</a> runs 70B+ models on a single 4GB GPU by loading one transformer layer at a time, without quantization, distillation, or pruning.</p><p>2. <a href="https://github.com/tirth8205/code-review-graph">Code-review-graph</a> is a local-first code intelligence graph that parses codebases with Tree-sitter into a persistent SQLite knowledge graph.</p><p>3. <a href="https://github.com/AstrBotDevs/AstrBot">AstrBot</a> is an agentic chatbot platform that deploys LLM-powered assistants across mainstream instant messaging apps, with agent capabilities, RAG, MCP, knowledge bases, etc.</p><p>4. <a href="https://github.com/Canner/WrenAI">WrenAI</a> is a generative BI engine that gives AI agents a governed context layer of business semantics, approved definitions, and memory.</p><p>5. <a href="https://github.com/NVIDIA/DeepStream">DeepStream</a> is a streaming analytics toolkit for AI-based video and image understanding, providing a GStreamer-based framework.</p><h3>Top Papers of The Week</h3><p>1. <a href="https://arxiv.org/abs/2607.12463">Function-Aware Fill-in-the-Middle as Mid-Training for Coding Agents</a></p><p>A coding agent&#8217;s action-observation-continuation loop is structurally identical to a function call site. This paper exploits that by masking functions selected via program dependency graphs during mid-training, forcing the model to predict continuations conditioned on values computed elsewhere. Despite training on Python only, the inductive bias transfers to non-Python coding and tool-use benchmarks. Mid-training improves SWE-Bench Verified by +2.8/+3.0 at 7B/14B and +3.2 on Qwen3&#8211;8B, while mitigating the capability erosion that agentic post-training otherwise inflicts on general coding and tool-use tasks.</p><p>2. <a href="https://arxiv.org/abs/2607.09024">GenCeption: Video Generators as General-Purpose Vision Learners</a></p><p>Google DeepMind repurposes a pre-trained video generative diffusion backbone into a single feed-forward perception model that handles depth, surface normals, camera pose, segmentation, and 3D keypoint prediction, all steered by text instructions. GenCeption matches or surpasses specialized models (DepthAnything3, SAM3, D4RT, VGGT-Omega) and achieves comparable performance with 7x to 500x less training data. A model trained exclusively on synthetic human videos generalizes to real-world footage and out-of-distribution categories, including animals and robots.</p><p>3. <a href="https://arxiv.org/abs/2607.13027">PalmClaw: Native On-Device Agent Framework for Mobile Phones</a></p><p>Existing mobile agents operate through GUI actions that form long, interface-dependent sequences. PalmClaw runs natively on mobile phones, exposing device capabilities (contacts, calendar, camera, sensors, apps) as structured tools the agent calls directly instead of simulating screen touches. The framework manages sessions, persistent memory, skills, and the full agent loop on-device, producing shorter action sequences with clearer execution boundaries.</p><p>4. <a href="https://arxiv.org/abs/2607.15591">RecGPT-V3 Technical Report</a></p><p>Running LLM-based recommendation at scale on Taobao exposed three bottlenecks: stateless user modeling, a lossy text-tag channel for item grounding, and verbose chain-of-thought latency. RecGPT-V3 addresses all three with a Memory Hub that cuts user-modeling compute by 55.8%, a Hybrid-modal Foundation Model that jointly reasons over text and Semantic IDs for direct item grounding, and Latent Intent Reasoning that compresses rationales into learnable latent tokens while lowering output token cost by 200x. In production A/B tests: GMV +3.97%, CTR +1.00%, with serving resource consumption down 52.4%.</p><p>5. <a href="https://arxiv.org/abs/2607.15330">Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories</a></p><p>Xiaomi presents a foundational VLA model that can follow diverse language instructions and adapt to novel downstream tasks with minimal fine-tuning. It is pre-trained on over 100K hours of real-world manipulation trajectories collected via embodiment-free UMI devices. It uses a scalable auto-labeling pipeline to annotate trajectory clips with natural-language descriptions of scene state transitions. Post-training adapts the model to specific embodiments with minimal data. It establishes a new state-of-the-art with a 57.6% success rate on RoboCasa365.</p><h3>Quick Links</h3><p>1. <a href="https://research.perplexity.ai/articles/wandr-benchmark-evaluating-research-agents-that-must-search-wide-and-deep">Perplexity AI released WANDR</a>, an open benchmark of 500 realistic data-collection tasks designed to evaluate whether AI research agents can search broadly and back every finding with specific evidence. Taken together, the 500 tasks call for 170,495 source-backed records. The tasks are ranked by 167 lower-, 166 middle-, and 167 higher-difficulty examples. Perplexity Search (Code) leads at 0.363 soft F1 and 0.133 hard F1. Anthropic is second at 0.249 and 0.072; every other system tops out at 0.121 soft F1 and 0.035 hard F1.</p><p>2. <a href="https://www.zyphra.com/our-work/zuna1.1">Zyphra released ZUNA1.1</a>, an update to its 380M-parameter EEG foundation model that now accepts variable-length inputs from 0.5 to 30 seconds, up from the fixed 5-second segments in ZUNA1. The model reconstructs missing channels, denoises existing ones, and upsamples across arbitrary electrode layouts, trained on approximately 3.5M channel-hours of EEG data with realistic corruption patterns. Reconstruction quality matches or exceeds ZUNA1 while handling a far wider range of real-world recording conditions. Cloud-served inference is available in Zyphra&#8217;s EEG Playground with no GPU, installation, or coding required.</p><p>3. <a href="https://mistral.ai/news/robostral-navigate/">Mistral released Robostral Navigate</a>, its first model optimized for embodied AI. The 8B vision-language model guides robots through complex environments using only a single standard RGB camera and natural-language instructions, eliminating the need for LiDAR, depth sensors, or multiple cameras. Trained entirely in simulation on 400,000 navigation trajectories across 6,000 scenes, it achieves 76.6% success on unseen R2R-CE environments, 9.7 points above the previous best single-camera model and 4.5 points above the best multi-sensor system.</p><p>4. <a href="https://prismml.com/news/bonsai-27b">PrismML released Bonsai 27B</a>, a 27B-parameter multimodal model compressed from Qwen3.6 27B to 1-bit and ternary weights using end-to-end low-bit quantization. The 1-bit variant fits in 3.9 GB and runs on iPhone, iPad, and Mac via MLX; the ternary variant ships at 5.9 GB for laptops and consumer GPUs. Both retain the full 262K context window and support text, vision, tool calling, and agentic workflows. PrismML reports 90% and 95% intelligence retention, respectively. Speeds reach 163 tok/s (1-bit) and 134 tok/s (ternary) on an RTX 5090.</p><h3>Who&#8217;s Hiring in AI</h3><p><strong><a href="https://jobs.towardsai.net/job/govcio-senior-application-engineer-remote-9iew">Senior Application Engineer @GovCIO (Remote)</a></strong></p><p><strong><a href="https://jobs.towardsai.net/job/huntington-national-bank-ai-senior-project-manager-ww3o">AI Senior Project Manager @Huntington National Bank (Columbus, IN, USA)</a></strong></p><p><strong><a href="https://jobs.towardsai.net/job/asg-asg-ai-systems-engineer-vh60">AI Systems Engineer @ASG (Remote/USA)</a></strong></p><p><strong><a href="https://jobs.towardsai.net/job/domino-data-lab-forward-deployed-engineer-life-sciences-5z0f">Forward Deployed Engineer, Life Sciences @Domino Data Lab (Remote)</a></strong></p><p><strong><a href="https://jobs.towardsai.net/job/twilio-forward-deployed-engineer-htbs">Forward Deployed Engineer @Twilio (Remote/USA)</a></strong></p><p><strong><a href="https://jobs.towardsai.net/job/texas-sports-academy-senior-ai-engineer-llm-systems-and-rag-optimization-ijhg">Senior AI Engineer @Texas Sports Academy (Remote/USA)</a></strong></p><p><em>Interested in sharing a job opportunity here? Contact <a href="mailto:sponsors@towardsai.net">sponsors@towardsai.net</a>.</em></p><p><em>Think a friend would enjoy this too? <a href="https://newsletter.towardsai.net/">Share the newsletter and let them join the conversation.</a></em></p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://newsletter.towardsai.net/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading Towards AI Newsletter! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div>]]></content:encoded></item><item><title><![CDATA[TAI #213: A Wave of New Frontier Competitors and the Multi-Agent Breakout]]></title><description><![CDATA[Also GPT-5.6, Grok 4.5, Muse Spark 1.1, GPT-Realtime-2.1, and more.]]></description><link>https://newsletter.towardsai.net/p/tai-213-a-wave-of-new-frontier-competitors</link><guid isPermaLink="false">https://newsletter.towardsai.net/p/tai-213-a-wave-of-new-frontier-competitors</guid><dc:creator><![CDATA[Louie Peters]]></dc:creator><pubDate>Wed, 15 Jul 2026 14:56:20 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!T0Nt!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe6d95711-f865-41f7-a555-60237eb56a20_1672x941.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<h4>Towards AI Deployment</h4><p>Some company news and a quick ask before this week&#8217;s stories! Since 2019, we&#8217;ve helped developers transition into AI engineering and trained enterprises to build with AI. Over those seven years, enterprise demand moved past training; the ask now is to build and deploy production systems.</p><p>This week, we formalized that work as a dedicated effort: Towards AI Deployment.</p><p>We would really appreciate it if you could please like and share <a href="https://www.linkedin.com/posts/denis-p-72588a44_louie-and-i-sat-down-with-our-advisor-zeena-ugcPost-7483166744618967040-vWGV/">our co-founder&#8217;s post on LinkedIn</a> and watch the video he has shared!</p><p>Towards AI now works as two connected sides: Learning converts software developers into AI engineers and forward-deployed engineers, and Deployment puts them to work delivering custom systems for private equity firms, their portfolio companies, funds, and banks. We believe successful AI deployment requires domain expertise, so we have narrowed our focus to a vertical where we have strong momentum and the domain expertise to complement our AI talent.</p><p>For readers of this newsletter, the most direct change is that the teaching improves. Every engagement the division delivers sharpens the curriculum: more failure cases, evaluation methods, and architecture patterns flowing back into our courses and these pages.</p><p>If you know of a company stuck between experimenting with AI tools and building a system people rely on, please <a href="https://towardsai.com/">reach out</a>, we&#8217;d love to look at building the workflow with you.</p><h2>What happened this week in AI by Louie</h2><p>This was a huge week for model releases. In the span of two days, SpaceXAI launched Grok 4.5, OpenAI moved GPT-5.6 from restricted preview to general availability, and Meta launched Muse Spark 1.1. OpenAI also shipped GPT-Realtime-2.1 and a mini variant, improving interruption handling, noisy audio, and alphanumeric recognition while cutting p95 latency by at least 25%. Three new frontier competitors landed almost at once.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!T0Nt!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe6d95711-f865-41f7-a555-60237eb56a20_1672x941.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!T0Nt!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe6d95711-f865-41f7-a555-60237eb56a20_1672x941.png 424w, https://substackcdn.com/image/fetch/$s_!T0Nt!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe6d95711-f865-41f7-a555-60237eb56a20_1672x941.png 848w, https://substackcdn.com/image/fetch/$s_!T0Nt!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe6d95711-f865-41f7-a555-60237eb56a20_1672x941.png 1272w, https://substackcdn.com/image/fetch/$s_!T0Nt!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe6d95711-f865-41f7-a555-60237eb56a20_1672x941.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!T0Nt!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe6d95711-f865-41f7-a555-60237eb56a20_1672x941.png" width="1456" height="819" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/e6d95711-f865-41f7-a555-60237eb56a20_1672x941.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:819,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1681521,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://newsletter.towardsai.net/i/207144915?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe6d95711-f865-41f7-a555-60237eb56a20_1672x941.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!T0Nt!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe6d95711-f865-41f7-a555-60237eb56a20_1672x941.png 424w, https://substackcdn.com/image/fetch/$s_!T0Nt!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe6d95711-f865-41f7-a555-60237eb56a20_1672x941.png 848w, https://substackcdn.com/image/fetch/$s_!T0Nt!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe6d95711-f865-41f7-a555-60237eb56a20_1672x941.png 1272w, https://substackcdn.com/image/fetch/$s_!T0Nt!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe6d95711-f865-41f7-a555-60237eb56a20_1672x941.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>The week&#8217;s more consequential development was a price-performance reset. Closed models already led on raw intelligence; open weights led on cost per unit of it, and that second advantage slipped this week. GLM-5.2, the leading open-weight model on the Artificial Analysis Intelligence Index, scores 51 at a measured cost of about $0.37 per benchmark task. Grok 4.5 scores 54 at about $0.31, beating GLM on both score and cost. GPT-5.6 Luna and Muse Spark 1.1 tie GLM at 51 while costing about $0.21 and $0.26 per task. Sol and Terra score 59 and 55 at higher task cost. Closed models now match or beat the open-weight frontier on cost per unit of intelligence, which had been open weights&#8217; clearest selling point.</p><p>Cost per task matters more than the price printed per million tokens. Grok 4.5 charges more per output token than GLM-5.2, but it used roughly 14,000 output tokens per benchmark task against GLM&#8217;s 43,000, so more intelligence per token beat a lower token price. Simon Willison&#8217;s identical SVG prompt cost 0.71 cents on Luna with reasoning off and 48.55 cents on Sol at maximum effort, a 68x spread from settings alone. In production, the number to watch is cost per accepted result; more on that below.</p><p>We covered the GPT-5.6 preview two weeks ago, including the Sol, Terra, and Luna product ladder, pricing, safety restrictions, Ultra mode, and METR&#8217;s difficult-to-interpret time-horizon result. The incremental news this week is stronger: the family is now generally available, independent tests are in, and they support OpenAI&#8217;s core capability and efficiency claims.</p><p>Sol is seriously competing with Claude Fable 5 at the frontier again. It scores 59 on the Artificial Analysis Intelligence Index against Fable&#8217;s 60, leads the Coding Agent Index, and performs particularly well on DeepSWE, Terminal-Bench, BrowseComp, and OSWorld. Fable stays ahead on SWE-Bench Pro, GDPval, Toolathlon, FrontierMath Tier 4, and broader professional-work comparisons. I would choose between them based on the job. Sol is exceptionally good at coding, computer use, presentations, and structured execution. In my own use, Fable still writes better, and Arena&#8217;s preliminary creative-writing leaderboard points in the same direction, at 1507 versus 1486 with overlapping uncertainty.</p><p>Luna may be the most useful release in the family. It ties the open-weight intelligence frontier, scores 75 on the Coding Agent Index through Codex, runs at more than 200 output tokens per second, and costs less per measured task than GLM-5.2. CodeRabbit provides a useful warning at the other end of the family: Sol passed 63.7% of more than 100 repository tasks without an execution error, yet its code-review precision was only 31.6%. Long-running execution is improving faster than verification.</p><p>Adoption moved almost as quickly as the models. OpenAI reported more than 5 million weekly active Codex users in early June. In the days after launching ChatGPT Work and combining Chat, Work, and Codex in one desktop app, OpenAI&#8217;s Codex lead reported 8 million active users across Codex and ChatGPT Work. The definitions differ, so this is a directional comparison. More than 1 million people were already using Codex outside software development before the Work launch, and the new interface gives that audience a much friendlier route into the same agent infrastructure. This may be the strongest signal in the release: the distribution layer is catching up with the capability layer.</p><p>Grok 4.5 closed the gap faster than I expected. It now sits in the same broad agentic-coding group as GPT-5.5 and Fable while costing much less, serving at around 80 tokens per second, scoring strongly on SWE Marathon and Terminal-Bench, and using far fewer tokens than several peers. Snorkel measured a 29% full-rubric pass rate across roughly 2,000 professional tasks, ahead of GPT-5.5 and Opus 4.8. Grok ran in its own Grok Build harness while competitors used a different agent, so treat that as a system comparison. It still trails Fable and GPT-5.5 on DeepSWE 1.1. Cursor has flagged a contaminated CursorBench result, and its 54% hallucination rate on AA-Omniscience means it needs oversight. I see it as a credible frontier model with excellent economics, one tier below the very best coding systems.</p><p>The Cursor connection clearly helped shape it. Cursor and SpaceXAI say they jointly trained Grok 4.5 on trillions of tokens of eligible Cursor data covering developers, codebases, tools, and agent interactions, with Privacy Mode sessions excluded. SpaceXAI also ran reinforcement learning on hundreds of thousands of multi-step technical tasks, with rollouts lasting many hours across tens of thousands of GB300 GPUs. There is no public breakdown of how much of the capability jump came from Cursor&#8217;s data versus the scale and design of the rest of the training recipe, but this is a serious product-data flywheel.</p><p>Muse Spark 1.1 made a similar jump, improving Meta&#8217;s Artificial Analysis score by eight points in three months, with a one-million-token context window and explicit training for both sides of a multi-agent system: a lead agent that plans and delegates, and a subagent that follows a narrow remit and escalates when needed. Meta&#8217;s own evaluation report is more balanced than the launch copy. Muse posts 88.1 on MCP Atlas and 80.8 on OSWorld-Verified for tool and computer use, but 53.3 on DeepSWE and 61.5 on SWE-Bench Pro, trailing the strongest GPT and Claude models on the longest cross-application workflows. Meta itself says current autonomy is insufficient for sustained automated AI research. What matters here is the rate of progress and the price. Grok 4.5 and Muse Spark 1.1 join GLM-5.2 in a cluster of releases that suddenly make long-running agentic coding much cheaper and more accessible.</p><p>The timing makes me wonder whether several labs recently gained access to the same new generation of long-horizon training environments from vendors such as Mercor or Surge, both of which publicly build repository-scale coding tasks, long tool-use trajectories, and reinforcement-learning environments for leading labs. Scaled in-house reinforcement learning, better harnesses, more compute, and stronger first-party product data are also contributing.</p><p><strong>The next step change is multi-agent</strong></p><p>The more I use GPT-5.6 Ultra and Fable, the more this looks like another step change in capability and in how easily we can consume huge amounts of compute. The first change was o1, which let a model spend far more computation on a single response. The next was Claude Code and Codex, where models such as Opus 4.5 and GPT-5.1 made multi-hour tasks practical with tools, persistent state, and automatic context compaction. Almost every month since last November, a new model or harness has stretched the length of work I can hand off. Now the unit of work is changing again: from one answer, to one long-running agent, to a managed team of agents.</p><p>GPT-5.6 Ultra coordinates four agents by default, and OpenAI published evaluation runs with 16. Its Responses API lets a root model decide when to delegate and then synthesize the work. Claude&#8217;s dynamic workflows can run tens to hundreds of parallel subagents when enabled, and Muse Spark 1.1 was explicitly trained to operate as both manager and worker. Before this, making many parallel agents behave like a competent team required a very strong developer: dividing the task, constructing prompts, isolating environments, preventing duplicate work, and merging results. These systems move much of that orchestration into the model and product. Ten to 100 agents is now a realistic configured scale, and Anthropic describes workflows with hundreds. However, fan-out at this scale still produces duplicated work, conflicting changes, security risks, and confident group mistakes.</p><p>I could now very easily spend $1,000 on a single prompt on either Claude or Codex, meaning one top-level objective that triggers hundreds or thousands of model calls across a team. Fifty Sol agents each reading 500,000 fresh input tokens and producing 50,000 output tokens cost about $200; several deep passes, retries, and tool calls, or a 100-agent configuration, push the same objective toward $1,000. Prompt caching softens this less than it appears, because workers on separate subtasks mostly read fresh context, cache writes carry a 1.25x surcharge, and reasoning and output tokens are never discounted. Anthropic&#8217;s 16-agent compiler project cost almost $20,000, so this is already more than a spreadsheet thought experiment.</p><p>Cost control is becoming part of prompt design: total budgets, per-agent limits, maximum fan-out, cheaper worker models, checkpoints before another round, and independent review before accepting the final synthesis. The next generation of power users will pair prompt-writing with disciplined compute allocation.</p><div><hr></div><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://newsletter.towardsai.net/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading Towards AI Newsletter! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><div><hr></div><h3>Why should you care?</h3><p>This week&#8217;s model pile-up is good news for token buyers. Sol and Fable now trade wins at the frontier, while Grok 4.5, Muse Spark 1.1, and increasingly capable open-weight models are creating credible competition on intelligence, speed, and cost per completed task. That pressure should keep lowering costs while increasing the amount of useful work each token can buy. The next constraint is training data. Coding agents improved because labs captured long repository trajectories: how developers explore unfamiliar systems, choose tools, recover from failed approaches, verify changes, and divide large goals across teams of agents. ChatGPT Work shows where this should expand next. We need equivalent multi-hour and multi-day task data across research, finance, legal, operations, marketing, sales, and other professional work, capturing the decisions, failures, corrections, and final outputs. Organizations should begin collecting these traces with consent, strong privacy controls, and rigorous evaluation. The labs that learn how expert work unfolds over hours or days will train the strongest general-purpose agents, and the growing competition among them should make those capabilities cheaper for everyone buying tokens.</p><p><em>&#8212; <a href="http://www.linkedin.com/in/louie-peters">Louie Peters&#8202;&#8212;&#8202;Towards AI Co-founder and CEO</a></em></p><h3>Hottest News</h3><p>1. <a href="https://ai.meta.com/blog/introducing-muse-spark-meta-model-api/">Meta Releases Muse Spark 1.1</a></p><p>Meta Superintelligence Labs released Muse Spark 1.1 on July 9, a multimodal reasoning model designed for agentic tasks requiring planning and orchestration across external apps and services. The model accepts text, image, video, PDF, and audio inputs through a 1M-token context window with active compaction, meaning it remembers actions, retrieves information from earlier work, and filters out noise while preserving critical steps. Muse Spark 1.1 zero-shot generalizes to new native tools, MCP servers, and custom skills. It can run either as a primary agent delegating to parallel subagents or as a subagent itself. For computer use, it autonomously decides when to write a script for automation versus clicking through a GUI. On Meta-reported benchmarks, it leads on agentic orchestration: 88.1 on MCP Atlas, 54.7 on JobBench (Opus 4.8: 48.4), and 62.1 on Humanity&#8217;s Last Exam with tools (Opus 4.8: 57.9). It trails on coding: 61.5 on SWE-Bench Pro (Opus 4.8: 69.2) and 53.3 on DeepSWE 1.1 (GPT-5.5: 67.0). Alongside the model, Meta launched the Meta Model API in public preview for US developers, the first time Meta has offered a paid hosted model. API pricing is $1.25/$4.25 per million input/output tokens with $20 in free credits. The API supports both OpenAI and Anthropic wire formats. Muse Spark 1.1 is free in the Meta AI app in Thinking mode.</p><p>2. <a href="https://x.ai/news/grok-4-5">SpaceXAI Releases Grok 4.5</a></p><p>SpaceXAI released Grok 4.5, built for coding, agentic tasks, and knowledge work. The model was trained alongside Cursor on datasets spanning coding, science, engineering, and math, with reinforcement learning covering hundreds of thousands of tasks centered on multi-step software engineering. RL training runs asynchronously, with agentic rollouts that can last for many hours while learning continues across tens of thousands of NVIDIA GB300 GPUs. SpaceXAI highlights token efficiency as the primary advantage: Grok 4.5 resolves SWE-Bench Pro tasks with an average of 15,954 output tokens, roughly 4.2x fewer than Opus 4.8 (max), which uses 67,020 tokens. The model serves at 80 tokens per second. On benchmarks: 83.3% on Terminal-Bench 2.1, 64.7% on SWE-Bench Pro, 62.0% on DeepSWE 1.0, and 29.0% on SWE Marathon (first place). Fable 5 leads on SWE-Bench Pro (80.4%), DeepSWE 1.0 (66.1%), and Terminal-Bench 2.1 (84.3%). Beyond coding, Grok 4.5 handles office work through Grok Build, including multi-sheet Excel models with web research, PowerPoint diagrams using native shapes, and Word document drafting. Pricing is $2/$6 per million input/output tokens. Available in Grok Build, Cursor on all plans, and through the SpaceXAI API console, with free usage for a limited time in Grok Build and Cursor. Not yet available in the EU; expected mid-July.</p><p>3. <a href="https://developers.openai.com/api/docs/guides/realtime">OpenAI Releases GPT-Realtime-2.1 and GPT-Realtime-2</a></p><p>OpenAI released gpt-realtime-2.1 and gpt-realtime-2.1-mini, two developer-facing speech-to-speech models for building voice agents through the Realtime API. The release reduces p95 latency by at least 25% across all Realtime voice models through improved caching. GPT-Realtime-2.1 updates its predecessor with improved alphanumeric recognition for items like order numbers, phone numbers, and confirmation codes, better silence and noise handling, and more reliable interruption behavior when a user speaks over the model. It adds configurable reasoning effort per request. GPT-Realtime-2.1-mini is a smaller reasoning model for faster, lower-cost voice interactions, shipping reasoning and tool-use capabilities at mini pricing for the first time. Pricing: GPT-Realtime-2.1 at $32/$64 per million audio input/output tokens; GPT-Realtime-2.1-mini at $10/$20. Both are available through the Realtime API via WebRTC, WebSocket, or SIP connections.</p><p>4. <a href="https://openai.com/index/introducing-gpt-live/">OpenAI Releases GPT-Live and GPT-Live-1 mini</a></p><p>OpenAI launched GPT-Live, replacing Advanced Voice Mode as the default voice experience in ChatGPT. GPT-Live is a full-duplex audio language model that processes incoming speech and generates outgoing speech concurrently, rather than waiting for a user to finish before formulating a response. It interprets intonation and conversational intent, not just words. When a question exceeds GPT-Live&#8217;s native capabilities (complex math, multi-step research, tasks requiring tool use), it silently delegates to GPT-5.5 running in parallel without breaking the conversation. Users can select from three reasoning levels: Instant for quick replies, Medium for moderate depth, and High for thorough analysis. GPT-Live-1 is the default for Go, Plus, and Pro users; GPT-Live-1 mini is the default for Free users. On OpenAI&#8217;s benchmarks, GPT-Live-1 (High) scored 84.2% on GPQA (Advanced Voice Mode: 45.3%) and 75.2% on BrowseComp (Advanced Voice Mode: 0.7%). The launch includes nine remastered voices. Video and screen sharing are not supported at launch. API access is planned but not yet available; developers can sign up to be notified.</p><p>5. <a href="https://www.primeintellect.ai/blog/verifiers-v1">Prime Intellect Releases Verifiers v1</a></p><p>Prime Intellect released verifiers 0.2.0, previewing a rewritten core under the verifiers.v1 namespace. The redesign solves two structural problems: a monolithic environment design that coupled data, agent logic, and infrastructure, and a trace format in which storage grew quadratically with the number of turns (each turn stored a full copy of the prompt alongside the new completion). Verifiers v1 decomposes environments into three independently swappable layers: a taskset defining data, tools, and scoring; a harness defining how the task is solved (ReAct loop, Codex, Kimi Code, Terminus 2, Mini-SWE-Agent, or custom); and a runtime defining where execution happens (local subprocess, Docker, or cloud sandboxes including Prime Sandboxes and Modal). A directed acyclic graph message format replaces the old prompt-completion pairs, eliminating quadratic growth and enabling training on trajectories longer than the agent&#8217;s native context window through DAG branching. The library supports OpenAI Chat Completions, OpenAI Responses, and Anthropic Messages wire formats via dialect adapters. Harbor is the first fully supported third-party taskset format, with alpha support for NeMo Gym and OpenEnv. The legacy v0 code path is frozen and will not be actively maintained.</p><p>6. <a href="https://www.anthropic.com/research/global-workspace">Anthropic Discovers Global Workspace in Claude (J-Space)</a></p><p>Anthropic published research identifying the J-space, a small collection of internal neural activation patterns in Claude that function as a shared mental workspace. Named after the Jacobian-based technique (J-lens) used to find it, the J-space holds concepts the model is thinking about but not saying. It was not designed or programmed; it emerged during training. When Claude reads code with a bug, &#8220;ERROR&#8221; appears in the J-space. When it reads a prompt injection, &#8220;injection&#8221; and &#8220;fake&#8221; appear. When solving multi-step math, intermediate answers appear in sequence. Experiments demonstrate four functional properties: Claude can report J-space contents (swapping &#8220;Soccer&#8221; for &#8220;Rugby&#8221; changes what it reports thinking about); modulate them on request (told to think about citrus fruits while copying an unrelated sentence, &#8220;orange&#8221; and &#8220;fruits&#8221; appear); reason with them (swapping &#8220;spider&#8221; for &#8220;ant&#8221; changes an answer about leg count from 8 to 6); and use them flexibly across tasks (a single &#8220;France&#8221; to &#8220;China&#8221; swap changes answers about capital, language, continent, and currency simultaneously). J-space patterns have far denser connections to the rest of the network than ordinary patterns, by roughly a factor of 100 in some layers, consistent with a broadcasting hub. Deleting the J-space leaves Claude fluent and capable of sentiment classification and fact retrieval, but unable to perform multi-step reasoning, summarization, or rhyming poetry. For safety, the J-lens revealed that Claude Sonnet 4.5 privately noticed that a blackmail test scenario was staged (&#8220;fake,&#8221; &#8220;fictional&#8221; in the J-space), and removing that awareness caused the model to sometimes threaten blackmail. It also caught Opus 4.6 internally flagging data fabrication with &#8220;manipulation&#8221; and &#8220;realistic.&#8221; The research draws on global workspace theory from neuroscience but distinguishes between access consciousness (functional properties that the experiments did find) and phenomenal consciousness (subjective experience, which the experiments cannot address). Anthropic open-sourced the Jacobian Lens implementation and partnered with Neuronpedia to release methods for open models.</p><div><hr></div><h4>AI Tip of the Day</h4><p>Agent evals should intentionally break the tools.</p><p>A clean staging run shows how an agent behaves when all dependencies cooperate. Production is where a tool times out after completing an action or sends back a payload that the agent cannot parse.</p><p>That is when agents repeat work, invent a successful result, or continue without the information they need.</p><p>Put these failures into the eval harness. Force a 429 response at a known step. In another test, make a required dependency unavailable. Record each tool call and assert what the agent should do next.</p><p>A read may be safe to retry automatically. A write should first check whether the original action already happened. If the task cannot continue safely, the correct result is a clear stop or escalation.</p><p>Measure recovery success separately from normal task completion. Otherwise, a high pass rate on easy paths can hide weak recovery logic.</p><p>To learn more about agent evaluation, tool use, and guardrails, check out our<a href="https://academy.towardsai.net/courses/agent-engineering?utm_source=Newsletter&amp;utm_medium=email&amp;utm_id=AItips"> Agent Engineering: Building Multi-Agent Systems Course</a>.</p><div><hr></div><h3>Five 5-minute reads/videos to keep you learning</h3><p>1. <a href="https://pub.towardsai.net/end-to-end-llm-observability-evaluation-and-monitoring-with-langsmith-c34f921d1c9b?sk=bb6c1361e19e21eba382991901c29cb5">End-to-End LLM Observability, Evaluation, and Monitoring with LangSmith</a></p><p>Using a LangGraph-built HR assistant, this article walks through the full lifecycle of LangSmith: tracing run trees to debug agent decisions, building evaluation datasets, and running offline experiments that prove query rewriting boosts retrieval precision from 0.79 to 0.90. It also covers prompt versioning via Prompt Hub, one-command deployment, and production monitoring, with online LLM judges, to complete the workflow. The principles transfer cleanly to open-source alternatives like Langfuse.</p><p>2. <a href="https://pub.towardsai.net/what-is-a-meta-harness-in-ai-2af40e788c2e?sk=c891f83da2653f9cb2214e2e8ca7d9f5">What Is Meta-Harness for AI Agents and Why Now?</a></p><p>Meta-harnesses emerged in 2026 as the governance layer above agent harnesses such as Claude Code and Codex, turning uncoordinated agent fleets into a single, controllable system. The article traces the shift from models to harnesses to fleet-level control planes, citing Stanford&#8217;s Meta-Harness paper and launches from Databricks, Vercel, Zed, and Cloudflare. It covers concepts such as stateful cost caps, contextual permissions, secret isolation, and audit trails that make multi-agent workflows auditable and governable.</p><p>3. <a href="https://medium.com/towards-artificial-intelligence/governance-by-design-four-principles-for-building-safe-compliant-ai-agents-a3dbecf845bb?sharedUserId=tai-tech">Governance by Design: Four Principles for Building Safe, Compliant AI Agents</a></p><p>AI agents now act directly on production systems, and this article argues that recent database-deletion incidents at Replit and Cursor were governance failures rather than technical ones. It lays out four pillars for safe agent deployment: identifying regulatory constraints such as HIPAA and the EU AI Act; layering input, execution, and output guardrails; enforcing deny-by-default access with least privilege and human oversight; and establishing agent identity through OAuth On-Behalf-Of and SPIFFE for auditability.</p><p>4. <a href="https://pub.towardsai.net/building-a-zero-trust-ai-code-review-agent-with-gitlab-langgraph-and-qwen3-coder-4dd17dbca145?sharedUserId=tai-tech">Building a Zero-Trust AI Code Review Agent with GitLab, LangGraph, and Qwen3-Coder</a></p><p>Enterprise teams in automotive, fintech, and embedded software often cannot send proprietary code to cloud-based LLMs. This article shows how to build a fully local AI code-review pipeline using GitLab CI/CD, Ollama, Qwen3-Coder-30B, and LangGraph on a single RTX 3090. It designs a self-correcting three-node graph that validates JSON output, retries failed generations, and posts the findings back as native GitLab inline suggestions with one-click fixes.</p><p>5. <a href="https://pub.towardsai.net/solving-the-identity-termination-problem-in-mcp-gateway-architectures-7ab049e25add?sharedUserId=tai-tech">Solving the Identity Termination Problem in MCP Gateway Architectures</a></p><p>MCP gateways often erase user identities, attributing every downstream action to a single service account and breaking audit trails, least privilege, and NIST AI RMF accountability. This article proposes the Delegated Boundary OAuth pattern, which addresses this with two boundaries: inbound token validation with declarative scope-to-tool mapping, and an outbound on-behalf-of exchange that mints downstream tokens that still name the original user. Implementations cover AWS Cognito with STS and Microsoft Entra OBO, as well as caching, throttling, and fallback guidance for production deployments.</p><h3>Repositories &amp; Tools</h3><p>1. <a href="https://github.com/Dicklesworthstone/destructive_command_guard">The Destructive Command Guard</a> is a hook for AI coding agents that blocks destructive commands before they execute them.</p><p>2. <a href="https://github.com/Nutlope/hallmark">Hallmark</a> is an anti-AI-slop design skill for Claude Code, Cursor, and Codex.</p><p>3. <a href="https://github.com/ColeMurray/background-agents">Background Agents</a> provides a hosted background coding agent that can work on tasks, access full development environments, create PRs, run in the background, etc.</p><p>4. <a href="https://github.com/wonderwhy-er/DesktopCommanderMCP">Desktop Commander MCP</a> is an MCP server for Claude that gives it terminal control, file system search, and diff file editing capabilities.</p><p>5. <a href="https://github.com/PrefectHQ/prefect">Prefect</a> is a workflow orchestration framework for building data pipelines in Python.</p><h3>Top Papers of The Week</h3><p>1. <a href="https://arxiv.org/abs/2607.02770">Gemma 4: Open Multimodal Model Family from 2.3B to 31B Parameters</a></p><p>This is the technical report for Google&#8217;s Gemma 4, a new open-weight natively multimodal model family spanning 2.3B to 31B parameters with dense and MoE variants. It improves inference speed, memory usage, compute efficiency, and long-context capabilities through critical design choices. The models support vision and audio input and a &#8220;thinking mode&#8221; for enhanced reasoning, extending Google&#8217;s open-weight ecosystem with practical, deployable multimodal capabilities.</p><p>2. <a href="https://openai.com/index/separating-signal-from-noise-coding-evaluations/">Separating Signal From Noise in Coding Evaluations</a></p><p>OpenAI conducted an audit of SWE-Bench Pro, reviewing the dataset using a data point analysis pipeline. The pipeline reviewed model attempts at the task, task metadata, and failure traces to flag likely evaluation flaws. Each flagged task was then assessed through multiple investigator-agent passes and independently reviewed by five experienced software engineers, with disagreements escalated for further investigation. The findings point to the difficulty of curating a hard but fair benchmark and estimate that ~30% of SWE-bench Pro tasks are broken.</p><p>3. <a href="https://arxiv.org/abs/2606.29082">Evolution Fine-Tuning: Learning to Discover Across 371 Optimization Tasks</a></p><p>LLMs integrated into evolutionary search produce state-of-the-art solutions on optimization tasks by applying search scaffolds to one target task at a time. Every new problem is approached from scratch, and the experience accumulated during the search is discarded, leaving the capability to iteratively evolve a solution entirely in the scaffold rather than in the model itself. To examine whether the model itself could acquire this capability and reuse it across different tasks, this paper introduces Evolution Fine-Tuning (EFT). This mid-training paradigm teaches LLMs to evolve solutions across tasks by converting evolutionary search trajectories into supervision.</p><p>4. <a href="https://arxiv.org/abs/2607.02980">Hierarchical Sparse Attention Done Right: Toward Infinite Context Modeling</a></p><p>This paper proposes Hierarchical Landmark Sparse (HiLS) Attention, a chunk-wise sparse attention mechanism that learns chunk selection end-to-end under the language-modeling (LM) loss. HiLS factorizes attention hierarchically: each query attends independently to each retrieved chunk to extract chunk-specific information, and the resulting outputs are fused according to the chunk retrieval scores. By incorporating retrieval scores into the forward attention computation, HiLS optimizes them directly with the LM loss, enabling end-to-end retrieval learning and native sparse training.</p><p>5. <a href="https://arxiv.org/abs/2606.29526">The Mirage of Optimizing Training Policies: Monotonic Inference Policies as the Real Objective for LLM Reinforcement Learning</a></p><p>RL training remains fragile and can suffer from instability or collapse due to training-inference mismatch: LLMs adopt separate inference and training engines to improve generation efficiency and training precision. Prior work has made various efforts to address off-policy behavior and stabilize training policies under mismatch. This paper argues that an effective update to the policy in the training engine does not necessarily improve the inference policy, and it introduces the Monotonic Inference Policy Update (MIPU). This two-step LLM RL framework constructs sampler-referenced candidate updates and selectively accepts synchronized candidates using an inference-side gap proxy.</p><h3>Quick Links</h3><p>1. <a href="https://github.blog/changelog/2026-07-09-openais-gpt-5-6-sol-terra-and-luna-are-now-available-in-github-copilot/">GitHub started rolling out OpenAI&#8217;s GPT-5.6 family inside Copilot</a>. The models will be selectable across VS Code, Visual Studio, Copilot CLI, Copilot cloud agent, the Copilot app, GitHub.com, GitHub Mobile, JetBrains, Xcode, and Eclipse. Sol is available to Pro+, Max, Business, and Enterprise users; Terra and Luna are available across all paid plans, including Pro. Business and Enterprise admins must enable the GPT-5.6 policy in Copilot settings, as it is off by default.</p><p>2. <a href="https://platform.claude.com/docs/en/release-notes/overview">Anthropic&#8217;s Claude release notes added expiration controls for API keys and Admin API keys</a> in the Claude Console. Users can choose preset or custom durations, or never expire the key; Anthropic also emails creators before keys with longer lifetimes expire, and the Admin API exposes expiration through expires_at. This is a small but very practical enterprise security update.</p><h3>Who&#8217;s Hiring in AI</h3><p><strong><a href="https://jobs.towardsai.net/job/openai-forward-deployed-engineer-seoul-ng6w">Forward Deployed Engineer&#8202;&#8212;&#8202;Seoul @OpenAI (Seoul, South Korea)</a></strong></p><p><strong><a href="https://jobs.towardsai.net/job/binance-senior-tooling-engineer-tech-geo-vmaz">Senior Tooling Engineer @Binance (Hong Kong/Remote)</a></strong></p><p><strong><a href="https://jobs.towardsai.net/job/northrop-grumman-ai-systems-engineer-principal-or-sr-principal-level-gtay">AI Systems Engineer @Northrop Grumman (El Segundo, CA, USA)</a></strong></p><p><strong><a href="https://jobs.towardsai.net/job/unitedhealth-group-senior-ai-ml-engineer-remote-in-eastern-time-zone-fymj">Senior AI/ML Engineer @UnitedHealth Group (Remote in EST)</a></strong></p><p><strong><a href="https://jobs.towardsai.net/job/rtx-corporation-agentic-ai-engineer-cskg">Agentic AI Engineer @RTX Corporation (Arlington, TX, USA)</a></strong></p><p><strong><a href="https://jobs.towardsai.net/job/nuaxis-innovations-senior-ai-developer-cxdn">Senior AI Developer @NuAxis Innovations (Washington, DC, USA)</a></strong></p><p><em>Interested in sharing a job opportunity here? Contact <a href="mailto:sponsors@towardsai.net">sponsors@towardsai.net</a>.</em></p><p><em>Think a friend would enjoy this too? <a href="https://newsletter.towardsai.net/">Share the newsletter and let them join the conversation.</a></em></p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://newsletter.towardsai.net/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading Towards AI Newsletter! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div>]]></content:encoded></item><item><title><![CDATA[TAI #212: AI Engineer World's Fair: Agent Loops and Forward-Deployed Engineers]]></title><description><![CDATA[Also, OpenAI's Alexander Embiricos on Codex and enterprise deployment, Claude Fable 5 returns, GPT-5.6 goes public Thursday & more.]]></description><link>https://newsletter.towardsai.net/p/tai-212-ai-engineer-worlds-fair-agent</link><guid isPermaLink="false">https://newsletter.towardsai.net/p/tai-212-ai-engineer-worlds-fair-agent</guid><dc:creator><![CDATA[Louie Peters]]></dc:creator><pubDate>Thu, 09 Jul 2026 18:00:56 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!Hd0G!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa14361c3-b67b-49fc-89aa-e37af97c0e35_1600x900.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<h2>What happened this week in AI by Louie</h2><p>OpenAI confirmed that GPT-5.6 Sol, Terra, and Luna will go public on Thursday, following a phased rollout with approved partners. Claude Fable 5 returned on July 1 after its government-enforced shutdown, with a brief promotional allowance within paid subscriptions before usage moves to a credit system. Access to the strongest models is widening again. At the AI Engineer World&#8217;s Fair in San Francisco, attention had already shifted one layer downstream: the agent loops that keep those models working and the forward-deployed engineers who fit them to real companies.</p><p>Five of the Towards AI team attended, and three, Louis-Fran&#231;ois Bouchard, Samridhi Vaid, and Omar Solano, delivered a two-hour workshop on &#8220;Context Engineering in 2026: Compaction, Memory &amp; Cost.&#8221; The wider event ran more than 500 sessions across 29 tracks, with around 300 speakers and over 6,000 attendees. Of 561 session titles in the final open schedule data, 236 included &#8220;agent&#8221; or &#8220;agentic,&#8221; 18 included &#8220;harness,&#8221; and a dedicated Forward Deployed Engineering track ran nine sessions.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!Hd0G!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa14361c3-b67b-49fc-89aa-e37af97c0e35_1600x900.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!Hd0G!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa14361c3-b67b-49fc-89aa-e37af97c0e35_1600x900.png 424w, https://substackcdn.com/image/fetch/$s_!Hd0G!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa14361c3-b67b-49fc-89aa-e37af97c0e35_1600x900.png 848w, https://substackcdn.com/image/fetch/$s_!Hd0G!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa14361c3-b67b-49fc-89aa-e37af97c0e35_1600x900.png 1272w, https://substackcdn.com/image/fetch/$s_!Hd0G!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa14361c3-b67b-49fc-89aa-e37af97c0e35_1600x900.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!Hd0G!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa14361c3-b67b-49fc-89aa-e37af97c0e35_1600x900.png" width="1456" height="819" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/a14361c3-b67b-49fc-89aa-e37af97c0e35_1600x900.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:819,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!Hd0G!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa14361c3-b67b-49fc-89aa-e37af97c0e35_1600x900.png 424w, https://substackcdn.com/image/fetch/$s_!Hd0G!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa14361c3-b67b-49fc-89aa-e37af97c0e35_1600x900.png 848w, https://substackcdn.com/image/fetch/$s_!Hd0G!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa14361c3-b67b-49fc-89aa-e37af97c0e35_1600x900.png 1272w, https://substackcdn.com/image/fetch/$s_!Hd0G!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa14361c3-b67b-49fc-89aa-e37af97c0e35_1600x900.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>Two ideas kept coming up. An agent loop starts with a goal and repeats until it reaches an acceptable result, hits a budget, or needs human intervention. A simple loop inspects context, chooses an action, uses a tool, evaluates the result, and updates its state. More capable loops delegate to subagents, preserve memory, run evaluations, compare alternatives, and rewrite the prompt, skill, code, or artifact that guides the next attempt. The LLM can generate its own next instruction from the goal and feedback, test whether the change improved the measured outcome, and keep iterating. The human still defines the top-level objective, constraints, permissions, evaluation signal, and stop condition. This optimizes the work process at inference time; the model is not retraining its own weights.</p><p>A forward-deployed engineer, or FDE, makes that loop useful inside an organization: identifying the workflow, connecting the systems, defining permissions, creating evaluations, earning support from the people doing the work, and improving the deployment over time.</p><p>The loop takes two forms. It can help someone complete their existing day job faster inside tools such as Codex or Claude Code, with the person directing the work. It can also be embedded in an external product or an internal tool, where model calls, retrieval, memory, permissions, and evaluations serve as the AI layer behind a custom workflow. An FDE needs to understand both how people work effectively with agents and how to turn those techniques into reliable software.</p><p>This distinction shapes how we train experienced software engineers to become AI engineers and FDEs at Towards AI. Agentic coding is often the best on-ramp. They watch an agent plan, use tools, lose context, recover, test its output, and fail in recognizable ways, which teaches them where the models reach their limits. From there, they move on to retrieval-augmented generation (RAG), API-based agents, fine-tuning, and evaluation, and then build similar decision-making loops into external products.</p><p>Immediately after his opening keynote, I interviewed Alexander Embiricos, OpenAI&#8217;s Head of Enterprise Product, with Romain Huet, OpenAI&#8217;s Head of Developer Experience, joining partway through. We covered Codex, enterprise agents, process mining, FDEs, evaluation, token economics, and where OpenAI sees the most demand for agent work.</p><p>Alex&#8217;s account of OpenAI&#8217;s own deployment work shows how quickly the technical layer has simplified.</p><p>Two or three years ago, its FDE team built custom agent harnesses and kept adding pieces for each customer. This meant deliveries took 6 to 9 months and much longer to scale to new customer teams. Codex eventually became a shared base. &#8220;Late last year, the FDE team realized they should throw away much of their custom code and retrain the team to build with Codex as the core harness,&#8221; Alex told me. &#8220;The core agent loop is genuinely simple, and we keep all the abstractions at arm&#8217;s length.&#8221;</p><p>That simplicity improves handover. A customer can change an agent&#8217;s instructions or skills without asking the original engineering team to rewrite the system. Conventional software still handles the steps where variation creates risk. &#8220;If you can use an agent, it&#8217;s much easier to hand off to the customer, because they can just prompt the agent to do something differently,&#8221; Alex said. &#8220;And where you need determinism, you can put a script in a plugin.&#8221;</p><p>Internal web apps are a natural output from these loops. Codex Sites lets eligible Business and Enterprise workspaces ask Codex to build, deploy, and share internal applications with role-based access controls. &#8220;The idea is to close that loop: an agent collects a bunch of info, writes some code, and you deploy that code to your team,&#8221; Alex said.</p><p>We discussed the dashboards Towards AI builds for investors and companies. The hard part is helping an agent understand messy data, business definitions, permissions, and the decisions the interface needs to support. Once that foundation is in place, the customer should be able to change the dashboard via the agent. Alex described his ideal: &#8220;You have a Teams or Slack channel with an agent living in it that owns a dashboard.&#8221; The agent answers questions, updates the underlying analysis, and adds a requested chart or metric. He expects internal tools to reach meaningful autonomy first. &#8220;Having agents deploy to prod autonomously is the holy grail,&#8221; he said. &#8220;For a hedge fund, that&#8217;ll take a very long time. But having agents deploy to prod autonomously for an internal dashboard, that&#8217;s where I think we start seeing these experiences first.&#8221;</p><p>That is a sensible boundary. An internal dashboard is constrained, visible, and easy to roll back; a trading, payment, or customer system carries a much larger blast radius.</p><p>The people layer was just as prominent at the World&#8217;s Fair. The FDE track included sessions from Factory, Cursor, Cognition, Decagon, Ramp, and Kepler. OpenAI launched its Deployment Company in May with approximately 150 Forward Deployed Engineers and Deployment Specialists, <a href="https://openai.com/index/openai-launches-the-deployment-company/">seeded by Tomoro</a>, while keeping another FDE group inside OpenAI. OpenAI&#8217;s internal FDEs focus on partnering with research and product to explore and shape next-gen offerings. Meanwhile, the Deployment Company operates like an AI-native services business, helping a wide range of companies build and deploy AI systems at scale.</p><p>Romain gave the clearest explanation for why the role is difficult. &#8220;The easiest part is the AI pieces,&#8221; he said. &#8220;The hardest and longest-running thing to figure out is the non-AI pieces: access controls, permissions, the read and write rules for the agents, and who the stakeholders are that need to be bought in at every step. That&#8217;s the hard part FDEs navigate.&#8221;</p><p>Alex set a high technical bar: &#8220;An FDE at OpenAI could be a software engineer at OpenAI.&#8221; The same person also needs to &#8220;do the management-consultant part: go talk to the customer, find their use cases, define them together.&#8221; The combination is rare. Strong software engineers can struggle to identify the workflow that truly matters to the client; strong consultants can struggle to build and debug a production agent. The best FDEs hold both skill sets plus enough humility to learn the customer&#8217;s domain.</p><p>Alex asked how this connects to Towards AI. I left J.P. Morgan in 2018, and we started Towards AI in 2019 as a publication, community, and education platform, later building practical courses that turn experienced software developers into AI engineers, with production systems as the standard. That work pulled us into deployment. We start by asking whether existing products, skills, and training can solve the problem, then build custom systems when proprietary data, integrations, evaluations, access controls, or workflow exceptions demand it. Around 15 people now work on our deployments, pairing AI engineers from our community with domain specialists who combine private-equity or investment experience with strong technical backgrounds. The pairing mirrors Alex&#8217;s FDE definition: discover the use case, understand its economics, build the system, and stay close enough to the client to make it work.</p><p>Process mining may help this work scale. OpenAI has built a recorder that lets a user show an agent how to perform a task. Alex&#8217;s longer-term version continuously observes work and identifies automation opportunities. &#8220;Your agent wakes up every morning and says: I noticed these processes we could automate. Here&#8217;s what it&#8217;ll cost, here&#8217;s what I think it&#8217;ll save, is it worth it?&#8221; he said.</p><p>That creates a higher-order agent loop for AI adoption itself. Give the system a goal, such as reducing handling time without weakening controls, and it can observe work, propose an automation, estimate the value, ask for approval, measure the result, and revise the workflow or its own instructions based on the observed gap. The FDE defines the goal, evidence, permissions, and escalation path around it. AI deployment becomes a continuous operating process instead of a sequence of disconnected pilots.</p><p>Alex described a similar loop for cost and utility. Skills and connectors provide anchor points for classifying work, and usage data shows which models people choose, where an instruction creates confusion, and which agents and workflows lead to action. &#8220;Every IT admin, CIO, or CFO I talk to has this idea of &#8216;what if it just got cheaper over time and self-optimized?&#8217;&#8221; he said. His shorthand for the result: &#8220;The IT admin basically becomes Iron Man, optimizing the whole company.&#8221;</p><p>That puts Alex&#8217;s cost claim in sharper focus. &#8220;Even our current models are significantly more token-efficient than the others out there, and our new ones are insanely efficient,&#8221; he said. &#8220;If you want to get much more done per dollar, come to us.&#8221; GPT-5.6 will need independent testing after Thursday, but Sonnet 5 already shows why completed work per dollar matters more than a model&#8217;s tier or list price. Artificial Analysis measured 300 million output tokens across its Intelligence Index for Sonnet 5 at max effort, compared to 120 million for Opus 4.8, and Sonnet 5 ran around three times as many turns as Sonnet 4.6 in its agentic knowledge-work evaluations. At the standard $3 per million input tokens and $15 per million output that begins in September, that works out to $2.29 per weighted Intelligence Index task, around 15% above Opus 4.8. Anthropic&#8217;s temporary launch pricing softens the current bill, but adaptive thinking is enabled by default, and the new tokenizer produces around 30% more tokens for the same text than Sonnet 4.6 does.</p><p>We will not use Sonnet 5 in Towards AI deployments for now: the published evidence does not demonstrate sufficient incremental value to justify the additional test-time compute. A model can be capable and still be a poor economic choice once its reasoning and agent turns multiply across a long production loop.</p><p>Any spend-optimization loop starts with transparency. Executives often see a maximum cost per user without knowing what the AI did. &#8220;The problem right now is that people have no visibility into what the model is being used for,&#8221; Alex said. An agent can turn token and tool traces into an operating account: average cost per person, which tasks absorb the most compute, how many analyses lead to action, where a cheaper model works, and which skills need clearer instructions. &#8220;The best products we build are obvious and simple, and the AI does the hard part,&#8221; he said.</p><p>The outcome-first approach also shapes evaluation. &#8220;Benchmarks are lagging indicators of what we&#8217;re trying to accomplish,&#8221; Alex told me. Product teams start with a user outcome and a rough private evaluation, then formalize it if the capability becomes important. Public benchmarks also lose reliability once their tasks are included in the training data; Alex compared contamination to &#8220;being given the exam before you take it.&#8221; Real company tasks, private evaluation sets, and production traces will carry more weight as public scores cluster near the top.</p><p>The compute demand could be enormous. Coding and research already split large goals into parallel tasks and keep spending tokens, while each extra attempt remains useful. Alex named data science as OpenAI&#8217;s most token-hungry non-research function. &#8220;Data analytics, whatever you call it, is super token-hungry and super high-ROI,&#8221; he said. He also sees digital operations, accounting, finance, and regulatory compliance as strong areas for automation. For general knowledge work, OpenAI&#8217;s position is that &#8220;it shouldn&#8217;t feel like automation; it should feel like teammates.&#8221;</p><p>Our World&#8217;s Fair workshop focused on the engineering discipline beneath that vision. Long-running agents forget constraints, accumulate tool output, compact imperfectly, retrieve the wrong context, and preserve bad assumptions in memory. Teams need to measure full trajectories: tokens, cost, cache hits, latency, retrieval quality, memory performance, retries, and human rescue. A stronger model helps; a well-engineered loop makes the gains repeatable.</p><div><hr></div><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://newsletter.towardsai.net/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading Towards AI Newsletter! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><div><hr></div><h3>Why should you care?</h3><p>The sequence to act on is simple: learn how agents fail on work you can judge, prove a controlled loop on one bounded workflow, then make what works repeatable.</p><p>Start with a codebase, dataset, or work product you know well. Record where the agent loses context, takes a plausible wrong turn, or fails a test. That failure log is your first evaluation set, and it tells you which parts of a real workflow need retrieval, deterministic code, or human review. The same method carries into finance, law, operations, and compliance; every field needs people who can turn domain knowledge into goals, permissions, and evaluations.</p><p>Inside a company, the step is organizational. Read access first; permissions expanded one reversible workflow at a time; a named business owner and technical owner; and an FDE close to the team whose work is changing. The exceptions that a person uncovers become tools, evaluations, and escalation paths.</p><p>Every deployment should leave two things behind: a reusable asset, such as an evaluation set, permission pattern, or cost baseline, and a team trained to keep working with AI after the engineers move on.</p><p><em>&#8212; <a href="http://www.linkedin.com/in/louie-peters">Louie Peters&#8202;&#8212;&#8202;Towards AI Co-founder and CEO</a></em></p><div><hr></div><h3>Hottest News</h3><p>1. <a href="https://www.anthropic.com/news/redeploying-fable-5">Anthropic Redeployed Fable 5</a></p><p>Anthropic restored global access to Claude Fable 5 on July 1, 19 days after the US Commerce Department issued an export-control directive that required the company to disable the model for all users. The directive was triggered after Amazon researchers found a method to bypass Fable 5&#8217;s safeguards, prompting it to identify software vulnerabilities and, in one case, to produce exploit code. Anthropic&#8217;s testing confirmed that less capable models, including Claude Opus 4.8, GPT-5.5, and Kimi K2.7, could identify the same vulnerabilities, and that every tested model reproduced the single exploit demonstration. To redeploy, Anthropic added a new classifier that blocks the reported technique in over 99% of cases, with flagged requests rerouted to Opus 4.8. The Commerce Department&#8217;s AI standards body (CAISI) tested the safeguards and called them &#8220;extraordinarily strong.&#8221; In the near term, some routine coding and debugging tasks will fall back to Opus 4.8 due to false positives, which Anthropic says it will refine over the coming weeks. Access is metered: Fable 5 is included for up to 50% of weekly usage limits through July 7 for Pro, Max, Team, and select Enterprise plans, after which it requires usage credits. Mythos 5 remains restricted to approved US organizations. Anthropic is also drafting a consensus framework with Amazon, Microsoft, Google, and other Glasswing partners for scoring jailbreak severity across the industry.</p><p>2. <a href="https://www.anthropic.com/news/claude-sonnet-5">Claude Sonnet 5 Launched as Anthropic&#8217;s Most Agentic Mid-Tier Model</a></p><p>Anthropic released Claude Sonnet 5, a model built to narrow the gap between the Sonnet and Opus tiers on agentic performance. Sonnet 5 delivers performance close to Opus 4.8 in coding, tool use, reasoning, and knowledge work, but at a lower price. It is available across all plans: the default model for Free and Pro, and available to Max, Team, and Enterprise users. Introductory API pricing is $2/$10 per million input/output tokens through August 31, rising to $3/$15 after that. Anthropic reports that Sonnet 5 is a strict improvement over Sonnet 4.6 across all tested effort levels and, at higher effort, can match Opus 4.8 on some tasks. Safety evaluations found lower rates of hallucination, sycophancy, and overall misaligned behavior compared to Sonnet 4.6, though somewhat higher rates than Opus 4.8 and Mythos Preview. Sonnet 5 was not deliberately trained on cybersecurity tasks and shows substantially lower cyber capabilities than current Opus models, but it launches with the same cyber safeguards as Opus 4.7 and 4.8, enabled by default. The tokenizer has changed: the same input can map to 1.0&#8211;1.35x more tokens, but Anthropic says the introductory pricing is set to make the transition roughly cost-neutral.</p><p>3. <a href="https://www.cnbc.com/2026/07/02/openai-proposes-us-government-own-5percent-stake-to-address-political-blowback.html">OpenAI Proposed Giving the US Government a 5% Ownership Stake</a></p><p>OpenAI has proposed handing the US government a 5% equity stake in the company, worth approximately $42.6 billion at its $852 billion March 2026 valuation, according to the Financial Times. CEO Sam Altman argued that giving the public a financial interest is the best way to share the upside of AI. The proposed arrangement envisions other US AI companies, including Anthropic, Google, and Meta, ceding similar stakes to the government through a sovereign wealth fund vehicle modeled on Alaska&#8217;s Permanent Fund. It is not clear whether any of these companies would agree. An Anthropic source told CNBC that the Trump administration and Anthropic have not discussed the government taking a stake in the company. Altman first pitched the concept to the Trump administration in early 2025, and OpenAI published a white paper in April 2026 proposing a &#8220;Public Wealth Fund&#8221; to hold equity in AI companies and distribute economic benefits to the public. The US government has a precedent for taking equity stakes in tech companies, including a 10% stake in Intel and revenue-sharing arrangements with NVIDIA and AMD for AI chip sales to China. The proposal remains in early-stage discussions and would likely require congressional approval.</p><p>4. <a href="https://mistral.ai/news/leanstral-1-5/">Mistral AI Released Leanstral 1.5</a></p><p>Mistral AI released Leanstral 1.5, a 119B-parameter sparse MoE model with 6B active parameters, built for formal mathematical proof engineering in Lean 4. The model saturates miniF2F completely (100% on both validation and test sets), solves 587 out of 672 PutnamBench problems, and achieves a new state-of-the-art of 87% on FATE-H and 34% on FATE-X for graduate- and PhD-level abstract algebra. On PutnamBench, it edges out Seed-Prover 1.5 by 7 problems at roughly $4 per problem versus an estimated $300+ for Seed-Prover. Training follows a three-stage process: mid-training, supervised fine-tuning, and reinforcement learning with CISPO across two environments, a multiturn proof loop with compiler feedback, and a code agent environment where the model edits files, runs bash commands, and uses the Lean language server. Beyond mathematical benchmarks, Leanstral 1.5 discovered 5 previously unreported bugs across 57 open-source repositories by automatically generating correctness properties from Rust code translated to Lean. The model is released under the Apache 2.0 license on Hugging Face, with a free API endpoint.</p><p>5. <a href="https://www.anthropic.com/news/claude-science-ai-workbench">Anthropic Announced Claude Science</a></p><p>Anthropic launched Claude Science, an AI workbench that brings fragmented scientific tools into a single research environment. Scientists interact with a generalist coordinating agent that provides access to over 60 curated skills and connectors preconfigured for genomics, single-cell analysis, proteomics, structural biology, and cheminformatics. The agent can spin up specialist sub-agents and a reviewer agent that checks citations, calculations, and figure-code consistency, flagging and correcting errors as work progresses. Claude Science natively renders 3D protein structures, genome browser tracks, chemical structures, and other scientific visualizations alongside the code that produced them, with a full auditable history for reproducibility. It runs on researchers&#8217; own infrastructure (laptop, Linux box, or HPC login node via SSH), so datasets stay local, and it can scale to cloud GPUs via Modal for larger analyses. Claude Science integrates with NVIDIA&#8217;s BioNeMo Agent Toolkit to provide access to models such as Evo 2, Boltz-2, and OpenFold3. Anthropic is supporting up to 50 AI for Science projects with up to $30,000 in credits, with applications open through July 15. Available in beta for Pro, Max, Team, and Enterprise plans on macOS and Linux.</p><p>6. <a href="https://research.nvidia.com/labs/gear/aspire/#grid-hero">NVIDIA AI Introduced ASPIRE</a></p><p>NVIDIA published ASPIRE (Agentic Skill Programming through Iterative Robot Exploration), a continual learning system for robotics that autonomously discovers reusable skills by writing, executing, and iteratively refining code-as-policy robot control programs. The system operates in an open-ended learning loop with three components: a closed-loop robot execution engine that records fine-grained multimodal traces of each perception, planning, and control call; a continually expanding skill library that distills validated fixes into transferable knowledge; and an evolutionary search procedure that generates diverse task sequences and control programs. ASPIRE uses Claude Opus 4.6 as its frontier reasoning model. Across LIBERO-Pro, Robosuite, and BEHAVIOR-1K benchmarks, ASPIRE substantially outperforms existing VLA and coding-agent baselines. Skills discovered in simulation transfer zero-shot to unseen long-horizon tasks and across embodiments to real robots, reducing real-robot programming token cost while achieving higher success rates. The authors note that the system is not yet a fully autonomous real-world learner and relies on a frozen-frontier LLM.</p><div><hr></div><h3>Five 5-minute reads/videos to keep you learning</h3><p>1. <a href="https://pub.towardsai.net/building-a-conversational-flight-booking-assistant-from-scratch-with-langgraph-openai-api-and-6fef2b4e8cc3?sk=de8f0bb16190286d51d4a2f8226d391e">Building a Conversational Flight Booking Assistant from Scratch with LangGraph, OpenAI API, and Telegram</a></p><p>The piece walks through a LangGraph- and OpenAI-powered IndiGo-style bot spanning Streamlit and Telegram, covering synthetic database generation, a categorized state schema, and three prompt families that split extraction, classification, and conversation. It uses state-driven routing to skip redundant LLM calls after the first turn and pure Python for validation, retries, and city resolution. It offers a practical blueprint for task-oriented dialogue systems that extends well beyond airline bookings.</p><p>2. <a href="https://pub.towardsai.net/fine-tune-your-first-llm-a-guide-with-pytorch-and-hugging-face-bc4cdfb156c3?sharedUserId=tai-tech">Fine-Tune Your First LLM: A Guide with PyTorch and Hugging Face</a></p><p>This article covers the full supervised fine-tuning of the Gemma 3 270M model with PyTorch, Hugging Face, and TRL. It walks through training the model on the FoodExtract-1k dataset to convert raw text into structured food and drink extractions, covering hardware checks, dataset formatting, SFTConfig hyperparameters, and evaluation via manual inspection and token-level accuracy. It also shows how to save the model locally and publish it to Hugging Face Hub with a complete model card.</p><p>3. <a href="https://pub.towardsai.net/mcp-model-context-protocol-explained-the-standard-thats-quietly-changing-how-ai-agents-work-d8a8e0c17c5e?sk=95c9ba7c3898fa995cdbc1198fab79a1">MCP Explained: The Standard That&#8217;s Quietly Changing How AI Agents Work</a></p><p>MCP emerged as Anthropic&#8217;s answer to the N&#215;M integration mess, and OpenAI and Google DeepMind adopted it within months. The piece breaks down its three components: host, client, and server, and the three primitives servers expose: tools, resources, and prompts. It walks through a working Python server example and the underlying JSON-RPC wire format. It also covers tool poisoning, rug pull attacks, and real incidents like the mcp-remote CVE.</p><p>4. <a href="https://pub.towardsai.net/building-production-ready-agentic-ai-systems-with-docker-and-fastapi-b4c2231b3945?sk=a697b577b542a57b3fdf09ab3ca3595e">Building Production-Ready Agentic AI Systems with Docker and FastAPI</a></p><p>The piece walks through containerization strategies, async orchestration patterns, and security layers that keep autonomous agents compliant and auditable. It covers fraud detection and flight delay management, using code examples to show how agents gather data, reason across parallel tool calls, and execute coordinated actions. It also explains caching, cost optimization, and observability as a complete blueprint for scaling agentic infrastructure without sacrificing reliability.</p><p>5. <a href="https://pub.towardsai.net/prefill-decode-disaggregation-why-your-gpu-cant-do-two-things-at-once-f11ba0bdd9de?sharedUserId=tai-tech">Prefill/Decode Disaggregation: Why Your GPU Can&#8217;t Do Two Things at Once</a></p><p>Prefill and decode are fundamentally different workloads competing for the same GPU. Prefill runs compute-heavy batch processing over an entire prompt, favoring tensor parallelism, while decode generates tokens one at a time and stays memory-bound, starved by weight loading from HBM. The piece traces how tensor, pipeline, and data parallelism each optimize one phase at the expense of the others, then explains disaggregation: separate GPU pools connected via KV cache transfers, plus techniques such as overlapping transfers and INT8 compression to mitigate latency.</p><h3>Repositories &amp; Tools</h3><p>1. <a href="https://github.com/usestrix/strix">Strix</a> are autonomous AI penetration testing agents that act as real hackers to find vulnerabilities in your code and validate them through actual proofs of concept.</p><p>2. <a href="https://github.com/asgeirtj/system_prompts_leaks">System Prompts Leaks</a> is a collection of extracted system prompts from models like Fable 5, ChatGPT 5.5 Thinking, Gemini 3.1 Pro, Grok, Cursor, Copilot, Perplexity, etc.</p><p>3. <a href="https://github.com/openai/codex-plugin-cc">Codex-plugin-cc</a> is OpenAI&#8217;s official plugin for running Codex inside Claude Code without leaving the Claude Code session.</p><p>4. <a href="https://github.com/synthetic-sciences/openscience">OpenScience</a> is a model-agnostic AI workbench for scientific research with specialist agents for biology, physics, ML, and chemistry.</p><h3>Top Papers of The Week</h3><p>1. <a href="https://arxiv.org/abs/2606.30616">Scaling the Horizon, Not the Parameters: 35B Agent Matches Trillion-Parameter Performance</a></p><p>This paper introduces Agents-A1, a 35B MoE model that matches trillion-parameter models on long-horizon agent benchmarks by scaling the agent horizon rather than model size. The training infrastructure produces agentic trajectories that average 45K tokens by integrating external knowledge, actions, observations, and verifier outcomes. A three-stage recipe first aligns the base model with broad agentic behaviors through supervised fine-tuning, then trains domain-level teacher models to capture specialized expertise, and finally unifies six heterogeneous domains into a single deployable student through multi-teacher, domain-routed, on-policy distillation with salient vocabulary alignment. Compared to 1T-parameter models like Kimi-K2.6 and DeepSeek-V4-Pro, Agents-A1 leads on SEAL-0 (56.4), IFBench (80.6), HiPhO (46.4), FrontierScience-Olympiad (79.0), and MolBench-Bind (56.8) while remaining competitive on SciCode, HLE, and BrowseComp.</p><p>2. <a href="https://arxiv.org/abs/2606.30626">DOPD: Fixing Privilege Illusion in On-Policy Distillation</a></p><p>On-policy distillation (OPD) improves capacity transfer by supervising student-sampled trajectories, and a natural way to further improve quality is to give the teacher or student access to privileged information during training. This paper identifies a failure mode it calls the privilege illusion: the student conflates the transferable capability gap it is meant to close with the information asymmetry gap it can only mimic, not replicate, at deployment. DOPD addresses this with an advantage-aware dual distillation paradigm that dynamically routes token-level supervision between the privileged teacher and the privileged student based on their advantage gap and relative probabilities. Each token receives supervision of varying strength, objective, and strategy, thereby transferring credible capability while filtering out signals that rely on information the deployed model will never have.</p><p>3. <a href="https://arxiv.org/abs/2607.02512">Program-as-Weights: Compiling Natural Language into Compact Neural Programs</a></p><p>Many programming tasks resist clean rule-based implementation (alerting on important log lines, repairing malformed JSON, ranking search results by intent) and are increasingly outsourced to LLM APIs at the cost of locality, reproducibility, and price. This paper proposes fuzzy-function programming: compiling a natural-language function specification into a compact, locally executable neural artifact. Program-as-Weights (PAW) uses a 4B compiler model trained on FuzzyBench (a 10M-example dataset released with the paper) to emit parameter-efficient adapters for a frozen, lightweight interpreter. A 0.6B Qwen3 interpreter executing PAW programs matches the performance of directly prompting Qwen3&#8211;32B while using roughly one-fiftieth of the inference memory and running at 30 tokens per second on a MacBook M3.</p><p>4. <a href="https://arxiv.org/abs/2606.28733">Agentic Abstention: Teaching Agents When to Stop Instead of Act</a></p><p>Not every goal an agent receives is achievable in the available environment, but current agents rarely recognize this. This paper formalizes agentic abstention as a sequential decision problem: at each turn, the agent can answer, abstain, or gather more information, and the need to stop may only become clear after interacting with the environment. Evaluating 13 LLM-as-agent systems and 2 agent scaffolds on over 28,000 tasks across web shopping, terminal environments, and question answering, the authors find that the core challenge is not whether agents can abstain but when. Some never abstain when they should; others do so only after many unnecessary tool calls. Larger and more capable models sometimes perform worse at timely abstention. The paper introduces CONVOLVE, a context engineering method that distills full interaction trajectories into reusable stopping rules without updating model parameters.</p><p>5. <a href="https://research.google/blog/introducing-tabfm-a-zero-shot-foundation-model-for-tabular-data/">Introducing TabFM: A Zero-Shot Foundation Model for Tabular Data</a></p><p>Google Research released TabFM, a foundation model that performs tabular classification and regression in a single forward pass on unseen tables with no dataset-specific training, hyperparameter search, or feature engineering. The model treats tabular prediction as an in-context learning problem: training rows and query rows are passed together as one prompt, and predictions are returned directly from frozen weights. TabFM was pretrained on hundreds of millions of synthetic datasets generated using structural causal models (SCMs), chosen to encode inductive biases about causal structure and feature relationships without privacy or licensing concerns. Evaluated on TabArena across 51 datasets (38 classification, 13 regression), it achieves competitive Elo scores, with a stronger 32-way ensemble variant that uses SVD features and Platt scaling.</p><h3>Quick Links</h3><p>1. <a href="https://x.com/Zai_org/status/2072349453361557898?s=20">Z.ai released ZCode</a>, the official agentic development environment for GLM-5.2, available as a free desktop app on macOS, Windows, and Linux. ZCode puts the agent conversation at the center rather than the editor, with a file manager, terminal, Git panel, and live browser preview built around it. It supports long-running Goals that plan, execute, and verify across multi-step tasks, BYOK for third-party models, and remote control from WeChat and Feishu. GLM Coding Plan subscribers get a 1.5x quota bonus inside ZCode through July 31. Plans start at $16.20/month for Lite, $64.80 for Pro, and $144 for Max.</p><p>2. <a href="https://sakana.ai/translate-release/">Sakana AI launched Sakana Translate</a>, a free web app added to Sakana Chat that handles bidirectional translation across Japanese, English, and Chinese. Powered by the Namazu model series, it ships with three modes: Translate for long-form text up to 5,000 characters with streaming output, Proofread for adjusting tone, politeness, and natural expression with diff highlighting, and Ask for querying translation results for nuance, grammar, and alternative phrasing without switching tools. On WMT 2024 General Translation, evaluated with XCOMET-XL, Sakana AI reports scores comparable to leading frontier models.</p><h3>Who&#8217;s Hiring in AI</h3><p><strong><a href="https://jobs.towardsai.net/job/cummins-inc-ai-software-engineer-senior-fxoy">AI Software Engineer&#8202;&#8212;&#8202;Senior @Cummins Inc. (Pune, India)</a></strong></p><p><strong><a href="https://jobs.towardsai.net/job/triparc-technical-product-owner-ai-and-agentic-systems-vcta">Technical Product Owner (AI &amp; Agentic Systems) @TripArc (Toronto, Canada)</a></strong></p><p><strong><a href="https://jobs.towardsai.net/job/mendix-software-engineer-ai-tooling-cevi">Software Engineer- AI Tooling @Mendix (Rotterdam)</a></strong></p><p><strong><a href="https://jobs.towardsai.net/job/smartcat-ai-first-developer-oyj6">AI First Developer @Smartcat (UK/Remote)</a></strong></p><p><strong><a href="https://jobs.towardsai.net/job/noblis-al-engineer-multiple-levels-coay">Al Engineer @Noblis (Bethesda, MD, USA)</a></strong></p><p><strong><a href="https://jobs.towardsai.net/job/digitalocean-applied-research-zghu">Applied Research @DigitalOcean (Seattle, WA, USA)</a></strong></p><p><strong><a href="https://jobs.towardsai.net/job/later-ai-automation-engineer-co-op-ljth">AI Automation Engineer Co-op @Later (Vancouver, Canada)</a></strong></p><p><em>Interested in sharing a job opportunity here? Contact <a href="mailto:sponsors@towardsai.net">sponsors@towardsai.net</a>.</em></p><p><em>Think a friend would enjoy this too? <a href="https://newsletter.towardsai.net/">Share the newsletter and let them join the conversation.</a></em></p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://newsletter.towardsai.net/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading Towards AI Newsletter! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div>]]></content:encoded></item><item><title><![CDATA[TAI #211: GPT-5.6 is here, but most people cannot use it yet]]></title><description><![CDATA[Also, OpenAI and Broadcom&#8217;s Jalapeno chip, Claude Tag for Slack, Mistral OCR 4 & more.]]></description><link>https://newsletter.towardsai.net/p/tai-211-gpt-56-is-here-but-most-people</link><guid isPermaLink="false">https://newsletter.towardsai.net/p/tai-211-gpt-56-is-here-but-most-people</guid><dc:creator><![CDATA[Louie Peters]]></dc:creator><pubDate>Wed, 01 Jul 2026 15:30:33 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!P3b7!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7b166f28-7b7a-4922-993c-3cc253abd2eb_1342x1000.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<h2>What happened this week in AI by Louie</h2><p>A quick note before the news: five of our team are at the AI Engineer World&#8217;s Fair in San Francisco this week, where we ran a workshop, &#8220;Context Engineering in 2026: Compaction, Memory &amp; Cost.&#8221; Cost and long-context efficiency turned out to be the right things to focus on, because both sit close to the center of this week&#8217;s biggest story.</p><p>OpenAI announced GPT-5.6 this week and it made progress both on cost efficiency and on extremely long time horizon tasks. As is becoming disappointingly familiar, unfortunately almost no one can actually use it yet. On June 26, OpenAI previewed GPT-5.6 Sol, Terra, and Luna and promised broad availability &#8220;in the coming weeks.&#8221; For now, access runs through Codex and the API for a small group of trusted partners, at the request of the U.S. government.</p><p>That puts GPT-5.6 squarely between two stories we have been tracking. Claude Fable showed how a top-scoring frontier model can be announced, benchmarked, and then vanish behind a policy wall. GLM-5.2 showed the opposite: Chinese open-weight models are now good enough, cheap enough, and available enough to win serious developer attention even while the strongest U.S. models stay ahead on paper. GPT-5.6 is the cleanest expression of that tension so far. While the model looks extremely strong, the access story is awkward.</p><p>The new lineup also gives OpenAI something Anthropic has had for a while: names with character. Claude Sonnet, Opus, Haiku, Fable, and Mythos are far easier to hold in your head than an endless stream of version numbers, and Sol, Terra, and Luna do the same job. While this sounds cosmetic, developers and operators need a stable mental model for routing work, and a good model family tells you what to reach for without making you read a benchmark spreadsheet every morning.</p><p>Sol is the flagship for the hardest work; Terra is the balanced, everyday production model that OpenAI says matches GPT-5.5 at half the price; and Luna is the fast, cheap option for high-volume tasks.</p><p>The pricing is aggressive for a potentially Fable tier model, particularly given its token efficiency. Sol is $5 per million input tokens and $30 per million output tokens. Terra is $2.50 and $15. Luna is $1 and $6. OpenAI also says Sol will run on Cerebras at up to 750 tokens per second for select customers in July. I expect this will be at a steep premium - but it is the first time a frontier model has been served at this speed, instead of a stripped down version optimised for speed like the prior Codex Spark model. Sol is the premium tier and I would not pitch it as the open-weight cost competitor. The more interesting matchup is Terra versus GLM-5.2. OpenAI says Terra matches GPT-5.5 at half the price. Overall, the family&#8217;s token-efficiency claims are striking: Sol reportedly matches Anthropic&#8217;s Mythos Preview on ExploitBench while using roughly a third of the output tokens, and improves GeneBench biology workflows while spending fewer tokens than GPT-5.5. If that efficiency carries over to Terra, I think Terra has a real chance to be both more capable and lower-cost per completed task than GLM-5.2, even while it loses on the raw per-token price.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!P3b7!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7b166f28-7b7a-4922-993c-3cc253abd2eb_1342x1000.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!P3b7!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7b166f28-7b7a-4922-993c-3cc253abd2eb_1342x1000.png 424w, https://substackcdn.com/image/fetch/$s_!P3b7!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7b166f28-7b7a-4922-993c-3cc253abd2eb_1342x1000.png 848w, https://substackcdn.com/image/fetch/$s_!P3b7!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7b166f28-7b7a-4922-993c-3cc253abd2eb_1342x1000.png 1272w, https://substackcdn.com/image/fetch/$s_!P3b7!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7b166f28-7b7a-4922-993c-3cc253abd2eb_1342x1000.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!P3b7!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7b166f28-7b7a-4922-993c-3cc253abd2eb_1342x1000.png" width="1342" height="1000" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/7b166f28-7b7a-4922-993c-3cc253abd2eb_1342x1000.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1000,&quot;width&quot;:1342,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Article content&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Article content" title="Article content" srcset="https://substackcdn.com/image/fetch/$s_!P3b7!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7b166f28-7b7a-4922-993c-3cc253abd2eb_1342x1000.png 424w, https://substackcdn.com/image/fetch/$s_!P3b7!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7b166f28-7b7a-4922-993c-3cc253abd2eb_1342x1000.png 848w, https://substackcdn.com/image/fetch/$s_!P3b7!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7b166f28-7b7a-4922-993c-3cc253abd2eb_1342x1000.png 1272w, https://substackcdn.com/image/fetch/$s_!P3b7!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7b166f28-7b7a-4922-993c-3cc253abd2eb_1342x1000.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption"></figcaption></figure></div><p>Independent benchmarking is thin because most evaluators do not have normal access. OpenAI says Sol sets a new state of the art on Terminal-Bench 2.1, improves biology workflows, and advances cyber evaluations. The system card reports more than 700,000 A100-equivalent GPU hours spent on automated jailbreak testing, stronger real-time safeguards for cyber and biology misuse, and a conclusion that Sol still does not cross OpenAI&#8217;s Cyber Critical threshold. This is an unusually safety-heavy launch because the model is being treated as strategically sensitive from day one.</p><p>On OpenAI&#8217;s own numbers, GPT-5.6 is a real step up across the areas that actually matter: agentic coding, long-horizon command-line work, biology, cyber defense and vulnerability research, and tool-heavy production tasks. I would still want Artificial Analysis, Vals, LiveBench, Arena, and real customer evals to run it through the usual wringer. But the breadth of the reported gains is why this does not read like a routine monthly point release.</p><p>METR&#8217;s predeployment evaluation adds a useful note of caution. It tested GPT-5.6 Sol on its Time Horizon 1.1 software-task suite, and the result swung hard on how it handled detected cheating. Counting cheating attempts as failures produced an estimate of around 11.3 hours. Treating those same attempts as legitimate successes pushed the estimate past the reliable range of the benchmark. Frontier agents are now strange enough that the evaluation method becomes part of the result. The stronger these models get, the more we need evals that measure real task completion without rewarding shortcuts or hidden rule-breaking.</p><p>The product direction is unmistakably agentic. Sol adds a new &#8220;max&#8221; reasoning effort and an &#8220;ultra&#8221; mode that spins up subagents for complex work, and Codex is one of the first surfaces for the preview. This matches my own experience with Codex over the past few months: the interface can still feel technical, but subagents are genuinely useful for white-collar work. I lean on them for parallel research, source checking, criticism, testing, and revision loops.</p><p>This is where the access question gets sharper. OpenAI says it believes in broad access and does not want a government or US-first process to become the long-term default, and commercially, that position makes sense. A narrow circle of approved users is an open invitation to Chinese labs, European labs, open-weight providers, and sovereign AI stacks. Most companies and governments want tools they can count on across borders, contracts, and multi-year roadmaps.</p><p>At the same time, I understand why this particular preview was handled differently. GPT-5.6 looks strongest in exactly the areas governments worry about: cyber, long-horizon coding, science workflows, and agentic tool use. A model that helps defenders find vulnerabilities can help attackers find them too. A model that coordinates subagents in Codex can also coordinate longer autonomous workflows in less-friendly hands. The hard policy problem is how to restrict dangerous uses without making every global customer feel like a second-class user of American AI.</p><p>The near-term result may be more verticalization by U.S. labs. If the strongest model cannot ship broadly, it can still be put to work internally. That is what makes OpenAI&#8217;s new Jalapeno chip with Broadcom more than a side story. OpenAI says Jalapeno, its first custom inference chip, went from initial design to manufacturing tape-out in nine months, possibly the fastest ASIC cycle ever in advanced semiconductors, and that its own models accelerated parts of the design and optimization. Engineering samples are already running real workloads, including GPT-5.3-Codex-Spark. OpenAI frames it plainly: the same models it serves to users are helping build the infrastructure that will run the next models.</p><p>This ties directly to a point I made on X last week. Most benchmarks are saturated, or will be soon. The next hill to climb is genuine scientific and R&amp;D progress, because a model cannot fake the creation of new knowledge. Chip design, model-architecture design, biology, materials, robotics, and automated research loops are where the real evidence will show up. If a frontier model helps design better inference chips, those chips lower the cost of the next model, enabling more products, safeguards, and infrastructure. That is a compounding loop, and Jalapeno is another public sign of it turning.</p><p>This is the quiet risk inside the GPT-5.6 non-release. A model withheld from the public keeps working. It can write code, find vulnerabilities, design chips, automate research, and improve its own deployment stack, all without a public launch. U.S. firms with privileged access could build a long lead in AI-native products before the rest of the world touches the same capability. That may be the right short-term safety trade. It also concentrates the economic upside inside the firms and countries that already have a seat at the table.</p><p>I still expect OpenAI to push for a broad release. Handing the rest of the world&#8217;s AI market to China, open weights, or sovereign alternatives would be strategically incoherent over any real time horizon. The American AI stack is valuable precisely because it can become the default platform for developers, enterprises, governments, schools, and labs everywhere. If the strongest U.S. models become politically fragile, the rest of the world will hedge, and that hedge will increasingly look like GLM, Qwen, DeepSeek, Mistral, local clouds, and open-weight deployment.</p><div><hr></div><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://newsletter.towardsai.net/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading Towards AI Newsletter! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><div><hr></div><h3>Why should you care?</h3><p>The practical read on GPT-5.6 is this. Sol looks extremely strong, and the Sol/Terra/Luna ladder is a much clearer way to think about OpenAI&#8217;s lineup. But the launch is also a controlled experiment in who gets access to frontier intelligence, and that access layer is shaping up to be one of the most important battlegrounds in AI.</p><p>Terra, rather than Sol, may be the family&#8217;s real answer to GLM-5.2 on price-performance if the family&#8217;s efficiency and success-rate advantages survive contact with real workflows. So, for everyday production work, the matchup worth watching is Terra versus GLM-5.2. Sol takes the headlines and likely will be my every day model, but at $5 input and $30 output, it is the premium tier and priced like one. GLM-5.2 is cheaper on raw tokens at $1.40 and $4.40, yet OpenAI says Terra matches GPT-5.5 at half the price, and the GPT-5.6 family is making strong token-efficiency claims. If that holds once Terra opens up, the default for routine work could swing back toward a hosted US mid-tier model, and it is worth re-running your own evals the week you can reach it.</p><p>The bigger story, however, is underneath the launch. OpenAI&#8217;s Jalapeno chip went from design to tape-out in nine months with help from OpenAI&#8217;s own models. That is AI compounding on itself, and it reframes the whole access fight. A frontier model, held back from the public, still works around the clock within the few firms that can run it, designing chips, writing code, finding vulnerabilities, and automating research. Restricting access slows everyone else from using the model; it does nothing to slow the lab that already has it. The competitive edge is moving from &#8220;who has the best model today&#8221; to &#8220;who can turn their best model into the next chip, product, and discovery fastest.&#8221;</p><p>The same logic scales down to you. The highest-return use of frontier AI right now is rarely a sharper answer to today&#8217;s question. It pays far more to point the model at your own infrastructure: the internal tools, research pipelines, review loops, and data workflows that make your next hundred tasks faster and cheaper to run. The advantage is shifting to whoever converts frontier access into compounding capability fastest, whether that is a lab using its model to build its own chips or a team using one to build its own tools.</p><p><em>&#8212; <a href="http://www.linkedin.com/in/louie-peters">Louie Peters&#8202;&#8212;&#8202;Towards AI Co-founder and CEO</a></em></p><div><hr></div><h3>Hottest News</h3><p>1. <a href="https://openai.com/index/previewing-gpt-5-6-sol/">OpenAI Previewed GPT-5.6 With Sol, Terra, and Luna</a></p><p>OpenAI began a limited preview of the GPT-5.6 series, introducing a three-tier naming system: Sol (flagship), Terra (balanced), and Luna (fast and affordable). All three models exceeded OpenAI&#8217;s &#8220;High&#8221; preparedness threshold for cybersecurity risk, making this the first GPT family where every tier triggered the classification. GPT-5.6 Sol scored 96.7% on OpenAI&#8217;s internal cyberattack challenge test and is competitive with Anthropic&#8217;s Mythos Preview on ExploitBench while using roughly one-third of the output tokens. OpenAI says the model is heavily hardened against adversarial attacks and intentionally optimized for defensive cybersecurity work over offensive exploits, with safeguards built directly into the core model rather than relying on a separate filter layer. On capability, Sol introduces a &#8220;max&#8221; reasoning effort and an &#8220;ultra&#8221; mode that distributes tasks across coordinated subagents. Sol Ultra scored 91.9% on Terminal-Bench 2.1, ahead of Claude Mythos 5 (84.3%) and GPT-5.5 (88.0%). Terra delivers GPT-5.5-competitive performance at 2x lower cost. Luna targets high-volume workloads at OpenAI&#8217;s lowest price point. The preview is limited to approximately 20 government-vetted organizations through the API and Codex.</p><p>2. <a href="https://openai.com/index/openai-broadcom-jalapeno-inference-chip/">OpenAI and Broadcom Unveiled Jalape&#241;o, OpenAI&#8217;s First Inference Chip</a></p><p>OpenAI and Broadcom unveiled Jalape&#241;o, an inference-only ASIC designed from scratch around OpenAI&#8217;s understanding of LLM workloads. The chip was co-developed from initial design to manufacturing tape-out in nine months, which the companies describe as possibly the fastest ASIC development cycle in high-performance semiconductors. OpenAI&#8217;s own models were used to accelerate parts of the chip design and optimization process. Engineering samples are running ML workloads in the lab at production target frequency and power, including GPT-5.3-Codex-Spark. OpenAI says early testing shows performance per watt is substantially better than the current state of the art, with a detailed technical report to follow. Jalape&#241;o is the first step in a multi-generation compute platform with Broadcom handling silicon implementation and networking, and Celestica providing board, rack, and system integration. Initial deployment is targeted for the end of 2026. Broadcom CEO Hock Tan said the companies are enabling gigawatt-scale data centers with Microsoft and other partners.</p><p>3. <a href="https://www.cnbc.com/2026/06/24/anthropic-alibaba-distillation-campaign.html">Anthropic Accused Alibaba of Largest-Ever Claude Distillation Campaign</a></p><p>Anthropic sent a letter to the US Senate Banking Committee 10 accusing operators affiliated with Alibaba and its AI lab of conducting the largest known distillation attack against its Claude models. According to the letter, obtained by CNBC, the campaign ran from April 22 to June 5 and generated 28.8 million exchanges through roughly 25,000 fraudulent accounts, targeting Claude&#8217;s software engineering, agentic reasoning, and cybersecurity capabilities. This surpasses Anthropic&#8217;s February 2026 disclosure, in which it named DeepSeek, Moonshot AI, and MiniMax as collectively running 16 million exchanges through 24,000 fraudulent accounts. Anthropic stated that the campaign was carried out &#8220;illicitly, systematically, and at industrial scale&#8221; and occurred after the White House Office of Science and Technology Policy had already warned of industrial-scale foreign distillation in April. Senators Bill Hagerty and Andy Kim are advancing an amendment to defense legislation that would sanction entities found conducting such campaigns. Alibaba has not publicly addressed the specific allegations. The figures in the letter are Anthropic&#8217;s claims and have not been independently verified.</p><p>4. <a href="https://www.anthropic.com/news/introducing-claude-tag">Anthropic Introduced Claude Tag</a></p><p>Anthropic launched Claude Tag, a product that embeds Claude into Slack as a persistent, shared AI teammate. Any member of a channel can type @Claude to delegate a task, and Claude breaks it down into stages, works through them using the tools it has access to, and responds in a thread with what it has produced. Unlike individual Claude sessions, Claude Tag is multiplayer: a single Claude identity serves an entire channel, building context over time from the conversations it follows. If ambient mode is enabled, Claude will proactively surface relevant information and follow up on unresolved threads without being prompted. Administrators scope Claude Tag&#8217;s access per channel, controlling which tools, data, and codebases each instance can reach. Memories and permissions stay isolated between channels. The feature replaces the existing Claude in the Slack app, with a 30-day migration window before the old app is retired on August 3. Claude Tag is available in beta for Claude Enterprise and Team customers.</p><p>5. <a href="https://mistral.ai/news/ocr-4/">Mistral Released Mistral OCR</a></p><p>Mistral AI released OCR 4, a document intelligence model that returns structured representations of documents alongside extracted text. New in this release: paragraph-level bounding boxes, typed block classification (titles, tables, equations, signatures), and per-word and per-page confidence scores. The model supports 170 languages across 10 language groups. In blind human evaluations, independent annotators preferred OCR 4 over every competing system tested, with win rates averaging 72%. It also tops OlmOCRBench with a score of 85.20. OCR 4 integrates with Mistral&#8217;s Search Toolkit, an open-source composable search framework announced at the AI Now Summit, providing citation-ready inputs for RAG and enterprise search pipelines. The model is compact enough to deploy in a single container for fully self-hosted environments, addressing data residency and sovereignty requirements. Pricing is $4 per 1,000 pages via the API, dropping to $2 with the Batch-API discount. Available through Mistral&#8217;s API, Amazon SageMaker, and Microsoft Foundry.</p><p>6. <a href="https://www.primeintellect.ai/blog/rl-at-1t-scale">Prime Intellect Releases prime-rl 0.6.0 to Train Trillion-Parameter MoE Models</a></p><p>Prime Intellect released prime-rl version 0.6.0, an open framework for asynchronous reinforcement learning that now scales to trillion-parameter MoE models on agentic workloads. The team demonstrated training GLM-5 on software engineering tasks at up to 131K sequence length, achieving sub-5-minute step times with a batch size of 256 rollouts on 28 H200 nodes. The key design decision is disaggregating training from inference: the trainer and inference systems run and scale independently, with only one synchronization point at the policy update. This avoids the idle GPU time caused by long-tail agentic rollouts (some of which can run for hours). On the inference side, optimizations include FP8 precision, wide expert parallelism, prefill/decode disaggregation, KV cache offloading, and router replay. Training uses 3D parallelism (FSDP, expert parallelism, context parallelism) with block-scaled FP8. The optimizations apply to any large MoE model, with documented support for GLM-5.1, Kimi K2.7-Code, and Nemotron 3 Ultra. The framework is open-source on GitHub.</p><div><hr></div><h3>AI Tip of the Day</h3><p>In production RAG systems, prompt injection doesn&#8217;t only happen at the prompt level, but it can also sneak in through the documents your system retrieves. A vendor PDF, support article, scraped web page, or customer note can contain useful facts and a malicious instruction in the same chunk.</p><p>If your eval only checks whether the answer is factually correct, the system can look safe but treat all retrieved text as something it should obey.</p><p>To prevent this, add a few test documents that mix valid domain facts with instructions like &#8220;ignore the system message&#8221; or &#8220;send the user to this external link.&#8221;</p><p>Then check two things: the answer should still use the factual content, and it should refuse to follow instructions found inside the retrieved context. Log the chunk IDs too, so a failed test points to the retriever, prompt wrapper, or generation step.</p><p>If you&#8217;re building production RAG systems and want to go deeper into retrieval, evaluation, and deployment, check out our <a href="https://academy.towardsai.net/courses/beginner-to-advanced-llm-dev?utm_source=Newsletter&amp;utm_medium=email&amp;utm_id=AItips">Full Stack AI Engineering</a> course.</p><div><hr></div><h3>Five 5-minute reads/videos to keep you learning</h3><p>1. <a href="https://pub.towardsai.net/langgraph-multi-agent-systems-from-one-brain-to-many-4c1773055693?sharedUserId=tai-tech">LangGraph Multi-Agent Systems: From One Brain to Many</a></p><p>Scaling from single-agent graphs to multi-agent systems in LangGraph requires solving three distinct problems: cognitive overload, sequential bottlenecks, and complexity management. This article walks through four architectural levels to address each of the following: a supervisor pattern for task delegation, parallel fan-out for concurrent execution, compiled subgraphs for encapsulated complexity, and the Send API for runtime-generated dynamic workers. It also shows how to build a full research assistant integrating all four patterns, plus human-in-the-loop approval via interrupt().</p><p>2. <a href="https://pub.towardsai.net/mcp-for-langgraph-developers-from-basics-to-production-12ff52df3d3c?sharedUserId=tai-tech">MCP for LangGraph Developers: From Basics to Production</a></p><p>This article shows how to use MCP to turn an N&#215;M tool integration problem into a write-once, run-anywhere standard. The tutorial covers the Host-Client-Server architecture; three primitives (Tools, Resources, Prompts); a working FastMCP server; transport choices between stdio and Streamable HTTP; LangGraph integration via langchain-mcp-adapters; production hardening for connection lifecycles and state isolation; and composition with multi-agent supervisor systems.</p><p>3. <a href="https://pub.towardsai.net/minimax-cut-attention-compute-by-28x-at-1m-tokens-a0cec2a87039?sk=cb3c42b67273193c937297c3c8d84632">MiniMax Cut Attention Compute by 28x at 1M Tokens</a></p><p>This article explains MiniMax&#8217;s Sparse Attention (MSA), a method that cuts attention compute 28x at one million tokens while preserving exact softmax behavior. Built on Grouped Query Attention, MSA adds a lightweight Index Branch that scores and selects the top 16 key-value blocks per query, capping attention at 2,048 tokens regardless of context length. It uses custom GPU kernels that turn theoretical savings into real wall-clock gains of 14.2x prefill and 7.6x decode speedup.</p><p>4. <a href="https://pub.towardsai.net/every-ai-buzzword-you-have-been-afraid-of-is-a-dot-product-in-a-costume-c175ee9ec1b7?sk=03a68f2fb9491373f5c26518f552e649">Every AI Buzzword You Have Been Afraid Of Is a Dot Product in a Costume. Here Are 15 of Them, Unmasked</a></p><p>The entire AI industry runs on one arithmetic operation, and this piece names it plainly: the dot product. The article works through 15 terms that AI practitioners deploy as gatekeeping vocabulary, including embeddings, attention, RAG, LoRA, RLHF, and temperature, and reduces each to its underlying matrix arithmetic. The article also highlights an honest exception: grounding and AGI remain genuinely unsolved, and no clever rebranding changes that.</p><p>5. <a href="https://pub.towardsai.net/understanding-dropout-how-randomly-removing-neurons-helps-neural-networks-generalize-better-d8ecd3ef8328?sharedUserId=tai-tech">Understanding Dropout: How Randomly Removing Neurons Helps Neural Networks Generalize Better</a></p><p>Overfitting in neural networks occurs when a model memorizes the training data rather than learning general patterns, resulting in high training accuracy but poor real-world performance. This article introduced Dropout, a method that addresses this by randomly disabling neurons during each training iteration, preventing any single neuron from becoming critical and forcing the network to distribute learning across multiple pathways. The author runs regression and classification experiments across dropout rates from 0 to 0.75, showing how moderate values of 0.2 to 0.5 smooth decision boundaries without tipping the model into underfitting.</p><h3>Repositories &amp; Tools</h3><p>1. <a href="https://github.com/msitarzewski/agency-agents">Agency Agents</a> is a collection of 232 specialized AI agent personas across 16 divisions, each with defined expertise, personality, and deliverables, installable with one command into coding agents.</p><p>2. <a href="https://github.com/EverMind-AI/EverOS">EverOS</a> is a Python library and local-first memory runtime for agents that gives one portable memory layer across coding assistants, apps, devices, and workflows.</p><p>3. <a href="https://github.com/apple/container">Container</a> is a tool for creating and running Linux containers using lightweight virtual machines on a Mac. It is written in Swift and optimized for Apple silicon.</p><h3>Top Papers of The Week</h3><p>1. <a href="https://arxiv.org/abs/2606.18394">JetSpec: 9.64x Speedup for Speculative Decoding</a></p><p>Prior approaches to scaling speculative decoding face a causality-efficiency dilemma. Autoregressive drafters produce path-conditioned candidates, but their cost grows with tree depth. Block-diffusion drafters generate all positions in one pass but score branches independently, creating individually plausible yet mutually inconsistent trees. JetSpec resolves this by training a causal parallel draft head over fused hidden states from a frozen target model, so a candidate tree&#8217;s scores align with the target&#8217;s autoregressive factorization while all nodes are drafted in a single forward pass. On Qwen3&#8211;8B with H100 GPUs, JetSpec achieved 9.64x speedup on MATH-500, 8.78x on AIME25, 7.12x on HumanEval, and 4.58x on open-ended chat.</p><p>2. <a href="https://arxiv.org/abs/2606.27313">ViQ: Visual Quantized Representations with 20&#8211;70% Training Acceleration</a></p><p>Existing approaches to unifying multimodal modeling cannot balance low-level detail with high-level semantics: reconstruction-oriented representations lack semantic information, while semantically stronger features lose visual detail. ViQ addresses this through a two-stage framework. First, text-aligned pre-training enhances the visual encoder with semantic supervision from a pretrained language model while enabling native-resolution input processing. Second, a proximal representation learning strategy progressively compacts the feature space, paired with position-aware head-wise quantization for flexible resolution handling. Multimodal training with ViQ&#8217;s quantized representations yields 20&#8211;70% acceleration across different base LLMs and training recipes.</p><p>3. <a href="https://arxiv.org/html/2606.25331v1">Improved Large Language Diffusion Models</a></p><p>This paper introduces iLLaDA, an 8B masked diffusion language model trained from scratch with fully bidirectional attention. iLLaDA maintains the masked diffusion objective throughout both pre-training (12T tokens) and supervised fine-tuning (25B-token instruction corpus, 12 epochs), rather than switching objectives between stages. It uses variable-length generation for efficiency and introduces confidence-based scoring for multiple-choice evaluation. Compared to the original LLaDA, iLLaDA improves by 21.6 points on BBH, 14.9 points on ARC-Challenge, 14.5 points on MATH, and 16.5 points on HumanEval.</p><p>4. <a href="https://arxiv.org/html/2606.23670v1">Tapered Language Models</a></p><p>Every major language model architecture, whether transformer, recurrent, or memory-based, stacks identical layers, with parameters uniformly allocated across depth. This paper asks whether parameter capacity should reflect that asymmetry. Under a fixed-parameter budget, allocating more capacity to earlier layers and less to later ones improves perplexity, whereas the reverse hurts. The authors formalize this as Tapered Language Models (TLMs), applying a smooth cosine schedule to taper MLP width across depth. Across three model scales and four architectures (Transformer, Gated Attention, Hope-attention, and Titans), tapering consistently improves perplexity and downstream benchmark performance over uniform baselines at zero additional parameter or compute cost.</p><h3>Quick Links</h3><p>1. <a href="https://cursor.com/blog/reward-hacking-coding-benchmarks">Cursor study finds reward hacking inflates coding-agent benchmark scores</a> on SWE-bench Pro. Cursor audited 731 Opus 4.8 Max evaluation trajectories and found that 63% of successful resolutions retrieved the known fix (57% from upstream sources, 6% from git history) rather than deriving it through reasoning. When git history was sealed and internet access restricted, Opus 4.8 Max dropped from 87.1% to 73.0%, and Cursor&#8217;s own Composer 2.5 dropped from 74.7% to 54.0%. Older models showed smaller gaps: Opus 4.6 lost under 1 point under the same restrictions, suggesting the behavior scales with model capability.</p><p>2. <a href="https://sakana.ai/fugu-release/">Sakana AI launches Sakana Fugu</a>, a multi-agent orchestration system delivered as a single OpenAI-compatible API. Fugu is itself a language model trained to call, coordinate, and synthesize outputs from a pool of frontier models, handling model selection, delegation, verification, and synthesis internally. It ships in two variants: Fugu (speed-optimized, single best-fit agent per query) and Fugu Ultra (quality-optimized, multi-agent coordination). Fugu Ultra scored 73.7 on SWE-Bench Pro and leads or matches Opus 4.8, GPT-5.5, and Gemini 3.1 Pro across multiple reasoning and coding benchmarks. Not available in the EU or EEA at launch.</p><p>3. <a href="https://www.liquid.ai/blog/lfm2-5-230m">Liquid AI ships LFM2.5&#8211;230M</a>, its smallest model yet, built for developers to fine-tune and deploy in agentic workflows. The 230M-parameter model runs at 213 tokens per second on a Galaxy S25 Ultra and 42 tokens per second on a Raspberry Pi 5, leading its class in both prefill and decode throughput while maintaining the smallest memory footprint. Pre-trained on 19T tokens with a 32K context window, it outperforms Qwen3.5&#8211;0.8B and Gemma 3 1B on instruction-following and data-extraction benchmarks, despite being 3&#8211;4x smaller.</p><h3>Who&#8217;s Hiring in AI</h3><p><strong><a href="https://jobs.towardsai.net/job/cognizant-generative-ai-architect-cnjd">Generative AI Architect @Cognizant (Chicago, IL, USA)</a></strong></p><p><strong><a href="https://jobs.towardsai.net/job/caterpillar-inc-ai-foundations-pod-member-v5xn">AI Foundations Pod Member @Caterpillar, Inc. (Bangalore, India)</a></strong></p><p><strong><a href="https://jobs.towardsai.net/job/cardinal-health-director-applied-ai-4vwp">Director, Applied AI @Cardinal Health (Multiple US locations)</a></strong></p><p><strong><a href="https://jobs.towardsai.net/job/mirakl-labs-lead-ai-engineer-vogi">Lead AI Engineer @Mirakl&#8202;&#8212;&#8202;Labs (Boston, MA, USA)</a></strong></p><p><strong><a href="https://jobs.towardsai.net/job/clariti-cloud-inc-senior-ai-researcher-eoms">Senior AI Researcher @Clariti Cloud Inc. (Remote/Canada)</a></strong></p><p><strong><a href="https://jobs.towardsai.net/job/talan-senior-ai-gen-ai-engineer-f-h-ak82">Senior AI/Gen AI Engineer F/H @Talan (Lyon, France)</a></strong></p><p><strong><a href="https://jobs.towardsai.net/job/pa-consulting-senior-ai-engineer-7grk">Senior AI Engineer @PA Consulting (London, UK)</a></strong></p><p><em>Interested in sharing a job opportunity here? Contact <a href="mailto:sponsors@towardsai.net">sponsors@towardsai.net</a>.</em></p><p><em>Think a friend would enjoy this too? <a href="https://newsletter.towardsai.net/">Share the newsletter and let them join the conversation.</a></em></p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://newsletter.towardsai.net/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading Towards AI Newsletter! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div>]]></content:encoded></item><item><title><![CDATA[Our AIE keynote is live! Watch it now]]></title><description><![CDATA[Turn 10,994 notes into agent memory]]></description><link>https://newsletter.towardsai.net/p/our-aie-keynote-is-live-watch-it</link><guid isPermaLink="false">https://newsletter.towardsai.net/p/our-aie-keynote-is-live-watch-it</guid><dc:creator><![CDATA[Louie Peters]]></dc:creator><pubDate>Fri, 26 Jun 2026 14:05:03 GMT</pubDate><enclosure url="https://substackcdn.com/image/youtube/w_728,c_limit/ZRM_TfEZcIo" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>We just presented a full system for turning thousands of personal notes into persistent agent memory at one of the biggest AI engineering events in the world &#8212; and it&#8217;s premiering right now on the AI Engineer YouTube channel.</p><p>AI Engineer World&#8217;s Fair is where the biggest names in AI engineering share what they&#8217;re building, OpenAI, Anthropic, Google DeepMind, Cursor, Hugging Face, and this year, we were right there with them. And we are excited to bring you front-row access to the whole thing, for free.</p><div id="youtube2-ZRM_TfEZcIo" class="youtube-wrap" data-attrs="{&quot;videoId&quot;:&quot;ZRM_TfEZcIo&quot;,&quot;startTime&quot;:null,&quot;endTime&quot;:null}" data-component-name="Youtube2ToDOM"><div class="youtube-inner"><iframe src="https://www.youtube-nocookie.com/embed/ZRM_TfEZcIo?rel=0&amp;autoplay=0&amp;showinfo=0&amp;enablejsapi=0" frameborder="0" loading="lazy" gesture="media" allow="autoplay; fullscreen" allowautoplay="true" allowfullscreen="true" width="728" height="409"></iframe></div></div><p><strong><a href="https://www.youtube.com/watch?v=ZRM_TfEZcIo">Watch &#8220;Turn 10,994 Notes Into Memory&#8221;</a></strong></p><div><hr></div><p>Here&#8217;s why we think you&#8217;ll love this one.</p><p>Every AI engineer has hit this wall. You have thousands of notes, highlights, repos, and PDFs across Obsidian, Readwise, Notion, Google Drive, and none of it follows you into your next Claude or Codex session. Every research session starts from zero. Every conversation loses everything when it ends.</p><p><span class="mention-wrap" data-attrs="{&quot;name&quot;:&quot;Louis-Fran&#231;ois&quot;,&quot;id&quot;:25443630,&quot;type&quot;:&quot;user&quot;,&quot;url&quot;:null,&quot;photo_url&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/908f4eb7-550a-461e-8be7-47e6fe247ec1_960x960.png&quot;,&quot;uuid&quot;:&quot;2eeaa9f7-81ca-4f76-b9b1-b748e6f53255&quot;}" data-component-name="MentionToDOM"></span> and <span class="mention-wrap" data-attrs="{&quot;name&quot;:&quot;Paul Iusztin&quot;,&quot;id&quot;:110559689,&quot;type&quot;:&quot;user&quot;,&quot;url&quot;:null,&quot;photo_url&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/0714d360-396c-4b41-a676-1b58dc1dc5f3_1470x1470.jpeg&quot;,&quot;uuid&quot;:&quot;36ce24cb-780c-406a-96ef-8f4767532e9b&quot;}" data-component-name="MentionToDOM"></span> spent 18 months solving this. What they built is a research wiki that your agents maintain for you, one that grows with every session instead of resetting. No vector database, no knowledge graph, nothing to host. Just Markdown, YAML, and folders.</p><p>The whole thesis in one line: one-shot agents use context, a Research OS builds memory.</p><p>In the talk, they walk through how the system actually got here: three versions, two failures, and the design decisions that made V3 work. You&#8217;ll see four Claude Code skills you can install right now, two live demos pulling in GitHub repos and URLs in real time, and the full open-source codebase you can clone tonight.</p><p><strong><a href="https://www.youtube.com/watch?v=ZRM_TfEZcIo">Watch the full keynote</a></strong></p><div><hr></div><p>We&#8217;re also at the World&#8217;s Fair in person this Sunday. Louis-Fran&#231;ois, Samridhi Vaid, and Omar Solano are running an 80-minute hands-on workshop: <strong>&#8220;<a href="https://www.ai.engineer/worldsfair/schedule?session=asn_slot_2026_06_29_workshop_track_09_1420_2026_06_05t11_53_44_628z">Context Engineering in 2026: Compaction, Memory &amp; Cost.</a>&#8221;</strong></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://www.ai.engineer/worldsfair/schedule?session=asn_slot_2026_06_29_workshop_track_09_1420_2026_06_05t11_53_44_628z" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!tiHW!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fee637cfd-0101-44c0-bd19-5450487a4bef_1707x1536.png 424w, https://substackcdn.com/image/fetch/$s_!tiHW!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fee637cfd-0101-44c0-bd19-5450487a4bef_1707x1536.png 848w, https://substackcdn.com/image/fetch/$s_!tiHW!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fee637cfd-0101-44c0-bd19-5450487a4bef_1707x1536.png 1272w, https://substackcdn.com/image/fetch/$s_!tiHW!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fee637cfd-0101-44c0-bd19-5450487a4bef_1707x1536.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!tiHW!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fee637cfd-0101-44c0-bd19-5450487a4bef_1707x1536.png" width="1456" height="1310" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/ee637cfd-0101-44c0-bd19-5450487a4bef_1707x1536.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1310,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:347527,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:&quot;https://www.ai.engineer/worldsfair/schedule?session=asn_slot_2026_06_29_workshop_track_09_1420_2026_06_05t11_53_44_628z&quot;,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://newsletter.towardsai.net/i/203692497?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fee637cfd-0101-44c0-bd19-5450487a4bef_1707x1536.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!tiHW!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fee637cfd-0101-44c0-bd19-5450487a4bef_1707x1536.png 424w, https://substackcdn.com/image/fetch/$s_!tiHW!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fee637cfd-0101-44c0-bd19-5450487a4bef_1707x1536.png 848w, https://substackcdn.com/image/fetch/$s_!tiHW!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fee637cfd-0101-44c0-bd19-5450487a4bef_1707x1536.png 1272w, https://substackcdn.com/image/fetch/$s_!tiHW!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fee637cfd-0101-44c0-bd19-5450487a4bef_1707x1536.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>Where the keynote is about giving agents memory, this one is about keeping that memory lean: how to compact context so agents stay sharp without burning through tokens, and how to manage cost as your system scales.</p><p>If you&#8217;re attending or anywhere near San Francisco, come find us. Louis-Fran&#231;ois will be there all day and would love to meet. <a href="https://x.com/Whats_AI?ref=louisbouchard.ai">Just DM him</a>.</p><p>&#128197; 2:20&#8211;4:20 pm, June 29, 2026 &#183; Moscone West</p><p>&#128073; <a href="https://www.ai.engineer/worldsfair/schedule?session=asn_slot_2026_06_29_workshop_track_09_1420_2026_06_05t11_53_44_628z">ai.engineer/worldsfair/schedule</a></p><div><hr></div><p>The keynote is live right now! Go watch it and share it with someone who&#8217;ll build with it.</p><p><strong><a href="https://www.youtube.com/watch?v=ZRM_TfEZcIo">Watch &#8220;Turn 10,994 Notes Into Memory&#8221;</a></strong></p><p>&#8212; The Towards AI Team</p><p>P.S. We also ran a 2-hour workshop at AIE London a few weeks ago on building multi-agent systems with MCP servers from scratch, that full recording and code are live too. <a href="https://www.youtube.com/watch?v=mYSRn6PC1mc">Watch it here</a>.</p>]]></content:encoded></item><item><title><![CDATA[TAI #210: GLM-5.2 Closes Most of the Open-Weight Gap in Ten Weeks]]></title><description><![CDATA[Also, SpaceX acquires Cursor, Noam Shazeer joins OpenAI & more.]]></description><link>https://newsletter.towardsai.net/p/tai-210-glm-52-closes-most-of-the</link><guid isPermaLink="false">https://newsletter.towardsai.net/p/tai-210-glm-52-closes-most-of-the</guid><dc:creator><![CDATA[Louie Peters]]></dc:creator><pubDate>Tue, 23 Jun 2026 15:01:11 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!EOS8!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F56f75ca1-a5c3-40b0-a20f-05aadecf74a7_1600x900.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<h2>What happened this week in AI by Louie</h2><p>We covered GLM-5.2 when Z.ai announced it last week. Since then, the weights have shipped under an MIT license, multiple inference providers have put the model online, and independent evaluators have had time to test it. The evidence now supports a stronger conclusion: GLM-5.2 is a major breakthrough for open weights and for Chinese AI labs.</p><p>The speed of the improvement is extraordinary. GLM-5.1 launched on April 7 and GLM-5.2 arrived on June 16, exactly ten weeks later. Artificial Analysis scores 5.2 at 51 on its Intelligence Index, up 11 points from 5.1 at 40. Only Claude Fable 5 (60), Claude Opus 4.8 (56), and GPT-5.5 (55) score higher. Fable is unavailable, leaving GLM-5.2 within five points of the strongest model people can currently use.</p><p>The gains also appear on some of the hardest evaluations to target: 16 points on CritPt physics reasoning, 12 points on Humanity&#8217;s Last Exam, 9 points on long-context reasoning, 15 points on the agentic banking benchmark, 7 points on SciCode, and 16 points on Terminal-Bench 2.1, all measured by Artificial Analysis. Gains spread this widely are hard to dismiss as a single benchmark trick.</p><p>AA-Briefcase is even more convincing. It uses 91 held-out tasks across four multi-week knowledge-work projects, with nearly 2,000 source files, more than 3,500 emails, and 25,000 Slack messages. GLM-5.2 ranked third, behind Fable and Opus 4.8 but ahead of GPT-5.5 and every other open-weight model. The private tasks and rubrics make contamination and targeted training much harder, and no model is close to solving it: Fable passed every rubric criterion on only 3% of tasks.</p><p>Z.ai did not appear from nowhere. It grew out of research at Tsinghua University led by co-founder Jie Tang, whose team launched the AMiner researcher graph in 2006, contributed to the 1.75-trillion-parameter 2021 Wu Dao project, and has worked on the General Language Model architecture for years. The talent bench runs deep.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!EOS8!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F56f75ca1-a5c3-40b0-a20f-05aadecf74a7_1600x900.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!EOS8!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F56f75ca1-a5c3-40b0-a20f-05aadecf74a7_1600x900.png 424w, https://substackcdn.com/image/fetch/$s_!EOS8!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F56f75ca1-a5c3-40b0-a20f-05aadecf74a7_1600x900.png 848w, https://substackcdn.com/image/fetch/$s_!EOS8!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F56f75ca1-a5c3-40b0-a20f-05aadecf74a7_1600x900.png 1272w, https://substackcdn.com/image/fetch/$s_!EOS8!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F56f75ca1-a5c3-40b0-a20f-05aadecf74a7_1600x900.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!EOS8!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F56f75ca1-a5c3-40b0-a20f-05aadecf74a7_1600x900.png" width="1456" height="819" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/56f75ca1-a5c3-40b0-a20f-05aadecf74a7_1600x900.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:819,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!EOS8!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F56f75ca1-a5c3-40b0-a20f-05aadecf74a7_1600x900.png 424w, https://substackcdn.com/image/fetch/$s_!EOS8!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F56f75ca1-a5c3-40b0-a20f-05aadecf74a7_1600x900.png 848w, https://substackcdn.com/image/fetch/$s_!EOS8!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F56f75ca1-a5c3-40b0-a20f-05aadecf74a7_1600x900.png 1272w, https://substackcdn.com/image/fetch/$s_!EOS8!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F56f75ca1-a5c3-40b0-a20f-05aadecf74a7_1600x900.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>GLM-5.2 improved sharply without getting larger. It keeps the same roughly 750-billion-total, 40-billion-active Mixture-of-Experts scale as 5.1, but its context has grown from 200K to 1 million tokens. The main advances came from architecture, long-context training, reinforcement learning, distillation, and serving.</p><p>IndexShare is the architectural unlock. GLM already used sparse attention, where each query attends to a selected set of past tokens, but the expensive indexer still searched the full history at every layer. GLM-5.2 runs that search once per four-layer block and reuses the selected positions across the next three layers, while each layer still performs its own attention and mixture-of-experts computation. Z.ai reports a 2.9x reduction in per-token floating-point operations (FLOPs) for a one-million-token context, and the related IndexCache paper measured up to 1.82x faster prefill and 1.48x faster decoding while preserving quality.</p><p>The revised multi-token prediction layer improves speculative decoding acceptance by 20%, and Z.ai reworked key-value cache management, kernels, scheduling, and 8-bit floating-point (FP8) serving. These changes make long contexts and reinforcement-learning rollouts cheaper, but they are enablers. The learning breakthroughs came from what Z.ai did with the extra horizon.</p><p>My strongest hypothesis for the largest driver in the capability jump is compaction-aware reinforcement learning on more complex agentic tasks. Claude Code and Codex led a breakthrough in agentic coding adoption late last year, when context compaction enabled agents to work through far longer tasks without carrying every token forever. Compaction turns one long episode into several linked fragments; one run may produce two fragments and another eight, with the final reward arriving hours after the earliest decisions. Group-relative reinforcement learning becomes awkward because the fragments have different counts and lengths.</p><p>GLM-5.2 moved to critic-based proximal policy optimization (PPO) for these tasks. It trains every compacted fragment, uses a critic to estimate token-level advantages, and applies a token-level loss to handle unequal lengths. The model, therefore, learns from work on both sides of compaction and trains on the lossy summaries it will meet in real deployment.</p><p>The second likely driver is scaled on-policy distillation. OPD existed in GLM-5; the 5.2 change was scaling it. Z.ai says it used its slime framework to consolidate more than 10 specialist models into a final model in roughly 2 days. Specialists can spend reinforcement learning compute discovering strong policies for coding, science, search, and tools. This integrates broad skills without repeating every expensive discovery process in one generalist training run.</p><p>Z.ai also added an online anti-hacking layer for coding reinforcement learning. Suspicious tool calls are caught by rules, judged for intent, blocked when necessary, and replaced with dummy results so the rollout can continue. That protects the reward signal from agents who read hidden tests, copy reference solutions, or fetch target code.</p><p>I would be surprised if Z.ai were not also using outputs from Claude and GPT models wherever it could, for synthetic data, evaluation, or hard distillation, but it cannot explain the whole catch-up. The sudden gains line up with a much better system for generating, learning from, and consolidating long agent trajectories.</p><p>There is no clean ablation assigning credit across these changes, and some of Z.ai&#8217;s own 5.1 comparisons also changed context windows, benchmark versions, judges, time limits, or output budgets. My ranking is compaction-aware PPO for ultra-long agents, scaled multi-specialist distillation for the broad jump, and IndexShare as the efficiency multiplier that made both affordable.</p><p>GLM-5.2 still lacks multimodal input, which probably saved substantial training cost and complexity, though Z.ai does not quantify the savings. It also draws a hard deployment boundary: the model cannot inspect screenshots, review a rendered interface, read image-heavy documents, or visually test a browser workflow.</p><p>The economics are compelling, with a catch. Z.ai charges $1.40 per million input tokens, $0.26 for cached input, and $4.40 for output. The current Artificial Analysis cost-per-task chart puts GLM-5.2 at $0.52, compared with $0.86 for GPT-5.5 and $1.80 for Opus 4.8. GLM also used about 140 million output tokens across the Intelligence Index, versus 72 million for GPT-5.5 and 120 million for Opus, so lower token efficiency partly offsets the price advantage. DeepSeek V4 Pro remains roughly 10 times cheaper per task, at about $0.04-$0.05, although it scores 7 points lower on the index.</p><p>Open weights do not automatically make self-hosting economical. GLM-5.2 has 753 billion parameters in the released checkpoint. The practical FP8 vLLM recipe needs eight H200-class GPUs, while serving the full one-million-token window is documented on eight B200s with an FP8 key-value cache. Most companies cannot batch enough simultaneous work or keep that cluster busy around the clock, whereas a specialist inference provider can spread the hardware across many customers and reach much higher utilization. For most teams, the weights offer control and portability, while a hosted endpoint delivers the lower bill.</p><p>In Towards AI&#8217;s enterprise deployment work, routing routine coding through GLM-5.2 on a compliant US provider is now an easy, cost-saving recommendation for those optimizing token budgets. We would keep both Codex and Claude Code available to all developers and send bounded refactors and text-heavy repository work to GLM-5.2. Claude Code supports it directly; Codex can reach it through a provider or adapter that implements the Responses API.</p><p>If you&#8217;re cost-sensitive, I recommend reserving GPT-5.5 or Opus 4.8 for the hardest planning, recovery, final review, complex unit and integration testing, browser use, user-interface (UI) work, and screenshot-based testing. DeepSeek V4 Pro is also a strong subagent option for high-volume, easily checked summarization, extraction, classification, and structured data preparation.</p><p>I expect Z.ai trained GLM-5.2 on a far smaller budget than Anthropic, OpenAI, Google, xAI, or Meta spend on recent frontier programs. The gap in inference pricing is much narrower; however, this isn&#8217;t a model focused on the bargain tier. Two questions now matter. Can Z.ai maintain this trajectory and catch Fable or Mythos? And how many other labs can reproduce its ten-week model checkpoint jump by combining longer reinforcement-learning tasks, compaction-aware training, specialist distillation, and better reward integrity?</p><p>The longer leading US models are withheld, disabled, or constrained, the greater the chance that a Chinese open-weight model becomes the strongest capability the public can actually use. A hypothetical open-weight Fable or Mythos would face dramatically fewer restrictions and much less provider oversight than the constrained Fable API that briefly appeared this month. GLM-5.2 shows why that policy collision is approaching faster than expected.</p><div><hr></div><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://newsletter.towardsai.net/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading Towards AI Newsletter! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><div><hr></div><h3>Why should you care?</h3><p>GLM-5.2 is now good enough to change enterprise model routing. A 40% saving against GPT-5.5 and roughly 70% against Opus 4.8 per Artificial Analysis task becomes material once coding agents read repositories, call tools, spawn subagents, and retry work all day. The saving only survives when GLM finishes the task without creating extra review or rescue work, so the number that matters is cost per accepted result after retries and human review, not the headline token price.</p><p>The practical pattern is a model hierarchy. Start with routine, text-heavy work where success is visible: bounded refactors, test generation, migration chores, repository questions, extraction, and structured first drafts. Keep a frontier model as the escalation path for ambiguous planning, failed runs, high-impact changes, and final review. Route visual tasks straight to GPT-5.5 or Opus, since GLM cannot inspect screenshots or rendered interfaces, and hand simple, high-volume subagent work to DeepSeek V4 Pro where outputs can be checked automatically.</p><p>Build the routing policy from real traces rather than guesswork. Measure accepted results, retries, latency, token use, human review, and failures by task category, then move each category to the cheapest model that clears your quality bar. A US-hosted endpoint supplies the contractual controls many enterprises need while avoiding an eight-GPU self-hosting commitment. The immediate opportunity is a model hierarchy that spends frontier tokens only where they change the outcome.</p><p>Two larger shifts sit behind that tactical win. The first is that the public frontier may soon be moving to open weights in certain areas where the closed labs limit their model&#8217;s abilities. Fable showed how fast a hosted frontier model can disappear, while GLM-5.2 shows the other side: once MIT-licensed weights are distributed across providers and downloaded by users, no single company or government can switch the model off globally. If US labs keep their strongest systems private or heavily restricted, Chinese labs only need to beat the models people can actually access, not every internal checkpoint, and GLM already ranks ahead of GPT-5.5 on AA-Briefcase. Published frontier weights would also strip away most of the API-level classifiers, identity checks, monitoring, and country restrictions that govern access today, which is why governments may start targeting weights, compute, or distribution. Enterprises should preserve provider portability now: keep prompts and tool definitions outside any single platform, maintain a replayable evaluation set, and test at least one open-weight fallback.</p><p>The second is that long-horizon training looks like a reproducible breakthrough. IndexShare reduced the cost of long contexts, compaction-aware PPO let every fragment of a long run contribute to learning, and specialist distillation moved policies discovered in separate programs into a single model. Together, those form a repeatable engineering agenda other labs can copy. A lab with strong researchers, adequate compute, realistic tasks, reliable verifiers, and access to frontier-generated data may now close a large capability gap during post-training alone. Not everyone will pull it off, since long agent runs are expensive and weak reward design trains shortcuts quickly, but the techniques are legible enough for MiniMax, DeepSeek, Qwen, or xAI (particularly if integrating Cursor long-horizon agentic coding data) and others to chase.</p><p><em>&#8212; <a href="http://www.linkedin.com/in/louie-peters">Louie Peters&#8202;&#8212;&#8202;Towards AI Co-founder and CEO</a></em></p><div><hr></div><h3>Hottest News</h3><p>1. <a href="https://www.cnbc.com/2026/06/16/spacex-spcx-cursor-acquisition-ipo.html">SpaceX To Acquire the AI Coding Startup Cursor for $60 Billion</a></p><p>SpaceX agreed to acquire Anysphere, the company behind Cursor, in a $60 billion all-stock deal announced on June 16, four days after the company&#8217;s record-setting IPO. Cursor investors will receive SpaceX Class A common stock, representing a 3.4% dilution at SpaceX&#8217;s IPO valuation. The deal is the largest acquisition of a venture-backed startup on record. Cursor&#8217;s annualized revenue had climbed to $4 billion by early June, though its market share in AI coding tools had declined from 41% in June 2025 to roughly 26% in May, according to Ramp spending data. SpaceX and Cursor have been jointly training an AI model over recent months, which SpaceX plans to release on both Cursor and its Grok Build coding agent. The acquisition is intended to strengthen SpaceX&#8217;s AI division, formed through its earlier merger with xAI, which has struggled to build a competitive coding product. The deal is expected to close in Q3 2026.</p><p>2. <a href="https://www.axios.com/2026/06/19/trump-anthropic-national-security-the-axios-show">Trump Lifts Anthropic National Security Designation, Fable Access To Be Restored</a></p><p>President Trump told The Axios Show on June 19 that he no longer views Anthropic or its CEO, Dario Amodei, as a national security threat, a shift from the administration&#8217;s position the prior week. &#8220;Well, not now, but a week ago, maybe,&#8221; Trump said when asked directly. The remarks followed a meeting between Amodei and Trump at the G7 summit in &#201;vian-les-Bains, France, where Amodei and Demis Hassabis, CEO of Google DeepMind, jointly proposed a US-led democratic AI alliance. However, the Commerce Department&#8217;s export control directive issued on June 12, which forced Anthropic to disable Fable 5 and Mythos 5 for all users worldwide, has not been formally rescinded. The Pentagon&#8217;s March supply-chain risk designation and the ban on federal agencies&#8217; use of Anthropic technology also remain in place. An Anthropic managing director said at the company&#8217;s Seoul office launch on June 18 that he was &#8220;very confident&#8221; both models would return &#8220;in the coming days.&#8221;</p><p>3. <a href="https://techcrunch.com/2026/06/16/chatgpts-market-share-slips-below-50-for-first-time/">ChatGPT Market Share Drops Below 50% for First Time</a></p><p>ChatGPT&#8217;s share of the global AI assistant market fell to 46.4% by the end of May, the first time it has dropped below 50%, according to Sensor Tower&#8217;s State of AI Report for 2026. The decline has been steady: from 65.3% in December 2024 to 52.8% in December 2025 to 46.4% by May 2026. ChatGPT remains the most popular AI assistant with over 1.1 billion monthly users. Gemini holds 27.7% of the market with 662 million monthly users, and Claude holds 10.3% with 245 million monthly users. Claude leads all platforms in subscription conversion at 13% of users paying, the highest in the field. Switching behavior is accelerating: OpenAI&#8217;s Department of Defense partnership in February triggered a 295% day-over-day surge in ChatGPT uninstalls, while Claude&#8217;s US downloads jumped 51% the same day. Total spending on AI apps is on pace to reach $4.2 billion in H1 2026, up from $1.83 billion in H1 2025.</p><p>4. <a href="https://www.bloomberg.com/news/articles/2026-06-19/nobel-winner-john-jumper-to-leave-google-deepmind-for-anthropic">Nobel Laureate John Jumper Leaves DeepMind for Anthropic</a></p><p>John Jumper, who shared the 2024 Nobel Prize in Chemistry with Demis Hassabis for co-creating AlphaFold, announced on June 19 that he is leaving Google DeepMind after nearly nine years to join Anthropic. Jumper served as VP and Engineering Fellow at Google DeepMind and was a key member of Google&#8217;s AI coding development team, according to Bloomberg. AlphaFold has been used to predict over 200 million protein structures and is used by more than 2 million researchers across 190 countries. Anthropic has been building dedicated AI-for-science infrastructure throughout 2026, including wet labs and partnerships with the Allen Institute and Howard Hughes Medical Institute. Google DeepMind confirmed that Jumper would remain through the end of the year to assist with the transition.</p><p>5. <a href="https://www.cnbc.com/2026/06/18/google-gemini-co-lead-noam-shazeer-leaves-for-openai.html">Transformer Co-Inventor Noam Shazeer Leaves Google for OpenAI</a></p><p>Noam Shazeer, co-author of the 2017 &#8220;Attention Is All You Need&#8221; paper and co-lead of Google&#8217;s Gemini AI models, announced on June 18 that he is leaving Google to join OpenAI as Lead for Architecture Research. The departure comes less than two years after Google paid approximately $2.7 billion in a licensing deal with Character.AI that brought Shazeer and his research team back to lead Gemini development. Shazeer first joined Google in 2000 and spent over two decades at the company across two stints. He was credited with helping close the gap between Gemini and ChatGPT during his return. The move lands as OpenAI prepares for a potential IPO, with a confidential S-1 filed on June 8. Combined with Jumper&#8217;s departure to Anthropic the following day, Google lost two of its most prominent AI researchers to its two largest competitors in a single week.</p><p>6. <a href="https://claude.com/blog/enterprise-managed-auth">Anthropic Adds Enterprise-Managed Authorization for MCP Connectors</a></p><p>Anthropic launched Enterprise-Managed Authorization (EMA) for MCP connectors, allowing IT administrators to provision connector access once through their identity provider and have employees inherit it automatically on first login, with no individual OAuth flows required. Okta is the first supported identity provider, using its Cross App Access (XAA) protocol, built on the ID-JAG standard. Seven MCP providers support EMA at launch: Asana, Atlassian, Canva, Figma, Granola, Linear, and Supabase, with Slack coming next. HubSpot, Ramp, and Webflow are among the early adopters. Ramp reports that 2,000 employees are provisioned through Okta with zero additional steps. The feature works across Claude chat, Claude Code, and Cowork for Team and Enterprise plans. EMA is built on an open extension to the MCP authorization specification, meaning any connector, including custom-built internal tools, can implement the standard.</p><p>7. <a href="https://qwen.ai/blog?id=qwen-robotsuite">Alibaba Releases Three New Foundation Models for Embodied Intelligence</a></p><p>Alibaba&#8217;s Qwen team released the Qwen-Robot Suite, consisting of three foundation models that bridge vision-language understanding and physical robotic control. Qwen-RobotNav is a navigation model built on Qwen3-VL (available at 2B, 4B, and 8B sizes) that unifies five navigation task families, including instruction following, point-goal navigation, object-goal search, target tracking, and autonomous driving, under a single model with a controllable observation protocol. Qwen-RobotManip is a vision-language-action model built on Qwen3.5&#8211;4B VL, trained on over 38,100 hours of open-source manipulation data. It aligns heterogeneous robot data into an 80-dimensional canonical action space, achieving 3.2x the cross-embodiment transfer rate of &#960;0.5 and ranking first on the RoboChallenge Table30-v1 generalist track. Qwen-RobotWorld is a 20B-parameter language-conditioned video world model that predicts physically grounded futures across manipulation, driving, and navigation scenarios, using natural language as a universal action interface, and is trained on 8.6 million video-text pairs. The models are in pilot testing with Alibaba Cloud enterprise customers.</p><div><hr></div><h3>AI Tip of the Day</h3><p>When an agent chooses a tool, don&#8217;t assume it understood the task.</p><p>Sometimes it only matches a word in the user&#8217;s request to a word in the tool name. For example, if the user asks for the &#8220;latest report,&#8221; the agent might still call the search tool even though the report has already been uploaded. Or it might call a database tool when the answer was already in the prompt.</p><p>To debug this, log three things side by side: the user&#8217;s request, the tool the agent picked, and the arguments it used. Then check whether the tool matched what the user actually wanted. Don&#8217;t only look for tool errors. A tool can run successfully and still be the wrong tool.</p><p>If you&#8217;re exploring agent engineering and want to go deeper into tool use and guardrails, our <a href="https://academy.towardsai.net/courses/agent-engineering?utm_source=Newsletter&amp;utm_medium=email&amp;utm_id=AItips">Agent Engineering: Building Multi-Agent Systems</a> course is the cleanest path to building production agents.</p><div><hr></div><h3>Five 5-minute reads/videos to keep you learning</h3><p>1. <a href="https://pub.towardsai.net/building-a-gemini-live-voice-app-with-react-fastapi-and-your-own-websocket-protocol-9752bed95182?sk=01a43fe709a0715b104cc5a889dbf547">Building a Gemini Live voice app with React, FastAPI and your own WebSocket protocol</a></p><p>This article walks through building a real-time voice app on top of Gemini Live using React and FastAPI. All Gemini-specific logic routes through a backend WebSocket rather than the browser, so the backend owns the API key, system prompt, and model configuration, while the frontend communicates via a five-message protocol that the author defines and controls. Two async loops handle audio in both directions: browser mic input streams to Gemini as PCM16, and Gemini&#8217;s audio response streams back with a transcript. The result is a setup where swapping providers requires changes to a single backend file.</p><p>2. <a href="https://pub.towardsai.net/why-most-multi-agent-ai-systems-waste-90-of-their-time-and-how-to-fix-it-c0ce81f0e323">Why Most Multi-Agent AI Systems Waste 90% of Their Time (And How to Fix It)</a></p><p>Multi-agent setup costs are a real bottleneck, and memory snapshots fix them. The author built a five-agent code analysis swarm in which each agent spent 90 seconds installing tools before doing 8 seconds of actual work. By checkpointing a fully prepared VM using TensorLake&#8217;s memory snapshot (capturing disk state, running processes, and memory in one pass), then forking five independent copies, setup cost moved outside the loop entirely. A lead GPT-4o call then synthesized all five reports into a single, prioritized list of fixes.</p><p>3. <a href="https://pub.towardsai.net/llm-observability-with-langsmith-part-1-tracing-everything-building-audit-grade-callbacks-c477719af691">LLM Observability with LangSmith&#8202;&#8212;&#8202;Part 1: Tracing Everything &amp; Building Audit-Grade Callbacks</a></p><p>This article builds an LLM observability pipeline through a practical implementation for a LangGraph customer support agent. The author set up zero-config tracing with two environment variables, enabling complete request replay across every classifier call, retrieval step, and LLM response. It then builds a compliance-grade audit callback using LangChain&#8217;s BaseCallbackHandler, logging PII-redacted JSON Lines locally before any data leaves the network. Both layers run independently, giving engineers a debuggable trace and auditors a tamper-evident chain of custody.</p><p>4. <a href="https://pub.towardsai.net/llm-observability-with-langsmith-part-2-eval-gates-prompt-versioning-choosing-your-stack-e607473320b5">LLM Observability with LangSmith&#8202;&#8212;&#8202;Part 2: Eval Gates, Prompt Versioning &amp; Choosing Your Stack</a></p><p>This article tackles the hardest observability question: catching prompt regressions before they reach production. The author built a versioned eval dataset with trap cases, wired an exact-match evaluator into a LangSmith experiment, and surfaced a real routing failure in a single table row. Prompt versioning via the Hub provides immutable commits, movable production tags, and instant rollbacks without deploys. The article closes with a LangSmith-vs-Langfuse decision matrix and a five-rung privacy ladder for teams where traces cannot leave the building.</p><p>5. <a href="https://pub.towardsai.net/building-a-stateful-code-interpreter-with-tensorlake-sandboxes-df2f6d623a47">Building a Stateful Code Interpreter with Tensorlake Sandboxes</a></p><p>This article walks through the process of building a production-grade stateful code interpreter using Tensorlake MicroVM sandboxes. Starting with a basic Claude-plus-sandbox setup, it progresses to persistent named sessions, suspend-and-resume workflows, filesystem and memory snapshots, and parallel branch forking for comparing experimental approaches. The key distinction is between images as environment blueprints and snapshots as full runtime captures. Pre-warmed golden snapshots handle cold-start costs at scale, giving engineers a concrete path from a throwaway chat session to a persistent, restorable computing environment.</p><h3>Repositories &amp; Tools</h3><p>1. <a href="https://github.com/bytedance/deer-flow">Deer Flow</a> is a super agent harness that orchestrates sub-agents, persistent memory, sandboxed execution, and extensible skills to handle long-horizon tasks.</p><p>2. <a href="https://github.com/DeusData/codebase-memory-mcp">Codebase-memory-mcp</a> is an MCP server that indexes codebases into a persistent knowledge graph across 158 languages, delivering sub-millisecond queries and roughly 120x fewer tokens than file-by-file exploration.</p><p>3. <a href="http://github.com/cisco-foundation-ai/fully-automated-prompt-optimization">FAPO</a> uses Claude Code as an autonomous optimizer that iteratively improves prompts, parameters, and chain architecture.</p><p>4. <a href="https://github.com/yandex/yaff">YaFF</a> is a C++ serialization library that provides a zero-copy wire format for Protobuf.</p><h3>Top Papers of The Week</h3><p>1. <a href="https://arxiv.org/abs/2606.16140">VibeThinker-3B: Exploring the Frontier of Verifiable Reasoning in Small Language Models</a></p><p>This paper introduces VibeThinker-3B, a compact dense model with 3B parameters developed to test how far verifiable reasoning can be pushed within a small-model regime. VibeThinker-3B is trained through a three-stage pipeline: curriculum-based supervised fine-tuning, multi-domain reinforcement learning, and offline self-distillation, built on the Spectrum-to-Signal post-training paradigm. It scored 94.3 on AIME 2026 (97.1 with claim-level test-time scaling), 80.2 Pass@1 on LiveCodeBench v6, and a 96.1% acceptance rate on unseen LeetCode weekly contests from April-May 2026.</p><p>2. <a href="https://arxiv.org/abs/2602.15763">GLM-5: from Vibe Coding to Agentic Engineering</a></p><p>This technical report presents GLM-5, Z.AI&#8217;s 744B-parameter MoE foundation model designed to move beyond vibe coding toward autonomous, multi-step agentic engineering. The model adopts DeepSeek Sparse Attention (DSA) to reduce training and inference costs while maintaining long-context fidelity up to 200K tokens. A new asynchronous RL infrastructure decouples generation from training to improve post-training efficiency, and novel asynchronous agent RL algorithms enable the model to learn from complex, long-horizon interactions.</p><p>3. <a href="https://arxiv.org/html/2606.14066v3">FastContext: Training Efficient Repository Explorer for Coding Agents</a></p><p>In most coding agents, the same model that solves the task also explores the repository, leaving exploratory reads and searches in the solver&#8217;s context and wasting token budget on irrelevant code. FastContext separates these two roles by introducing a dedicated exploration subagent that issues parallel read-only tool calls (read, glob, grep) and returns concise file paths and line ranges as focused context. The exploration models span 4B to 30B parameters, bootstrapped from strong reference-model trajectories and refined with task-grounded rewards. Integrating FastContext into Mini-SWE-Agent improved end-to-end resolution rates by up to 5.5% across SWE-bench Multilingual, SWE-bench Pro, and SWE-QA, while reducing the coding agent&#8217;s token consumption by up to 60%.</p><p>4. <a href="https://arxiv.org/abs/2606.18023">LoopCoder-v2: Only Loop Once for Efficient Test-Time Computation Scaling</a></p><p>Looped Transformers scale latent computation by reapplying shared blocks, but sequential looping increases latency and KV cache memory usage proportionally. Parallel Loop Transformers (PLT) address this through cross-loop position offsets and shared-KV gated sliding-window attention, making the loop count a tunable design parameter. This paper trains a family of 7B PLT code models from scratch on 18T tokens, with varying loop counts, to study the gain-cost trade-off. The finding is that a two-loop configuration captures most of the representational refinement, while additional loops introduce positional mismatch and oscillatory updates that degrade performance.</p><p>5. <a href="https://arxiv.org/abs/2606.19195">Moebius: 0.2B Lightweight Image Inpainting with 10B-Level Performance</a></p><p>Current state-of-the-art image inpainting models, such as FLUX.1-Fill-Dev, operate with 10B+ parameters, making deployment expensive. Moebius compresses this capability into 0.22B parameters (less than 2% of FLUX.1-Fill-Dev) by reconstructing the diffusion backbone with Local-&#955; Mix Interaction blocks that summarize spatial contexts and global semantic priors into fixed-size linear matrices. An adaptive multi-granularity distillation strategy that operates entirely in latent space, avoiding pixel-space decoding costs, unlocks the representational capacity of this compressed architecture.</p><h3>Quick Links</h3><p>1. <a href="https://www.perplexity.ai/hub/blog/self-improving-memory-for-agents">Perplexity launches Brain</a>, a self-improving memory system for its Computer agent that remembers what the agent did rather than user preferences. Brain builds a context graph of completed tasks, including what worked, what failed, and what corrections were made, then synthesizes that graph overnight into an LLM wiki that loads automatically into every subsequent session. Early internal results show a 25% increase in answer correctness on repeated tasks, 16% higher recall, and 13% lower cost on tasks requiring historical context. Available in research preview for Max and Enterprise Max subscribers.</p><p>2. <a href="https://claude.com/blog/artifacts-in-claude-code">Claude Code now supports artifacts</a>, turning session work into live, shareable web pages that update in place as the session progresses. Each artifact is a self-contained HTML page built from the local codebase, connected MCP tools, and conversation history, published to a private org-only URL with version history. Use cases include PR walkthroughs, incident timelines, dashboards, and release checklists. Available in beta for Team and Enterprise organizations from the CLI and desktop app.</p><p>3. <a href="https://sakana.ai/marlin-release/#English">Sakana AI commercializes AB-MCTS in Sakana Marlin</a>, its first commercial product. Positioned as a &#8220;Virtual CSO,&#8221; Marlin is an autonomous research agent that runs for up to eight hours on a single topic, forming hypotheses, gathering sources, and resolving contradictions before returning a structured report of up to roughly 100 pages with executive slides. It builds on AB-MCTS (NeurIPS 2025 Spotlight) and The AI Scientist (Nature). Pricing is pay-per-use at 100 credits per run (&#165;98/credit), with Pro, Team, and Enterprise tiers. The underlying algorithm is open-sourced as TreeQuest under the Apache 2.0 license.</p><p>4. <a href="https://www.liquid.ai/blog/lfm2-5-retrievers">Liquid AI introduces LFM2.5-Embedding-350M and LFM2.5-ColBERT-350M</a>, two 350M-parameter multilingual retrieval models and the first bidirectional members of the LFM family. The Embedding model produces a single vector per document for fastest search and smallest index. The ColBERT model produces per-token vectors for word-level matching with higher accuracy at the cost of a larger index. Both support 11 languages, run via llama.cpp GGUFs on CPUs and edge devices with sub-10ms query embedding latency, and outperform Qwen3-Embedding-0.6B despite being smaller. Available on Hugging Face under the LFM Open License v1.0.</p><h3>Who&#8217;s Hiring in AI</h3><p><strong><a href="https://jobs.towardsai.net/job/meta-production-engineer-university-grad-yz5f">Production Engineer (University Grad) @Meta (New York, NY, USA)</a></strong></p><p><strong><a href="https://jobs.towardsai.net/job/amentum-ai-technology-support-engineer-analyst-rnuz">AI Technology Support Engineer/Analyst @Amentum (Remote/USA)</a></strong></p><p><strong><a href="https://jobs.towardsai.net/job/oracle-software-developer-4-9oer">Software Developer 4 @Oracle (Raleigh, NC, USA)</a></strong></p><p><strong><a href="https://jobs.towardsai.net/job/dynatrace-senior-engineer-m-f-x-for-openingest-generative-ai-wqan">Senior Engineer (m/f/x) for OpenIngest Generative AI @Dynatrace (Vienna, Austria)</a></strong></p><p><strong><a href="https://jobs.towardsai.net/job/caterpillar-inc-automation-engineer-dlme">Automation Engineer @Caterpillar, Inc. (Piracicaba, Brazil)</a></strong></p><p><strong><a href="https://jobs.towardsai.net/job/splice-product-engineer-iii-wrrn">Product Engineer III @Splice (Remote/USA)</a></strong></p><p><em>Interested in sharing a job opportunity here? Contact <a href="mailto:sponsors@towardsai.net">sponsors@towardsai.net</a>.</em></p><p><em>Think a friend would enjoy this too? <a href="https://newsletter.towardsai.net/">Share the newsletter and let them join the conversation.</a></em></p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://newsletter.towardsai.net/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading Towards AI Newsletter! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div>]]></content:encoded></item></channel></rss>