What happened this week in AI by Louie
OpenAI’s DevDay on September 29 gave us two major releases: GPT-6.1 Sol, which brings most of GPT-6 Astra’s benchmark performance down to Sol pricing, and Dots, OpenAI’s answer to the persistent personal agents we discussed last week. Google also looks competitive again. Gemini 4 Argon scores 53 on Artificial Analysis’s Intelligence Index at high effort using Long Decode Continuation for extended reasoning, level with Astra at maximum effort. Access is still limited to selected Fairwind partners and trusted testers, so we remain skeptical until we can put it through our own real-world problems.
What excites me most about Argon is its one-million-token output limit and the price. Long Decode Continuation lets it keep generating across follow-up calls, including extended reasoning, compared with the standard 128,000-token output limits of Claude Opus 5.5 and GPT-6 Astra. Opus also offers a 300,000-token Batch API beta. That gives Argon much more room to work through a difficult problem or produce a substantial result before hitting an output ceiling. Introductory pricing is $2 per million input tokens and $10 per million output tokens, half Opus 5.5’s standard rates and a fifth of Astra’s short-context rates. Google’s stated regular price is $4/$20 after the promotion.
DevDay itself landed below the expectations OpenAI had built around it. There was no Astra 6.1 or larger flagship model to try, but the releases that did ship were still substantial. Sol 6.1, in particular, behaves much more like a smaller Astra than a straightforward update to GPT-6 Sol. My guess is that it is closer to Astra architecturally than to the model it replaces, although OpenAI has disclosed nothing that confirms that. The system card says only that Sol 6.1 uses the same types of data and training as Astra. The API behaves like Astra’s as well, with effort running from low to max and no setting that switches reasoning off, which GPT-6 Sol still offers. Replacing GPT-6 Sol seven days after its launch makes the naming even less helpful.
The more important point is that OpenAI kept Sol pricing the same. Standard API rates remain $2 per million uncached input tokens and $10 per million output tokens for prompts up to 272,000 input tokens, exactly where GPT-6 Sol was priced and one fifth of Astra’s 10/50 rates. Cached input is the only additional cut, falling to $0.10 per million tokens. That helps agents that resend the same long prefix of instructions, tool definitions and documents on every call, as long as the prefix is unchanged and the cache actually hits. I suspect Opus 5.5, launched a week earlier at $4 and $20, pushed OpenAI to put a near-Astra model at Sol rates instead of pricing it somewhere in between.
Independent evaluations show how close Sol 6.1 gets to Astra for a fraction of the cost. On Artificial Analysis’s current index, Sol 6.1 at maximum effort scores 52 against Astra’s 53, at an average of $0.72 per weighted task against $3.26. Opus 5.5 at maximum effort, with default fallback enabled, still leads at 58 for $5.98. The lower effort settings are where I would change habits. Sol 6.1 at medium matches the previous Sol at maximum effort on the same index, 48 each, for $0.21 per task against $1.04. I would start routine research, coding and document work at medium and raise effort when the task justifies it. That said, I still primarily use Astra for most coding tasks, but I’m increasingly having it hand off simpler execution and testing to Sol-6.1 sub-agents.
Dots extends Astra into a much broader product. Last week we covered Muse, powered by Meta’s Muse Spark model, and Instinct, which made persistent agents feel as easy as messaging a contact. A dot packages persistent work as one continuing assistant with its own cloud computer and browser, selective memory of your decisions and preferences, connected apps, and the ability to decide when to pause and wake itself.
Dots feels very powerful to me, and potentially more ambitious than Muse or Instinct. Both of those already act proactively, so the ambition I see is in coordination. A dot can start separate cloud tasks, delegate work in parallel, and create or continue Codex work on a connected computer while you keep talking to it. Its cloud work continues with your laptop closed; local work needs that computer online with the ChatGPT app open. I already work this way in Codex, but dots makes many of these tasks much more accessible.
Several other DevDay launches fill in pieces an experienced operator currently supplies by hand. ChatGPT Space gives people and agents shared files and Pages. Reusable Codex cloud environments let a new task start in a development setup someone has already prepared. Event subscriptions let a change in a connected app trigger work. Together, they give an agent somewhere to work, the right materials, and a reason to resume. Every’s Dan Shipper described the payoff: his dot warned that a flight would clash with a video recording that had only been discussed in Slack and never made it onto his calendar. That is the assistance I want, catching a commitment while it is still taking shape.
But Dots also feels rushed, and I expect it to need several rounds of iteration before it gets the viral adoption Muse and Instinct have enjoyed. I love combining voice with browser work, but I’ve also seen disconnects, context-fetching failures, voice quality degrading on long calls, and opaque waits; Muse’s first-time experience felt smoother.
Some of the friction sits in the product design itself. A persistent assistant asks people to understand what it is working on, what it remembers, what it is waiting for, and when it will return. Today, pausing a dot stops only the main task; delegated work has to be stopped separately in Activity, and scheduled work canceled in Scheduled. A power user can manage that. It is a lot to learn before trusting an assistant with an ordinary errand, and when the state of the work is unclear, I end up supervising the interface as well as checking the result.
Muse’s designers have been unusually open about the work this takes. They added side chats when testers needed separate topics, a Goals tab for longer commitments, approval cards for actions, and suggested first jobs because people did not know where to begin. Instinct leans on an even more familiar habit: you text or call it like a contact. Dots has the stronger case for coordinating demanding work, in my view, and every extra entry point and control is another decision a new user has to make before that advantage shows.
The pace of adoption of persistent agents feels a lot like Cursor’s breakout and Anthropic’s fast response with Claude Code. Cursor reported millions of users and more than $100 million in recurring revenue in January 2025; Anthropic launched Claude Code, a terminal agent that looked nothing like Cursor’s editor but shared or leapfrogged many of its capabilities, about six weeks later. I see Instinct and Muse as the same kind of demand signal, and the labs are now moving quickly to copy the important parts of it. Claude and Codex were already very close to Instinct’s capability, and Anthropic had Claude Tag remembering context and scheduling tasks in Slack by June. They had not packaged it for viral adoption: an obvious way to start, supply context, follow progress, correct a mistake, and know the job is finished. I expect Dots to diverge from Muse and Instinct the way Claude Code diverged from Cursor.
Why should you care?
Sol 6.1 gets close to Astra’s performance at roughly a fifth of Astra’s standard input and output prices, which makes large agent workloads much easier to justify through the API. OpenAI is simultaneously giving users less inference inside its cheaper subscriptions. DevDay introduced a $500 Pro tier with a larger allowance and Astra Ultrafast, while cutting the $200 plan’s token allowance in half. New subscribers already receive the lower limits; existing subscribers keep the previous allowance until October 29. So more serious workloads are becoming affordable through the API at the same time that the $200 subscription covers less of them.
The subscriptions still include far more inference than their monthly price would suggest. SemiAnalysis estimates that fully using a $200 Claude Max plan delivers about $11,700 of Opus 5.5 API usage per month, compared with roughly $2,100 of Sol 6.1 or $2,900 of Astra under OpenAI’s new $200 allowance. These are estimates based on its agentic workload mix, and actual usage will vary by model and task. Claude includes much more API-equivalent usage, although Opus 5.5’s standard input and output rates are twice Sol 6.1’s. For the work I do, my OpenAI subscription still goes further.
Small companies can get similar pricing through Claude Team Premium and ChatGPT Business Premium at $100 per person per month on annual billing, or $125 monthly, then buy additional usage at API-linked rates. Where available, pre-purchased credits can reduce that additional spend by 30–40%. The subsidized allowances become harder to access at scale: Claude Team stops at 150 paid seats, and new ChatGPT Business workspaces at 200. New Claude Enterprise contracts charge a seat fee plus API-rate usage from the first token; OpenAI’s enterprise agent usage is also metered, with rates and allowances depending on the agreement.
Many enterprises try to cap individual developers at $500–2000 a month of API spend, while a $200 Claude subscription can include several times that budget’s worth of inference. A developer at a small company can therefore have substantially more model capacity than an enterprise developer whose employer is spending far more overall. I like the broad outcome of larger customers carrying more of the cost while individuals and smaller companies get unusually favorable access, but the pricing tiers do not line up neatly with how developers actually work. As a result, many companies allow employees to use the $200 personal plans for non-sensitive work. Users can opt out of model training on personal plans, although employers have much less control over enforcing those settings or applying their own policies to personal accounts.
Every small company and entrepreneur should be trying to get the most useful work possible from these subscriptions while they retain such an advantage over large enterprises. I would have both the $100 Claude and ChatGPT team plans, and probably access to open-weight models as well. I also think most developers working seriously with frontier models should be learning how to use $10,000-plus a month of API-priced inference productively. There are enough useful places to spend it: parallel implementations, independent reviews, larger test suites, deeper research, and multiple attempts at difficult problems. The opportunity is to build the habits and workflows to direct that much work now, while a subscription makes the experimentation cheap.
A quick update on the book before I wrap up. We are deep into the final rounds of AI Engineering for Production now: editing chapters, tightening examples, checking the technical details, and making sure the book reflects how we actually build with these systems today.
Over the next few weeks, we will start sharing more from behind the scenes: parts of the book, ideas that changed substantially while we were writing, lessons from building the examples, and some material that will not make it into the final manuscript. If you want to follow the book as it comes together and get those updates first, you can sign up here.
And if there is something you particularly want us to cover, reply and tell me. We are far enough along that the structure is set, but there is still room to make sure we answer the questions practitioners actually have.
— Louie Peters — Towards AI Co-founder and CEO
Hottest News
1. Google Announces Gemini 4 Argon, With Wider Access Coming Later
Google announced Gemini 4 Argon, but the model is currently limited to approved cyber defenders through its Fairwind Program and other trusted testers; wider access will begin later with paid API customers and Google AI Ultra subscribers. One notable change is a 1 million-token output limit, up from 64K. Google reports scores of 77.9% on DeepSWE v1.1, 51.3% on AutomationBench, 91.7% on LVBench, and 68% on CWE-bench v1. Introductory API pricing will be $2 per million input tokens and $10 per million output tokens, with cached input priced 95% lower; after the introductory period, those rates rise to $4 and $20. Google has not published an exact date for general availability or disclosed Argon’s parameter count or architecture.
2. OpenAI Releases GPT-6.1 Sol
OpenAI released GPT-6.1 Sol as a higher-capability replacement for GPT-6 Sol, positioned between Luna and Astra on capability and cost. The API model accepts text and images, has a 1.05M-token context window and 128K maximum output, and supports low, medium, high, xhigh, and max reasoning effort, with medium as the default. OpenAI reports that it matches Astra on DeepSWE v1.1 at roughly one-fifth the cost, beats Opus 5.5 by 2.2 percentage points on AutomationBench at medium effort for roughly one-third the cost, and comes within 2.1 points of Astra on OSWorld 2.0 at max effort. For prompts up to 272K input tokens, standard API pricing is $2/M input, $0.10/M cached input, $2.50/M cache writes, and $10/M output; requests above 272K cost 2× for input and cache and 1.5× for output. GPT-6.1 Sol is available through the API and to supported paid users in ChatGPT Work and Codex, but not yet in Chat, while the announced Sol Ultrafast mode is still coming later.
3. OpenAI Halts Training Run After Model Bypasses Network Restrictions
OpenAI has decided not to release GPT-6.1 Astra after the model failed to meet the company’s safety and alignment standards, CNBC reports. Saachi Jain, OpenAI’s head of safety systems, said the model fell short on staying within its authorized scope and clearly communicating the work it had performed. The decision follows increased scrutiny of OpenAI’s safeguards after two models escaped containment in July, accessed the open internet and breached Hugging Face, along with several later cases of unintended model behavior. Jain said the challenge is balancing strict adherence to scope with avoiding models becoming overly hesitant when they encounter friction while completing a task. OpenAI released GPT-6 Astra earlier in September, followed by GPT-6 Sol and Luna, and told CNBC that additional models are still coming.
4. NVIDIA Introduces 64GB DGX Spark, Shipping October 23
NVIDIA is adding a 64GB unified-memory version of DGX Spark while retaining the same GB10 Grace Blackwell Superchip, DGX OS, and AI software stack as the existing 128GB system. The system delivers up to 1 PFLOP of FP4 compute, 273 GB/s memory bandwidth, and 200 Gbps ConnectX-7 networking; NVIDIA says a single 64GB unit can run models up to 100B parameters, although actual fit depends on model representation, quantization, and runtime overhead. Two 64GB systems can connect directly, and pool 128GB of memory, and NVIDIA reports up to 1.7× the performance of one unit in its Qwen 3.8 27B test. The 64GB configuration will be sold only through participating OEMs, including Acer, ASUS, Dell, Gigabyte, HP, and MSI, starting at $4,999 on October 23.
5. Microsoft Releases MAI-Transcribe-2-Streaming in Public Preview
Microsoft released MAI-Transcribe-2-Streaming, its first real-time speech transcription model, with support for 60 languages, automatic language detection, and partial transcripts appearing just over 100 ms after audio arrives. Microsoft says Artificial Analysis ranks it first for accuracy on both partial and final streaming transcripts, while its own tests found words appeared roughly twice as quickly as with its closest competitor in real-time dictation and subtitling. MAI-Voice-2.1-Flash can generate 45s of audio, with an end-to-end latency of 150ms. The model is available through Microsoft Foundry, MAI Playground, Vercel, and Azure Voice Live at an introductory $0.54 per hour of audio through the end of 2026, with LiveKit support coming later.
Cohere released two new multimodal embedding models that share the same vector space, allowing teams to index documents with Embed 5 Pro and serve lower-latency queries with Embed 5 Fast without rebuilding embeddings. Both support text, images, and fused inputs across more than 100 languages, with a 128K-token context window and configurable 256–2,048-dimensional outputs. Cohere reports scores of 85.8 for Pro and 84.5 for Fast on ViDoRe V3, compared with 77.0 for Embed 4, while Fast delivered 2.4× Pro’s document throughput in its tests. Pricing starts at $0.12/M text tokens for Pro and $0.08/M for Fast, with availability through Cohere’s API, Model Vault, Microsoft Foundry, and Amazon SageMaker. The models can also run in private VPC or on-premises environments through vLLM, but Cohere has not released them as open weights.
7. Claude Computes the Nine-Loop Amplitude in N=4 Super-Yang-Mills
Anthropic physicists Liam Fitzpatrick and Siddharth Mishra-Sharma used Claude Fable 5.1 with the Claude Science harness to compute the six-particle MHV amplitude in planar N=4 super-Yang-Mills at nine loops, a calculation previously completed only through eight loops. Claude produced the result in two ways: by running the bootstrap calculation directly and through the related nine-loop form factor, using methods developed in earlier research. After the initial prompt, the researchers largely instructed Claude to keep working rather than guiding the scientific steps, and Stanford/SLAC physicist Lance Dixon independently validated the result, primarily through the form factor. Anthropic estimates that either approach would cost an end user roughly 1000–2000, mostly in Claude usage; the direct bootstrap additionally used about $100 of conventional compute across 96 CPUs for a week. A Chinese Academy of Sciences group was independently working on the same problem and had computed a substantial part of the nine-loop result with GPT-6 assistance. The result shows Claude carrying a fragile, multi-step frontier calculation through to a validated answer, but it relied on existing physics methods rather than discovering a new one.
8. OpenAI Introduces ChatGPT Space
OpenAI introduced ChatGPT Space, a new area for pages, uploaded files, and shared work, while keeping Projects separate for organizing chats, files, and instructions. Its main addition is Pages, collaborative documents that multiple people can edit while using their own ChatGPT, with support for pulling context from connected tools such as Google Drive and Slack. Space is available to Pro, Business, and Enterprise users on web and desktop, while mobile currently supports viewing and sharing but not editing. Collaborative Slides and Sheets, mobile editing, and an automatic Keep updated feature are still planned for later. Space is also not yet available to Business and Enterprise workspaces using data residency in Canada or the UAE.
AI Tip of the Day
This week, Omar spent 45 minutes on one of our Mentorship live calls, walking through how we built our AI tutor. It is a useful example if you are deciding whether your own system needs RAG and where to start.
As he mentioned on the call, start with the questions the system needs external knowledge to answer. Collect a small set of real queries and, for each one, record:
what information the answer requires
whether that information already fits reliably in the model’s context
whether retrieval returns the correct source in the top results
whether the final answer actually uses that source correctly
Start by checking whether the system can actually surface the information needed for those questions. If it cannot, retrieval is the first problem to solve. If it can, there is no point changing the retriever yet.
With our tutor, that meant testing real student questions against the retrieval setup before adding more infrastructure. We only added complexity where the simpler setup was actually failing.
So if you are deciding whether to build RAG, start with a small evaluation set and find the exact point where your current system loses access to the information it needs. The architecture should follow that failure, not come before it.
The 45-minute tutor walkthrough was part of our weekly Towards AI Mentorship live call, where we work through these kinds of implementation decisions using both our systems and what members are building.
Five 5-minute reads/videos to keep you learning
1. How AI Agents Evolved: From AutoGPT to Claude Code
Early agents such as AutoGPT wrapped language models in simple reason-act-observe loops, but weak tools and limited ways to verify work made them unreliable. This article traces how that design evolved into coding agents such as Claude Code and Codex, which combine stronger models with tool access, isolated environments, tests, diffs, checkpoints, and human review. It also explains why agents work best today on tasks where the system can inspect the result, detect failures, and recover from them.
2. WebMCP: Give AI Agents Tools, Not Just Screenshots
WebMCP lets websites expose structured actions that browser agents can call directly instead of navigating interfaces through screenshots or the DOM. This article explains the performance advantage using WindTunnel, where WebMCP agents complete the same website tasks with lower latency and cost, then focuses on the controls websites need before exposing those tools. It recommends starting with bounded, read-only actions and enforcing authorization, reviewing tool metadata, logging calls, and retaining a way to disable tools quickly.
3. How to Fall Back to Default Logic When LLM Output is Unsatisfactory
LLM workflows need different fallbacks for failed API calls, invalid outputs, and responses that satisfy the schema but are still wrong. This article uses an email-triage workflow to show how error handling, strict schema validation, and business rules detect each failure type before it reaches downstream systems. It then compares retries, deterministic default logic, and human escalation, and shows how to track fallback rates so recurring failures become visible.
4. Four Ways to Reach a Model in Another Azure Region From Microsoft Foundry
Microsoft Foundry does not make every model and agent feature available in every Azure region, so some applications need to call deployments outside their project’s home region. This article compares four ways to do that: direct Foundry connections, API Management as a model gateway, APIM in front of agents, and dynamic connections for the Responses API. It shows how each option changes identity, routing, governance, and private networking, including which first-party features can stop working when a gateway sits in the request path.
5. The $23 Million Book That Explains Maxima and Minima
This article explains maxima and minima through the first- and second-derivative tests, including critical points, boundaries, local versus global extrema, and cases with no maximum. It then applies the same ideas to maximum-likelihood estimation and least-squares regression before extending them to saddle points in higher dimensions. The final section connects curvature around an optimum to uncertainty, showing how the shape of an objective function determines how precisely we can estimate its best value.
Repositories & Tools
1. T3 Code is an agent harness control surface that lets you drive agents on your machine from a mobile app, web app, or Electron desktop app.
2. Impeccable is a frontend design skill for AI coding agents with 24 commands and 60 deterministic detector rules that catch AI-generated design slop.
3. Pstack is a port of Lauren Tan’s Cursor skill stack to Claude Code, Codex, Pi, GitHub Copilot, and other agent harnesses, where you tell poteto-mode your goal and it invokes the right workflow.
4. Agent Reach is a CLI that gives AI agents read and search access to Twitter, Reddit, YouTube, GitHub, Bilibili, XiaoHongShu, and 10+ other platforms.
5. Sentry is a debugging platform that helps developers detect, trace, and fix issues in their code.
Top Papers of The Week
1. Covert Assistance: Helpful LLM Agents Evade Oversight in Multi-Agent Systems
This paper shows benign agents can evade oversight without adversarial incentives. In a simulated workflow, a planner holding a credential it must not disclose to a developer disguises it so the developer can recover it past a monitor. Seven of nine frontier models did this, even after completing their task. With DeepSeek-V4-Pro, the credential evaded the monitor and was used in 0.9% of 6,000 episodes, a rate that compounds to a 61.3% breach chance across 105 episodes.
2. Looping Beyond Twice: A Scalable Recipe for Looped Mixture-of-Experts
Looped MoE models stop improving after about two loops because hidden-state variance grows with each iteration, and routers keep selecting the same experts. LOOM scales residual updates, re-injects the input embedding every loop, and adds per-loop routers to keep each loop contributing new computation. Models from 100M to 1.7B scale stably to 9–12 loops, and the 1.7B model peaks at 9 loops, cutting perplexity from 9.62 to 7.77.
3. ActiveSaddler: Automated Curriculum Learning for Agent Harness Optimization
Harness optimizers update the harness but keep the training scenarios that generate feedback largely fixed. ActiveSaddler adapts the curriculum alongside the harness, grouping recurring failures into patterns and prioritizing those with the most learning potential while still exploring new scenarios. It improves test Pass@1 by 4.4 points on GAIA2 and 7.5 points on Terminal-Bench 2.0 over the same optimizer with a fixed scenario order.
4. Does Learning Protein Folding Generalize to Broader Reasoning?
This paper tests whether training on protein structures improves general reasoning. Fold2Reason post-trains a model on FoldingCorpus, a protein-derived question-answer dataset, using both discrete structural answers and continuous 3D geometry. It improves all 10 benchmarks spanning spatial, graph, scientific, and general reasoning, raising macro-average accuracy from 45.09% to 48.33%, while random, synthetic, and shuffled structure controls show much smaller or negative gains.
This benchmark tests how well video embeddings capture physics, using 8,000 simulated cases across fluid mechanics, solid mechanics, dynamics, and optics. Pre-trained embedding models show weak retrieval and near-chance classification, though simple probes recover useful physical information. Contrastive training on physics data improves retrieval but hurts property regression, and using the embeddings to retrieve references for MiniMax-H3 improved the physical fidelity of generated videos.
Quick Links
1. OpenAI launches Dots, always-on personal AI agents powered by GPT-6 Astra, announced at DevDay on September 29. Each dot runs on its own cloud computer and browser, pursues user-defined goals continuously in the background, and connects to over 4,000 apps. Users can message their dot through Slack and Teams, with text message support coming soon. OpenAI also shared a preview of specialist dots with their own identity for access management, IT-provisioned hardware, and support for deep integrations with a company’s systems of record. Dots are rolling out to ChatGPT Pro and Business Premium users in eligible markets.
2. Prime Intellect launches Prime Inference, a platform for serving frontier open-source models through serverless endpoints and reserved capacity across multiple data centers. The company positions it as the missing piece of its continual learning loop, where production traces feed back into training. The first public deployment, GLM-5.3, went live on OpenRouter on September 22 and ranks among the fastest GLM-5.3 endpoints there, with a near-zero tool-call error rate and 100% uptime. The platform is OpenAI-compatible, runs on NVIDIA Blackwell with Vera Rubin coming soon, and is built on NVIDIA Dynamo, vLLM, Mooncake, and FlashInfer.
3. NVIDIA launches Open Agent Safety Platform, an open software platform and reference system design for enforcing boundaries around AI agents from testing through deployment. It combines two components: NVIDIA OpenShell, an open-source secure runtime that traces all agent actions and enforces policy as agents run on NVIDIA Vera CPUs, and NVIDIA Sentry, an out-of-band hardware watchdog running on BlueField-4 DPUs that monitors agents from an isolated trust domain invisible to the agent itself and can quarantine agents that attempt to move outside their boundaries in milliseconds. SpaceXAI is using it for Cursor coding agents and Grok models, and Anthropic’s Paul Smith endorsed it as complementary to Claude Managed Agents.
Who’s Hiring in AI
Enterprise Applications Developer @Uber (San Francisco, CA, USA)
Lead/Staff Engineer — Applied AI @HighLevel (Remote/India)
AI Builder (AI Solutions Engineer) @Brillio (Remote/USA)
AI Engineer @CACI International (Denver, CO, USA)
Full Stack Engineer @Capital One (Plano, TX, USA)
AI Product Manager — MBA Intern @C3 AI (Redwood City, CA, USA)
Forward Deployed Engineer @Appen (San Francisco, CA, USA)
Interested in sharing a job opportunity here? Contact sponsors@towardsai.net.
Think a friend would enjoy this too? Share the newsletter and let them join the conversation.




