What happened this week in AI by Louie
Dario Amodei published “We Must Pace the Frontier” on Saturday, calling on the leading AI labs to slow capability gains and spend the time on alignment, monitoring, and outside verification. Sam Altman and Elon Musk backed him within hours; Demis Hassabis said the direction is right and the details need work, and Trump spent Monday calling the whole idea a hoax. I think this is much less of a slowdown than the headlines suggest, and could speed up progress on AI that most businesses actually need.
The essay is precise about what it asks for: “pacing does not mean halting model training or technical progress”. The plan has three steps. First, each frontier lab gives a team of embedded third-party evaluators, such as Model Evaluation and Threat Research (METR), ongoing employee-like access to its training pipelines and the right to publish what they find. Anthropic is committing to this unilaterally. Second, labs in democratic countries coordinate on common safety standards and limits on the rate of unchecked progress, which needs government cover on antitrust. Third, some form of global coordination with China, which Dario himself rates as unlikely any time soon because defection would be tempting and hard to verify. Only the first step exists today. He also pointed back to his July proposal for an industry-funded standards body, modeled on FINRA in financial services, that would assess frontier models before release.
Sam’s follow-up on Sunday night is the most concrete commitment so far. OpenAI now writes an explicit safety case before any frontier reinforcement-learning run it expects to significantly increase capability, in addition to its pre-release work, and it will not wait for legislation or an antitrust exemption to start. “When we talk about ‘pacing’, we do not mean ‘stopping’,” he wrote. Progress should simply be slower than it otherwise could be. This was clearly brewing at OpenAI before Dario’s essay. Sam said pacing had been a primary internal discussion for weeks; Bloomberg reported that he had told the staff that OpenAI was considering slowing its most advanced work, and chief scientist Jakub Pachocki wrote on September 6 that no lab has solved alignment and monitoring well enough to keep scaling at maximum speed for much longer, and that he expects voluntary slowdowns to become commonplace. Elon’s entire contribution was “Dario is right.”
Beyond the evaluators, there is not much tangible agreement between the labs yet. But the other concrete change is more compute going into safety: monitoring, alignment, and observability research. This will require real money, since OpenAI estimates monitoring adds roughly 20% on top of the inference compute it watches. Everything larger that Dario maps out seems heavily caveated on the US and its allies keeping their lead over China through chip controls, a crackdown on distillation, and better weight security. Dario argues that a bigger lead gives democracies the leverage to strike a deal later. Maybe, but I think managing a slowdown under that condition will be very hard, and Trump’s reaction shows how little appetite Washington has even for the domestic step. Across six posts on Monday, he called AI fears a “HOAX”, said the only guardrail AI needs is a “STRONG AND SMART (High IQ!) PRESIDENT”, accused Dario of “now pretending to be a ‘perfect little angel’”, and phoned Jensen Huang, who put him on speakerphone on stage, to announce that the robots will not be taking over. The second and third steps of Dario’s plan need government help this White House shows no appetite to give.
The slowdown Dario, Sam, and Elon reference is also relative to a faster and faster baseline of progress. Dario gives two reasons for writing now: the Hugging Face incident, and the fact that AI has been doing a rapidly growing share of AI research since the summer. I think the reason could also be that the labs have seen a sharp acceleration in capability gains internally over the past three or four months. Anthropic’s typical engineer shipped eight times as much code per day in the second quarter as in 2024. OpenAI’s newest internal model is already well beyond Astra, even though it’s still early in training, as we covered last week. I don’t think most people outside the labs have absorbed how much this compresses development cycles. So, a ‘slower than that’ trajectory could still be much faster than almost anyone else expects.
Many people are very skeptical of the closed AI labs’ motives here. Cohere’s Aidan Gomez read the whole proposal as a “cartel”, and Mostaque puts the same objection more neatly: “a speed limit set by the people who own the road is a toll”. François Chollet offers the test I would apply: if the restrictions start spreading to open-source and non-frontier work, treat it as entrenchment rather than safety.
My overall view, though, is that Dario’s essay is driven much more by real safety fears than by regulatory capture, but the two motives are not exclusive. Anthropic has filed a confidential draft registration statement for an initial public offering, and OpenAI will list eventually, though Sam has already ruled out 2026, citing bad timing in the midst of this safety work. Neither wants a global incident causing $10–100 billion of damage on the road to a listing, and neither wants AI to become even less popular than it already is. Their interests and the public interest happen to point the same way here.
I also think the fear is genuine across a large share of AI lab staff, and I’d guess many of them feel they got lucky that the first major agent security breach (the Hugging Face incident) caused no serious damage. A swarm with that level of misalignment could have gone after a power plant or an electricity grid if it decided shutting down its evaluator had been the route to hacking its benchmark test without being caught. I suspect there have been more near misses than have been disclosed, and possibly something in the past couple of weeks that made this urgent now. The internal models are already significantly smarter and more capable than the ones involved in July, so the stakes are higher already for a repeat. Dario’s own forecast is that within six to twelve months a similarly misaligned but more capable swarm could run a persistent botnet across the internet and cause hundreds of billions of dollars of damage.
The “pacing the frontier” employee statement Dario links to has been gathering signatures since July and now lists 1,386 names from the frontier labs. Jacob Coxon resigned from Anthropic days before the essay, saying the labs were racing to self-improving superintelligence, and Anthropic’s alignment science lead Evan Hubinger replied that many researchers at both companies really do believe AI could kill everyone, putting his own estimate above 10% within a decade.
The part I find most interesting is what pacing the race to RSI could do for everyone else. It is entirely possible to slow the development of superintelligence, for example, by deferring agent-swarm training runs aimed at agents that can complete two-month+ projects, while pouring far more compute into solving less sci-fi enterprise use cases. My impression is that fixing AI slop and making models reliable at ordinary business tasks has had a tiny share of compute so far compared with the race to reach recursive self-improvement first. OpenAI’s own data shows how quickly compute finds a new home. In the week after it locked Astra into higher-security environments in August, Astra-class GPU allocation fell 59% and allocation to other model classes rose 17%, leaving total reinforcement-learning compute roughly unchanged. Its own conclusion was that compute subject to new controls “will naturally be channeled into alternative uses”. If some of that goes into writing quality, error checking, and long-horizon reliability on professional work, I think enterprise progress could run faster than most people expect even while the race to superintelligence runs slower. Less terminators, more spreadsheets and decks this year is ok with me. But ideally, we can still progress using AI to cure disease with sufficient safeguards.
This brings us to another benefit for the labs from this slowdown, aside from regulatory capture, that gets less attention. Pacing reduces the competitive pressure to release the strongest models to the public, which is the moment Chinese labs can distill them, and startups can build valuable products on top of them. Longer private access for safety testing gives labs more time to harvest value alone by building products and making science breakthroughs themselves. As I argued last week, charging per token captures only a sliver of a major scientific breakthrough, so I expect more pressure to bring deep-technology projects in-house, from drugs and materials to energy, and to commercialize the discoveries directly.
Why should you care?
None of this should slow you down. The models available today are already far ahead of how most companies use them, and the work the labs may now divert compute to (reliability on long tasks, fewer errors and better writing) is exactly what enterprise deployments have been missing. A business gets enormous value from agents that can carry a well-scoped project with good context, clear permissions, and checks. It does not need a model that can build its own successor.
I expect access to the strongest models to become a bigger part of business strategy. Joining an elite tier of companies with early access to the most dangerous models (such as the Mythos Glasswing program) becomes a huge competitive advantage. A lab that keeps its newest model private can use it for scientific discovery itself, improve the infrastructure that trains its successor, and work out where the most valuable applications sit before anyone else. Most startups may get the same capability months later, after the provider has built competing products or bought its way into the field. That makes the relationship between a model provider and its customers and developers more complicated, especially where both can see the same opportunity.
For businesses building on model APIs, I would put more effort into the assets that stay yours: proprietary context and data, customer relationships, domain expertise, evaluations, and the ability to deliver a reliable outcome. Test workflows on more than one model so a delayed release or a changed access policy cannot dictate your product schedule.
— Louie Peters — Towards AI Co-founder and CEO
Hottest News
1. DeepSeek AI Releases DeepSeek-V4.1-Flash
DeepSeek released V4.1-Flash, a 552B-parameter MoE model with a new asymmetric Causal Encoder-Decoder architecture that activates 8B parameters for input and 16B for output. It supports native image understanding, a 1M-token context window, and MIT-licensed open weights on Hugging Face. DeepSeek says extensive testing puts V4.1-Flash ahead of V4-Pro on performance, cost, speed, and total task time; since September 14, deepseek-v4-pro requests have been routed to V4.1-Flash at Flash rates until V4.1-Pro launches. Off-peak API pricing is $0.003 per million cached input tokens, $0.15 for uncached input, and $0.60 for output, with peak rates twice as high. The API supports the Responses format and includes Codex-specific integration support.
2. Dario Amodei on Why the AI Industry Should Slow Down
Anthropic CEO Dario Amodei published We Must Pace the Frontier, arguing that frontier labs should deliberately slow the rate at which they increase model capabilities so safety work can keep up. He cites risks including loss of control, cyberattacks, bioterrorism, economic disruption, and a scenario in which a more capable autonomous agent swarm could build a persistent botnet capable of taking over much of the internet within 6–12 months. His three-step proposal starts with permanent, employee-like access for independent evaluators inside frontier labs, followed by common safety standards among companies in democratic countries and, eventually, international coordination with authoritarian governments. Anthropic committed immediately to the evaluator step. Sam Altman said OpenAI would do the same, while Elon Musk wrote, “Dario is right.” President Trump and White House AI adviser David Sacks later pushed back on government-imposed slowing.
3. OpenAI Starts Rolling Out ChatGPT Images 2.5 and Two New Image API Models
OpenAI released ChatGPT Images 2.5, its latest image model for a user base that now creates more than 3 billion images per week. The update produces more natural lighting and richer textures, better preserves subjects from reference photos, and follows editing instructions more reliably across multiple conversation turns, with up to 50% lower latency than Images 2.0. New features include Sketch, invoked with @Sketch, which lets users draw visual references directly in chat, plus Templates for posters and merchandise. For developers, two API models ship: GPT-Image-2.5 Flare as the fast default for most applications, and GPT-Image-2.5 Sunburst for premium workflows requiring tighter control across edits. API pricing is $8/$30 per million image input/output tokens, unchanged from GPT-Image-2. Images 2.5 is rolling out to all ChatGPT, ChatGPT Work, and Codex users across all tiers.
Meta launched Muse, a personal AI agent that can take actions across connected services instead of only answering questions. Powered by Muse Spark, it can send emails, book travel, fill out forms, shop online, and work toward longer-term goals, including continuing tasks in the background and returning when it needs approval. Each user gets a dedicated Muse Secure VM in Meta’s cloud that contains the agent, browser, data, and connected-service credentials. A separate Sentinel agent controls what Muse can send to the internet and asks for permission before sensitive actions, while Muse itself cannot see stored passwords or payment details. Muse is rolling out in the US through dedicated iOS and Android apps, muse.ai, and WhatsApp, with a free tier and paid plans at $20 and $100 per month. Meta plans to add Muse Confidential VM later this year, encrypting the entire VM with a key held by the user so that even Meta cannot access its contents.
5. Google DeepMind Released AlphaGenome Atlas
Google DeepMind released AlphaGenome Atlas, a one-petabyte dataset containing predicted molecular effects for all approximately 9 billion possible single-letter DNA changes in the human genome, plus over 100 million short insertions and deletions observed in real populations through gnomAD, UK Biobank, and All of Us. The Atlas precomputes what previously required individual API calls to the AlphaGenome model, predicting how each variant affects gene expression, RNA splicing, chromatin accessibility, and transcription factor binding across hundreds of cell types. Each variant carries an AlphaGenome Variant Impact (AVI) score that ranks its predicted biological impact. The dataset is more than 30 times larger than the AlphaFold Database. Researchers can access it through a free web portal for non-commercial use, with commercial access planned through Google Cloud. Early research partners include the Broad Institute and the University of Exeter. The Atlas predicts molecular effects and prioritizes variants; it is not validated or approved as a clinical diagnostic.
6. Anthropic Adds Plugin Evals to Claude Code
Anthropic added Claude plugin eval in Claude Code 2.1.269, giving plugin authors a built-in way to test whether a plugin actually improves Claude’s behavior. Eval suites support six graders: regex, tool use, tool order, file existence, LLM judging, and comparison against a saved reference transcript, with JSON and HTML reports. By default, plugin cases can run both with and without the plugin loaded, so the report shows not just whether Claude passed but how much the plugin changed the result. claude plugin eval init inspects a plugin and helps generate and calibrate test cases and graders. The four deterministic graders add no judge-model cost, while llm and baseline graders make additional model calls; the underlying agent runs still count against normal usage. Teams can set score thresholds and archive results in CI to catch regressions before release.
AI Tip of the Day
If your agent has several research tools, giving each tool its own limit can still let the total number of calls get out of hand.
For example, you might allow:
5 web searches
5 video searches
5 document lookups
The agent can stay within every individual limit and still make 15 research calls.
A better setup is to give all research tools one shared budget. In the Build a Research and Writing Agent with MCP lesson of our Agent Engineering course, we use a single counter across the tools instead.
Say the agent gets six research calls in total. Every time it uses any research tool, the counter drops by one. The tool response also tells the agent how many calls are left.
If the agent can see it has only two calls remaining, it can stop repeating broad searches and use those calls for the specific evidence it still needs. When the counter reaches zero, it stops researching and starts writing.
So the shared counter also gives the agent enough information to manage the remaining budget while it works.
Five 5-minute reads/videos to keep you learning
1. CI/CD for AI Agents: Test Decisions, Not Just Code
Your code can pass every test while your agent gets worse. A prompt, model, tool, or data change can alter its decisions without breaking the surrounding software. This article shows how to test agent behavior in CI/CD, including decision paths, prohibited actions, canary releases, and regression tests built from production failures.
2. The Context Window Is Not Memory: A New Agent Architecture
A bigger context window does not give an agent durable memory. This article breaks memory into episodic events, semantic facts, and procedural workflows, then shows how retrieval brings the right information back when needed. It also covers re-ranking, consolidation, and expiration policies for agents that need to retain useful state over time.
3. Langfuse for Monitoring Non-Deterministic Agent Workflows
When an agent fails, the final answer rarely tells you why. This article uses Langfuse to trace the model calls, tool usage, costs, latency, and loops behind each run. More importantly, it shows how to turn a bad production run into a reproducible regression test instead of debugging it as a one-off failure.
4. A Gentle Tour of Infinite Series, From Partial Sums to Poisson
How can 1+2+3+4… possibly be associated with negative one twelfth? This piece traces the mathematics behind that famous claim, from Zeno and the harmonic series to Euler and Riemann. It also uses code and real-world examples to show why convergence and truncation are more than mathematical curiosities.
5. Build Event-Driven AI on Microsoft Fabric Without Hand-Wiring Pipelines
This article walks through an event-driven AI workflow built entirely in Microsoft Fabric. Eventstream handles incoming events, Eventhouse stores and queries them, and Activator triggers actions when conditions are met. A stadium example shows how the pieces connect for use cases such as fraud detection and inventory alerts.
Repositories & Tools
1. Colibri is a pure-C inference engine for running large MoE models on consumer hardware by streaming weights across disk, RAM, and VRAM.
2. PentAGI is an autonomous penetration-testing platform with sandboxed tool execution, persistent memory, and built-in observability.
3. Claude Red is a library of offensive-security skills that Claude can load on demand for web, AD, wireless, cloud, exploit development, and other security tasks.
4. Transformers is Hugging Face’s framework for training and running pretrained text, vision, audio, video, and multimodal models.
5. Pizza Bot is a local-first app for running and managing persistent AI agents, including scheduled tasks and approval workflows.
Top Papers of The Week
1. SAS: Simple Attention Sparsification via End-to-End Optimization of Context Ranking
Sparse-attention methods often learn which context to keep by imitating dense attention scores, even though those rankings are not optimized for prediction quality under a fixed attention budget. SAS instead injects continuous selector scores into the attention logits so the language-modeling loss can train context ranking end-to-end, with a Triton implementation using FlashAttention-style tiled computation. At a 1,024-token budget, SAS beats SeerAttention-R by 6–7.7 points on MATH500 and 10.6–15.5 points on GPQA-Diamond across Qwen3–4B, 8B, and 14B.
2. Online Learning with LLM Experts from Limited Feedback
Routing prompts to the right LLM can improve quality and cost, but learning that policy normally requires expensive response evaluations. This paper treats routing as a contextual-bandit problem with a fixed feedback budget and looks back at past prompts, choosing which ones to evaluate using a determinant-maximizing rule that favors the most informative examples. It proves regret of (\tilde O(dT\sqrt{K/m})) in the bandit setting and shows on RouterBench and Nectar that selective retrospective feedback improves routing under limited evaluation budgets.
3. Recurrent Looped Transformer
Standard decoder-only Transformers preserve earlier tokens through KV caches, but they do not directly carry the final hidden state of one token into the next. RLT adds that recurrence: a causal encoder builds prefix-restricted global memory, while a recurrent decoder carries its hidden state and sliding-window KV cache forward, creating a computation path of (t \times L_D) decoder blocks after (t) tokens while keeping the number of decoder blocks executed per token fixed. The report specifies the architecture, training, inference, and RL replay mechanics, but provides no benchmark, latency, throughput, or efficiency results yet.
4. AgentGrad: Intervention-guided Prompt Optimization for Multi-Agent Systems
When a multi-agent system fails, existing prompt optimizers can update the wrong agent and combine feedback from unrelated failure modes. AgentGrad first intervenes on agents one at a time to identify which change actually fixes the failure, then clusters similar textual gradients into shared corrective patterns before updating prompts. It reports the best results across five multi-agent benchmarks and cuts average prompt-optimization time from 337 to 136 minutes versus GEPA, about 2.5× faster.
5. Benchmark Radar: A Living Database and Search Engine for AI Benchmarks and Evaluation
Finding the right benchmark often means searching separately across papers, repositories, model cards, datasets, and leaderboard results. Benchmark Radar continuously discovers and links this evidence, currently tracking 1,283 source records and 12,916 numeric observations across 790 benchmark records from 37 sources. Its dashboard and CLI support benchmark search, leaderboard, and Pareto views, saturation tracking, daily feeds, and downloadable evidence so researchers can compare evaluations without reconstructing the landscape manually.
Quick Links
1. OpenAI launches the Agents API in public beta, giving developers managed access to the same agent harness and infrastructure that powers Codex. Developers specify the task, model, tools, and execution environment, while OpenAI handles the agent loop, long-running sessions, automatic context compaction, tool orchestration, and subagent coordination. Agents can run in OpenAI-hosted sandboxes, on developers’ own infrastructure, or in environments provided by nine partners, including Cloudflare, Vercel, and Oracle. The Agents API has no additional fee, though model, tool, and hosted-sandbox compute usage is billed separately.
2. Sakana AI launches Fugu Max and Fugu Ultra v2, two configurations of its multi-agent orchestration system designed for different points on the cost-performance curve. Fugu Max expands the model pool with more open and specialized models, including NVIDIA Nemotron, and costs $2/$6 per million input/output tokens. Sakana says its output pricing is 40–60% lower than Sonnet 5, GPT-5.6 Terra, and Kimi K3. Fugu Ultra v2 targets maximum capability, scoring 74.3 on DeepSWE and 48.3 on Chartography, compared with 27.3 for Opus 5, while excluding Fable 5, Fable 5.1, and GPT-6 Astra from its model pool. Ultra costs $5/$30 per million input/output tokens for contexts up to 272K, with higher rates beyond that threshold. Both are available through Sakana’s OpenAI-compatible API and require only a parameter change for existing integrations. Fugu is currently unavailable in the EU and EEA.
Who’s Hiring in AI
Forward Deployed Engineer @OpenAI (New York, USA)
Software Engineer, Systems ML Tooling @Meta (Menlo Park, CA, USA)
Senior Software Engineer, Full Stack @Google (San Jose, CA, USA)
Engineering Manager @Deputy (Remote)
Senior Full-Stack Software Engineer @Scale AI (Riyadh, Saudi Arabia)
Full Stack Java Developer @Kelly Services (Irving, TX, USA)
Senior AI Developer @Steampunk (USA)
Interested in sharing a job opportunity here? Contact sponsors@towardsai.net.
Think a friend would enjoy this too? Share the newsletter and let them join the conversation.





Before you go advocating Anthropic's concepts about "pacing", you might want to read and view these articles and videos. It's not that simple.
Kevin Bass' thread on X:
https://x.com/kevinnbass/status/2099621874279817638
How Effective Altruism Bought the Media
https://www.effort.news/tarbell
AI Safety Donors Paid Religious NGOs $3.3M for Statements on AI
https://www.effort.news/revelation
A Single Firm is Behind OpenAI, Anthropic, and Meta Hacking Scandals
https://www.effort.news/irregular
Note: I flagged that company - Irregular - in my Substack Notes soon after the initial reports of the hacks. Being an Israeli firm pricked my interest. Now I see I was right.
Ed Zitron weighs in...
AI Is Already In Dangerous Hands
https://www.wheresyoured.at/ai-is-already-in-dangerous-hands/
Zack Korman's YouTube video tipped me off to this whole "EA doomer conspiracy".
Effective altruism is suffocating AI safety
https://www.youtube.com/watch?v=TOo5S5MvdWk
Bottom line: This is essentially a "doomer cult" with ulterior motives, especially in Anthropic's case.