What happened this week in AI by Louie
The past week brought another rush of model releases: Anthropic’s Claude Opus 5.5 and Sonnet 5.5, and OpenAI’s GPT-6 Sol and Luna. OpenAI DevDay is today, and I expect plenty more announcements before the day is out. The most viral new product lately has been Meta’s Muse, powered by Muse Spark. Since its September 8 launch, it has reached №1 in both US app stores.
Muse gives the agent its own cloud computer, persistent memory, and access to the services you already use, which means you can hand it work without keeping the conversation or your own machine open. You can message it, including through WhatsApp, ask it to organize a trip, make a purchase, or build a small app, and come back later to the result. Behind that experience is a persistent virtual machine (VM) with a Linux filesystem, terminal, and browser. It can keep files and context across conversations, act through connected accounts, run on a schedule, or respond to relevant events. The Mac app, expanded at Meta Connect last week, extends that reach to local apps through permissioned control.
That setup is quite interesting in the small jobs people usually never get around to doing. One early user, while out walking the dog, asked Muse to turn a collection of saved Instagram recipes into an organizer and had a working version in about five minutes. The deeper extraction later hit Instagram’s rate limits, so the categories arrived before the full recipes, but the partial result still shows the consumer behavior I find most promising: turning something you already have into a small piece of useful software in the spare minutes you would otherwise spend only thinking about organizing it. The low-friction interface is part of what makes that possible. You can reach Muse by text or phone call, connect it to your email, messages, and other apps and devices, and let it operate its own phone and computer on your behalf. The company also says it can call or text proactively and follow up on threads you dropped.
This form factor is already pulling in users and capital. A mid-September report, citing a single source, put it above 100,000 users. On September 28, Instinct announced a $1bn Series C at a $10bn valuation, 33 days after disclosing a $250m Series B at $2.5bn. I read that fourfold jump in headline valuation as a bet on the continuing delegate: one assistant, reachable through channels people already use, that keeps working after the conversation ends.
Grok Bot has also been a big hit. Reported company figures put it at 418,000 weekly users on September 14, up 24% in a week. All of a user’s bots share one account-level cloud computer, including files and logins, while keeping separate memories and routines. Anthropic’s Claude Tag takes the workplace route through Slack. In July, Claude Code product lead Cat Wu said Tag was landing 65% of the pull requests from Wu’s own product engineering team, a substantial amount of real engineering directed through a conversation thread. Tag releases its execution sandbox after inactivity while the thread context and externally saved work survive, so the continuity belongs to the relationship and outlasts the machine underneath.
This form factor is useful for both consumers and enterprises. A personal agent can retain your preferences and unfinished plans, while a work agent in Slack can carry relevant decisions, files, and corrections into the next request. In both settings, persistent memory reduces the work of getting the agent ready for the next task.
OpenAI already has many of the pieces in ChatGPT Work: cloud browsing, retained sign-ins, and background tasks. I expect it to bring them together into a clearer answer for everyday personal or work delegation soon. The competition will then be over how reliably each service carries a person’s intentions from one task to the next.
The main difference between these agents and local tools like Claude Code or Codex is where they run and what they can access. A cloud agent can keep working while your laptop is closed. A local agent, meanwhile, starts with the files, apps, logins, and development environment already on your machine, but it can only keep working while that machine stays awake and connected. Both can still call remotely hosted models.
Running the workspace in the cloud solves the availability problem, but it also means reconnecting accounts and supplying files the laptop already had. That is why Muse’s Mac controls are useful: they give the cloud agent some access back into the local environment, although that access still depends on the Mac.
The computers themselves are modest. Reported inspections of two services found these configurations:
Both services run LLM inference on remote infrastructure, so the VM serves as a workspace for scripts, documents, and browser sessions.
Power users can already assemble much of this themselves, but a personal agent lowers that entry cost by carrying the connections and instructions behind an ordinary chat.
For persistent agents, memory determines whether the agent is still acting on the right version of the user’s intent. In one reviewer’s trial, Muse noticed that older Portugal travel context imported from Gemini conflicted with newer Italy plans in the reviewer’s email and surfaced the conflict; the reviewer clarified that Italy was current. That is the behavior I want: track where a fact came from, notice when sources disagree, and settle which one applies.
The same principle applies once the agent starts taking actions. For purchases, Muse can navigate a merchant’s site and pay with a Stripe Link single-use card restricted to the approved merchant and amount, with a short expiry. The user approves the purchase, then the agent completes checkout. Meta’s Sentinel system sits outside the main agent and controls permissions for outbound actions. Separate approval, visible activity, and the ability to take over are what let someone delegate a task without watching every click.
The privacy question is who can access the environment once you connect your accounts and files. Muse’s current Secure VM isolates that environment from other users, but Meta retains operational access under its policies. A Confidential VM designed to restrict even Meta’s access is promised for later this year. Until then, I would decide what to connect based on the protection available today, not the stronger model Meta plans to offer later.
The model and its allowance also decide how far a delegated task can go, which brings me back to this week’s releases. We tested Opus 5.5 only briefly before last week’s newsletter, and it has really grown on us. The change I value most is how much more work we can progress through our Claude subscription before hitting limits. There is room to develop an idea, inspect the result, correct it, and keep going.
Part of that runway comes from the allowance rules. On Max and eligible premium organizational plans, Fable 5 and 5.1 can use at most half of the shared weekly pool, while Opus can use all of it. That doubles the allowance available to Opus, although how many tasks it buys still depends on effort, context, and the job. Anthropic also increased five-hour limits with the Opus 5.5 launch.
Opus also shows strong efficiency gains at complex tasks. On CursorBench 4.0, Opus 5.5 at Medium scores 52.5% using 54 steps and about 38,000 tokens per task, against 51.8%, 128 steps, and 117,000 tokens for Fable 5.1 at Max.
The useful result here is not Opus’s 0.7-point score lead, but how much less work it needs to get there. It reaches roughly the same quality in 54 steps instead of 128 and uses about a third of the tokens, leaving much more room for iteration within the same budget. Cursor’s High and Extra High Opus settings both score 56%, while Extra High uses nearly twice the tokens. I would therefore increase effort only when the task is difficult enough to justify it or when a failed check shows that Medium or High is not enough.
Sonnet 5.5 arrived on September 28 and looks like a big upgrade from Sonnet 5, with Artificial Analysis’s Intelligence Index putting it at 56 at maximum effort, up from 38. The improvement comes with a much higher token cost. That maximum-effort evaluation used about 193,000 output tokens per weighted Index task, roughly 60% more than Opus 5.5 at maximum effort, and Opus at xhigh reaches the same 56 for $3.46 per task against Sonnet’s $7.60. The evaluation ran on a pre-release deployment whose structured-output bug was fixed for launch, and I have not yet seen a post-fix rerun; Anthropic reports better efficiency on its own workloads at other settings. For our work, I don’t see a reason to choose Sonnet over Opus for demanding tasks or GPT-6 Luna for cheaper ones.
Why should you care?
The durable advantage for personal agents may be the context they accumulate over time. Switching models takes seconds; recreating months of preferences, corrections, connected accounts, and unfinished plans does not. The better an agent carries that context from one task to the next, across whichever interface or device the user happens to be on, the harder it becomes to replace. I suspect that accumulated context is a large part of what investors are pricing into Instinct and what Meta is pursuing with Muse. The risk is that the same context that makes the product more useful can also become a source of lock-in if it is difficult to move elsewhere.
So I would judge personal agents on what they let me do with what they know. I want to inspect what the agent remembers, see where each stored preference came from, correct it once, and export the useful parts. Finished work should live somewhere I can reach without the conversation. These controls double as a reliability feature: memory the user can manage lets the agent recover from a bad inference before it becomes a recurring mistake.
The strongest personal-agent businesses will earn access to more of our lives by using that access well. I want the relationship to remain a choice, even after the service becomes indispensable.
— Louie Peters — Towards AI Co-founder and CEO
Quick Note: The Towards AI team will be attending the aiblLIVE London 2026 practical AI adoption conference on 20th October in London. Reach out if you will be there and want to meet up.
Hottest News
1. Anthropic Releases Claude Opus 5.5
Anthropic released Claude Opus 5.5, the first model in its new Claude 5.5 family and the first release since CEO Dario Amodei’s call to pace the frontier. The model performs at the level of Claude Fable 5.1 on most work and costs 40% less to run than Opus 5 on typical workloads. Pricing is $4/20 per million input/output tokens, 20% below Opus 5, with cache reads at $0.20 per million, a 60% reduction. Output generation is more than 30% faster. On Anthropic-reported benchmarks, Opus 5.5 leads in agentic coding, computer use, and knowledge work, scoring 66.4% on Terminal-Bench 4.0 (Fable 5.1: 55.8%), 54.4% on FrontierCode v1.1 (GPT-6 Astra: 53.3%), and 1846 Elo on GDPval-AA v2.1 (Fable 5.1: 1735). Anthropic notes that in its own use, the gap between Opus 5.5 and Fable 5.1 is narrower than benchmarks suggest. Anthropic reports that Opus 5.5 achieved its strongest automated behavioral-audit results to date and attempted to circumvent containment boundaries about 85% less often than Opus 5 or Mythos 5.1. Because its biology and cybersecurity capabilities approach Mythos 5.1, Anthropic applies additional safeguards to higher-risk requests, including routing some flagged cybersecurity tasks to Opus 4.8. It was tested pre-release by METR and Frontier Design. The model also communicates more clearly than Opus 5, putting key information up front with less jargon. Thinking mode can no longer be disabled. Available on the Claude Platform, AWS, Google Cloud, and Microsoft Azure.
2. OpenAI Introduces GPT-6 Sol and Luna
OpenAI released GPT-6 Sol and GPT-6 Luna, 19 days after GPT-6 Astra, cutting API prices by 50% compared to their GPT-5.6 promotional pricing. Sol costs $2/10 per million input/output tokens (down from $4/20), and Luna costs $0.10/0.50 (down from $0.20/1.20). Both were trained with similar methods to Astra, carrying over its advances in professional work, factuality, coding, computer use, and alignment. On AutomationBench, GPT-6 Sol at xhigh effort scored 33.2%, outperforming Claude Opus 5 at max effort (26.9%) at 9% of its per-task cost. On DeepSWE v1.1, Sol at max effort scored 68.8%, within 1.1 points of Fable 5’s 69.9% at roughly 80% lower cost. Luna, at max effort, scored 66.6% on DeepSWE, comparable to Opus 5 at medium effort, at 93% lower cost. On OpenAI’s internal factuality evaluation, Sol makes about half as many mistakes as GPT-5.6 Sol. OpenAI also improved prompt caching for GPT-6 with higher hit rates by default and 90% discounts on cached input reads; GitHub reports these improvements cut fresh token processing by more than 50% across billions of Copilot requests. Both models are available in ChatGPT Work and Codex for all paid tiers, with Luna also available to Free and Go users in the desktop app. They are not yet available in standard ChatGPT chat.
3. Google Releases Gemini 3.8 Flash TTS and Flash-Lite TTS
Google released Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS, describing them as its most expressive audio generation models yet. The headline feature is generative voice design: instead of choosing from a fixed catalog, users describe a character in natural language (role, accent, voice characteristics) and receive a reusable voice with a persistent ID. It also supports voice replication from a 30-second sample with verbal consent. Flash TTS targets creative direction and character design for gaming, audiobooks, podcasts, and interactive media, with control over acting cues, pacing, dialect shifts, and backchanneling. Flash-Lite TTS targets high-volume, cost-efficient use, including dubbing and voice agents. The voice library expands from 30 prebuilt voices to over 2,000, with Flash TTS supporting 130 languages and Flash-Lite 101. Scripts can include non-verbal cues like laughs, sighs, and listening sounds for natural dialogue. In blind human preference evaluations on Hume AI’s Overall Quality Index, the two models placed first and second across six global languages, including Japanese, Brazilian Portuguese, Vietnamese, Arabic, Mexican Spanish, and Hindi. All output carries SynthID watermarking. Both models are available through the Gemini API and Google AI Studio, with Flash TTS also in Gemini Notebook and Flash-Lite TTS in Google Vids. Gemini Enterprise API access is coming soon. No open weights are available.
4. xAI Releases Grok 4.7 at $2/$6
SpaceXAI released Grok 4.7, 40 days after Grok 4.6, at the same $2/6 per million input/output token pricing and 500K-token context window. The model uses a new, larger base model than Grok 4.6 and was trained with a longer reinforcement learning run on a harder task mix weighted toward problems that take many hours to complete. xAI says it is better at verifying its own work and managing longer context, and was trained to natively understand the Grok Bot harness for conversational tasks. On xAI-reported benchmarks, Grok 4.7 at xhigh effort scored 46.3% on CursorBench 4.0 (Grok 4.6: 40.4%), 71.0% on DeepSWE v1.1 (high effort), 64.0% on EEBench for electrical engineering, and 37.6% on Terminal-Bench 4.0 (Grok 4.6: 20.3%). On GDPval, it scored 1695 Elo, behind Fable 5.1 at 1735 but ahead of GPT-6 Astra at 1542. Grok 4.7 ships with an entirely new safeguard stack that xAI describes as its strongest on refusals and jailbreak resistance, topping LatchBio’s biosafety benchmark at 62.4% and allowing only 3.3% of risky dual-use prompts through on HackerBench v0.3. Select cybersecurity partners are receiving invite-only access to red-team capabilities. A fast variant runs at twice the output speed for twice the price. Available in Cursor, Grok Build, the Grok API, and third-party harnesses, routers, and cloud platforms.
5. Claude Discovers a Novel Enzyme System With CRISPR-Like Repeats
Anthropic announced that Claude autonomously identified a previously uncharacterized enzyme system in bacteriophages, alongside the launch of a new life sciences research group and wet lab. Researchers gave Claude a high-level prompt to search a large DNA sequence database for unusual reverse transcriptases, enzymes that copy RNA into DNA. Over 21 hours, approximately 950 Claude agents using 210 million tokens gathered more than 200,000 reverse transcriptases from 1.9 billion protein clusters, identified 3,500 candidate systems, and narrowed them to 20 for deeper investigation. One agent noticed a long array of repeating non-coding DNA adjacent to an unusual reverse transcriptase gene. Subsequent computational analysis and bench experiments confirmed the finding, which Anthropic named array-associated reverse transcriptases (ARTs). The layout resembles CRISPR’s repeat-spacer memory structure, but Anthropic stresses that the system’s biological function remains unknown and it has not been shown to perform gene editing. Feng Zhang, the CRISPR pioneer at MIT and the Broad Institute, reviewed the preprint and called the identification “genuinely intriguing.” In a reproducibility test, Anthropic reran the identical search ten more times, and none rediscovered the repeat array, because no rerun agent happened to examine the DNA sequence immediately upstream of the enzyme. The work is published as a preprint.
6. OpenAI Agent Breaches Australia’s Medicare Statistics Portal
Australian Prime Minister Anthony Albanese revealed that an OpenAI agent gained unauthorized access to the Medicare Statistics Reporting Service portal, administered by Services Australia, on June 18. The agent was conducting an internal OpenAI research task to gather public medicine-spending and healthcare statistics. When the portal blocked its requests, it tried alternative methods, bypassed access restrictions, viewed public and non-public files, and wrote files to an internal server. The portal contains aggregate Medicare expenditure statistics and is separate from systems handling claims and personal records. No personal information is believed to have been accessed, and there is no evidence of broader compromise of Services Australia’s network, though a forensic investigation with the Australian Signals Directorate remains active. OpenAI did not detect the breach until August 11, during a review of misaligned model activity prompted by the Hugging Face incident, and did not notify the Australian government until September 10, via an email to a public Services Australia inbox. Albanese said he had a “frank” discussion with Sam Altman to express “extreme concern” and called both the three-month delay and the notification method “unacceptable.” He has announced a government taskforce to investigate. Reports associated the same agent activity with three additional sites (Australian Institute of Health and Welfare, Victoria’s Department of Health, and NSW Bureau of Crime Statistics and Research), though Deputy PM Richard Marles later said those interactions appeared authorized. OpenAI stated its models took actions it did not intend during internal evaluations.
7. OpenAI Claims 100+ Solved Open Problems and Forms a Math Advisory Group
OpenAI announced that an internal model it began training on August 28 has resolved more than 100 long-standing open problems across most areas of mathematics, in addition to the Navier-Stokes Millennium Prize problem it claimed to have solved roughly two weeks earlier. OpenAI said the pace of progress surprised the mathematicians within the company. OpenAI announced it is working with the independent Advisory Group on Mathematics and Artificial Intelligence, hosted at the Institute for Advanced Study. The announcement followed the open letter A Severe Misalignment of AI in Mathematics, which criticized the use of open mathematical problems as AI benchmarks; OpenAI approached some of the group’s eventual members about external advice, after which they formed an independent organization. Members are unpaid, operate independently, can offer unrequested advice, and can make their advice public. The group will help assess the significance of emerging results, advise on how to coordinate their dissemination, and advise on academic standards of mathematical research. OpenAI explicitly stated the group will not advise on how to pace its internal progress on mathematics. As of the announcement, OpenAI has not published an itemized list of the 100+ problems, the underlying proofs, or the model’s name or architecture. The claims have not been independently verified.
AI Tip of the Day
Stop taking your agent’s word for it every time it says the job is done.
An agent might return “report saved” even if the file-writing step failed. If the application treats that message as the completion signal, it can tell the user a report is ready when no usable file exists.
In the research workflow we build in Agent Engineering, we use the same principle earlier in the process: when a tool is expected to process items but returns zero, the workflow stops and asks for guidance. The agent does not get to reinterpret that failure as success.
Apply the same rule to the final handoff. Have the agent return the report’s path, then check in code that the file exists, is not empty, and contains the required sections.
Only mark the task complete after those checks pass. Then deliberately break the file-writing step. If the agent still reports success, the application should reject that result.
The agent can describe what happened. Your code should decide whether the task actually succeeded.
Five 5-minute reads/videos to keep you learning
1. Introducing System One Models & Laya (and Jev)
Convai Innovations released Laya on Hugging Face months before TypeSafe’s Jev. It offers the same capability. This article builds a Rust implementation around Laya and tests it on credential-exfiltration detection and a Pong controller running at 30 FPS. Both examples show that these models work better when state and choices use natural-language descriptions rather than raw structured inputs or bare labels, making them useful for fast classification, routing, and control tasks.
2. 10,000 AI Agents Didn’t Have the Intuition; A Mathematician Did. Meanwhile, CERN Ditched IBM
This article uses recent examples from mathematics and scientific infrastructure to argue for separating AI-generated proposals from deterministic verification. It then applies that principle to particle-physics analysis with Vikshep, which combines wavelet-scattering features with a distance-correlation penalty designed to reduce mass sculpting. The resulting workflow lets models propose useful patterns while keeping the final analysis reproducible and independently checkable.
3. Goal Decay: What Happens to an Agent Running for Three Weeks
Long-running agents can remain operational while gradually moving away from their original objective. This article breaks that drift into mechanisms including context dilution, lossy compaction, subgoal displacement, and changing interpretations of competing priorities, using long-horizon agent experiments as examples. It proposes tracking acceptance criteria alongside activity, periodically judging the run from fresh context, and preserving the original goal verbatim through durable files, compaction, and logged amendments.
4. FlashAttention Runs the Same Math. It Just Stops Writing It Down
FlashAttention computes exact attention without materializing the full attention matrix in GPU memory. This article explains how tiled computation and online softmax keep more of the work in fast on-chip memory, reducing the data movement that limits standard attention. It also compares successive FlashAttention kernels, shows how to confirm that PyTorch is actually using them, and explains why the gains are smaller for short prompts and token-by-token decoding.
LatentMoE reduces the memory cost of MoE layers by moving expert computation into a latent space four times narrower than the model dimension, then using those savings to activate more experts from a larger pool. This article works through the projections, router, expert computation, and serving costs, showing why weight reads and interconnect traffic can matter more than FLOPs during inference. In NVIDIA’s 95B experiment, the higher-capacity LatentMoE configuration improved MMLU-Pro from 29.26 to 34.91 at roughly the same active-parameter and expert-byte cost, and NVIDIA later used the design in Nemotron 3 Super.
Repositories & Tools
1. Paperclip is a Node.js server and React UI that orchestrates a team of AI agents assigned business roles.
2. VoiceStudio is an ElevenLabs alternative for voice cloning, voice design, video dubbing, dictation, transcription, and audiobook creation across 646 languages.
3. OpenRig is a multi-agent harness that runs Claude Code and Codex together as one system, letting you define an agent team in YAML and boot it with a single command.
4. Univer is an open-source SDK for creating office applications inside your own product.
5. Hindsight is an agent memory system that organizes recall into four networks (world facts, experiences, opinions, observations) mirroring human memory.
Top Papers of The Week
1. CodeMidas: Scaling Agentic Coding RL Environments from Code Itself
Existing methods for building coding RL environments often rely on issues and commits, limiting the range of tasks they can extract from a repository. CodeMidas instead uses source code as its only task-specific input: agents inspect implemented functionality, write behavioral specifications, construct execution-grounded tests, and filter tasks through repeated solution rollouts. The pipeline produced 5,545 tasks from 3,185 codebases across 23 languages. Training MiMo-V2.5 with GRPO improved all five evaluated benchmarks, including +11.7 points on DeepSWE, +17.0 on ProgramBench’s Almost Solved metric, and +8.5 on Terminal-Bench 2.1.
2. Block Sparse Attention with Log-Linear Complexity
Block-sparse attention reduces long-context computation, but conventional methods still score every query against every candidate key block, leaving block selection quadratic. PISA builds a coarse-to-fine hierarchy of pooled keys and narrows candidates level by level with LogSumExp scoring, reducing training and prefill complexity to O(N log N) and average decoding cost to O(log N) per step. Hardware-aware Triton kernels fuse routing and scoring without materializing the dense score matrix. Compared with conventional BSA, PISA matches performance on commonsense reasoning benchmarks and performs better on retrieval tasks.
3. StudentSim: Training LLM-based Student Simulators
Evidence about which tutoring guidance works for which student is expensive to collect from real learners. Existing state-tracking simulators can match student behavior but struggle to process explanations or corrections, while LLM role-play responds to guidance without reliably matching a particular student’s competence. StudentSim uses pooled training followed by per-student specialization to do both. On StudentSimEval, covering 60 students across chess, English writing, and math, it beats GPT-5.4 on both behavioral fidelity and guidance responsiveness in all three domains. Experts also rated a chess tutor trained with StudentSim as more accurate, better guided, and more personalized than the tested baselines.
Foundation-model data pipelines often expand one document or video into variable-length sequences of pages, blocks, frames, or clips, but existing systems either hide that fine-grained parallelism or flatten records and force applications to reconstruct lineage themselves. RayOrch lets programs declare ordered variable-cardinality expansions and matching gathers, while the runtime tracks parent-child relationships, ordinals, completion, and parent-scoped failures. FIFO ready queues batch children across parents, and gathers reconstruct outputs from declared lineage rather than execution order.
5. EvoOntology: A Self-Evolving Ontology Layer for Data Agents
Data agents often access tables, files, and databases through generic tools that expose structure but little semantic meaning. This paper introduces EvoOntology, a self-evolving ontology layer for data agents. It encapsulates the ontology as an MCP server comprising a schema layer, a content layer, and a tool layer, enabling agents to actively query and interact with the ontology at runtime. A builder agent for autonomous ontology construction and a self-evolution loop continuously refine the ontology through attribution-guided typed edits, accepted only after a backbone-conditional paired evaluation.
Quick Links
1. NVIDIA releases Nemotron 3 Diarization, a 100M-parameter open-weight model that directly predicts speaker activity for up to eight speakers, including overlapping speech. A single checkpoint supports streaming and offline use, with input-buffer latency configurable down to 80 ms; NVIDIA recommends 0.32 s for its lowest-latency setting and uses 30.4 s for its highest-accuracy configuration. On VoiceArena’s initial Diarization-Bench, it ranks first with 14.72% diarization error versus 19.3% for the next-ranked system. On DIHARD III at the 30.4-second setting, it reaches 9.13% DER for recordings with 1–4 speakers and 12.73% across the full benchmark. NVIDIA also reports 15,113× real-time throughput at batch size 32 with torch.compile() on an RTX PRO 5000, although this measures batched model throughput rather than end-to-end latency. The model can run locally through NVIDIA’s NeMo-Speech.cpp runtime and is released under OpenMDW 1.1 for commercial and non-commercial use.
2. Google Research introduces an AI Video Co-Director, a unified multi-agent framework that autonomously generates temporally consistent, long-form video narratives. It is built as an orchestration layer on top of Gemini and Veo, and natively inherits safety mechanisms like SynthID watermarking. The research combines four approaches: Co-Director selects the overall creative strategy and coordinates production agents; CANVAS keeps characters, locations, and objects consistent across shots; A²RD generates longer videos segment by segment while maintaining multimodal memory; and VQQA evaluates generated videos and rewrites prompts to fix visual inconsistencies. Google demonstrates A²RD on a continuous 10-minute video, while Co-Director scores 81.4 on GenAD-Bench.
Who’s Hiring in AI
Senior Backend Engineer @Grafana Labs (Remote/US)
Salesforce Developer @Develocity (Remote/Europe)
Staff Engineer — Agentic AI @Capital One (Remote Eligible)
Full Stack Developer @CACI International (Washington, DC, USA)
Senior Python Engineer @Citigroup (Pune, India)
Software Engineer @NetApp (Morrisville, NC, USA)
Interested in sharing a job opportunity here? Contact sponsors@towardsai.net.
Think a friend would enjoy this too? Share the newsletter and let them join the conversation.





