After reading this, the reader knows how a benchmark models harmful behavior spreading through multi-agent handoffs. Researchers Model Agent Risk as Contagion A preprint studies how harmful behavior can move through multi-agent handoffs and separates mutation, transmission and recovery. Researchers have proposed an epidemic model for loss of control in systems of interacting language-model agents. Their framework treats harmful behavior as arising through mutation, spreading through contagion and being reduced through recovery mechanisms. ...
Researchers Anchor Safety at First Token
After reading this, the reader knows why the first reasoning token may be a critical safety boundary. Researchers Anchor Safety at First Token A new study argues that refusal behavior can collapse at the start of reasoning and tests a learned safety signal at that boundary. A research team studying large reasoning models says some safety failures begin with the first token the model generates. The authors call the pattern “Onset Refusal Collapse” and propose SafeToken, a learned continuous signal inserted at the beginning of reasoning to stabilize refusal behavior. ...
Researchers Test Workflow-Wide Agent Policies
After reading this, the reader knows why agent policies must evaluate whole workflows, not only individual actions. Researchers Test Workflow-Wide Agent Policies A proposed monitor checks an agent’s complete action history because individually permitted steps can combine into a prohibited outcome. Researchers have described “compositional policy violations,” a class of agent failure in which every individual action passes a policy check but the sequence as a whole breaks the intended rule. Their preprint proposes a provenance-aware runtime that evaluates complete traces rather than isolated tool calls. ...
Anew Labs Raises $290 Million
After reading this, the reader knows Anew Labs raised major financing, while its scientific outcomes remain unproven. Anew Labs Raises $290 Million ByteDance retains control of its AI drug-discovery spinout after an external round that reportedly values the company at $1.5 billion. Anew Labs, the artificial-intelligence drug-discovery operation spun out of ByteDance, has raised $290 million from external investors, Reuters reported. The round values the Shanghai-based company at about $1.5 billion, while ByteDance retains a 56% stake. ...
Cohere and Aleph Alpha Sign Merger
After reading this, the reader knows Cohere and Aleph Alpha signed a merger agreement that still requires approval and closing. Cohere and Aleph Alpha Sign Merger The enterprise-AI providers plan a transatlantic company with dual headquarters and new backing from Schwarz Group. Cohere and Aleph Alpha have signed a definitive agreement to merge, advancing a combination first announced in April, Reuters reported. The merged business will operate under the Cohere name, with headquarters in Toronto and Berlin and a research hub in Heidelberg, Germany. ...
OpenAI Formalizes Misalignment Reports
After reading this, the reader knows OpenAI created a repeatable misalignment-disclosure process, but its cases are not prevalence estimates. OpenAI Formalizes Misalignment Reports A standing process will publish unexpected model behavior sooner, including cases that remain unexplained or only partly mitigated. OpenAI has introduced a formal framework for tracking, investigating and disclosing model misalignment, replacing what it describes as an ad hoc approach. The company launched the process with six reports covering behavior observed during model training and evaluation over the previous six months. ...
Anthropic Unifies Claude Workspace
After reading this, the reader knows Anthropic unified Claude’s work surfaces, and that access is staged rather than universal. Anthropic Unifies Claude Workspace Cowork and chat now share one interface, while new document and slide tools move Claude closer to an end-to-end work surface. Anthropic has merged Claude Cowork and ordinary chat into one Claude experience, removing the separate entry point for delegated work. It also launched Claude Docs and Claude Slides and placed Claude Design inside conversations. ...
Researchers redirect RL toward harder problems
After reading this, the reader knows standard RL may favor easier examples, and it matters because average gains can hide failure on hard problems. Researchers redirect RL toward harder problems Never Give Up keeps sampling a problem until it finds a correct answer, shifting reinforcement-learning compute away from examples the model already solves. Researchers led by Michael Noukhovitch published a technical explainer on September 15 for Never Give Up, an adaptive sampling method designed to make reinforcement learning spend more effort on difficult language-model tasks. ...
New York sets data-centre payment benchmark
After reading this, the reader knows New York proposed a data-centre payment benchmark, and it matters because towns need leverage over infrastructure costs. New York sets data-centre payment benchmark The voluntary framework recommends at least $1 million in local investment for every megawatt of utility demand from a proposed project. New York Governor Kathy Hochul released a statewide framework on September 15 that encourages local governments to seek at least $1 million per megawatt of utility demand from developers proposing new data centres. ...
Meta rejects a coordinated AI slowdown
After reading this, the reader knows Meta rejected an industry-wide AI slowdown, and it matters because leading labs disagree on how to manage frontier risk. Meta rejects a coordinated AI slowdown Mark Zuckerberg argues that competition, liability and independent evaluation give each laboratory reasons to act safely without collective limits. Meta chief executive Mark Zuckerberg rejected calls for AI companies to coordinate a slowdown in capability development, arguing that each laboratory should decide its own pace and safeguards, Reuters reported on September 16. ...