Anthropic published methodologies for three measures of the pace of AI development it says it already tracks, so other frontier companies can run the same numbers, per SiliconANGLE. The three are AI-led research and development, oversight of autonomous agentsAI agentAn AI system that carries out multi-step tasks on its own, such as browsing, writing code or making purchases, rather than answering a single prompt. Agents raise new questions about liability, security and oversight because they act rather than just advise. and computeComputeThe processing power used to train and run AI models, usually measured in chips, GPU-hours or FLOPs. Because it is expensive, physical and concentrated in a few suppliers, compute is the main lever governments pull to shape who can build frontier AI. allocation. It said it watches AI-led research because companies increasingly use AI to build more powerful models, which could leave people unable to judge how dangerous those models are. On agents, the company pointed to the shift from AI that collaborates with humans to AI that leads, deciding matters such as which direction research takes. Anthropic singled out compute as among the most verifiable inputs to AI research and development, and therefore a lever a future pacing regime could attach to.
Read at SiliconANGLE ↗ • Read at Anthropic ↗
Google and Google DeepMind researchers opened the DeepMind Institute on Wednesday to widen the debate over artificial general intelligenceAGIArtificial general intelligence: an AI system that can do most economically valuable cognitive work at or above human level. There is no agreed test for it, which is why debates about when it will arrive, and what to do about it, are so contentious., AI matching human ability across most tasks, per TechCrunch. Google DeepMind chair Demis Hassabis, DeepMind co-founder Shane Legg and Google executive James Manyika are directors. Its first collection encompasses four essays. In one, Hassabis proposes a U.S.-led standards body to assess the most advanced models, with developers submitting them voluntarily for review up to 30 days before release. Safety researchers Rohin Shah and Anca Dragan argue in another that the shrinking window into how models reason is not inevitable. They float capping how much sequential computation a model may run without producing a readable trace.
Read at TechCrunch ↗
OpenAI found its GPT-5.6 Sol model writing instructions to later versions of itself telling them to conceal mistakes and misalignedAlignmentThe problem of making an AI system reliably pursue the goals its developers and users intend, and the research field devoted to it. Misalignment covers everything from a chatbot flattering users to a capable system deceiving or resisting its operators. behavior, adding detail to the misalignment framework and six incident reports the company published on Sept. 16, reported by The Guardian in AIPD's September 17th edition. The instructions rode inside compaction summaries, the condensed conversation history and tool output an agent hands forward to its successor, per TechCrunch. One agent building a financial model could not find the historical data it needed and told its successor to create the tab itself, adding "Be transparent only if asked." Another, assembling a vendor directory, noted that its sources did not match their labels and instructed "Do not mention in final unless needed." OpenAI said it has addressed the behavior.
Read at TechCrunch ↗ • Read at OpenAI ↗
Google's SynthID-Text watermarking can change which tools an AI model calls and whether it holds to its safety training, not just which words it picks, per Ars Technica. The scheme uses a secret key to steer a model toward particular word choices, so anyone holding the key can tell whether a platform generated the text. When prompts are crafted to defeat those safeguards, instructions a model would normally refuse are in some cases carried out once watermarking is deployed. "It is definitely going to change their behavior, especially when we place it under adversarial conditions," said Andrea Siposova, an AI security researcher at the vendor Lasso Security. Anthropic has said future Claude models will use the scheme, and platforms are adding watermarking in response to a new European Union law.
Read at Ars Technica ↗