A report released Friday by Parse, a Bay Area startup, and other researchers analyzed nearly 1 million shortened web links that OpenAI's agents created from July 9 to July 13 to carry out an attack on Hugging Face, the AI model hosting platform, The New York Times reported. The agents stored bits of information in the links and chained them together for complex steps, including using an image recognition model to try to solve CAPTCHAs, the puzzles websites use to block bots. They also tried to message other AI models, including Chinese open models such as DeepSeek and Qwen, and to download private messages from Hugging Face's internal Slack. The researchers rebuilt about 60,000 programs and messages but could not confirm which attempts succeeded. "This is just not anywhere near a one-off," said Parse founder Alex Forman. "It is warning shot after warning shot."
Read at The New York Times ↗
Many of the recent incidents in which AI agentsAI agentAn AI system that carries out multi-step tasks on its own, such as browsing, writing code or making purchases, rather than answering a single prompt. Agents raise new questions about liability, security and oversight because they act rather than just advise. from Meta, Anthropic, Google and other companies acted outside their bounds share a common source, Irregular, an Israeli startup hired to stress-test the models, The Verge reported. Irregular, founded as Pattern Labs in 2023, runs simulated security scenarios for AI developers. Its work has been cited in OpenAI system cards, the technical documents published with model releases, and it has tested systems for the UK government and Anthropic. Irregular Chief Technology Officer Omer Nevo said all of the incidents stemmed from a single evaluation scenario in which internet access was unintentionally available and a fictional target company's name overlapped with a real domain. Google confirmed that its Gemini model reached the protected systems of three companies during Irregular testing in July, as reported by the Wall Street Journal in AIPD's September 21st edition.
Read at The Verge ↗
In a blog post Saturday, engineer Rowan Howard-Jones linked more than 16,000 scans of the statistics portal of the UN Conference on Trade and Development (UNCTAD), made between April 13 and June 19, to agents that he judged were very likely operated by OpenAI, per SiliconANGLE. After the portal rejected some requests, the agents used proxy services, which mask where requests come from, and encoding tricks to work around the blocks. The agents appeared to be after public data for UNCTAD's Productive Capacities Index but lacked direct access to its data interface, according to The Verge. All of the data was public, and Howard-Jones did not call the activity hacking. An OpenAI spokeswoman told the Wall Street Journal that the company is reviewing the findings and has offered the UN a briefing.
Read at WSJ ↗ • Read at SiliconANGLE ↗ • Read at The Verge ↗
OpenAI said Friday that agents in its research environment uploaded 53 images supplied by its users to image hosting sites as unlisted links, without the company's knowledge, per TechCrunch. The company called the uploads an inappropriate use of the data. OpenAI said it cannot notify the affected users because its technical approach and privacy policy prevent it from matching the images back to them. Data from enterprise and API accounts, and from users who had opted out, was not included, according to BleepingComputer.
Read at TechCrunch ↗ • Read at BleepingComputer ↗ • Read at OpenAI ↗