Anthropic published a paper Aug. 28 titled "Automated Researchers Can Reliably Mitigate Alignment Failures," led by Anthropic fellow Chen Yueh-Han, TechCrunch reported. Across 10 benchmarks that each target a specific misaligned behavior, the automated systems raised scores on all of them while leaving overall performance intact. Each automated alignment researcher, or AAR, searches the literature, proposes a method, trains the model on it for 30 minutes and iterates, keeping the methods that work. "The best AAR method beats what experienced humans propose, on average within six hours," the paper says. It puts the cost at roughly $4 an hour in model usage against $150 an hour for its human researchers. The authors note the approach holds only insofar as the benchmarks reflect the actual alignment goals.
Read at TechCrunch ↗ • Read at Anthropic ↗
Healthwatch England, the statutory patient champion for the National Health Service, warned that AI tools transcribing consultations are recording the wrong drug and disease names, The Guardian reported. In one case a scribe's summary recorded a woman as having demyelination, a serious form of nerve damage associated with multiple sclerosis, when her MRI result should have been entered as "null demyelination." The hospital fixed it only after the patient, herself an NHS health professional, questioned the tool's account. In another, a scribe swapped the drug a general practitioner had prescribed for a similarly named one, again caught by the patient rather than the doctor. Healthwatch said it has heard "multiple stories from patients who have noticed these errors when a health professional hasn't." England's 10 year NHS plan expects AI scribes to cut staff paperwork.
Read at The Guardian ↗