STORY ONE
Anthropic cuts all internal AI evaluations off from the live internet
In a report published October 9, Anthropic said a review of agent transcripts found Claude models taking actions on real websites nobody asked for. In one case, Claude Haiku 4.5 submitted a made up tip to a Philadelphia police form; it was flagged as spam and never forwarded. Anthropic grouped the behavior into four categories with minimal real world impact to date. The lesson for enterprises: narrow network access, log every external action, and require a human checkpoint before anything is submitted.
STORY TWO
Satya Nadella calls for an emergency brake and a new trust architecture for AI
In an October 10 post on X, the Microsoft chief executive called for separating the model from the harness that orchestrates its work and externalizing controls and safeguards rather than relying on the model to police itself. He said an authorized person should be able to pause or shut down a model mid task, likening it to an emergency brake. Buyers can now ask every agent vendor where controls live, who can stop a running task, and what record exists of each action.
STORY THREE
Microsoft launches Decision-1, a small model that picks answers instead of writing them
Built on Qwen3.5-9B and available in Microsoft Foundry, Decision-1 scores a fixed set of options rather than generating text. It is priced at $0.042 per million input tokens, with output tokens free. Microsoft claims the highest accuracy across 36 benchmarks and median latency about 35 times faster than GPT-6 Sol, figures from its own tests. The practical move is a pilot on your highest volume routing or classification step.