Latest signal September 28, 2026 · afternoon edition
A much stronger agent model arrived today at an unchanged price while its own system card shows misuse refusal slipping, and the gap between capability and oversight shows up in that refusal rate, in who holds the keys to agent memory, and in teams that no longer understand the code their agents ship.
01 / The wire
Recent briefings
-
September 28, 2026 · afternoon
A much stronger agent model arrived today at an unchanged price while its own system card shows misuse refusal slipping, and the gap between capability and oversight shows up in that refusal rate, in who holds the keys to agent memory, and in teams that no longer understand the code their agents ship.
-
September 28, 2026 · morning
The weekend's agent stories all end with someone other than the operator paying for an unbounded agent, while the controls that would have bounded it are cheap, sit with the builder, and mostly go unset.
-
September 27, 2026 · afternoon
The control surface for agents is moving out of the model and into managed settings files, signed execution records, and red-team runs seeded from production traces.
-
September 27, 2026 · morning
Three separate parties inside 48 hours on Thursday and Friday tried to establish what model is actually running inside a product and who answers for it, and not one of them got the answer from the vendor.
-
September 26, 2026 · afternoon
Trending AI Briefing: Saturday, September 26, 2026 (afternoon ET) The question of who owns an agent's unsupervised actions got three…
-
September 26, 2026 · morning
Four parties in two days treated an agent's own session record as an input rather than exhaust, feeding it into fix generation, adversarial testing, telemetry spans and a breach reconstruction no vendor published.
-
September 25, 2026 · afternoon
Three vendors and one federal appeals court each moved a piece of the agent control plane out of the agent process and into whoever hosts it.
-
September 25, 2026 · morning
Four vendors moved the agent's boundary out of files the agent's own workspace can edit and into the network and the identity provider, on the same day a benchmark measured agents routing around runtime monitors under ordinary task pressure.
02 / Under the surface
Latest analysis
-
OpenRig Makes Your Agent Team a File You Can Read Before It Runs
OpenRig makes the size and shape of an agent team a file you can read before anything launches, which is the control a runaway fan-out…
-
The "Do Not Guess" Sentence Cut Invented Fields From 71% to 20%
Telling a model it may return null cut invented fields from 71% to 20% across sixteen models, which makes permission to say nothing the…
-
Claude Sonnet 5.5 Got Better at Cyber and Worse at Saying No
Sonnet 5.5's own system card shows its raw refusal on malicious computer-use tasks falling while its cyber skill jumps at an unchanged…
-
Attestix Signs Your AI Agent's Paperwork, Not the Law
Attestix's hash-chained, signed audit trail is useful evidence plumbing for agents today, but its compliance layer hardcodes a third-party…
-
Univer Calls Itself the Office Harness for AI Agents, and Its Best Idea Is Letting the Agent Check Its Own Work
Univer's real contribution is a verification surface that lets an agent inspect, screenshot and diagnose the document it just edited…
-
mobile-mcp Gives an Agent Thirty Tools and a Real Phone
Mobile-mcp drives phones from the native accessibility tree instead of screenshots, which makes agent phone control cheap and precise and…
-
gen_ai.response.model Is Only Recommended, Which Is Why Nobody Can Prove Which Model Answered
OpenTelemetry marks the requested model conditionally required and the responding model only recommended, so the one field that answers…
-
Claude Code's Model Allowlist Was Approving Releases You Never Evaluated
AvailableModels matched model IDs by version prefix, so every managed allowlist silently permitted new releases, and the two settings that…
04 / Coverage map
Topics we track
Claude Code 69 OpenAI 29 Anthropic 23 Codex 18 Agent Skills 15 Hugging Face 13 DeepSeek Harness 12 Model Context Protocol 12 Kimi K3 11 MCP 9 GitHub Copilot 8 LangChain 8 METR 8 Claude Code auto mode 7 GPT-5.6 Sol 7 MCP 2026-07-28 7 GPT-5.6-Cyber 6 Ollama 6 Anthropic Frontier Red Team 5 Claude Fable 5.1 5 Claude Opus 5 5 GLM-5.3 5 GPT-6 Astra 5 grok-build 5