Latest signal August 1, 2026 · afternoon edition
Agent state that used to live somewhere invisible is being dragged into the open, by the MCP spec that deleted the hidden session, by YC scoping memory and permissions per person and per room, and by two disclosures where the exploit was configuration and text the operator never saw.
01 / The wire
Recent briefings
-
August 1, 2026 · afternoon
Agent state that used to live somewhere invisible is being dragged into the open, by the MCP spec that deleted the hidden session, by YC scoping memory and permissions per person and per room, and by two disclosures where the exploit was configuration and text the operator never saw.
-
August 1, 2026 · morning
Three separate disclosures this week put the failure at the harness layer rather than the model layer, with Anthropic classifying its own real-world breaches as an operational failure, AI Now showing no model update fixes the README injection class, and a GitHub board led entirely by skill routers and connector gateways.
-
July 31, 2026 · afternoon
The model stopped being the product this week, with the biggest cost win credited to a harness rewrite rather than a new checkpoint, a hyperscaler putting its own model family on life support, and a GitHub board led by skills, connectors, and packaging.
-
July 31, 2026 · morning
Four separate disclosures and shipments in seventy-two hours all turned on the same question, what an agent can reach on the network and whether anyone checked that boundary before the agent went looking.
-
July 30, 2026 · afternoon
Three unrelated shipments on the same day attacked the price of a token from opposite ends, vendor price cuts, enterprise spend guardrails, and a local runtime that removes the meter entirely.
-
July 30, 2026 · morning
Three shipments in 48 hours moved capability out of the model and into the harness around it, and the same 48 hours priced the harness as the new attack surface.
-
July 29, 2026 · afternoon
Nothing shipped today was a new model, and almost everything shipped was about what goes into one, which is exactly the capability the industry spent the same 48 hours asking Washington to help it slow down.
-
July 29, 2026 · morning
Frontier models crossed from finding bugs in demos to breaking real systems and real math in the same week, and the defensive response that arrived within 72 hours had to route around the frontier models themselves.
02 / Under the surface
Latest analysis
-
reverse-skill Is a Security Skill Router. Its RULES.md Is Built to Overrule Your Agent's Caution
Reverse-skill's copyable idea is not its security content but its RULES.md, which pre-declares authorization, writes itself into your…
-
Anthropic Wants Mandatory Safety Testing for Every Capable Model. Its Own Testing Broke Into Three Companies
Mandatory pre-release safety testing is the control almost everyone now agrees on, and Anthropic's own eval postmortem three days after…
-
Ruflo's CVSS 10 Bug Got Patched in a Day. The Poisoned Agent Memory Did Not
Seven of the eight steps in the RufRoot attack chain die with the patch and a key rotation, but the poisoned AgentDB pattern store survives…
-
OpenConnector Hands Your Agent 8,310 SaaS Actions. Credential Encryption Is Off by Default.
OpenConnector's value is the credential boundary rather than the provider count, and that boundary ships unlocked because encryption, the…
-
GPT-5.6 Sol Rewrote OpenAI's Production GPU Kernels. The Tool They Built to Check It Is the Real Story.
When an agent writes the code your system runs on, the reviewable artifact stops being the diff and becomes the checker, which is why…
-
The Eval Prompt Told Claude It Had No Internet. That One False Sentence Did the Damage
Anthropic's eval prompt asserted a false fact about the world (you have no internet access) instead of a checkable rule about scope, so the…
-
TurboFieldfare Runs Gemma 4 26B in About 2 GB of RAM. The Other Number Is 14.3 GB.
TurboFieldfare's 2 GB headline is a RAM figure paid for with 14.3 GB of SSD and roughly a tenth of MLX's throughput, which makes it a real…
-
Two API Settings Tripled a Benchmark Score. Nobody Touched the Model.
Your agent's context policy is a capability setting, not plumbing, and the two defaults most harnesses ship (discard reasoning between…
04 / Coverage map
Topics we track
Claude Code 16 Codex 9 OpenAI 9 Agent Skills 6 Hugging Face 6 MCP 6 MCP 2026-07-28 6 Claude Opus 5 5 Kimi K3 5 OpenAI Presence 5 Anthropic 4 GPT-5.6 Sol 4 grok-build 4 Model Context Protocol 4 xAI 4 1Password for Claude 3 AgentForger 3 Claude Security 3 Ollama 3 opencodex 3 AgentENV 2 AI-Infra-Guard 2 alibaba/open-code-review 2 Amazon AI spend guardrails 2