01 / The wire
Recent briefings
-
August 25, 2026 · afternoon
The measurement layer stopped being a bolt-on and became the shipped product, with LangChain releasing three separate agent-grading systems in one day while OpenAI's CFO priced the whole stack in cost per successful result.
-
August 25, 2026 · morning
The harness stopped being plumbing and became the thing being engineered, with the top two papers on Hugging Face this morning both being agent harnesses and a Microsoft benchmark landing the same day to say no harness is reliable twice in a row.
-
August 24, 2026 · afternoon
Every layer of the agent stack now ships a vendor-neutral version, from the local inference engine to the orchestrator to the ruleset, while precision measurement shows the substrate underneath those layers is not interchangeable at all.
-
August 24, 2026 · morning
Four separate shipments this weekend attack the same broken assumption, that a human sits in a browser to approve what software does, and the replacement being built is per-task consent plus a distinct identity for the agent.
-
August 23, 2026 · afternoon
Running many agents at once stopped being a technique this weekend and became infrastructure, and almost everything shipped around it is about supervision and cost rather than capability.
-
August 23, 2026 · morning
Across protocol, infrastructure, tooling and research this weekend, the same move keeps repeating, replacing a stated claim with a mechanically checkable one.
-
August 22, 2026 · afternoon
Model weights sat still this week while nearly every notable release moved capability into the scaffolding around the model, and the scaffolding is now learning to rewrite itself.
-
August 22, 2026 · morning
The expensive part of running an agent is not the model, it is the context the agent keeps re-deriving, and three of today's top projects attack that waste from three different layers.
02 / Under the surface
Latest analysis
-
OpenWiki 0.4.0 Proves Its Claims Against Your Code. The Claims It Can't Pin Look Exactly the Same
OpenWiki 0.4.0's grounded claims deterministically re-verify every fact it could pin to a repository file, and the facts it could never pin…
-
Headlong Gives Your Team One Agent With One Memory, and No Wall Between You
Headlong's single thought stream is exactly what makes a shared agent feel like a colleague instead of a service, and it is also why every…
-
Codex Deprecated Its MCP Server, Not MCP. The Direction of That Cut Is the Story
Codex stopped serving MCP while expanding its MCP client support in the same release, and that one-directional cut marks the real boundary…
-
One Success Isn't Reliability: The Agent Number Almost Nobody Reports
Running an agent workflow once and watching it succeed measures almost nothing, because success collapses under repetition and the failures…
-
x64dbg-MCP Server Gives an Agent 71 Debugger Tools and Ships Listening on 0.0.0.0
X64dbg-MCP Server proves agentic reverse engineering works today, and its hand-rolled static bearer token sent in cleartext to a default…
-
FreeToken Runs a 753B Model on One Workstation GPU. The Real Trick Is That Your VRAM Split Moves at Runtime.
FreeToken's headline parameter counts matter less than its elastic runtime reallocation of VRAM between expert cache and KV memory, which…
-
Cloudflare's Optional OAuth Scopes Make Partial Grants Normal, and Most Agents Will Break On Them
Cloudflare's optional OAuth scopes turn partial grants into a routine outcome, so every agent and MCP server that assumes it received the…
-
Agent Skills Compose Right Up Until Two of Them Disagree. Then Nothing Decides Who Wins.
Agent skills are sold as composable but the format defines no precedence and no scope, so when two installed skills govern the same…
04 / Coverage map
Topics we track
Claude Code 33 Codex 15 OpenAI 11 DeepSeek Harness 10 Agent Skills 9 Kimi K3 9 Model Context Protocol 9 MCP 8 GPT-5.6 Sol 7 Hugging Face 7 MCP 2026-07-28 7 Anthropic 6 Anthropic Frontier Red Team 5 Claude Code auto mode 5 Claude Opus 5 5 GLM-5.3 5 GPT-5.6-Cyber 5 OpenAI Presence 5 Claude Code self-hosted environments 4 FreeToken 4 GLM 5.2 4 GPT-5.6 Luna 4 grok-build 4 LangSmith Preview Builds 4