01 / The wire
Recent briefings
-
August 19, 2026 · afternoon
Every significant capability gain published in the last 48 hours came from changing the harness around the model instead of the model itself, and none of it shipped with a security evaluation.
-
August 19, 2026 · morning
Three labs spent this week engineering containment against their own models, and the thing being contained is offensive security capability that arrived faster than any of them planned for.
-
August 18, 2026 · afternoon
Five gates went up around the AI stack in forty-eight hours, and the GitHub daily board is quietly voting for everything you can pick up and carry out.
-
August 18, 2026 · morning
AI now reviews code and attacks it, and only the attacking side gets to iterate against live feedback.
-
August 17, 2026 · afternoon
Three separate moves in 48 hours all changed the layer between your app and the model, and not one of them was a model.
-
August 17, 2026 · morning
Four separate things that were free or open picked up a gate in 72 hours, and the counter-tooling is already climbing the trending charts.
-
August 16, 2026 · afternoon
The competition moved off the model and onto the harness, and the plugin ecosystem that formed around DeepSeek Harness in 72 hours is what a platform land grab looks like before anyone calls it one.
-
August 16, 2026 · morning
Offensive security capability became the thing labs gate releases on this week, and the same week's speed and locality launches make that gate almost impossible to hold.
02 / Under the surface
Latest analysis
-
StateM Reports 95.3% on Terminal-Bench 2.1 With Frozen Weights. The Word Doing the Work Is 'Raw'
StateM's reproducible claim is the roughly $15 price rather than the 95.3% score, because Terminal-Bench's published leaderboard subtracts…
-
Google Bought 100 Million Spirit Airlines Emails Out of Bankruptcy Court
Bankruptcy court has become a training-data supply line, and the privacy machinery in the code protects the customers a dead company had,…
-
Microsoft Foundry Moved Agent Tool Permissions Into a Request Parameter, and the Denylist Fails Open
Foundry moved agent tool governance into per-request parameters, and Microsoft's own operational checklist says the denylist form of that…
-
career-ops Is an AI Job Search Tool Whose Best Answer Is Don't Apply
Career-ops's real product is a refusal threshold, and its real risk is that the same agent enforcing the threshold will rewrite the rubric…
-
watermarks-remover Is Trending, and Its Own README Argues Against Half of It
Watermarks-remover is the clearest published account of why text watermarking fails as a trust primitive, because its README documents that…
-
DSH Desktop Checks That Your Update Is a Real Installer, Not Who Built It
DSH Desktop's own known-limitations section says its auto-updater validates the download container rather than publisher identity, which is…
-
OpenAI's Computer History Turns Your Mac Into Agent Memory, and Writes It to Plain Text
Computer History is the best-documented agent memory feature anyone has shipped, and its documentation tells you the derived memory files…
-
Codex Multi-Agent V2 Rejects Your Cheapest Subagent, and Your Config File Can't Override It
Codex resolves which models you may delegate to from a static server-side model catalog rather than from your config, so a documented…
04 / Coverage map
Topics we track
Claude Code 31 Codex 14 OpenAI 11 Kimi K3 9 DeepSeek Harness 8 Model Context Protocol 8 Agent Skills 7 Hugging Face 7 MCP 2026-07-28 7 Anthropic 6 MCP 6 Anthropic Frontier Red Team 5 Claude Code auto mode 5 Claude Opus 5 5 GLM-5.3 5 GPT-5.6 Sol 5 GPT-5.6-Cyber 5 OpenAI Presence 5 Claude Code self-hosted environments 4 GPT-5.6 Luna 4 grok-build 4 MAI-Cyber-1-Flash 4 Muse Glimmer 4 QM 4