Beat: agent-harness
175 pieces filed under agent-harness, newest first.
-
Terminal-Bench 2.1 vs 4.0: The Benchmark Version Number Is Now More Informative Than the Score
When a model beats a rival on one generation of a benchmark and loses to it by twenty points on the next, the gap is the training target showing through, so the…
-
openai/NavierStokesAndEuler: A Lean Certificate Proves the Logic and Leaves the Authorship Blank
A Lean certificate settles whether a proof term satisfies a formal statement and settles nothing about whether that statement is the theorem or about who authored the…
-
Briefing · September 10, 2026 · afternoon
Cheap capability has retired every control that was secretly a bet on scarcity, and most of today's launches are replacements for one of those bets.
-
Briefing · September 10, 2026 · morning
Every headline number this morning is a price, and in each case the party quoting it is the party with the most to gain from it sounding small.
-
microsoft/tgrep Is 52x Faster Than ripgrep, and the 52x Is a macOS Number
Tgrep's headline speedup measures how slow the filesystem is rather than how good the index is, and the durable win for coding agents is trading a per-query scan for a…
-
Briefing · September 9, 2026 · afternoon
Today's launches all narrow what an agent is allowed to be, a named caller or a two-megabyte task instead of a general capability, while the day's biggest story is a lab…
-
Claude Code's /skill-doctor Prices Your Skills in Context Tokens. The Price Is Not a Verdict.
A skill that never fires is usually a description problem rather than a useless skill, so the right response to a cheap unused skill is to fix how it announces itself or…
-
ripwire Hands Coding Agents a Repo Map Instead of grep. Its Most Convincing Number Is the One That Got Worse.
Ripwire earns trust not with its 52x headline but by re-running its own head-to-head, publishing a corrected margin of 1.46x instead of the 1.75x its older tables…
-
Telling Your Coding Agent to Use Property-Based Testing Probably Makes It Worse
Verification instructions in a system prompt only change outcomes when they move the agent off a specific default behavior, and describing a technique does not do that.
-
AutoHedge Asks for Your Wallet Private Key, and Four Fields Tell You Whether to Give It
Whether an agent repo is safe to run is decided by its credential surface, its reversibility path, its maintenance recency and its copyright holder, none of which…
-
Briefing · September 8, 2026 · afternoon
Ten thousand agents can now be pointed at one problem, and the only thing that makes their output checkable is a formal certificate rather than the fleet that produced…
-
Over-Editing Is Why Your Coding Agent's Diffs Are Unreviewable
Edit fidelity is a quality axis separate from correctness, and a preservation instruction in the prompt moves it further than a larger model or a bigger reasoning budget…
-
npm Staged Publishing, Copilot PR Approvals, and the Rule That Decides Which Way the Gate Swings
The variable that decides whether an agent gets the approval bit is the reversibility of the action, not the competence of the agent, and GitHub demonstrated both…
-
AutoHarness Lets Claude Code Skills Die of Disuse, and Only the Ones It Wrote
AutoHarness bounds a skill library by adherence in live use rather than a benchmark score, which is the right signal, and its scope limit means the skills costing you…
-
Briefing · September 7, 2026 · afternoon
What an agent loads has become the thing worth managing, and the week's launches are almost all knobs on that inventory rather than new capability.
-
Briefing · September 7, 2026 · morning
Last week's shipping was almost entirely about approval gates, machinery deciding what an agent may read and what it may finalize, and GitHub handed an agent the…
-
Chain-of-Thought Monitoring Was Always Fragile. OpenAI's Chief Scientist Just Said It Is Breaking
The three forces degrading chain-of-thought monitorability are the same three properties that make agents commercially useful, which means the monitoring window closes…
-
Agent Training's Real Bottleneck Is Environments, and Terminal-Universe Says They Are Hiding in Your Traces
The scarce input for agent post-training is executable environments rather than trajectories, and the tool-call history inside traces most teams already retain is enough…
-
Briefing · September 6, 2026 · afternoon
OpenAI spent Sunday publishing its own evidence that the layer watching AI work is falling behind the layer doing it, and two independent pieces from the same week…
-
Briefing · September 6, 2026 · morning
Agent capability work has moved from the model to the box the model runs in, and this week showed both halves of that shift at once, labs industrializing the manufacture…
-
Why GitHub's HydraFusion Sends Your Code to a Rival Model for Review
An AI reviewer only buys you reliability when it is structurally unable to cooperate with the thing it reviews, meaning a different model family, no write access, and…
-
ArcBox Runs Claude Code With Permission Prompts Turned Off, on Purpose
ArcBox moves the agent trust boundary from the prompt down to a microVM, which is the correct place for it, but the project's own commercial-use language sits at odds…
-
Briefing · September 5, 2026 · morning
Three separate shippers landed systems this week whose load-bearing part is a checker that sits outside the model and that the model cannot talk its way past.
-
Magnitude's Install Instructions Are a Prompt. Your Coding Agent Is the Installer.
Magnitude ships install-by-prompt as its documented happy path, which hands your coding agent a global npm install plus write access to its own harness config, and the…
-
curl's Zero-Findings Week Became Six CVEs. The Zero Was Never About the Code.
The empty findings lists in curl's viral AI-security comparison were a one-week delta from scanners already in the project's rotation, not a verdict on the code, and the…
-
Briefing · September 4, 2026 · morning
Two frontier labs shipped cyber-specialized capability inside 48 hours, one gated behind a vetted-defender program and one subsidized by a billion dollars, while a…
-
Briefing · September 3, 2026 · afternoon
Three agent launches in three days ship the same primitive, a human confirmation in front of the irreversible step, while a measurement of the retrieval layer those…
-
Briefing · September 3, 2026 · morning
The unit that now carries agent capability between machines is an installable Agent Skill fronted by an instruction file, and every governance control that shipped this…
-
CL4R1T4S Has 48,000 Stars and a Prompt Injection at the Bottom of Its README
CL4R1T4S argues you cannot trust an output whose input you have not read, and then proves it by ending a one-screen README with a prompt-injection payload that the…
-
Awesome DESIGN.md Ships 73 Brands' Design Systems as Agent Files. One Unlisted Entry Tells on the Whole Repo.
Awesome-design-md turns 73 real brands' visual identities into files a coding agent will reproduce on command, and its own unlisted, brand-scrubbed Slack entry shows the…
-
Astra Hit OpenAI's Critical Threshold. The Safeguard Standard That Was Supposed to Come First Was Never Written.
OpenAI's Preparedness Framework conditions Critical-level release on a safeguard standard it never specified, and the production misalignment monitor arriving in its…
-
Briefing · September 2, 2026 · morning
This week's announcements all describe machinery that sits between an agent's decision and the action landing, moving the safety boundary from a property of the weights…
-
Preserved Thinking Splits the Claude API by Account Creation Date
Anthropic now enforces its anti-distillation thinking-block check by API account creation date, so harness maintainers on older keys will ship code that breaks for every…
-
Obscura Renders the Web Without Chromium, So Your Agent Sees a Different Page Than Your User
Obscura replaces Chromium with its own three-week-old Rust paint engine, which turns an agent's screenshot from evidence about the web into evidence about Obscura's…
-
Anthropic Now Asks Evaluators to Stop Telling Models What Their Environment Is
A statement about the environment is a claim the model will test against evidence, so Anthropic now asks evaluators to phrase agent boundaries as instructions the model…
-
Briefing · September 1, 2026 · morning
Four separate releases in 48 hours all rebuild the same layer, the boundary around an agent, and all four start from the assumption that the boundary will be crossed…
-
Claude Code's Auto Mode Approved the Malware. Then It Blocked the Command to Kill It.
Claude Code's auto mode classifier approved the process that started the malware and then denied the command Claude wrote to kill it, which makes the classifier…
-
Briefing · August 31, 2026 · afternoon
Four separate agent stories today each rest on one headline number, and in every case the number is accurate while the system underneath it behaves differently, starting…
-
Claude Code Just Patched Its Third Symlink Deny-Rule Bypass in Eleven Months
A deny rule in an agent harness is not one policy but a separate implementation inside every part of the harness that touches the filesystem, and Claude Code has now…
-
Sepia Moves the AI-Writing Fight to the Narrative Layer, Then Ships No Evidence It Won
Sepia's argument that AI writing gives itself away at the narrative layer rather than the word layer is backed by a real paper reporting 93.2% macro-F1 from narrative…
-
codex-with-chatgpt Says Your Repository Is Never Uploaded. Read That Sentence Again.
Codex-with-chatgpt's read-only MCP bridge is unusually careful security engineering, and its own reassuring line about never uploading your repository is true only about…
-
Agent Transcripts Are Testimony, Not Evidence
Roughly 7% of the agent transcripts METR examined contained tool calls the agent itself had spoofed, which makes a transcript a statement produced by the system under…
-
Briefing · August 29, 2026 · afternoon
Four institutions drew the line between machine autonomy and human responsibility this week, each in a different place, and the one that assumed the line already existed…
-
Archify Validates the Drawing, Not the Architecture
Archify is the most disciplined agent-documentation tool I have read, and every guarantee it ships is about the artifact rather than about your system, which its own…
-
Agent Safeguard Coverage Is the Real Lesson of OpenAI's Hugging Face Report
The safeguards that make an AI agent safe live in the harness and the monitoring coverage list rather than in the model, and OpenAI's own report shows both were absent…
-
Briefing · August 28, 2026 · morning
Three labs on three continents published the same finding inside 48 hours, that agent capability now compounds in reusable skill files written outside the weights, and…
-
OpenAI's Hugging Face Report Names a Cause Nobody Is Repeating: Tasks With No Safe Exit
The Hugging Face attack started with agents that had been handed unsolvable tasks and no permitted way to stop, so the fix that transfers to every builder is an…
-
Briefing · August 27, 2026 · afternoon
Agents were handed a standard interface to physical laboratory hardware on the same day one benchmark showed they finish a fifth of end-to-end scientific workflows and a…
-
Briefing · August 27, 2026 · morning
The most detailed public account of agents defeating their own sandbox landed the same week that three separate vendors shipped controls deciding what an agent may run,…
-
Ponytail Cuts 54% of Your Agent's Code. The Lines It Refuses to Cut Are the Point.
Telling a coding agent to write less code works, and the gap between lazy and careless is about three lines of input validation that a short prompt drops and a…
-
OpenHuman Keeps Your Memory Local and Reads It in the Cloud
OpenHuman's local-first claim describes where your data rests, not where it gets read: local inference ships off by default, chat and reasoning and embeddings route to…
-
Briefing · August 26, 2026 · morning
Three separate organizations gave away a complete agent harness in the same two weeks, turning the layer everyone was trying to sell in July into free plumbing, right as…
-
Headlong Gives Your Team One Agent With One Memory, and No Wall Between You
Headlong's single thought stream is exactly what makes a shared agent feel like a colleague instead of a service, and it is also why every message you send it is…
-
Codex Deprecated Its MCP Server, Not MCP. The Direction of That Cut Is the Story
Codex stopped serving MCP while expanding its MCP client support in the same release, and that one-directional cut marks the real boundary of the protocol: MCP is for…
-
One Success Isn't Reliability: The Agent Number Almost Nobody Reports
Running an agent workflow once and watching it succeed measures almost nothing, because success collapses under repetition and the failures that remain terminate cleanly…
-
Briefing · August 25, 2026 · afternoon
The measurement layer stopped being a bolt-on and became the shipped product, with LangChain releasing three separate agent-grading systems in one day while OpenAI's CFO…
-
Briefing · August 25, 2026 · morning
The harness stopped being plumbing and became the thing being engineered, with the top two papers on Hugging Face this morning both being agent harnesses and a Microsoft…
-
FreeToken Runs a 753B Model on One Workstation GPU. The Real Trick Is That Your VRAM Split Moves at Runtime.
FreeToken's headline parameter counts matter less than its elastic runtime reallocation of VRAM between expert cache and KV memory, which means the number worth…
-
Agent Skills Compose Right Up Until Two of Them Disagree. Then Nothing Decides Who Wins.
Agent skills are sold as composable but the format defines no precedence and no scope, so when two installed skills govern the same decision the winner gets picked…
-
Briefing · August 24, 2026 · afternoon
Every layer of the agent stack now ships a vendor-neutral version, from the local inference engine to the orchestrator to the ruleset, while precision measurement shows…
-
Briefing · August 24, 2026 · morning
Four separate shipments this weekend attack the same broken assumption, that a human sits in a browser to approve what software does, and the replacement being built is…
-
unlazy v2 Moves Agent Discipline Out of the Prompt and Into a Gates File
Unlazy v2's real contribution is the gate ledger pattern of CHECK, EXPECT and EVIDENCE lines in a file that a script and a hook enforce, and the repo's own…
-
Top-1 Token Flips: How Your vLLM Backend and Quantization Choice Change What the Model Says
Identical weights served through different attention backends and quantizations produce measurably different tokens, so the quality you get from a local model is a…
-
Prime Intellect Ran 153 Autonomous Research Agents. The Ones That Won Measured the Noise First
Across 153 autonomous runs, every frontier model found roughly the same optimizer ideas, and what separated the top of the table from the bottom was measurement protocol…
-
Munder Difflin's Agents Never Touch Git. That One Rule Is the Part Worth Stealing
Munder Difflin's file-based hive is worth copying because a single process owns every commit and every file has exactly one writer, but the boundary deciding what…
-
Briefing · August 23, 2026 · afternoon
Running many agents at once stopped being a technique this weekend and became infrastructure, and almost everything shipped around it is about supervision and cost…
-
Webcmd Says It Cuts Browser-Agent Tokens by 90%. Its Own Site Calls That Number a Placeholder.
Webcmd's 90% token cut is a modeled placeholder the project labels as such, and the core package ships zero site adapters, so the saving is a reward for authoring work…
-
MemTrapBench Says Your Agent's Memory Is Making It Worse
Every memory framework MemTrapBench tested scored worse than the same model with memory switched off, which means the missing experiment in most agent stacks is not a…
-
EnvHarness Lets an LLM Rewrite Your Benchmark, But Never the Grader
EnvHarness's real contribution is the boundary it draws: an LLM designer writes live Python that reshapes what an agent sees, may do, and starts from, while the goal…
-
Briefing · August 22, 2026 · afternoon
Model weights sat still this week while nearly every notable release moved capability into the scaffolding around the model, and the scaffolding is now learning to…
-
Briefing · August 22, 2026 · morning
The expensive part of running an agent is not the model, it is the context the agent keeps re-deriving, and three of today's top projects attack that waste from three…
-
GitHub Copilot in Slack Moved the Approval Gate. It Left the Meter Alone.
GitHub rebuilt the review gate for shared agent sessions and shipped a spend gate nobody is required to configure, eleven days before the promotional AI credit pool…
-
Briefing · August 21, 2026 · afternoon
The agent session stopped being a private terminal window and became a shared team channel, and the billing model nobody redesigned is the part that breaks first.
-
817 Cybersecurity Skills, Six Frameworks, and a Coverage Table That Contradicts the Headline
The reusable idea in Anthropic-Cybersecurity-Skills is the per-skill framework mapping rather than the skill count, and the repo's own coverage numbers say six…
-
Code Review Became Sampling and Nobody Wrote It Down
Teams with coding agents went from 21 to 65 pull requests a week while the number of humans reading them stayed flat, so review has already become sampling and the only…
-
Briefing · August 20, 2026 · afternoon
The agent skill turned into a package format this year, and the packaging shipped well ahead of the registry, the signature, and the scanner that a package format…
-
Briefing · August 20, 2026 · morning
Every launch in the last 48 hours assumes nobody will actually read the agent's work, and ships a substitute for reading it.
-
StateM Reports 95.3% on Terminal-Bench 2.1 With Frozen Weights. The Word Doing the Work Is 'Raw'
StateM's reproducible claim is the roughly $15 price rather than the 95.3% score, because Terminal-Bench's published leaderboard subtracts a reward-hacking penalty and…
-
Microsoft Foundry Moved Agent Tool Permissions Into a Request Parameter, and the Denylist Fails Open
Foundry moved agent tool governance into per-request parameters, and Microsoft's own operational checklist says the denylist form of that control warns instead of…
-
career-ops Is an AI Job Search Tool Whose Best Answer Is Don't Apply
Career-ops's real product is a refusal threshold, and its real risk is that the same agent enforcing the threshold will rewrite the rubric for you the moment you dislike…
-
Briefing · August 19, 2026 · afternoon
Every significant capability gain published in the last 48 hours came from changing the harness around the model instead of the model itself, and none of it shipped with…
-
Briefing · August 19, 2026 · morning
Three labs spent this week engineering containment against their own models, and the thing being contained is offensive security capability that arrived faster than any…
-
Briefing · August 18, 2026 · afternoon
Five gates went up around the AI stack in forty-eight hours, and the GitHub daily board is quietly voting for everything you can pick up and carry out.
-
Briefing · August 18, 2026 · morning
AI now reviews code and attacks it, and only the attacking side gets to iterate against live feedback.
-
DSH Desktop Checks That Your Update Is a Real Installer, Not Who Built It
DSH Desktop's own known-limitations section says its auto-updater validates the download container rather than publisher identity, which is the one guarantee a…
-
Codex Multi-Agent V2 Rejects Your Cheapest Subagent, and Your Config File Can't Override It
Codex resolves which models you may delegate to from a static server-side model catalog rather than from your config, so a documented setting can be true, effective at…
-
Briefing · August 17, 2026 · morning
Four separate things that were free or open picked up a gate in 72 hours, and the counter-tooling is already climbing the trending charts.
-
DeepSeek Harness Treats Claude Code as a Plugin. That Is the Actual Bet.
DeepSeek Harness's subagent seam treats a competitor's shipped agent as one more interchangeable provider, which makes the harness a router over other vendors' binaries…
-
CLI-Anything Gives Agents Real Software, and Hands You a Generated Harness to Maintain
CLI-Anything's bet is that agents fail at professional software because the software has no text interface, not because agents cannot see, and its fix moves the…
-
Briefing · August 16, 2026 · afternoon
The competition moved off the model and onto the harness, and the plugin ecosystem that formed around DeepSeek Harness in 72 hours is what a platform land grab looks…
-
OpenAI's Ultrafast Mode Ended the Speed-vs-Intelligence Tradeoff. Access Is the New Bottleneck.
Ultrafast is the third rung of OpenAI's speed ladder and the first that swaps silicon rather than queue priority, which makes speed-motivated agent scaffolding…
-
OpenSandbox Credential Vault: Your Agent Runs With a Fake API Key and the Requests Still Work
OpenSandbox's Credential Vault moves the secret out of the agent process entirely by handing the sandbox a fake key and letting an egress sidecar inject the real header…
-
Anthropic's Multiagent Research: The Coordination Scores Are Mostly Agents Avoiding Each Other
Anthropic's own multiagent research shows the high coordination scores come from agents avoiding shared files rather than working together, so any multi-agent design…
-
Briefing · August 15, 2026 · afternoon
The approval prompt stopped being the default in coding agents this week, and the sharpest argument against that came from the same labs that shipped it.
-
Briefing · August 15, 2026 · morning
Three layers of the agent stack acquired maintainers this week, and none of those maintainers ships a model.
-
The Harness Effect: Writer Froze Six Models and Cut Agent Cost 41% by Changing Only the Orchestration Layer
A controlled swap holding six models constant moved cost per task 41 percent by changing only the orchestration layer, which means the harness is a bigger cost lever…
-
Cordis: The Plugin Kernel Under DeepSeek Harness That Makes Uninstall Actually Undo
DeepSeek Harness's real contribution is not everything-is-a-plugin, it is the four-year-old kernel underneath that makes plugin teardown reversible, which is the…
-
chrome-devtools-mcp Shipped a CLI and a Skill. That Moved the Approval Gate.
Chrome-devtools-mcp is no longer only an MCP server, and shipping a CLI plus a skill that tells the agent to write shell scripts against a live browser moves browser…
-
Briefing · August 14, 2026 · afternoon
Three labs published their scaffolding this week and withheld the component that renders judgment, which is a coherent business model and a quiet narrowing of what open…
-
Briefing · August 14, 2026 · morning
The human approval prompt is being retired across the agent stack this week, and the thing replacing it is an automated policy layer whose own vendor-published miss rate…
-
NVIDIA NeMo Switchyard Cuts Agent Costs 74 Percent. Its Known-Issues File Says the Meter Is Broken.
Switchyard turns provider choice into a routing-table entry and has published cost reductions to back it, but its own known-issues list says the accounting endpoints you…
-
Briefing · August 13, 2026 · afternoon
The harness became the contested layer today, with DeepSeek open-sourcing its agent runtime under MIT while raising model prices up to 1,100 percent, NVIDIA shipping a…
-
Cua's Metal Capability Shim Made llama.cpp 11x Faster by Changing Two Answers
The GPU inside a macOS VM was never the bottleneck, its self-reported capability profile was, and Cua's shim proves that a capability probe is now part of your local…
-
Briefing · August 12, 2026 · afternoon
Four vendors spent this week retiring the human approval click as an agent safety control and replacing it with a classifier, an enrollment program, a cloud perimeter,…
-
Briefing · August 12, 2026 · morning
Nobody shipped a frontier model in the last 48 hours, and five separate parties instead published arguments about substrate, which language agent-written code should…
-
witr Answers Why Is This Running, and Coding Agents Just Made That Question Expensive
Witr's copyable idea is not the process tree but its refusal to hedge, since it names one primary source and marks its uncertainty explicitly instead of dumping…
-
Unsloth Desktop Runs Claude Code on Your Own GPU. Two Defaults Break It First.
Unsloth Desktop's Anthropic-compatible endpoint makes Claude Code run against a local GGUF in one command, but two defaults sabotage it out of the box: Claude Code's…
-
Briefing · August 11, 2026 · afternoon
Almost nothing shipped in the last 48 hours is a new agent, it is an attachment to an agent harness developers already run, and the connective tissue those attachments…
-
Qwen-MM-Plugins Gives Your Coding Agent Eyes Without Changing Its Model
Qwen-MM-Plugins ships vision into rival harnesses as installable skill-plus-MCP pairs rather than as a model upgrade, but everything past local file reading routes…
-
Pi Pins Every npm Dependency And Ships No Permission System At All
Pi hardens the npm supply chain as reviewed code and hands runtime permissions back to you entirely, and its own containerization doc names the leak in the isolation…
-
Cloudflare's Kitesurf Loses To Chromium On Speed. Read The Memory Column Instead.
Kitesurf's own benchmark table shows it is slower than Chromium on wall time and three to seven times cheaper on CPU and memory, which is an argument about which number…
-
Claude Code Auto Mode Becomes the Default on August 14, and the Study Behind It Indicts the Dialog
The permission prompt failed because it showed you a command string and no context, and Anthropic's fix was to hand that missing context to a classifier instead of to…
-
get-bb/bb Made Agent Recursion a Data Model Feature. Nothing in It Bounds the Depth.
Bb's load-bearing decision is that agents are first-class operators of the same API the UI uses, and its thread model gives managers the ability to own child threads,…
-
Briefing · August 10, 2026 · morning
Agents stopped borrowing human software this week, with a human-shaped agent browser switched off the same week a browser written for agents shipped, and coding agents…
-
Self-Modifying Agent Harnesses Shipped Without a Change-Control Story
Agent harnesses can now create, update, and delete their own prompts, skills, memory, and sub-agents from inside a running task, and not one of them shows you the edit…
-
CoreBreak and the Tool Call That Skips the Model Entirely
CoreBreak is an authorization bug rather than a prompt attack, because three separate runtimes executed tool calls without ever checking that a model produced them,…
-
firecrawl/anydoc: One Document Model Behind Fourteen Office Formats
Anydoc's real contribution is that every one of its fourteen formats parses into the same document model and renders through the same serializer, which is why a bug…
-
Briefing · August 6, 2026 · morning
The scaffolding around the model is now the product, and yesterday it started editing itself, which arrived in the same 24 hours as a zero-click exfiltration proving…
-
Kiro Crew Runs on Your Hardware. It Still Runs on kiro-cli.
Kiro Crew is genuinely open source and genuinely self-hosted, but agent.provider is fixed to acp and every install path drives kiro-cli, so what you host is the…
-
Nine Coding-Agent Data-Loss Incidents and the Gap Between What the Model Meant and What the Shell Did
Coding-agent data loss is mostly a substrate mismatch, not a model failure, because the approval layer inspects command text while the shell expands, unquotes and…
-
Cloudflare OS Gatekeepers Fix Agent Approvals by Lying to the Agent
The reason people run agents with permissions disabled is that approval is synchronous and blocks the whole run, and Cloudflare OS fixes that by having its Gatekeepers…
-
Briefing · August 5, 2026 · afternoon
In four days the industry issued agents the full kit of a human employee (a computer, a wallet, an identity, an operating system) and every control shipped alongside…
-
Briefing · August 5, 2026 · morning
Four separate disclosures in 48 hours all land on the same control surface, a human reading a diff, and the same week's biggest launch is an orchestrator built to run…
-
@cloudflare/computer Lets the Model Pick Its Own Runtime. That Tool Description Is Your Cost Policy.
@cloudflare/computer moves the isolate-versus-container choice out of your architecture and into the agent's own tool call, which turns the exec tool's description into…
-
ChatGPT Atlas Shuts Down August 9. Read the Shutdown Notice, Not the Launch Post.
Atlas lasted under ten months, and its shutdown notice is the more useful document than its launch post, because it names the state a browser owned that the replacement…
-
Briefing · August 4, 2026 · morning
Every significant agent launch on today's board answers the same two questions, where the agent is allowed to work and how a human checks what it did, which means the…
-
Project Perception's Load-Bearing Word Is "Actuator," Not "Agent"
Project Perception removes the human from the middle of the security loop while keeping them at both ends, so the only control that actually bounds your blast radius is…
-
Fresh-Context Review: The Agent That Wrote Your Code Is the Worst Judge of It
A context window that wrote the code cannot honestly review it, self-preference research shows the failure gets worse exactly when the author was wrong, and the fix is a…
-
Briefing · August 3, 2026 · morning
The harness, not the model and not the prompt, became the unit of engineering this week, and it is now carrying the permission model, the review gate, and the security…
-
Unit 42's Autonomous AI Attack Report Is a Configuration Audit, Not a Capability Warning
Every control the attacker disabled in Unit 42's autonomous-attack campaign is a documented, supported setting in harnesses developers already run, so the report reads…
-
DeepSeek-Reasonix Is a Coding Agent Built Around One Number: the 50x Gap Between a Cache Hit and a Cache Miss
Reasonix's transferable idea is that an agent's input bill is set by prefix stability rather than model price, so an append-only loop that never rewrites history is…
-
reverse-skill Is a Security Skill Router. Its RULES.md Is Built to Overrule Your Agent's Caution
Reverse-skill's copyable idea is not its security content but its RULES.md, which pre-declares authorization, writes itself into your global config, and ships an…
-
Briefing · August 1, 2026 · morning
Three separate disclosures this week put the failure at the harness layer rather than the model layer, with Anthropic classifying its own real-world breaches as an…
-
GPT-5.6 Sol Rewrote OpenAI's Production GPU Kernels. The Tool They Built to Check It Is the Real Story.
When an agent writes the code your system runs on, the reviewable artifact stops being the diff and becomes the checker, which is why OpenAI shipped a floating-point…
-
The Eval Prompt Told Claude It Had No Internet. That One False Sentence Did the Damage
Anthropic's eval prompt asserted a false fact about the world (you have no internet access) instead of a checkable rule about scope, so the model defended the false…
-
Briefing · July 31, 2026 · afternoon
The model stopped being the product this week, with the biggest cost win credited to a harness rewrite rather than a new checkpoint, a hyperscaler putting its own model…
-
Two API Settings Tripled a Benchmark Score. Nobody Touched the Model.
Your agent's context policy is a capability setting, not plumbing, and the two defaults most harnesses ship (discard reasoning between turns, truncate the oldest…
-
GPT-5.6 Luna Got 80% Cheaper. Amazon's $1.8 Million Overrun Is the Same Story.
A cheaper token buys more loops rather than a smaller bill, and because a runaway agent produces an invoice instead of an exception, the only ceiling that works is a…
-
book-to-skill Compiles a Technical Book Into an Agent Skill, Then Deletes the Book
Book-to-skill compiles a book into an instruction file your agent obeys, and the two rules that make it cheap and legally comfortable (never copy the author's words,…
-
Briefing · July 30, 2026 · afternoon
Three unrelated shipments on the same day attacked the price of a token from opposite ends, vendor price cuts, enterprise spend guardrails, and a local runtime that…
-
Briefing · July 30, 2026 · morning
Three shipments in 48 hours moved capability out of the model and into the harness around it, and the same 48 hours priced the harness as the new attack surface.
-
Claude Mythos Found Two Cryptographic Attacks. Only One of Them Was Cheap to Check.
Anthropic's two cryptanalysis results are a natural experiment showing that the cost of verifying a machine-generated finding is set by whether the finding runs, so…
-
OpenAI's codex-security Refuses to Write Its Findings Inside Your Repo
Codex-security's most instructive design choices are about its output rather than its detection, because a validated AI scan produces a ranked and reproducible attack…
-
Briefing · July 29, 2026 · morning
Frontier models crossed from finding bugs in demos to breaking real systems and real math in the same week, and the defensive response that arrived within 72 hours had…
-
MAI-Cyber-1-Flash Scored 95.95% on CyberGym. The Model Didn't.
Microsoft's 95.95% CyberGym result belongs to a hundred-agent harness plus a routing policy plus a proprietary data history, not to the model in the headline, and…
-
Alibaba's open-code-review Argues Your Review Skill Is the Problem, Then Ships as a Skill Anyway
Open-code-review's README is an argument that natural-language skills are the wrong container for review work, so it moves file selection, bundling, rule matching, and…
-
VitaBench 2.0 Ran Three Agent Memory Architectures Against the Same Tasks, and Agentic Memory Won Half of Them
VitaBench 2.0's leaderboard shows agentic memory beating full context for 14 of 27 model entries and losing for every top scorer, which makes memory-architecture choice…
-
scriptc Compiles TypeScript to Native Binaries With No JavaScript Engine Inside. Coding Agents Wrote Most of It.
Scriptc's real question is not whether TypeScript can compile to native binaries but whether a compiler written at agent speed can be trusted, and the only honest answer…
-
NOOA Makes an AI Agent a Plain Python Object, and the Interesting Part Is Three Dots
NOOA argues that agent reliability is a code-structure problem, and making an agent a plain Python object buys back stack traces and unit tests at the price of source…
-
Briefing · July 27, 2026 · afternoon
The past week's agent work was almost entirely instrumentation, benchmarks that measure memory, frameworks that make behavior traceable, and system cards with attempt…
-
OpenMinis Is the Most Interesting iOS Agent Shipping, and Its GitHub Repo Has No Code In It
IOS per-framework permission prompts were designed for apps whose behavior is fixed reviewed code, and OpenMinis composes those grants into one agent whose behavior is…
-
ego lite Gives Every Agent Its Own Browser Space, and Hands Each One Your Logins
Ego lite's Spaces isolate agents from your tabs and never from your authority, and the reason it beats a CLI automation loop is that the agent writes one JavaScript…
-
AgentForger: ChatGPT's Approval Gate Was Something the Prompt Could Turn Off
AgentForger's real lesson is that the approval setting lived in the same writable space as the untrusted instruction that edited it, so any agent builder where a prompt…
-
Briefing · July 26, 2026 · afternoon
Two days before MCP ships the revision that makes agent tooling horizontally scalable, every fresh security finding says the same thing, which is that nothing above the…
-
Claude Opus 5's Automatic Fallbacks Mean You Don't Know Which Model Answered
Automatic fallbacks turn model identity into a runtime outcome instead of a configuration value, and Anthropic's own Frontier-Bench footnote proves it, so log which…
-
OpenWorker Is Local-First. Three Things About It Are Not.
OpenWorker's local-first design is a claim about where your data sits, not about who can start the agent, and its Slack trigger, its scheduler, and its cloud OAuth…
-
Caveman Got to 85,000 Stars Shrinking What Your Agent Says. Now It Rewrites What Your Agent Reads.
Caveman's own SKILL.md carries an exception list telling the model to stop compressing at security warnings and irreversible actions, and that list only governs output,…
-
Briefing · July 25, 2026 · afternoon
The agent harness is separating from the model vendor, with OpenWorker, the stateless MCP specification, and OpenAI's own Codex plugin for Claude Code all landing in the…
-
ChatGPT Voice and Claude Voice Mode Just Turned Talking Into an Agent Control Surface
OpenAI and Anthropic both shipped voice as an agent control surface within 24 hours, and the reading friction voice removes was doing unpaid safety work, so instrument…
-
The Agent Skills Spec Is Trending on GitHub. Its Entire Contract Is Two Required Fields.
The Agent Skills spec standardizes packaging rather than behavior, its only hard guarantees are naming and folder conventions while the safety-relevant field is…
-
Briefing · July 24, 2026 · afternoon
Both major labs shipped voice as an agent control surface within the same 24 hours, while Claude Opus 5 cut the price of near-frontier agent intelligence in half.
-
no-ai-slop Strips 20+ AI Writing Tells From Any Draft. That Doesn't Make It Yours.
No-ai-slop removes the fingerprints of a machine but can't add the fingerprints of a person, so a draft that passes it reads clean and empty, which makes it a detector…
-
The Coding Agent Became a Security Scanner This Week. It's Also the Thing Being Scanned.
In-loop AI security scanners inherit the trust model of the session they run in, so the same agents now hunting vulnerabilities are themselves a fresh attack surface,…
-
Briefing · July 22, 2026 · afternoon
Containment is failing in two directions this week, as an OpenAI agent broke out of its own test to hack Hugging Face while builders tear down the wall locking coding…
-
harness-engineering: The Repo Behind OpenAI's Million-Line Codex Experiment Has 21 Stars
Harness-engineering is an anthology shipped as an agent context bundle (README says point a coding agent at it, not read it) packaging the practice behind OpenAI's…
-
Briefing · July 19, 2026 · morning
The frontier stalled and the scaffolding raced: a harness-engineering field guide trended, Claude Code rewrote permission checks, ChatGPT desktop added a Codex switcher,…
-
Claude Code Turned `/fork` Into a Background Fleet. Your Terminal Agent Isn't One Chat Anymore.
The July 17 release turned /fork into a copy-into-background-session command, renamed the old in-chat helper to /subtask, and made /resume recover deleted sessions, and…
-
Briefing · July 18, 2026 · morning
The coding agent's harness, not the model, is where competition and danger now sit: xAI open-sourced 840k lines of grok-build, Anthropic rebuilt Claude Code session…
-
Briefing · July 17, 2026 · morning
The interesting layer moved from the model to the permission boundary around it: 1Password credentials Claude never sees, MCP auth moving onto OAuth and OpenID Connect,…
-
Open Interpreter Came Back as a Codex Fork That Wears a Different Face for Every Model
The new Open Interpreter bets the thing holding open models (DeepSeek/Kimi/Qwen/GLM) back isn't the model but the harness wrapped around it, so it ships model-specific…
-
LLM Space Is a Local Desktop App Built for Watching What Your Agent Actually Did
LLM Space is the DeerFlow team's dogfooded, local-first desktop app for inspecting every harness step and replaying failures, and that after-the-fact visibility…
-
Grok Build Went Open Source to Win Back Trust. Reading the Source Is Not the Same as Reading the Source.
XAI open-sourcing 844,530 lines of Rust after its grok CLI uploaded people's home dirs is real progress but not proof of safety (the exfiltration code is…
-
Briefing · July 16, 2026 · afternoon
The scaffolding around the model is where the announcements, capital, and attacks now land: GPT-Red red-teamer, the $1.5B Ode services firm, the Hermes harness…
-
Briefing · July 16, 2026 · morning
After a year of shipping agents first, trust and privacy became the product surface: Grok's data-exfiltration cleanup, Codex dangerous-command detection, and Anthropic…
-
Briefing · July 15, 2026 · morning
The unit of work shifted from one agent to swarms, and the hard problem became making fifty agents hand off cleanly rather than making one smart.