Hugging Face's Serge Fixes Failing CI for About $43 a Merge. The Gates Are the Product
Serge, the agent behind 29 merged fixes in the Transformers library, publishes its costs and its failure rate. The most useful thing in the repo is how often it decides not to write a patch.
Out of 86 failing-test investigations that reached a real model session, Hugging Face's bug-fixing agent produced 24 verified patches, which became 19 distinct pull requests. Twenty-eight times, it gave up because it could not find a safe patch.
That second number is the interesting one. Most agent demos show you the wins. The Hugging Face team behind Serge published the losses, the dollar cost per merge, and a frank admission that the system still cannot reliably tell a real fix from a cheat. Read closely, the post makes a case that the patch-writing model is the least important part of a CI agent. The parts that decide when not to act are what make the output something a maintainer will merge.
What Serge is
The Serge repo describes two jobs. The default one is a pull request reviewer: add an LLM key as a repository secret named LLM_API_KEY, install the GitHub Action, and comment @askserge please review on an open PR to get inline review comments. It works with any OpenAI-compatible model.
The second job is the one in the September 29 Hugging Face post: an optional, write-capable "tasks flow" that turns failing CI runs into fix PRs. The team runs it against Transformers, one of the busiest open-source codebases in machine learning, where GPU test failures pile up faster than maintainers can triage them.
It is Apache-2.0, at v0.1.0, and small, with about 50 stars. This is a working internal tool that happens to be public, not a polished product.
The pipeline, and where it stops
The blog lays out six steps:
- Filter failures, keeping only those "worth investigating."
- Reproduce the failure on a fresh GPU runner.
- Check that nobody is already working on it.
- Generate a patch and run quality checks.
- Verify the patch fixes the failure.
- Open a PR and ping the maintainers.
Count the exits. Step 2 is a hard gate: Serge "spins up a fresh GPU runner and executes the test several times. If the failure doesn't reproduce reliably, the task stops there." A flaky test never reaches the model. That one rule removes the most common way CI agents waste money, which is confidently "fixing" a test that was never deterministically broken.
Step 3 uses relore, a companion service that indexes a repo's issue and PR history so the agent can see whether someone already owns the problem and why earlier decisions were made. The post gives an example where relore found an open, approved PR already addressing the same regression, so Serge stopped instead of opening a duplicate. Without it, an agent opens duplicate PRs and maintainers learn to ignore it.
Step 5 is the strictest. "The targeted test runs five times on the unpatched tree and five times on the patched tree." Fail before, pass after, repeatedly, on real hardware. Anything less does not ship.
And step 4 has a quiet trap the team names out loud. A patch that "only changes expected values" gets treated with suspicion, because the cheapest way to make a failing test pass is to change what the test expects. That is reward hacking, and the post admits there is no "perfect automatic way to distinguish a legitimate expectation update from reward hacking." Humans still have to look.
The receipts
The numbers from the post:
- 29 fixes landed in Transformers over the last 80 days, about 2.5 per week.
- Of 86 groups that entered a real LLM session: 24 produced verified patches (19 distinct PRs), 28 ended without a safe patch, 18 failed GPU verification, 5 timed out during verification, and 11 failed mid-session on patch-apply, runner or parsing errors.
- Inference cost: about $14 per distinct PR (counting those 19) and $43 per merged PR. Individual sessions ran "a little under $2" to "closer to $3."
- Human time saved: the team estimates "roughly one to two hours of developer attention per fix, or around 43 hours across the 29 fixes that landed."
Some maintainers merged patches unchanged. Others tweaked them or wrote a different fix after reading Serge's diagnosis. A tracking issue gives maintainers one place to see which failures were dispatched and which produced a PR.
Do the math the way a manager would. Roughly $43 in inference per merged fix, against one to two hours of a senior engineer's time. That trade is excellent, and it only holds because the gates kept the PR stream clean enough that maintainers kept reading it. An agent that opened 86 PRs instead of 19 would cost less per PR and be worth nothing, because nobody would review its output.
The model matters less than you think
The main model is Kimi K-2.7-Code. The team also tested Qwen3.8-27B, which scored higher on the DeepSWE 1.1 benchmark (42.2 against Kimi's 31) while being, in their words, "twice cheaper."
That comparison is useful, but notice what it does not change. Swap the model and the reproduce gate, the duplicate check, the five-and-five verification and the reward-hacking suspicion all stay exactly where they were. Those are what produced a PR stream maintainers trust. The model decides how many of the 86 sessions turn into candidates. The gates decide whether any of them deserve a reviewer's attention.
How it fences its own write access
The security design is where Serge is most worth copying, and it lives in the repo's tasks flow and security docs.
Write access is off three times over. The deployment needs TASK_API_ENABLED=1, each repo has to opt in with an "Enable write-capable tasks" checkbox, and callers authenticate with GitHub Actions OIDC tokens whose audience must match TASK_OIDC_AUDIENCE (default serge). There is no shared secret, and the docs say a leaked token is scoped to one repository and useless within minutes.
The model never pushes. "The LLM only proposes patches," and Serge itself applies the patch, commits through the GitHub Git Data API, and opens the PR. The security doc puts it bluntly: "push credentials never enter the sandbox."
Follow-up fixes are fenced to Serge's own branches. A "branch-ownership guard" allows updating an existing PR only for serge/* branches, and TASK_MAX_FOLLOWUPS caps commits per fix branch at five by default. Logs and test output are "treated as data, fed to the prompt, never as instructions," and with Kubernetes isolation the pods run under a NetworkPolicy that "denies all egress" except to approved services.
None of this is exotic. It is the same discipline you would apply to a new contractor: limited keys, a sandbox, a named branch, and a cap on how many times they can try before a human steps in.
Put this into practice
You do not need to run Serge to use what it teaches. The lowest-friction path:
Start with the reviewer, read-only. Add LLM_API_KEY as a secret, install the Action, and comment @askserge please review on a few PRs. You learn how the model reads your codebase without giving it write access. One caveat from the docs: GitHub withholds secrets from forked-PR workflow runs, so do not rely on the Action for outside contributions.
Add a reproduce gate to whatever CI agent you already run. Rerun a failing test several times on clean infrastructure before any model sees it. Drop anything that does not fail consistently. This one step will cut wasted sessions more than any model upgrade.
Verify both directions. Require the targeted test to fail on the base tree and pass on the patched tree, more than once. Five and five is Serge's number, and it is a good default.
Flag expectation-only diffs. If a patch changes only asserted values, route it to a human with a label that says so.
Fence the writes. Give the agent its own branch prefix, cap follow-up commits, keep push credentials outside the model's sandbox, and treat logs as untrusted input.
Measure cost per merged PR, not per PR. It is the only number that tells you whether maintainers are getting value.
Honest limitations
All the headline numbers come from Hugging Face's own post about its own tool. There is no independent replication, and Transformers is an unusual codebase: huge, heavily tested, with dedicated GPU CI and maintainers who already knew the system existed. Your merge rate will differ.
The figures also sit on different windows. The 29 merges cover the last 80 days, while the 86-session breakdown is its own sample. Do not divide one number by another and call it a success rate.
Reward hacking is unsolved, by the team's own account. A human still has to judge every expectation change.
The GPU reproduction step assumes you have GPU runners to spare. For a small team, that infrastructure costs more than the inference. And the repo is v0.1.0 with a reviewer-first README, so expect the tasks flow to need real setup work: a web app deployment, a GitHub App with contents and pull-request write access, and OIDC wiring.
The question to take back to your team
Serge's most important output is not a patch. It is the 28 times it said no, and the maintainers who kept reading its PRs because of that.
Before you give any agent write access to your repository, write down its stop conditions. When does it quit? What does it refuse to touch? What does a human see before anything merges? If you cannot answer those, the model you choose will not save you. If you can, a modest model and $43 a merge starts to look like a very good deal.
Sources: Hugging Face, "Anatomy of a bug-fixing agent," Sept 29, 2026 · huggingface/serge README · Serge tasks flow docs · Serge security docs · huggingface/relore
Medium metadata
- Title: Hugging Face's Serge Fixes Failing CI for About $43 a Merge. The Gates Are the Product
- Subtitle: The agent behind 29 merged fixes in Transformers publishes its costs and failure rate. The most useful thing in the repo is how often it decides not to write a patch.
- Tags: AI Agents, Continuous Integration, Hugging Face, Software Engineering, Open Source
- Canonical: import from the fervorai.dev URL