If you run parallel coding agents and your collision check is "do the two PRs touch the same files?", you are missing about half your conflicts. In a new MergeGym study, 47.3% of the textual merge conflicts (79 of 167) happened between pull requests whose authored file sets did not overlap at all.
That number reframes how I think about fan-out. I had been treating file-level disjointness as the safety guarantee that makes it fine to give three sub-agents three tickets. The paper says that guarantee is much weaker than it feels.
I have been guilty of this myself. Splitting a feature into non-overlapping tickets feels like engineering discipline, and when it fails I usually blame the agent for wandering outside its lane. This paper suggests the lane was never as clean as the diff made it look.
Disjoint diffs, conflicting merges#
The cause is mundane. A PR diff is computed against a PR-specific base. A merge is computed from both heads against their common merge base. If one branch carries earlier commits the other has never seen, the real merge surface includes files neither PR "authored" in its own diff. Your agent's change looks isolated and the merge disagrees.
The authors built 715 stratified replay-labeled pairs (167 conflicts) from the AIDev-pop dataset: 33,596 PRs across 2,807 repositories. They deliberately oversampled cross-agent pairs, 16.1% versus a 7.38% population share, so interactions between different agents show up.
The framing matters for tooling too. Most dashboards for parallel agents show per-PR diffs, because that is what the review UI shows. Those views are structurally blind to the exact failure described here, so whatever your orchestrator displays and whatever it gates on are two different things.
A lineage rule beats file overlap by 18 AUROC points#
Their simplest predictor checks whether the history paths from the merge base to each head overlap. That standalone rule reaches held-out AUROC 0.877. The baselines it beats:
File overlap alone: 0.694
Metadata risk (title, body, timing): 0.743
LLM given only the file lists: 0.732
Learned models add a little: logistic regression at 0.882 AUROC (0.661 PR-AUC), an untuned random forest at 0.902 (0.701 PR-AUC). With a replay budget of one third of pairs, the logistic and random forest models recovered 81.2% and 82.9% of held-out conflicts.
Note what that says about the LLM baseline. Reading the file lists with a language model did no better than a metadata heuristic. The signal is in the git graph, not in anyone's judgment about the task.
Just run the merge#
The paper's most practical line is also its bluntest:
When exact replay is cheap, verify everything.
And it is cheap. Local textual merge replay took a median of 0.02 s, with a 90th percentile of 0.38 s and a 95th of 0.50 s. A prediction model is only worth building when the check it replaces is expensive. Here the check costs less than the model call you would use to guess.
In a Claude Code style setup with an orchestrator and worktree-isolated sub-agents, that translates directly. Before an agent opens or updates a PR, do a trial merge of its head against the current target and every other open agent branch, and gate on the result. A dry-run `git merge-tree` is the kind of thing I mean. It is a few lines of orchestrator code, not a research project.
MergeGym also frames this as three tasks: forecasting from open-time signals whether PRs will share authored files, resolving conflicts against observed reconciliations, and scheduling across 97 repository streams (19,397 PRs) to minimize scope-collision exposure. I would not take the scheduling track as a recipe yet. But the forecasting results are a useful warning against the obvious idea of asking a model to predict collisions from ticket text: title, body and timing signals were the weakest family in the comparison.
Zero-overlap conflicts are often fixable#
Among the 79 zero-overlap conflicts, patch reconstruction worked on 48 and all 48 became clean, none remained conflicting. The other 31 were inconclusive because of patch conflicts, missing paths or network failures. So roughly 61% of the surprising conflicts were mechanical, which suggests the right response is automatic rebase or replay, not escalating to a human or burning an agent turn on conflict resolution.
There is a second-order benefit. If every branch is verified against every other open branch continuously, conflicts surface while each agent still has its task in context. A conflict found minutes after a push is a cheap rebase. The same conflict found after three more dependent commits is an expensive one, and it lands on whoever, human or agent, has to untangle history they did not write.
What this does not show#
Be careful with the scope. Everything here is textual mergeability. A clean merge says nothing about whether the build passes or the semantics hold, and two agents can still produce perfectly merging code that breaks each other. The held-out set is small (497 pairs, 117 conflicts) and enriched for cross-agent pairs on purpose, so the rates will not transfer to your repo unchanged. The data is also real-world agent PRs, not controlled orchestration where one scheduler sees all branches from the start, so the 47.3% may overstate or understate what a single orchestrator sees.
One more practical point: the cross-agent oversampling means some of these conflicts come from different agent products working the same repository with no shared coordinator. If your team mixes tools, say one agent in the IDE and another running unattended in CI, the only shared state between them is the repository itself. That makes the git graph the one honest source of truth, and it is exactly what the lineage rule reads.
What I'm changing#
I'm dropping "disjoint files" as a scheduling criterion and replacing it with an actual trial merge on every push, plus a test run on the merged result, since the paper explicitly does not cover that half. My prediction: orchestrators that schedule by file ownership will quietly pile up integration failures, and the fix will be boring git plumbing rather than smarter planning.