What happened

GitHub shipped a changelog update on September 11 changing how Copilot’s code review feature checks and manages its own feedback on pull requests.

Three changes are live now, described in the post using present tense with no mention of a preview period. First, auto-resolution: “when you push a commit that addresses a Copilot code review comment, Copilot now resolves that comment during its rereview,” instead of leaving you to close out threads by hand. Second, when you accept one of Copilot’s suggested code changes, it now writes a commit message based on what the suggestion actually changes, rather than a generic one. Third, Copilot’s review agent now uses, in GitHub’s words, “the full set of shell tools from the Copilot SDK,” including running build commands, running tests, executing targeted scripts, and pulling information from other available tools and APIs, to validate its findings before it posts them.

Screenshot of a GitHub Copilot code review comment thread showing a note that Copilot automatically resolved the comment

GitHub also says reviews run at its lower-cost “Lite” effort tier now use an ensemble of multiple review agents working together instead of one. Internally, GitHub reports this raised the average number of addressed comments per review by 47% for high-severity findings, 31% for medium, and 11% for low, while cutting review cost by about 8%. The post doesn’t say which plans or repositories have access to any of this, or whether a setting exists to turn auto-resolution off.

Why it matters

The auto-resolution feature changes what “no open comments” means on a pull request. Previously, an open Copilot comment stayed open until you closed it, so the list of open threads was a list you controlled. Now Copilot decides, on its own judgment, whether your latest commit actually addressed what it flagged. If it judges that incorrectly, a comment that still matters can disappear from your open-threads view marked as resolved, not because the underlying issue is fixed, but because Copilot believed a later commit fixed it.

The shell-tool change is worth watching for a different reason. Copilot’s review agent isn’t just reading your diff anymore. It’s running build commands, tests, and “targeted scripts” against your repository to check its own findings. GitHub frames this purely as a quality improvement, and for most projects that’s exactly what it is. But if any of your repository’s build or test scripts do something beyond compiling code and asserting on results, such as calling a real external API (another service’s software, reached over the internet), writing to a shared database, or triggering a webhook (an automatic notification sent to another system when something happens), an automated review agent invoking those scripts on its own schedule is a different risk profile than a script only a human or your CI, the automated pipeline that already runs your build and tests on every push, runs deliberately.

Who should care

Anyone using Copilot code review who treats a clean, all-resolved comment thread as their signal that a pull request is ready to merge. Also anyone whose repository’s build or test scripts have side effects that reach outside an isolated test environment, since those scripts can now run as part of an automated review rather than only under your own CI setup.

What builders should do next

Run a bounded check before you trust the new behavior blindly. Pick two pull requests where Copilot leaves a review comment. On one, push a commit that fully fixes what it flagged. On the other, push a commit that only partially addresses the comment, or that touches a nearby but different line. Compare two things: whether Copilot correctly leaves the second comment open instead of resolving it, and whether the resolution note it leaves on the first (the one shown in the screenshot above) is specific enough that you can tell, without re-reading the diff yourself, that it actually verified the fix rather than closing the thread on any new commit. If it resolves the partial fix too, don’t treat an empty comment list as your merge signal until GitHub tightens that behavior.

Separately, if your team’s build or test scripts reach outside the repository, check with whoever owns those scripts about whether they’re safe to run unattended by a review agent, not just by your own CI.


End of article