GitHub pushed stacked Pull Requests into public preview on July 31, and followed up on August 4 with an engineering blog describing a complete workflow for turning giant AI-generated diffs into reviewable chains. This is not just a new feature—it is GitHub's first direct admission that when an agent generates 1000+ lines of code in one shot, the traditional "one PR, one conversation" model breaks down.
How it works
The technical form is straightforward: a stacked PR is an ordered, linked set of PRs within the same repository, where each PR targets the branch of the PR below it, forming a bottom-up merge chain. The bottom layer merges first; upper layers automatically rebase to follow.
- Native UI + stack map on GitHub.com for visualizing the chain
- CLI support via the
gh-stackextension (requires GitHub CLI 2.90.0+ and Git 2.20+) for lifecycle management - AI agent integration through the
gh-stackskill, letting Copilot and similar agents invokegh stackcommands to split changes automatically - Branch protection and required checks apply across the whole chain—stacking bypasses no merge gates, which is GitHub tying "safety governance" directly to the workflow upgrade
- Splitting is not free—when lower layers change, upper layers must rebase and resolve conflicts; CI cannot be skipped.
- Tests must attach to every layer, not run once after the whole chain is done.
- Upper-layer code is not independently releasable until it lands on the main branch.
- The gh-stack skill's auto-splitting depends on prompt quality—it isn't magic; reviewers still need to inspect each layer's actual changes.
The motivating case
GitHub's blog example splits a 1000+ line AI-generated diff into four layers: L1 (data model) → L2 (API) → L3 (wiring) → L4 (UI), each assigned to a different reviewer. The bottom layer is reviewed and merged first; upper layers build on it and continue review on top.
This matches the real-world experience of WHOOP engineer Mayank Saini: "In the past, a large change meant one giant PR nobody wanted to review. Now it becomes a stack of small PRs—reviewers can actually follow the reasoning, and the entire chain merges in one pass."
Why now
AI coding has driven down the cost of *generating* code, but the cost of *reviewing* it hasn't changed. An engineer reviewing a 1000-line PR won't be much faster than reviewing 200 lines—yet an agent can produce 1000 lines in minutes. The bottleneck has shifted from writing to reading, and the entire development pipeline's throughput is choked at the review stage.
Stacked PRs are not just "encourage small PRs" discipline talk. They make small, incremental changes the platform default: agents decide how many layers, how big each layer is, and who reviews. The stack map lets reviewers browse layers in parallel while still seeing dependencies clearly.
Four real-world caveats
Assessment
Stacked PRs, OpenAI Codex's Sol/Luna layering (August 3 topic), and Cloudflare's ADLC (August 4 topic) form the same wave: when AI compresses the "implementation" stage to near zero, every stage of engineering management is forced to be rewritten.
GitHub chose stacking because it stays within the "PR as the unit of collaboration" tradition without inventing new workflow terminology—the path of least resistance and best ecosystem compatibility. Teams already using Graphite, Sapling, or git-branchless may find GitHub's native stacking doesn't fully replace mature third-party tools. But GitHub's entry means stacking shifts from an "advanced technique" to a "default action." By GA, native stacking will very likely become the default workflow for new repositories.
Original post: https://github.blog/engineering/turn-one-giant-ai-generated-pull-request-to-a-reviewable-stack
Docs: https://docs.github.com/pull-requests/get-started/stacked-prs-quickstart