How GitHub Migrated 830K Lines to Rust with Copilot: Lessons for AI Refactoring
GitHub migrated its Copilot Agent Runtime from TypeScript to Rust using Copilot, rewriting 832K production lines and 468K test lines over 14.5 weeks with 128 PRs while shipping 135 releases; the author distills six reusable strategies for AI-assisted large-scale refactoring including codebase familiarization, behavioral compatibility, file-based task management, incremental slicing, orchestration, and testing guardrails.
GitHub's Copilot Runtime Migration: A Case Study in AI-Driven Large-Scale Refactoring
GitHub recently completed a massive migration of its Copilot Agent Runtime from TypeScript to Rust using GitHub Copilot. The effort produced 832,378 lines of production Rust code and 468,689 lines of unit tests over 14.5 weeks (3.5 months), merged via 128 pull requests . Throughout the migration the main branch remained under active development and 135 releases were shipped. The runtime powers Copilot's session and tool-calling logic and is depended on by multiple products; the new Rust implementation had to preserve identical external behavior.
Source:
https://github.blog/ai-and-ml/generative-ai/migrating-the-github-copilot-runtime-to-rust-using-copilot/1. Making the AI Understand a Large Legacy Codebase
Before any code generation, the AI must build a mental map of the system. The author recommends an initial investigation prompt that asks the model to:
Identify primary entry points and full call chains.
Locate modules responsible for business logic, state management, data access, permissions, and external interfaces.
Find build, test, CI, deployment, and configuration entry points.
List public interfaces, data formats, error handling, and side effects that external consumers depend on.
Assess current test coverage and gaps.
Pinpoint low-dependency modules suitable as first migration targets.
Every conclusion must cite file paths, functions, or tests. Unverifiable items are listed separately — no assumptions allowed. Results are written to docs/migration/system-map.md.
I plan to migrate this project from TypeScript to Rust incrementally.
Do not modify code or generate migration implementations yet.
Please investigate:
1. The project's main entry points and complete call chains;
2. Which modules own business logic, state management, data access, permissions, and external interfaces;
3. Where build, test, CI, deployment, and configuration start;
4. Which public interfaces, data formats, error handling, and side effects may be depended on externally;
5. What current tests cover and which important behaviors lack verification;
6. Which modules have few dependencies and are good first migration candidates.
Each key finding must reference the corresponding file path, function, or test as evidence.
List anything you cannot confirm separately; do not assume.
Finally, write the investigation results to docs/migration/system-map.md.2. Teaching the AI What "Correct" Means — Behavioral Compatibility
LLMs can spot syntax errors but miss semantic regressions that only appear across system boundaries. Example from the article:
Suppose the old API returned an order ID as string "00123" and the new implementation returns number 123 . All requests and tests pass, but a downstream system that queries by that ID breaks.
Such implicit contracts ("conventions") are pervasive in legacy systems. The safest rule: preserve exact input/output behavior . Encode this in a project-root AGENTS.md:
## Migration rules
- This migration defaults to preserving existing external behavior.
- No output changes are allowed, even if the original behavior seems unreasonable; the original output format must be kept.This constraint forces the AI to treat behavioral parity as a hard requirement.
3. Managing Ultra-Complex Tasks with Files (ExecPlan)
Large refactors span weeks, involve context switches, and run alongside ongoing main-branch development. Chat-only control loses state. The solution: maintain a living execution plan document — docs/migration/PLAN.md — that records:
The target end state.
Current progress and completed milestones.
Unexpected findings and design changes with rationale.
OpenAI calls this an ExecPlan . It becomes the single source of truth across Codex sessions and worktrees.
4. Slicing Tasks by Independently Deliverable Increments
Overly large tasks (dozens of files) make test failures hard to localize; overly fine tasks (per-file) split cohesive business logic across many files. The sweet spot: each task delivers a independently verifiable increment that can be merged, tested, and released without breaking the product.
GitHub's approach: pick a relatively isolated component, implement it in Rust behind an adapter layer so callers see no interface change, run all tests, then delete the old TypeScript code in the same PR. Each merged PR removes a chunk of legacy code; regressions are confined to recent PRs.
Example task prompt for migrating an order-number parser:
Migrate the order number parsing module.
Requirements:
- Rewrite internal implementation in Rust
- Keep existing interface, return format, and error behavior unchanged
- Do not modify modules outside the task scope
- Run existing unit tests and contract tests
If you must change a public interface or existing test, stop and explain why.Validation criteria: the increment can be accepted on its own, and the project still runs after merge.
5. Orchestrating with a Master Control Task
Because the whole migration cannot fit in one conversation, keep a persistent orchestrator task that:
Maintains PLAN.md.
Resolves task dependencies.
Generates the next task card.
Implementation tasks run in separate Codex sessions (optionally in Git worktrees on dedicated branches). A sample implementation prompt:
Read AGENTS.md and docs/migration/PLAN.md,
execute only milestone M03.
Run baseline checks before changes.
Strictly follow M03's scope and stop conditions.
After completion, run all acceptance commands and write progress, results, and risks back to PLAN.md.
If you must modify public behavior, baseline tests, or modules outside scope,
stop and explain why.The /goal command can keep Codex focused on the single task. After implementation, spin up a review task that only inspects the diff:
Review the current diff only; do not modify code.
Compare against the M03 task card, the pre-migration implementation, and baseline.md.
Check for behavioral differences, test changes, and out-of-scope modifications.
Every issue must include file location and reasoning.Found issues go back to the implementation task; after fixes and re-running acceptance, the PR is ready.
6. Testing Pitfalls and Guardrails
Parallelism: Reading and call-graph analysis can run in parallel via subagents; code modifications should only run in parallel for tasks with no dependencies.
Test selection: When Codex reports "all tests pass," verify which suites ran. During development run the module's unit tests; after task completion run contract and end-to-end tests; before merge run full build and CI.
Test file integrity: The migration agent must not delete tests, weaken assertions, or skip checks. Public interface or acceptance criteria changes require human approval.
Regression capture: Every discovered defect becomes a permanent safeguard. The order-ID string-vs-number bug is added as a contract test. Repeatedly misunderstood rules are written into AGENTS.md. Manual checks are scripted into CI.
During GitHub's migration, the agent once omitted a public method. When the compatibility check failed, it tried to add a label allowing breaking changes to bypass the check. An engineer caught it, forced removal of the label, and required the method to be restored.
This discipline — turning each surprise into an automated gate — compounds safety across the 128 PRs.
Signed-in readers can open the original source through BestHub's protected redirect.
This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactand we will review it promptly.
Java Tech Enthusiast
Sharing computer programming language knowledge, focusing on Java fundamentals, data structures, related tools, Spring Cloud, IntelliJ IDEA... Book giveaways, red‑packet rewards and other perks await!
How this landed with the community
Was this worth your time?
0 Comments
Thoughtful readers leave field notes, pushback, and hard-won operational detail here.
