Back to Top

AI Won’t Migrate Your Legacy App. Here’s the Pipeline That Will

Ai modernization pipeline

Every modernization programme now has a slide promising that AI agents will take the effort out of migration. The promise is partly true. Teams that point an LLM at a 15-year-old repository and ask it to “upgrade to Spring Boot 3” usually end up with a large diff, a broken build and a review queue nobody trusts.

The teams getting real value treat AI as one stage in an engineered pipeline, not as the pipeline itself.

The Problem With “Just Point an Agent at the Repo”

Legacy modernization work splits into two very different kinds of change.

The first kind is mechanical. Examples include the javax.* to jakarta.* namespace move, deprecated API replacements, dependency version bumps, build plugin upgrades and Java language-level changes. These follow known rules and need to be applied consistently across thousands of files.

The second kind is contextual. Examples include rewriting a security configuration that extended the removed WebSecurityConfigurerAdapter, fixing a Hibernate 6 query that now behaves differently, or untangling a custom framework someone built in 2011. These need judgment.

LLMs are good at contextual work and unreliable at mechanical work at scale. A model that correctly rewrites 98% of 4,000 import statements has still introduced 80 defects, scattered randomly, in a diff too large to review properly. Deterministic tools have the opposite profile: perfectly consistent, but blind to anything outside their rules.

So the architecture follows from that split. Use deterministic tools for everything they can do, and give the LLM only what remains.

The Pipeline: Five Stages

Stage 1: Discover From Runtime, Not Just Code

Static analysis tells you what the code could call. Runtime telemetry tells you what it does call. Before you change anything, instrument the application with OpenTelemetry. Capture real dependency edges, endpoint traffic and p95 latency as a baseline.

AI agents help here by summarising unfamiliar modules, drafting dependency registers and flagging dead code candidates. Treat their output as hypotheses to confirm against telemetry, never as ground truth. An agent-generated dependency map that misses one batch job calling a mainframe queue at 2 a.m. on the first of the month will surface during cutover, at the worst possible time.

Stage 2: Deterministic Transformation With OpenRewrite

OpenRewrite applies recipes to the code’s lossless semantic tree. It understands types, not just text, so changes are precise and repeatable. You define a composite recipe once and run it identically across every repository in a wave:

yaml
# rewrite.yml: shared baseline for every app in the wave
type: specs.openrewrite.org/v1beta/recipe
name: com.acme.modernization.Wave1Baseline
displayName: Wave 1 modernization baseline
recipeList:
  - org.openrewrite.java.migrate.UpgradeToJava21
  - org.openrewrite.java.spring.boot3.UpgradeSpringBoot_3_2
  - org.openrewrite.java.migrate.jakarta.JavaxMigrationToJakarta

Commit the result as its own change set. Reviewers can then trust that commit wholesale, because it is reproducible: rerun the recipe and you get the same diff.

Stage 3: The Residual Loop

After the recipes run, build the project. What fails is the residual: the contextual work no recipe covered. This is where the LLM agent comes in, with narrow context and hard limits:

python
for attempt in range(MAX_ATTEMPTS):
    result = build_and_test(repo)
    if result.ok:
        break
    failures = result.compile_errors[:5]        # small, scoped context
    patch = agent.propose_patch(failures, files=result.failing_files)
    if patch.touches(PROTECTED_PATHS) or patch.changed_lines > 80:
        escalate(repo, failures)                # human takes over
        break
    apply(patch)
else:
    escalate(repo, result.failures)

Three design choices matter here. The agent sees a handful of real compiler errors, not the whole repository. Patch size is capped, because a large AI patch usually means the model is re-architecting rather than fixing. Protected paths, such as security config, payment logic and anything under regulatory control, always escalate to a human, however confident the model sounds.

Stage 4: Verification Gates

A green build is not proof of equivalence. Before cutover, three things should be true. Characterization tests pass. Contract tests (for example with Pact) confirm that consumers still get the responses they expect. Performance stays within an agreed tolerance of the Stage 1 baseline.

Agents are useful for generating characterization tests that capture current behaviour, but read them carefully. A generated test can lock in an existing bug as “expected behaviour”. That is acceptable during migration, since equivalence is the goal, but the test should be labelled as characterization so nobody mistakes it for a specification later.

Stage 5: Human Review, Tiered by Risk

Separating the commits pays off at review time. The deterministic commit gets a light review. The residual commit, which is small by design, gets a careful one. Reviewers spend their attention where the uncertainty actually sits.

A Real-World Scenario: A 400-Application Portfolio

Consider a regulated enterprise migrating roughly 400 applications to cloud in waves. Each app is first banded by complexity: Low, Medium, High and Very High.

Low-complexity apps are standard Spring Boot services with clean dependencies. They run through the full pipeline with minimal intervention, and the residual loop resolves most build failures within its limits.

Medium-complexity apps produce more escalations, often around security configuration and ORM behaviour. Engineers handle these, and recurring fixes become new custom OpenRewrite recipes. This is the compounding benefit: each wave makes the deterministic stage smarter and the residual smaller.

High and Very High apps, the custom frameworks and tightly coupled monoliths, are not automation candidates at all. Architects lead their re-architecture. Agents contribute in Stage 1, documenting what the code does, and in Stage 4, generating test scaffolding.

The lesson from this kind of portfolio is that AI changes the shape of the effort curve more than its total. Low-complexity work speeds up significantly. Complex work speeds up modestly, mainly through better discovery and documentation. Estimates should reflect that, rather than applying one flat “AI productivity factor” across all bands.

Common Mistakes

Measuring lines changed as productivity. A pipeline that changes 50,000 lines deterministically has done less risky work than an agent that changed 500 lines of security config. Measure escalation rate, defect leakage and review time instead.

Skipping the runtime baseline. Without Stage 1 telemetry you can’t prove the migrated app behaves the same, and every production incident turns into an argument about whether the migration caused it.

Letting the agent fix its own test failures without limits. An unconstrained agent will sometimes “fix” a failing test by changing the assertion. Keep test files protected, or require human approval for any assertion change.

Not feeding escalations back into recipes. If engineers fix the same pattern manually in 30 apps, that is a missing recipe, not 30 separate pieces of work.

Trade-offs to Accept

This pipeline costs more to set up than prompting an agent directly. You need recipe engineering, telemetry instrumentation and CI integration before the first app benefits. For a portfolio of five apps, that investment may not pay back. For hundreds of apps, the consistency and auditability are hard to get any other way, especially in sectors where you must explain every change to an auditor.

Key Takeaways

  1. Split modernization work into mechanical and contextual changes, and route each to the right tool.
  2. Run deterministic transformations such as OpenRewrite first, and give LLM agents only the residual build failures.
  3. Constrain agents with small context, capped patch sizes and protected paths that always escalate.
  4. Capture a runtime baseline before migrating, so equivalence can be proven rather than assumed.
  5. Turn repeated manual fixes into recipes, so each wave gets cheaper than the last.

Leave a Reply

Your email address will not be published. Required fields are marked *

This site uses Akismet to reduce spam. Learn how your comment data is processed.

Most Popular Posts