ORIGINAL RESEARCH • OCTOBER 2026

AI involvement and code drift

AI involvement and code drift

AI involvement
and code drift

A provenance-labeled comparison of drift density and severity across nine production-readiness dimensions.

A provenance-labeled comparison of drift density and severity across nine production-readiness dimensions.

Jonathan Gordon · info@reweaver.ai

More AI involvement, more drift.
And it lands where design intent lives.

More AI involvement = more drift.
And it lands where design intent lives.

More AI involvement, more drift.
And it lands where design intent lives.

Our first study measured freshly generated components in isolation. This one asks what AI involvement does to real, shipping code. We scanned 21 open-source repositories, labeled by provenance before scanning, on identical instrumentation, and compared drift across human-led, AI-assisted, and AI-led code.

Our first study measured freshly generated components in isolation. This one asks what AI involvement does to real, shipping code. We scanned 21 open-source repositories, labeled by provenance before scanning, on identical instrumentation, and compared drift across human-led, AI-assisted, and AI-led code.

Our first study measured freshly generated components in isolation. This one asks what AI involvement does to real, shipping code. We scanned 21 open-source repositories, labeled by provenance before scanning, on identical instrumentation, and compared drift across human-led, AI-assisted, and AI-led code.

sampling

SAMPLING

Real code on both sides of the ledger.

Real code on both sides of the ledger.

Real code on both sides of the ledger.

Greenfield and brownfield code differ in age, review history, and accumulated debt, so comparing fresh AI output with mature human code mixes up provenance and maturity. Here every repository is a real codebase, scanned at a fixed commit by ReWeaver’s deterministic drift engine across nine production-readiness dimensions.

Greenfield and brownfield code differ in age, review history, and accumulated debt, so comparing fresh AI output with mature human code mixes up provenance and maturity. Here every repository is a real codebase, scanned at a fixed commit by ReWeaver’s deterministic drift engine across nine production-readiness dimensions.

Greenfield and brownfield code differ in age, review history, and accumulated debt, so comparing fresh AI output with mature human code mixes up provenance and maturity. Here every repository is a real codebase, scanned at a fixed commit by ReWeaver’s deterministic drift engine across nine production-readiness dimensions.

21

REPOSITORIES

Open-source TypeScript projects, each scanned at a fixed commit.

Open-source TypeScript projects, each scanned at a fixed commit.

21

REPOSITORIES

From user experience and accessibility to maintainability and AI code governance.

9

From user experience and accessibility to maintainability and AI code governance.

9

9

DIMENSIONS

DIMENSIONS

From user experience and accessibility to maintainability and AI code governance.

REPOSITORIES

21

Open-source TypeScript projects, each scanned at a fixed commit.

92

92

UX, accessibility, and
design-consistency rules

examined one by one.

UX, accessibility, and design-consistency rules

examined one by one.

RULES

92

RULES

UX, accessibility, and
design-consistency rules

examined one by one.

Human-led (n=6)

AI-assisted (n=10) AI-led (n=5), labeled before scanning.

Human-led (n=6)
AI-assisted (n=10)
AI-led (n-5), labeled before scanning.

Human-led (n=6)
AI-assisted (n=10)
AI-led (n=5)
labeled before scanning.

PROVENANCE GROUPS

PROVENANCE GROUPS

3

3

3

MEASURES

MEASURES

Drift Density

Distinct drifted lines divided by scanned lines. Density does not saturate, so it is the clearest comparator between groups.


Distinct drifted lines divided by scanned lines. Density does not saturate, so it is the clearest comparator between groups.


Production Drift Ratio (PDR)

Drift frequency weighted by estimated remediation cost, from 0 (no drift) to 1 (severe, costly drift). A PDR of 0.30 or less is considered production ready.

Drift frequency weighted by estimated remediation cost, from 0 (no drift) to 1 (severe, costly drift). A PDR of 0.30 or less is considered production ready.

RESULTS

What changed as AI involvement rose.

RESULTS

RESULTS

What changed as AI involvement rose.

What changed as AI involvement rose.

Drift Density

MEASURES

Distinct drifted lines divided by scanned lines. Density does not saturate, so it is the clearest comparator between groups.


Production Drift Ratio (PDR)

Drift frequency weighted by estimated remediation cost, from 0 (no drift) to 1 (severe, costly drift). A PDR of 0.30 or less is considered production ready.

RESULTS

100%

DRIFT IS ENDEMIC

DRIFT IS ENDEMIC

Every repository in every group drifted, including mature, widely used human-led projects. All three group means sit well above the 0.30 production-ready line.

Every repository in every group drifted, including mature, widely used human-led projects. All three group means sit well above the 0.30 production-ready line.

AI-led repositories averaged 10.15% drifted lines, about 1 in 10, against 4.42% (1 in 23) for the human-led baseline.

2.3×

2.3×

MORE DRIFT IN AI-LED CODE

MORE DRIFT IN AI-LED CODE

100%

DRIFT WAS ENDEMIC.

Every repository in every group drifted, including mature, widely used human-led projects. All three group means sit well above the 0.30 production-ready line.

AI-led repositories averaged 10.15% drifted lines, about 1 in 10, against 4.42% (1 in 23) for the human-led baseline.

6.1×

MORE DRIFT DENSITY IN AI-LED CODE.

AI-led repositories averaged 10.15% drifted lines, about 1 in 10, against 4.42% (1 in 23) for the human-led baseline.

UX drift density rose sharply in AI-led code, with PDR up 0.34. Accessibility (2.3×) and design consistency (2.6×) followed.

2.3×

UX drift density rose sharply in AI-led code, with PDR up 0.34. Accessibility (2.3×) and design consistency (2.6×) followed.

MORE USER EXPERIENCE DRIFT

MORE USER EXPERIENCE DRIFT

6.1×

MORE USER EXPERIENCE DRIFT

6.1×

UX drift density rose sharply in AI-led code, with PDR up 0.34. Accessibility (2.3×) and design consistency (2.6×) followed.

MORE USER EXPERIENCE DRIFT

MAINTAINABILITY PDR GAP

MAINTAINABILITY PDR GAP

The largest source of human drift was no costlier to fix in AI-led code. Hygiene debt looks endemic to software, not specific to AI.

MAINTAINABILITY PDR GAP

0.00

0.0

The largest source of human drift was no costlier to fix in AI-led code. Hygiene debt looks endemic to software, not specific to AI.

The largest source of human drift was no costlier to fix in AI-led code. Hygiene debt looks endemic to software, not specific to AI.

Density and PDR increased
with AI involvement.

Density and PDR increased with AI involvement.

Both metrics rise with AI involvement. PDR is already near the top of its range, so its gap looks modest (+0.06). Density, which does not saturate, shows AI-led code carrying roughly 2.3× the drift of the human baseline.

Density and PDR increased with AI involvement.

Both metrics rise with AI involvement. PDR is already near the top of its range, so its gap looks modest (+0.06). Density, which does not saturate, shows AI-led code carrying roughly 2.3× the drift of the human baseline.

Density and PDR increased with AI involvement.

Both metrics rose with AI involvement. At >0.8, PDR is already near the top of its range, so its gap looks modest (+0.06). Density, which does not saturate, shows AI-led code carrying roughly 2.3× the drift of the human baseline.

Two bar charts showing PDR and drift density
Two bar charts showing PDR and drift density

Added drift was design-facing.

Added drift was
design-facing.

Added drift was design-facing.

The increase concentrates in user experience, accessibility, and design consistency. Structural dimensions are flat or slightly lower. Testability is the exception: its density rose 3.5×, more than any dimension but UX.

The increase concentrates in user experience, accessibility, and design consistency. Structural dimensions are flat or slightly lower. Testability is the exception: its density rose 3.5×, more than any dimension but UX.

The increase concentrates in user experience, accessibility, and design consistency.
Structural dimensions are flat or slightly lower. Testability is the exception: its density
rose 3.5×, more than any dimension but UX.

Orange and gray bar chart on a black background

Missing code, not malformed code.

Missing code, not
malformed code.

Missing code, not malformed code.

Most of the increase came from code that wasn’t there: absent empty states, forms that fail without feedback, modals that strand keyboard focus, and inconsistent spacing patterns. None of it announces itself in a passing demo. Hardcoded values ran the other way: they were most common in human-led code and absent from AI-led code.

Most of the increase came from code that wasn’t there: absent empty states, forms that fail without feedback, modals that strand keyboard focus, and inconsistent spacing patterns. None of it announces itself in a passing demo. Hardcoded values ran the other way: they were most common in human-led code and absent from AI-led code.

Most of the increase came from code that wasn’t there: absent empty states, forms that fail without feedback, modals that strand keyboard focus, and inconsistent spacing patterns. None of it announces itself in a passing demo. Hardcoded values ran the other way: they were most common in human-led code and absent from AI-led code.

A signal, not a settled effect.

A signal,
not a settled effect.

A signal, not a settled effect.

LIMITATIONS

The sample is small

Twenty-one repositories across three groups support directional findings, not formal statistical tests. Six were added by hand.

Labels are the main threat

An independent authorship detector agreed with all six human-led labels but disagreed with 14 of the 15 AI-group labels. A configuration file shows an AI tool was used, not how much of the code it wrote.

The groups differ in kind

Human-led repositories were large, long-lived projects; most AI-group repositories were small front-end apps. Some of the design-facing increase may reflect that difference.

The sample is small

Twenty-one repositories across three groups support directional findings, not formal statistical tests. Six were added by hand.

Labels are the main threat

An independent authorship detector agreed with all six human-led labels but disagreed with 14 of the 15 AI-group labels. A configuration file shows an AI tool was used, not how much of the code it wrote.

The groups differ in kind

Human-led repositories were large, long-lived projects; most AI-group repositories were small front-end apps. Some of the design-facing increase may reflect that difference.

The sample is small

Twenty-one repositories across three groups support directional findings, not formal statistical tests. Six were added by hand.

The groups differ in kind

Human-led repositories were large, long-lived projects; most AI-group repositories were small front-end apps. Some of the design-facing increase may reflect that difference.

Labels are the main threat

An independent authorship detector agreed with all six human-led labels but disagreed with 14 of the 15 AI-group labels. A configuration file shows an AI tool was used, not how much of the code it wrote.

Read the pattern as a well-defined hypothesis. A companion field study, now in preparation, tests it across 2,825 repositories with a measured AI-exposure gradient.

Read the pattern as a well-defined hypothesis. A companion field study, now in preparation, tests it across 2,825 repositories with a measured AI-exposure gradient.

Highest rate in each row highlighted. Full results for all 92 rules are in Appendix A of the study.

Read the pattern as a well-defined hypothesis. A companion field study, now in preparation, tests it across 2,825 repositories with a measured AI-exposure gradient.

The question isn’t whether code drifts. It’s where.

The question isn’t whether code drifts.
It’s where.

The question isn’t whether code drifts. It’s where.

What AI involvement added was concentrated in the parts of a product that encode design intent: the states around the nominal case, the affordances that make an interface accessible, and adherence to the design system. These are also the parts that functional tests and demos are least likely to catch.

What AI involvement added was concentrated in the parts of a product that encode design intent: the states around the nominal case, the affordances that make an interface accessible, and adherence to the design system. These are also the parts that functional tests and demos are least likely to catch.

The practical response is not to slow AI adoption. It is to make design intent explicit and checkable, treating missing states, focus management, and token adherence as first-class review criteria rather than polish to add later.

The practical response is not to slow AI adoption. It is to make design intent explicit and checkable, treating missing states, focus management, and token adherence as first-class review criteria rather than polish to add later.

DOWNLOAD THE FULL STUDY