ORIGINAL RESEARCH • OCTOBER 2026
AI involvement and code drift
AI involvement and code drift
AI involvement
and code drift
A provenance-labeled comparison of drift density and severity across nine production-readiness dimensions.
A provenance-labeled comparison of drift density and severity across nine production-readiness dimensions.
Jonathan Gordon · info@reweaver.ai
More AI involvement, more drift.
And it lands where design intent lives.
More AI involvement = more drift.
And it lands where design intent lives.
More AI involvement, more drift.
And it lands where design intent lives.
Our first study measured freshly generated components in isolation. This one asks what AI involvement does to real, shipping code. We scanned 21 open-source repositories, labeled by provenance before scanning, on identical instrumentation, and compared drift across human-led, AI-assisted, and AI-led code.
Our first study measured freshly generated components in isolation. This one asks what AI involvement does to real, shipping code. We scanned 21 open-source repositories, labeled by provenance before scanning, on identical instrumentation, and compared drift across human-led, AI-assisted, and AI-led code.
Our first study measured freshly generated components in isolation. This one asks what AI involvement does to real, shipping code. We scanned 21 open-source repositories, labeled by provenance before scanning, on identical instrumentation, and compared drift across human-led, AI-assisted, and AI-led code.
sampling
SAMPLING
Real code on both sides of the ledger.
Real code on both sides of the ledger.
Real code on both sides of the ledger.
Greenfield and brownfield code differ in age, review history, and accumulated debt, so comparing fresh AI output with mature human code mixes up provenance and maturity. Here every repository is a real codebase, scanned at a fixed commit by ReWeaver’s deterministic drift engine across nine production-readiness dimensions.
Greenfield and brownfield code differ in age, review history, and accumulated debt, so comparing fresh AI output with mature human code mixes up provenance and maturity. Here every repository is a real codebase, scanned at a fixed commit by ReWeaver’s deterministic drift engine across nine production-readiness dimensions.
Greenfield and brownfield code differ in age, review history, and accumulated debt, so comparing fresh AI output with mature human code mixes up provenance and maturity. Here every repository is a real codebase, scanned at a fixed commit by ReWeaver’s deterministic drift engine across nine production-readiness dimensions.
21
REPOSITORIES
Open-source TypeScript projects, each scanned at a fixed commit.
Open-source TypeScript projects, each scanned at a fixed commit.
21
REPOSITORIES
From user experience and accessibility to maintainability and AI code governance.
9
From user experience and accessibility to maintainability and AI code governance.
9
9
DIMENSIONS
DIMENSIONS
From user experience and accessibility to maintainability and AI code governance.
REPOSITORIES
21
Open-source TypeScript projects, each scanned at a fixed commit.
92
92
UX, accessibility, and
design-consistency rules
examined one by one.
UX, accessibility, and design-consistency rules
examined one by one.
RULES
92
RULES
UX, accessibility, and
design-consistency rules
examined one by one.
Human-led (n=6)
AI-assisted (n=10) AI-led (n=5), labeled before scanning.
Human-led (n=6)
AI-assisted (n=10)
AI-led (n-5), labeled before scanning.
Human-led (n=6)
AI-assisted (n=10)
AI-led (n=5)
labeled before scanning.
PROVENANCE GROUPS
PROVENANCE GROUPS
3
3
3
MEASURES
MEASURES
Drift Density
Distinct drifted lines divided by scanned lines. Density does not saturate, so it is the clearest comparator between groups.
Distinct drifted lines divided by scanned lines. Density does not saturate, so it is the clearest comparator between groups.
Production Drift Ratio (PDR)
Drift frequency weighted by estimated remediation cost, from 0 (no drift) to 1 (severe, costly drift). A PDR of 0.30 or less is considered production ready.
Drift frequency weighted by estimated remediation cost, from 0 (no drift) to 1 (severe, costly drift). A PDR of 0.30 or less is considered production ready.
RESULTS
What changed as AI involvement rose.
RESULTS
RESULTS
What changed as AI involvement rose.
What changed as AI involvement rose.
Drift Density
MEASURES
Distinct drifted lines divided by scanned lines. Density does not saturate, so it is the clearest comparator between groups.
Production Drift Ratio (PDR)
Drift frequency weighted by estimated remediation cost, from 0 (no drift) to 1 (severe, costly drift). A PDR of 0.30 or less is considered production ready.
RESULTS
100%
DRIFT IS ENDEMIC
DRIFT IS ENDEMIC
Every repository in every group drifted, including mature, widely used human-led projects. All three group means sit well above the 0.30 production-ready line.
Every repository in every group drifted, including mature, widely used human-led projects. All three group means sit well above the 0.30 production-ready line.
AI-led repositories averaged 10.15% drifted lines, about 1 in 10, against 4.42% (1 in 23) for the human-led baseline.
2.3×
2.3×
MORE DRIFT IN AI-LED CODE
MORE DRIFT IN AI-LED CODE
100%
DRIFT WAS ENDEMIC.
Every repository in every group drifted, including mature, widely used human-led projects. All three group means sit well above the 0.30 production-ready line.
AI-led repositories averaged 10.15% drifted lines, about 1 in 10, against 4.42% (1 in 23) for the human-led baseline.
6.1×
MORE DRIFT DENSITY IN AI-LED CODE.
AI-led repositories averaged 10.15% drifted lines, about 1 in 10, against 4.42% (1 in 23) for the human-led baseline.
UX drift density rose sharply in AI-led code, with PDR up 0.34. Accessibility (2.3×) and design consistency (2.6×) followed.
2.3×
UX drift density rose sharply in AI-led code, with PDR up 0.34. Accessibility (2.3×) and design consistency (2.6×) followed.
MORE USER EXPERIENCE DRIFT
MORE USER EXPERIENCE DRIFT
6.1×
MORE USER EXPERIENCE DRIFT
6.1×
UX drift density rose sharply in AI-led code, with PDR up 0.34. Accessibility (2.3×) and design consistency (2.6×) followed.
MORE USER EXPERIENCE DRIFT
MAINTAINABILITY PDR GAP
MAINTAINABILITY PDR GAP
The largest source of human drift was no costlier to fix in AI-led code. Hygiene debt looks endemic to software, not specific to AI.
MAINTAINABILITY PDR GAP
0.00
0.0
The largest source of human drift was no costlier to fix in AI-led code. Hygiene debt looks endemic to software, not specific to AI.
The largest source of human drift was no costlier to fix in AI-led code. Hygiene debt looks endemic to software, not specific to AI.
Density and PDR increased
with AI involvement.
Density and PDR increased with AI involvement.
Both metrics rise with AI involvement. PDR is already near the top of its range, so its gap looks modest (+0.06). Density, which does not saturate, shows AI-led code carrying roughly 2.3× the drift of the human baseline.
Density and PDR increased with AI involvement.
Both metrics rise with AI involvement. PDR is already near the top of its range, so its gap looks modest (+0.06). Density, which does not saturate, shows AI-led code carrying roughly 2.3× the drift of the human baseline.
Density and PDR increased with AI involvement.
Both metrics rose with AI involvement. At >0.8, PDR is already near the top of its range, so its gap looks modest (+0.06). Density, which does not saturate, shows AI-led code carrying roughly 2.3× the drift of the human baseline.




Added drift was design-facing.
Added drift was
design-facing.
Added drift was design-facing.
The increase concentrates in user experience, accessibility, and design consistency. Structural dimensions are flat or slightly lower. Testability is the exception: its density rose 3.5×, more than any dimension but UX.
The increase concentrates in user experience, accessibility, and design consistency. Structural dimensions are flat or slightly lower. Testability is the exception: its density rose 3.5×, more than any dimension but UX.
The increase concentrates in user experience, accessibility, and design consistency.
Structural dimensions are flat or slightly lower. Testability is the exception: its density
rose 3.5×, more than any dimension but UX.

Missing code, not malformed code.
Missing code, not
malformed code.
Missing code, not malformed code.
Most of the increase came from code that wasn’t there: absent empty states, forms that fail without feedback, modals that strand keyboard focus, and inconsistent spacing patterns. None of it announces itself in a passing demo. Hardcoded values ran the other way: they were most common in human-led code and absent from AI-led code.
Most of the increase came from code that wasn’t there: absent empty states, forms that fail without feedback, modals that strand keyboard focus, and inconsistent spacing patterns. None of it announces itself in a passing demo. Hardcoded values ran the other way: they were most common in human-led code and absent from AI-led code.
Most of the increase came from code that wasn’t there: absent empty states, forms that fail without feedback, modals that strand keyboard focus, and inconsistent spacing patterns. None of it announces itself in a passing demo. Hardcoded values ran the other way: they were most common in human-led code and absent from AI-led code.

A signal, not a settled effect.
A signal,
not a settled effect.
A signal, not a settled effect.
LIMITATIONS
The sample is small
Twenty-one repositories across three groups support directional findings, not formal statistical tests. Six were added by hand.
Labels are the main threat
An independent authorship detector agreed with all six human-led labels but disagreed with 14 of the 15 AI-group labels. A configuration file shows an AI tool was used, not how much of the code it wrote.
The groups differ in kind
Human-led repositories were large, long-lived projects; most AI-group repositories were small front-end apps. Some of the design-facing increase may reflect that difference.
The sample is small
Twenty-one repositories across three groups support directional findings, not formal statistical tests. Six were added by hand.
Labels are the main threat
An independent authorship detector agreed with all six human-led labels but disagreed with 14 of the 15 AI-group labels. A configuration file shows an AI tool was used, not how much of the code it wrote.
The groups differ in kind
Human-led repositories were large, long-lived projects; most AI-group repositories were small front-end apps. Some of the design-facing increase may reflect that difference.
The sample is small
Twenty-one repositories across three groups support directional findings, not formal statistical tests. Six were added by hand.
The groups differ in kind
Human-led repositories were large, long-lived projects; most AI-group repositories were small front-end apps. Some of the design-facing increase may reflect that difference.
Labels are the main threat
An independent authorship detector agreed with all six human-led labels but disagreed with 14 of the 15 AI-group labels. A configuration file shows an AI tool was used, not how much of the code it wrote.
Read the pattern as a well-defined hypothesis. A companion field study, now in preparation, tests it across 2,825 repositories with a measured AI-exposure gradient.
Read the pattern as a well-defined hypothesis. A companion field study, now in preparation, tests it across 2,825 repositories with a measured AI-exposure gradient.
Highest rate in each row highlighted. Full results for all 92 rules are in Appendix A of the study.
Read the pattern as a well-defined hypothesis. A companion field study, now in preparation, tests it across 2,825 repositories with a measured AI-exposure gradient.
The question isn’t whether code drifts. It’s where.
The question isn’t whether code drifts.
It’s where.
The question isn’t whether code drifts. It’s where.
What AI involvement added was concentrated in the parts of a product that encode design intent: the states around the nominal case, the affordances that make an interface accessible, and adherence to the design system. These are also the parts that functional tests and demos are least likely to catch.
What AI involvement added was concentrated in the parts of a product that encode design intent: the states around the nominal case, the affordances that make an interface accessible, and adherence to the design system. These are also the parts that functional tests and demos are least likely to catch.
The practical response is not to slow AI adoption. It is to make design intent explicit and checkable, treating missing states, focus management, and token adherence as first-class review criteria rather than polish to add later.
The practical response is not to slow AI adoption. It is to make design intent explicit and checkable, treating missing states, focus management, and token adherence as first-class review criteria rather than polish to add later.
DOWNLOAD THE FULL STUDY