JULY 2026

Measuring the frequency and severity of drift

IN AI-GENERATED AND HUMAN-AUTHORED CODE

REWEAVER AI

info@reweaver.ai

ORIGINAL RESEARCH · REWEAVER AI · JULY 2026

AI writes more code. It doesn’t write production-ready code.

The promise was clear: AI coding tools would accelerate development. They have. What velocity charts do not show is what accumulates underneath: bugs, incidents, rework, and drift from the standards the code was supposed to follow.

All drift is not created equal.

This controlled comparison measured the frequency and cost of drift across eight production-readiness dimensions in five leading AI coding tools and a human-authored reference.

5

AI TOOLS

Cursor · Claude Code · Lovable · VS Code · Figma Make

42

PROMPTS

Identical component prompts run across every tool.

6

HUMAN REPOS

Collected on GitHub and verified as human-authored.

8

DIMENSIONS

Production-readiness dimensions measured per file.

Drift Frequency

Percentage of lines containing at least one drift occurrence. Counts what went wrong.

Production Drift Ratio (PDR)

Drift frequency weighted by remediation cost.

HEADLINE FINDINGS

The numbers that change how you think about AI code quality.

100%

DRIFT IS ENDEMIC

Both humans and all tools produced drift. No tool won across all dimensions; each showed a different drift profile.

6 of 8

NO MEANINGFUL FREQUENCY GAP

Raw frequency alone misses the story. Only Security & Privacy and Testability showed meaningful frequency differences.

6.5×

MORE UX DRIFT

AI tools produced 6.5× more costly UX drift than humans, 5× more costly Accessibility drift, and 4× more costly Design Consistency drift.

22×

MORE COSTLY SECURITY & PRIVACY DRIFT

AI tools produced three times the human drift frequency—but a cost that was 22× higher once remediation effort was measured.

EVIDENCE

Frequency understates the risk.
Here’s the proof.

The Production Drift Ratio exposes the remediation burden hidden behind familiar measures of output and raw defect frequency.

MAIN CONCLUSIONS

Drift is endemic to AI generation.

Every tool tested—regardless of benchmark performance or commercial positioning—produced meaningful drift across all dimensions. Tool selection does not substitute for post-generation verification.

The actionable conclusion is not which tool to use; it is that every tool requires a verification layer capable of catching what generation leaves behind.

REFERENCES

Becker, J., Rush, N., Barnes, E., & Rein, D. (2025). Measuring the impact of early-2025 AI on experienced open-source developer productivity. ArXiv, 2507.09089.

CodeRabbit. (2025). State of AI vs. human code generation report.

Faros (2026). AI Engineering Report: The Acceleration Whiplash.

Peng, S., Kalliamvakou, E., Cihon, P., & Demirer, M. (2023). The impact of AI on developer productivity: Evidence from GitHub Copilot. arXiv preprint arXiv:2302.06590.

Sabra, A., Schmitt, O., & Tyler, J. (2025). Assessing the quality and security of AI-generated code: A quantitative analysis. arXiv preprint, arXiv:2508.14727.