JULY 2026
Measuring the frequency and severity of drift
IN AI-GENERATED AND HUMAN-AUTHORED CODE
REWEAVER AI
info@reweaver.ai
ORIGINAL RESEARCH · REWEAVER AI · JULY 2026
AI writes more code. It doesn’t write production-ready code.
The promise was clear: AI coding tools would accelerate development. They have. What velocity charts do not show is what accumulates underneath: bugs, incidents, rework, and drift from the standards the code was supposed to follow.
All drift is not created equal.
This controlled comparison measured the frequency and cost of drift across eight production-readiness dimensions in five leading AI coding tools and a human-authored reference.
5
AI TOOLS
Cursor · Claude Code · Lovable · VS Code · Figma Make
42
PROMPTS
Identical component prompts run across every tool.
6
HUMAN REPOS
Collected on GitHub and verified as human-authored.
8
DIMENSIONS
Production-readiness dimensions measured per file.
Drift Frequency
Percentage of lines containing at least one drift occurrence. Counts what went wrong.
Production Drift Ratio (PDR)
Drift frequency weighted by remediation cost.
HEADLINE FINDINGS
The numbers that change how you think about AI code quality.
100%
DRIFT IS ENDEMIC
Both humans and all tools produced drift. No tool won across all dimensions; each showed a different drift profile.
6 of 8
NO MEANINGFUL FREQUENCY GAP
Raw frequency alone misses the story. Only Security & Privacy and Testability showed meaningful frequency differences.
6.5×
MORE UX DRIFT
AI tools produced 6.5× more costly UX drift than humans, 5× more costly Accessibility drift, and 4× more costly Design Consistency drift.
22×
MORE COSTLY SECURITY & PRIVACY DRIFT
AI tools produced three times the human drift frequency—but a cost that was 22× higher once remediation effort was measured.
EVIDENCE
Frequency understates the risk.
Here’s the proof.
The Production Drift Ratio exposes the remediation burden hidden behind familiar measures of output and raw defect frequency.

MAIN CONCLUSIONS
Drift is endemic to AI generation.
Every tool tested—regardless of benchmark performance or commercial positioning—produced meaningful drift across all dimensions. Tool selection does not substitute for post-generation verification.
The actionable conclusion is not which tool to use; it is that every tool requires a verification layer capable of catching what generation leaves behind.
REFERENCES
Becker, J., Rush, N., Barnes, E., & Rein, D. (2025). Measuring the impact of early-2025 AI on experienced open-source developer productivity. ArXiv, 2507.09089.
CodeRabbit. (2025). State of AI vs. human code generation report.
Faros (2026). AI Engineering Report: The Acceleration Whiplash.
Peng, S., Kalliamvakou, E., Cihon, P., & Demirer, M. (2023). The impact of AI on developer productivity: Evidence from GitHub Copilot. arXiv preprint arXiv:2302.06590.
Sabra, A., Schmitt, O., & Tyler, J. (2025). Assessing the quality and security of AI-generated code: A quantitative analysis. arXiv preprint, arXiv:2508.14727.