Wily: High-Performance Complexity Gated-Feedback for AI Coding Agents
Wily: Высокопроизводительная система обратной связи с учётом сложности для агентных AI-инструментов программирования
2026-05-22
SCID: 54.1/8f2mpfh4
Discuss with AI
Claude Opus 4.6Cyclomatic ComplexityGPT-5.3-CodexMaintainability IndexSWE-benchWilycomplexity deltacomplexity-gated feedbackgit history analysislogical lines of code
Figures from the paper
Abstract (AI)
AI coding agents can now resolve a large fraction of real-world software engineering tasks, yet the patches they produce are consistently more complex and less maintainable than those written by skilled human engineers. Left unchecked, widespread adoption of such agents will accelerate technical-debt accumulation in production codebases. We present Wily, a high-performance code-analysis tool that computes Cyclomatic Complexity (CC), Maintainability Index (MI), and logical lines of code (LLOC) across an entire git history at over 570 000 LOC/s and is up to 55 × faster than existing alternatives. We integrate Wily as a complexity-gated feedback signal (backpressure) inside an agentic SWE-bench harness: after each candidate patch the agent receives a complexity delta report and may iterate to reduce regressions. In a controlled study of 500 SWE-bench Verified tasks with two frontier models—Claude Opus 4.6 and GPT-5.3-Codex—the Wily-feedback condition reduces mean complexity growth by 10–\(27\%\) and LLOC growth by 8–\(23\%\) relative to a control that receives only test-pass/fail feedback, while maintaining comparable resolution rates (< 1.3 pp drop). Both AI conditions remain substantially above the human baseline across both models, confirming a persistent complexity gap that motivates continued research into quality-aware agent training.
Key Findings
1
Both AI agent conditions (with and without Wily feedback) remain substantially above the human baseline in complexity, indicating a persistent complexity gap.
2
In a controlled study of 500 SWE-bench Verified tasks with Claude Opus 4.6 and GPT-5.3-Codex, Wily-feedback reduced mean complexity growth by 10–27% compared to test-pass/fail-only feedback.
3
In the same study, Wily-feedback reduced LLOC growth by 8–23% while maintaining comparable resolution rates (less than 1.3 percentage point drop).
4
Wily is a high-performance code-analysis tool that computes Cyclomatic Complexity (CC), Maintainability Index (MI), and logical lines of code (LLOC) across entire git histories at over 570,000 LOC/s.
5
Wily is integrated as a complexity-gated feedback signal (backpressure) inside an agentic SWE-bench harness, giving agents a complexity delta report after each candidate patch.
6
Wily is up to 55× faster than existing alternatives for computing code complexity metrics across histories.
Research Object
Wily code-analysis tool integrated as a complexity-gated feedback mechanism in an agentic software-engineering benchmarking harness
Research Subject
Effect of providing Cyclomatic Complexity, Maintainability Index, and logical lines-of-code (LLOC) delta feedback (complexity-gated backpressure) on AI coding agents' patch complexity growth, LLOC growth, and task resolution rates
Publication Details
Publication Date
2026-05-22
Journal
Publisher
ISSN
Cited by
0
Open access PDF
Access Type
Author Information
Download PDF
Subscribe to digest