Wily: High-Performance Complexity Gated-Feedback for AI Coding Agents

Wily: Высокопроизводительная система обратной связи с учётом сложности для агентных AI-инструментов программирования
Amin Beheshti, Anthony Shaw
2026-05-22

Claude Opus 4.6Cyclomatic ComplexityGPT-5.3-CodexMaintainability IndexSWE-benchWilycomplexity deltacomplexity-gated feedbackgit history analysislogical lines of code
AI coding agents can now resolve a large fraction of real-world software engineering tasks, yet the patches they produce are consistently more complex and less maintainable than those written by skilled human engineers. Left unchecked, widespread adoption of such agents will accelerate technical-debt accumulation in production codebases. We present Wily, a high-performance code-analysis tool that computes Cyclomatic Complexity (CC), Maintainability Index (MI), and logical lines of code (LLOC) across an entire git history at over 570 000 LOC/s and is up to 55 × faster than existing alternatives. We integrate Wily as a complexity-gated feedback signal (backpressure) inside an agentic SWE-bench harness: after each candidate patch the agent receives a complexity delta report and may iterate to reduce regressions. In a controlled study of 500 SWE-bench Verified tasks with two frontier models—Claude Opus 4.6 and GPT-5.3-Codex—the Wily-feedback condition reduces mean complexity growth by 10–\(27\%\) and LLOC growth by 8–\(23\%\) relative to a control that receives only test-pass/fail feedback, while maintaining comparable resolution rates (< 1.3 pp drop). Both AI conditions remain substantially above the human baseline across both models, confirming a persistent complexity gap that motivates continued research into quality-aware agent training.
1
Both AI agent conditions (with and without Wily feedback) remain substantially above the human baseline in complexity, indicating a persistent complexity gap.
2
In a controlled study of 500 SWE-bench Verified tasks with Claude Opus 4.6 and GPT-5.3-Codex, Wily-feedback reduced mean complexity growth by 10–27% compared to test-pass/fail-only feedback.
3
In the same study, Wily-feedback reduced LLOC growth by 8–23% while maintaining comparable resolution rates (less than 1.3 percentage point drop).
4
Wily is a high-performance code-analysis tool that computes Cyclomatic Complexity (CC), Maintainability Index (MI), and logical lines of code (LLOC) across entire git histories at over 570,000 LOC/s.
5
Wily is integrated as a complexity-gated feedback signal (backpressure) inside an agentic SWE-bench harness, giving agents a complexity delta report after each candidate patch.
6
Wily is up to 55× faster than existing alternatives for computing code complexity metrics across histories.

Wily code-analysis tool integrated as a complexity-gated feedback mechanism in an agentic software-engineering benchmarking harness

Effect of providing Cyclomatic Complexity, Maintainability Index, and logical lines-of-code (LLOC) delta feedback (complexity-gated backpressure) on AI coding agents' patch complexity growth, LLOC growth, and task resolution rates

Publication Details
Publication Date
2026-05-22
Journal
Publisher
ISSN
Cited by
0
Access Type
Author Information
Authors
Amin Beheshti
Anthony Shaw
Explore further
Open the scid.ai AI chat with a ready-made request: it will find papers on a similar topic and help build a literature review.
Find similar papers in the chat
Make a presentation
100%