Why AI Code Volume Lies: Measuring Real Efficiency via Task Acceptance and Rework
The article argues that AI programming efficiency should be measured by task acceptance, rework, and defects—not code volume or tokens—and proposes a practical framework tracking human interventions, shared improvements, and maintenance costs to assess true team productivity gains.
An agent keeps running, code keeps increasing, and tokens are consumed heavily. Yet developers still repeatedly explain requirements, correct implementations, remind about tests, and finally take over rework. In this situation it is hard to judge efficiency gains merely by "code written fast."
The author focuses on a frontline question: delivering a compliant feature, how much human effort is actually required? And how much of that effort is spent re-solving problems already solved? The author references InfoQ's summary of Patrick Debois's talk, which offers two observation angles: human interventions needed to complete a task, and shared benefits that one system improvement brings to other members. This shifts attention from individual code generation to the team's ability to consistently do things right.
The previous article discussed turning corrections into reusable improvements. This article continues: after improvements are made, how to judge whether they are truly effective?
Look at Delivery Results First, Then Resources Spent
Token usage helps analyze model call costs; code volume describes change scale. Neither alone represents productivity. Fewer tokens may mean better context organization or skipped necessary checks; more code may mean more features completed or duplicate implementations introduced. Judgment must return to results.
A hypothetical case illustrates: a management system needs batch data import. Business allows partial success; retries must not duplicate successful records; failure reasons must be queryable. The agent quickly produces an upload page and import API. But if humans must then add retry behavior, modify error messages, and investigate duplicate records, the "first version generation time" covers only a tiny slice of the work.
The author prefers tracking from task start through agreed acceptance, then observing subsequent defects. This includes delivery time, design, implementation, review, and rework human time. Wait times for environments or approvals should be recorded separately, otherwise process waits get attributed to coding ability.
Before comparing efficiency, confirm the same task type and the same acceptance criteria. A version that skips exception handling cannot be directly compared for speed with a fully validated version.
Human Intervention Worth Counting, But Not Just Counting Frequency
Continuing the import task: the agent forgot an agreed convention, human reminded once; found wrong test entry, human corrected once; claimed completion but didn't verify retry, human required补测 once. These repeated interventions are worth reducing.
If effective designs are easier to find, test entries directly usable, and completion conditions clear, similar tasks can avoid such detours. Here "Harness" refers to the tools, context, rules, and verification environment that support the agent.
But not all human participation counts as rescue.
Repeated reminders of confirmed rules, correcting same execution errors → Prioritize reduction, check common causes.
Confirm partial success meets business requirements → Necessary decision, focus on timely completion.
Review key designs, approve high-risk operations → Retain responsibility, improve basis and wait time.
Take over erroneous implementation and large-scale rework → Record time and impact, cannot count as just "once".
A single two-hour human takeover is not necessarily lighter than five brief confirmations. Counting only frequency misses the continuous background monitoring time.
Therefore the author observes "repeated rescue count" and "corresponding human duration" together, recording causes. The statistical unit must be consistent: one continuous handling of the same problem counts as one intervention, not by number of messages sent.
The goal is not to make the number drop every time, but to observe: under similar tasks and same quality requirements, whether repeated rescues gradually decrease.
If numbers are made to look good by underreporting issues or canceling checks, the result is only later-exposed risks.
Can One Improvement Let Others Avoid Detours?
Personal proficiency improvement has value. But if everyone still individually finds designs, tries commands, and adds the same checks, the team's repeated investment hasn't disappeared.
Still using the import task: one developer finds the correct test entry, solving only the current task. If the team organizes the entry so subsequent similar tasks default to retry checks, other members no longer need to re-explore—this investment starts yielding shared benefits.
The benefit depends not on how many files sit in a shared directory, but on how many applicable tasks actually benefit. Shared context—the business and project background needed for the task—must be findable, understandable, and correctly used; putting it in does not equal making it work.
Hypothetical ledger: using records of similar tasks as a baseline, estimate an improvement reduces 10 hours of repeated search and rework across a batch of tasks, while organizing, integrating, and maintaining actually took 6 hours. The estimated saved human effort for this phase is 4 hours. Saved time relies on a comparison baseline, not a directly measured value, but this hypothetical ledger provides more judgment basis than "ten people used it, so efficiency times ten."
This is only a human-time account. Model and infrastructure costs must be viewed separately. Delivery quality cannot be offset by saved hours. If rework time is already recorded in task human investment, do not add it again in shared benefits to avoid double counting.
Some improvements require large upfront investment and may not pay back short term. As long as applicable tasks keep appearing and maintenance costs are controllable, they are worth continued observation; if reuse scope is narrow and each integration requires the original author to modify, they are not suitable for rapid rollout.
Sharing also amplifies errors. If the import check hardcodes "allow partial success," it cannot be directly used for projects requiring full rollback. The more adopters, the more need to confirm benefits come from correct usage, not from everyone collectively missing a class of issues.
Small Team Just Needs a Task Table
No need to build a complex measurement platform initially. Pick a recurring, relatively stable task scope, keep current records, then improve one shared step, observe changes.
For import-type tasks, record the following:
Task scope, acceptance criteria, model and tool versions → Is there a before/after comparison basis?
Start, pass acceptance, and wait times → Did delivery get faster? Did the bottleneck just shift?
Human investment, repeated rescue causes and durations → What repetitive work did humans actually reduce?
Acceptance results, rework, and subsequent defects → Was quality traded for speed?
Shared improvement version, integration and maintenance investment → Do team benefits cover the new burden?
Observe a batch of similar tasks first, then adopt improvements on a small scale. Record not only successful tasks, but also failed, abandoned, and human-taken-over tasks; subsequent defects should use a consistent observation period—cannot compare just-completed tasks with long-running ones directly.
If both old and new approaches meet quality and safety baselines, they can be interleaved in similar tasks to reduce time-variation interference. Approaches already confirmed to leak checks or cause errors should be fixed first, not continued for comparison. When task volume is low, don't rush to calculate a precise "efficiency percentage"; first check whether repeated problems and long takeovers decrease.
Before/after comparison only provides clues. Model upgrades, developer proficiency increases, task simplification—all may affect results. If multiple conditions change simultaneously, clarify it's an overall work-style change, not attributable entirely to one guideline or tool.
Finally, see where saved time goes. Coding faster but results piling up in review queues means local improvement hasn't translated to overall delivery improvement; staff using freed time to add important tests may show quality improvement first, not quantity increase.
Summary
The author endorses observing AI programming via human intervention and shared benefits, but won't turn them into two scores everyone must chase. They are better suited to help teams discover: where repeated rescues still happen, which investments truly make subsequent work easier.
AI efficiency is not just one person completing the current task faster; it's about the team, next time on a similar task, under quality and responsibility guardrails, having fewer repeated rescues.
References
InfoQ, Fu Yuqi, Tina: "Stop fixing code, fix the system: DevOps father says organizational change in Agent era harder than technology", 2026-08-06.
https://www.infoq.cn/article/hLA2I6DD1v0ou0sE8KKB
This article borrows the two observation angles of human intervention and shared benefits from that piece. Measurement calibers, the import case, and the hours ledger are personal analysis and hypothetical examples, not completed team effectiveness evaluations.
Signed-in readers can open the original source through BestHub's protected redirect.
This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactand we will review it promptly.
Data Bricklaying Diary
Records practices, thoughts, and pitfalls on the data grunt-work journey, sharing content on data platforms, data analysis, data processing, data governance, knowledge graphs, and more. Less theory, more hands‑on, making complex data technologies simple.
How this landed with the community
Was this worth your time?
0 Comments
Thoughtful readers leave field notes, pushback, and hard-won operational detail here.
