close

DEV Community

Cover image for The Epistemology of Quality
Peter Harrison
Peter Harrison

Posted on

The Epistemology of Quality

Should we expect human software developers to review code written by AI? For many the answer is clearly yes: how else will we maintain accountability? Few have challenged the thinking, but many of the advantages of past code review practices have evaporated in modern software development.

Underneath the question is another one: how do we know code is correct? For a simple function we can sometimes get close to proof. The inputs are bounded, the behaviour is specified, and we can test exhaustively or reason about it formally. A real system is different. It has users, requirements that were never fully written down, dependencies, concurrency and failure modes nobody anticipated. We cannot prove it correct. What we have is evidence, and the question is how much each kind of evidence is worth.

In this article I trace how code review evolved, what evidence each stage actually provided, and where it fell short. Then I set out the kinds of evidence available in a modern development process and the limits of human review among them.

In Person Code Review

Code review is not a single practice. It has changed over time, and perhaps not for the better. When I worked for an insurance company they had a stringent process where code was printed out for each developer and marked up with red pen. The review was conducted in a meeting room with all the developers, with one section of code under review.

This let all the developers see what was being written. However, the reviews were more about syntax checking and coding standard enforcement than teaching. It felt like a hostile interrogation where every last standard violation would be raised.

The benefit was that we got into a room and discussed code and cross-pollinated ideas. Sometimes a developer would contribute a better implementation. What we didn't have was unit tests, or any discussion of the purpose of the software and whether it met requirements. A separate testing team did that.

The evidence: the review told us the code conformed to the standards and that several people had read it. It said little about whether the code was correct. Human readers are good at spotting style violations and poor at finding defects that only appear when the code runs.

In later projects I decided this was too time intensive. Without a separate test team we made developers responsible for quality, including ensuring their code was free of defects. We introduced unit tests, automated coverage analysis, and compulsory review by a second developer before commit. In practice the reviewer sat next to the developer and reviewed the code on the developer's machine, which was possible because we all worked in the same office.

Developers also paired up to discuss and review code long before commit. Collaboration was the norm. Pair programming takes this further, making it continuous: less a pre-commit audit and more active, ongoing feedback.

A review in this model had the developer explain the feature, show it working, show the unit tests covering it, and only then show the code. Reviewing the code was one part of the review.

The evidence: this was much stronger, and the reason is that it combined several different kinds of evidence. A person watched the feature run, tests passed, coverage was measured, and a second pair of eyes read the code. Reading the code was the weakest of those signals for finding defects. When done by someone who understood the architecture, it could add judgment about structure that tests don't see. Review as ordinarily practised rarely did.

Git and Geographic Isolation

Open source created a challenge for source control. People were no longer co-located, and sitting next to a colleague before commit was not possible when contributions came from around the world.

Git encouraged separate branches that could be merged, and the branch became the unit of review. The Pull Request, where a branch is merged into main, became the place review occurred. A diff identified the changes and was easily displayed in web interfaces such as GitHub.

What began in open source has become the near-exclusive approach in the corporate world too. Tools automated certain gates: running unit tests to catch regressions, coverage reports to maintain test coverage, and systems like FindBugs to find common defects.

These tools addressed a failure in manual review. Eyeballing code was good for developer education and alignment, but hideous at finding defects. Human review was already being superseded by automated tooling.

The PR model also broke something. It removed the side-by-side discussion of implementations, which I consider perhaps more important than review as a quality gate, because it educated developers about what quality looks like. It also became too easy to glance at a PR and approve it without checking that the code would even compile.

The evidence: the PR gave us a diff, which is the least informative view of a change. The real evidence moved to CI/CD: build, unit tests and static analysis on every commit. In effect we were already relying on many forms of evidence other than human eyeballs.

AI and the Review Crisis

So why, in the age of AI, are we told that human review is so critical? Why do AI governance approaches specify mandatory human code review?

Part of the reason is trust. LLMs hallucinate, meaning they make things up, or less charitably, lie. I have been working with AI coding for some time and have seen its failure modes. The most concerning is that, given deterministic tests, it will write code that passes the tests but breaks the spirit of what the tests intended. This is overfitting: code written to pass the test rather than achieve the goal.

If there are ten tests and they are all green it looks like a working system, and you might ship it, only for it to be obviously broken for anyone using it.

I have addressed this in two ways. The first is heretical: manually running the app and interacting with it. It is the proof-is-in-the-pudding test. It is not rigorous or objective, but it is a minimum bar. The second is more diverse unit tests. Usually you would have one test per function, but several tests with different data, even where the program flow is identical, make it harder for the code to cheat a single test.

Overfitting reveals the deeper problem. If the same model writes the code and the tests, a misunderstanding of the requirement can end up in both. The tests pass because they encode the same mistake as the code. The green result looks like evidence when it is really one opinion stated twice. This is the central issue in AI-assisted development: not that AI makes more mistakes, but that its mistakes can be correlated with the checks meant to catch them.

What Counts as Evidence

Any single check is evidence, not proof. A passing unit test says the code behaves correctly for the cases the test covers. A clean static analysis says it avoids the patterns the analyser knows about. Neither says the software does what the user needs. So the question is not whether any one check is reliable, but what makes a collection of checks a reasonable basis for belief.

My answer is independence. Two checks that fail for the same reason add little. If they fail for different reasons, a defect has to slip past all of them at once, and confidence grows quickly with each one added. The object is to employ enough independent safeguards to have reasonable epistemic belief that the requirements have been met and the solution is reliable.

Independence is not binary, and it comes in three kinds. There is independence of the author: a different model, or a different person. There is independence of the method: running the application rather than reading its code. And there is independence of the oracle, the account of what "correct" means that the check is measured against. The third matters most. A perfectly independent test of the wrong requirement still checks the wrong thing.

Here are the kinds of evidence available in a development process, and what each can and cannot tell you.

  • Build and type checks. The code is syntactically valid and internally consistent. Nothing about behaviour.
  • Unit tests. Behaviour is correct for the cases covered. Weak against overfitting and shared misunderstanding if the same author wrote code and tests.
  • Coverage analysis. Shows which code the tests exercise, not whether they check anything meaningful. High coverage with weak assertions is common.
  • Functional and integration tests. Components work together. Stronger than unit tests, but still written against someone's understanding of the requirement.
  • Manually running the application. Catches the obviously broken. Not repeatable or exhaustive, but fails in a completely different way from automated tests.
  • Static analysis (FindBugs and similar). Known defect patterns are absent. Blind to anything outside the rules.
  • Cyclic dependency and complexity analysis. Objective measures of structure: coupling, cycles, cyclomatic complexity, duplication. They say how hard the code will be to change, not whether it works.
  • Vulnerability and security analysis. Known classes of weakness are absent. Specialised, and a different failure mode from functional tests.
  • Load tests. Behaviour under stress. Nothing else on this list covers it.
  • AI code quality review. Useful for simplicity and readability, which are hard to measure. Only independent if it doesn't share the blind spots of whatever wrote the code, so a different model and different prompt help.
  • AI acceptance testing. An agent exercises the application as a user would, for instance through a browser. Strong because it checks behaviour from outside the code, but only as good as the acceptance criteria it is given.
  • Production observation. Telemetry, error reporting, canary releases and rollback. Some evidence can only be gathered after deployment, so quality is a continuing process, not a judgment made at the PR boundary.
  • Human code review. Covered below.

Quality needs the same treatment as correctness. Tests say the code works today and nothing about how hard it will be to change tomorrow. The most objective test I know of is a change test: give a fresh agent a realistic new requirement and measure how it goes. Measure the effort needed to make the change correctly, the regression rate, and how much unexpected coupling the change exposes. The result partly reflects the agent as well as the code, so the test is best used comparatively, running the same change against alternative designs. The structural metrics above support this, and unlike readability they are not matters of taste.

The Limits of Human Review

This is not a case against human involvement with source code. Even if code is human authored, I think we should have objective measures of quality rather than subjective review. By all means humans can interrogate the code. The question is whether mandatory human review is a systematically valid way of ensuring quality, because as those responsible for delivering systems we are accountable for the result.

Do humans find defects in review?

Senior developers review junior developers' work because new developers need to earn trust, and review has been part of that. But AI agents are no longer mainly caught out by obvious errors. They still make them, but they now also routinely produce code that looks plausible, compiles and passes tests while containing a subtle conceptual error. Obvious mistakes are cheap to catch. Plausible wrong ones are exactly what a quick read of a diff misses.

Eyeballing a diff on GitHub is not conducive to finding subtle defects, or to understanding the effect on the software as a whole. To be sure code is correct, the reviewer must understand the requirement, the overall solution, and then the detailed implementation. That can take nearly as long as writing the code.

Telling reviewers they are responsible for the defects in code they approve misunderstands how hard certainty is. The practice delivers a single, weak form of evidence, and it holds one person accountable for a judgment they lack the time and means to make.

What is the harm?

AI is meant to make development faster and produce more code. Putting human review in the loop creates a bottleneck: AI generates at machine speed while humans cannot act as a meaningful quality gate at the same rate. Expecting people to review AI work as fast as it is created is unreasonable. It adds stress, and in my experience the defects it finds rarely justify the time, compared with spending that effort on acceptance tests and on exercising the running system.

Where Humans Fit

Every check above verifies the code against a reference: a test, a standard, a specification. None of them can establish that the reference itself is right. What the software is supposed to do isn’t a fact waiting to be discovered, it’s a decision, and it can’t be proven. It can only be stated, agreed, and then tested against real users and real conditions. Something has to supply that reference, and it can’t be the generator checking its own work.

That is properly a human responsibility. People define what "correct" means through requirements and acceptance criteria, and they set the standards the other checks enforce. This is a different job from reading a diff, and a better use of scarce attention, because it is the one check no tool can perform for us.

Retaining human responsibility doesn't require them to read every line. It requires them to own the decisions. Human responsibility gets clearer once it moves from approving diffs to determining the quality criteria. This is a matter of authority and accountability, not superior technical ability.

The checks should also run as early as possible. With AI development, the agent should have the reviewers available as tools while it works: linters, tests, dependency analysis and security scans run after every change, so failures are fixed before a human sees anything. The PR becomes the final gate, a summary of what each independent reviewer found, not the first time anyone looks.

Traditional review confuses accountability with evidence of correctness. An approved PR says nothing about what the reviewer understood, what they checked, or which defects they could reasonably detect. Only organisations treat the approval as proof the code is satisfactory. AI exposes the weakness because it pushes code volume beyond what anyone can meaningfully inspect.

The alternative is a different shape of accountability. Instead of a person signing off on code they cannot realistically verify, the organisation stands behind a set of independent checks and a human-defined account of what correct means. Accountability moves from the person who clicked Approve to the people who design and maintain the assurance process. When something fails, the question is which check should have caught it and why it didn't, and that question can be answered and acted on.

Top comments (0)