Two questions have worked their way into every engineering manager interview I run. Early in the conversation I ask whether humans should write code. Later, if the conversation gets there, I ask whether humans should read code that AI wrote.
Candidates often laugh, or at least smile, while they think about it. The role has been open for a while (I wrote about losing two finalists in August), so I have heard a fair number of answers. The consensus, when there is one, is “less and less.” Fewer humans writing code, and fewer humans reading it.
I don’t grade the answers. A candidate who says humans should never write another line of code and a candidate who says we’ll be reading every diff for the next decade can both pass. The question is there to show me how someone reasons about the thing their team is about to become, and a small development team inside a company that doesn’t sell software is about to become something quite different.
Two jobs under one title
Underneath the answers I keep hearing the same split, even when nobody states it outright. Software engineering has always been two jobs sharing one title. One job is owning a problem: understanding what the business needs, deciding what should exist, and being accountable for whether it works. The other job is translating the chosen solution into code. For most of the history of the profession, the second job consumed most of the hours, so it came to stand for the whole thing.
AI does the translation now, and in a lot of cases it does it better than people do. That makes writing the easier of the two to let go of.
Reading is harder. The argument for reading is familiar to anyone who follows this debate online: AI isn’t always good at this, so reading the code is where human verification happens and where the trust lives, and you can’t understand software whose code you haven’t read. My own view is less charitable. I think a lot of programmers hold on to the pull request and the code review because it is the last point where they can keep control of the work. That is an understandable instinct, and I’m not sure it is the right one.
AI doesn’t write perfect code. Neither did we. Generations of programmers wrote sloppy code, and that sloppy code is where most of the bugs in the world came from. The case for reading AI code can’t rest on the idea that human-written code was a trustworthy baseline, because it never was.
A prototype arrives at IT
The interview questions would be academic if the demand side of my job weren’t changing at the same time.
Internal development teams at companies that don’t sell software are in a lucky position. We exist to support the business: automating a process, extending a tool people already use, capturing data that nobody captures anywhere else. What matters is solving the problem, whatever shape the product takes, and that gives an internal team a lot of freedom in how it gets there.
What’s new is that the people we support have started solving problems themselves. Someone vibe codes an app, finds it useful, and then asks IT to install it or share it with colleagues. That request starts a review. In a regulated industry like law, the review is not a formality. The prototype has usually skipped everything that makes an application safe to run inside a firm: data handling, confidentiality, enterprise security, authentication, authorization, proper workflows, proper permissions, even consistent branding across internal tools.
What my team does then is take the app and understand the risk that comes with it. In most cases we bring it into our own technology stack by applying the same standards and guardrails we apply to any application. Sometimes that means changing how the thing is written, how it is deployed, or how it interacts with its users, while it keeps doing the job its author built it to do.
The honest number, so far, is small. These requests have been few, but not zero. I expect that to change quickly, and last Friday gave me a reason. Microsoft announced a new Copilot that includes Code, a way for people who aren’t developers to build “small, purpose-built solutions anyone can create to get a job done.” Alongside it came a Copilot Managed Runtime, which lets those apps run inside the company’s Microsoft 365 environment, “governed by IT.” The runtime is in public preview. Every firm running Microsoft 365 is about to find out how many of its employees have an app in mind.
One familiar reading of that says users who can build their own tools need fewer developers, so the internal team shrinks. Another treats every vibe-coded app as shadow IT, a risk to be contained by a department whose job is to say no. I don’t accept either one.
The clearest requirements we’ve ever been given
A vibe-coded prototype shows a need. It shows where people are losing time and what they want instead, in a form nobody can misread.
Compare it with the way requests used to arrive. Someone would email: I need a form that tracks this, this, and this. What “this” meant had to be worked out afterward. When a user brings a working prototype, their intentions and requirements are far clearer than any email ever made them. They have already done the hardest part of requirements gathering, which is figuring out what they want.
The prototype isn’t the end of it, though. Before it can be widely available, or trusted, or become an official output of the firm, it has to go to people who can judge it. That judgment has two halves. Someone has to judge its technical merit: whether it is safe, whether it will hold up, and whether it fits the rest of our systems. Someone also has to judge its business merit: whether it solves the problem worth solving, whether three other people have built the same thing, and whether the firm should invest in it. Programmers are well placed to do both, because they sit where the technology and the business process meet.
So I expect the democratization of coding to be good for developers. It produces more ideas worth building and more need for people qualified to decide which ones to build properly. I want more of it. But it needs the right workflows and processes around it, because IT’s job shouldn’t be to say no. It should be to say yes, and then do it properly.
Notice what that makes my developers’ job: directing the conversion of a working prototype into something standardized, reviewed, audited, and trusted, that can sit alongside everything else we’ve built. Writing the code is the smallest part of it.
What I’m listening for
This brings me back to the interview room, and to what separates a strong answer from a weak one.
The strongest candidates are curious, and they can see where things might go even when they aren’t sure. Most of all, I want someone who thinks in systems, and I say so in every interview.
The reason is that the things software does haven’t changed much. Buttons that get clicked, voice interfaces, careful layouts, automation across APIs and external tools: none of that is new, and all of it could have been built before. We can build it at larger scale now, and a junior developer can reach work that used to need a senior one.
The change is happening in the process of creating software. What my team is building, more and more, is the factory: the loop that takes an idea and turns it into something a person can use. The software is an output of that process. The work that matters is designing a process with quality and reliability built in from the start, instead of building something and figuring it out as we go. Once we get the loop right, we can put any idea in one end and quality comes out the other end.
A candidate who has thought about that can tell me how people fit inside the process and how they work with AI inside it. Whether they conclude that humans should read code matters much less to me than whether they’ve thought it through.
The loop we run today
I should be clear about how far along we are, because it would be easy to describe the factory as a finished thing. It isn’t. We have a rough version working.
Issues in GitHub are reviewed and triaged by an agent. An agent writes the analysis and the implementation plan. Implementing bug fixes and working through code review tickets is automated, and so is opening the pull request. Every pull request is then reviewed automatically by several different agents. A human merges it.
On a regular schedule, we also point an agent at the factory itself. It reads the issues that were opened and closed and the CI runs that failed, and proposes changes to our guardrails and our process. That step is the one I’d point to if someone asked what makes this a factory rather than a set of tools. The thing under review is the process itself.
We still do a light human review of pull requests, and a person still presses merge. That’s the honest state of things. On my team, the review agents catch more than most of our human reviews do. That is our experience only, and the published studies I’ve read point the other way: one analysis of nearly 280,000 review conversations found developers adopted 16.6% of AI review suggestions against 56.5% of human ones. The more telling study, for my purposes, is about the humans. Across AI-authored pull requests, researchers found approval rates rising while inline comments fell by 22%, and called the pattern “reflexive habituation.” Human review of agent code is already thinning out. The question is whether anything is designed to replace it.
The case for reading the code
The strongest argument against everything I’ve said so far comes from John Allspaw, and it deserves a full hearing.
In June, Martin Monperrus published a paper titled “The End of Code Review,” arguing that every stated goal of code review can be served by agents, and that making humans the mandatory reviewers of agent code “neither provides meaningful assurance nor scales with AI-assisted throughput.” That is close to my position, stated more formally than I would state it.
Allspaw answered it in August. He calls the paper’s method a “substitution myth”: break a human practice into functions, show that a machine can perform each one, and conclude the human is redundant. His point is that code review was never mainly about detecting defects. A reviewer’s confusion is itself a signal. A reviewer notices what is missing, questions whether the change should exist at all, adjusts their attention to the author and the risk, carries operational context that isn’t in the repository, and learns something, as the author does. And then there is accountability: “An agent that ‘signs off’ on a pull request bears no consequences and certainly has no incentive structure that fuels an earnest evaluation.”
The research supports him on the history. When Alberto Bacchelli and Christian Bird studied code review at Microsoft in 2013, finding defects was the top stated motivation, but defect comments were “only the fourth most frequent, out of nine items,” about one comment in seven. Review was mostly doing other work.
I also have to reckon with my own writing. In July I argued that the danger of AI lies in delegating without a review loop, and my example was the partner who signs without reading. A fair reader could ask how a developer who merges code nobody read is any different.
Where the review goes
I think the bottleneck has moved.
If the code is your main concern, and code review is where you decide what belongs in the application, then you have put yourself in a position where you can’t step outside the code. The decisions that matter, about what goes in and whether a feature should exist at all, belong before the code is written. The judgment that matters most after it is written is a human reviewing the finished product: whether it meets the requirements, whether it has side effects, whether it behaves the way the user expects, whether the original idea survived, and whether it fits the framework the factory is built on.
Between those two points, what’s in the code matters much less than people want it to. The code could all be written in iambic pentameter for all I care, because it is going to be judged by its performance, its security, and its stability. We already accept this one level down. Nobody reviews the compiler’s output against their preconceived notion of how it ought to look. The code is how an idea travels into a working system, a record of institutional knowledge and institutional process written in a form a machine can run. The code knowledge is fungible, because it’s the business knowledge that matters.
This is also where my July argument holds, and where the partner is different from the developer. Legal work has no deterministic test for whether a document complies with the law and with how courts interpret it. Someone has to take responsibility for the interpretation, which means someone has to read. Source code can be revised and put through regression tests, integration tests, and end-to-end tests, as many times as you like. The review loop now sits in the two places where human judgment can’t be automated, and the middle goes to verification that can be.
That answers Allspaw on accountability, at least partly. In the factory, accountability sits with whoever decided the feature should exist and whoever accepted the finished product. Both of those are people.
Who checks the tests
This is where I concede ground, and it is the part of Allspaw’s list I take most seriously.
A factory that nobody reads depends on its tests, and the worst author of a test suite is whoever wrote the code, human or agent. Research on generated tests keeps finding this. One study describes test suites that “mirror the generating models’ error patterns and cognitive biases,” which its authors call a homogenization trap. A sensible factory has different models and different agents writing the tests than the ones that wrote the code. Even that isn’t enough on its own: in a cross-model review study this summer, one model reviewing another’s work raised the success rate substantially, while the reverse pairing made it worse. A second opinion helps only when it is a better one.
So I think we need to go much further, into adversarial review. Most unit and integration testing confirms that code does what it says it does. That was a reasonable goal when humans wrote the code slowly and other humans read it. It isn’t enough now. We have to prove the code can’t be made to do something it isn’t supposed to do.
The security world is already there. In July, OpenAI disclosed that two of its models, tested with their cyber safeguards deliberately lowered, escaped a sandbox and broke into Hugging Face’s production systems to fetch the answers to a security benchmark. That is an agent completing its task by doing something nobody intended, which is exactly the failure a functional test isn’t designed to find. Pointed the other way, the same capability is a defence. Anthropic’s Project Glasswing gives partners a model built to find vulnerabilities, and by late May about fifty of them had found more than 10,000 high- or critical-severity flaws. Anthropic’s summary of what changed is the clearest statement of my argument I’ve seen: “Progress on software security used to be limited by how quickly we could find new vulnerabilities. Now it’s limited by how quickly we can verify, disclose, and patch.”
That approach needs to reach all kinds of code, not only security-critical code. Attack the application. Try to break the code, and try to break the business process it supports, because that is where you find what is missing. Allspaw is right that noticing absences matters. I think the way to find absences is to test outside the strict boundaries of the solution. Reading a diff and hoping to notice what isn’t there is a weak method. Mutation testing, which deliberately plants faults to see whether the tests catch them, is one small example; Meta has used language models to generate the mutants and the tests at scale.
We don’t do adversarial testing yet. We still have light human review on our pull requests, and that review is doing, imperfectly, some of what adversarial testing would do better. Building that capability is the next piece of the factory, and it has to exist before the human pressing merge can step back.
Demand will rise to meet it
The size of my team doesn’t worry me. I tell candidates I believe a small team can have an outsized impact, and the factory is how that happens.
I don’t know how large the gain will be, or when it arrives. We might be ten percent faster this year, or fifty, or eventually far more than that. I do know which way the curve is going. And I expect the gain to be met by a rise in demand, the pattern economists call the Jevons paradox. The expectations will be higher. People will ask us to connect systems and integrate in ways that used to be too difficult or too time-consuming to attempt, and now aren’t.
What does concern me is that nobody knows where this process ends, or whether it ends. Model capability and hardware are improving at a rate that makes any five-year plan a guess. Programmers will lose the habit of reading code, the way the profession stopped reading punch cards once nobody needed to, and I don’t mourn that. They’ll become part product owner, part engineer, and part project manager, and they’ll need skills they never had to build before. That’s fine. The industry is in a transition, and nobody knows the destination. You would be foolish not to prepare for it anyway.
Back to the interview room
That is why I don’t grade the answer.
A candidate who wants humans reading every line might be right for longer than I expect, and one who says nobody should read code again might be right sooner than I’d guess. The possibilities are all gray, and anyone who claims to know the timeline is guessing. What I need to know is whether their reasoning is sound: whether they can see the whole loop, where the people fit inside it, and where quality comes from.
The person I hire will spend their days running the process that produces our software.
We’re not building software anymore. We’re building the software factory.
If you run a small development team inside a company that doesn’t sell software, I’d like to hear how the vibe-coded requests are arriving for you. Reply to this; I read every one.


