A senior engineer opens a PR from a more junior teammate and finds code that does the right thing the wrong way. Too many useEffects fighting React’s lifecycle, that kind of thing. They know exactly what is wrong, write the review comment, and then have to decide whether it is worth writing again next week.
Code that does the right thing the wrong way is genuinely wrong, and it does require a human to intervene. The question is where.
Where to Intervene
The volume of plausible-looking-but-wrong code is running ahead of what any reviewer can keep up with, and reviewing faster does not fix that. So the intervention has to move upstream as well, into getting the agents to build things better in the first place.
The Guardrail Stack
Each piece of senior judgment belongs somewhere on a stack, and picking the right rung is most of the work:
- Linters, at the cheap end, for anything mechanically detectable.
- CLAUDE.md and Cursor rules, in the middle, for conventions an agent should carry into every task.
- Specialized sub-agents, for judgment a lint cannot express, where a pattern has legitimate uses but telling those from the rest takes reading the situation.
- PR checks, as hard gates, for the things that must never merge.
- Scaffold-level constraints baked into the agent harness itself, so the wrong shape is hard to produce at all.
Two things seem to decide the rung: how much of the call is mechanical versus contextual, and how bad it is to get wrong. The useEffect example sits in the awkward middle, because a linter probably catches some of it and not all of it, which is precisely the case for encoding it as a sub-agent or a rule.
Rolling Out a New Rule
This is not about juniors reading the CLAUDE.md. They are not reading it.
The pattern is two-part. The senior ships a PR with the new rule, then sends a message to the team: here is the thing I keep seeing, here is what it should be instead, here is the rule I added to encode it, and here is the backing evidence for why the new way is better. The PR is the durable enforcement. The announcement is the transferable explanation, and it names the pattern and the why in a place people actually read.
You need both halves. Merge without announce gets you enforcement without buy-in. Announce without merge gets you buy-in without enforcement.
Security Already Had This Argument
The security discipline has been through this and largely converged on shift-left, off the same pain shape: a small group of specialists holding a manual gate at the end of the SDLC, with more code arriving than any gate review could absorb and no practical way to assess whether it was safe.
My read is that you stop being the gate and embed the thinking throughout the lifecycle instead. Slash commands. Plugins and skills developers run locally. Branch reviews, PR checks, ticket-level checks. Supplementing wide-scale security review, so that by the time something reaches the gate most of the obvious problems are already gone. Senior engineering judgment looks like it is on the same arc, one cycle later.
What Stays Human
We are effectively training the AIs instead of training the individuals, and that is a little soul-destroying if you got into this to develop people.
The counterweight is that the human contribution I would argue matters most is building the right thing and making sure it has the right behavior. The judgment that decides whether something has the right architecture is still the human’s, and that is not going away. What changes is how that judgment reaches the code.
On style specifically, I do not think it goes by the wayside in the next six months, because humans still read this code at review time. But just like JavaScript compiles down to assembly and nobody looks at the assembly anymore because the compiler is trustworthy, I think we are going to be somewhere similar in the next one to two years, where if it works, it works, and people do not read the code as closely. Style survives in that world as a guardrail concern, because consistency reduces the hallucination surface for the next agent that reads the file.
Some senior judgment is not encodable. Architecture calls that need whole-system context are still going to come out of a human’s head. Maybe most recurring code-review judgment is encodable and the rest stays human, and the uncodable part is not a reason to skip the encodable part.
This is the IC-side version of a shift I have written about elsewhere, where engineering’s job moved up a level from tasks to epics. Here the move is from PR-reviewer to guardrail-author, and the reason to make it is arithmetic: catching a mistake in review scales with reviewer hours, and encoding it once scales with generation volume.