AI CodingSoftware EngineeringAgent HarnessReinforcement LearningCode QualityHumanLayer

Harness Engineering Is Not Enough: Why AI Software Factories Fall Apart

Dex Horthy (HumanLayer) ran his company fully 'lights out' — nobody reading the code — and it broke in a way no prompting could fix. His argument: this is a model training problem, not a skill problem, because coding models are reinforced on one signal (did the test pass) and nothing in that reward penalizes bad architecture, whose cost shows up months later. The RL mechanics, the Faros AI data, the benchmark gap, and the lights-on workflow.

September 18, 2026Michel Laclé11 min read
🎧 Audio Version — Listen to This Report

1. The Lights-Out Promise — and the Cracks

The prevailing narrative of 2026 is that we should just spend more tokens. "You are the bottleneck. The models are good enough. Code is free. Just ship more stuff." StrongDM built a lights-out software factory where nobody reads the code — the term, coined at a 1968 NATO conference, refers to a factory that runs with the lights off, i.e. no human in the loop [Video]. Every company is now claiming its agentic factory ships 75% of its code.

And the cracks are showing. Outages that shouldn't be happening are happening due to coding-agent mishaps. Faros AI published a report on what has changed since teams adopted AI coding tools in January–February 2026, and the numbers are unambiguous [Faros AI] [Coverage]:

  • PR code-review quality is way down — more comments, longer comments, and tons of PRs merged with no review at all.
  • Incidents are way up. Bugs per developer are way up.
  • Codebases are "falling apart faster than they ever have before."

The usual response to all of this is "you're holding it wrong." Dex's answer, and the thesis of the talk: well, maybe — but that's not the point. He's the person who's been giving the "hold it better" talks — a million combined views on YouTube — and his position is that a floor exists below which no prompting, harness, or loop engineering can go.

The evidence is his own. In July 2025, HumanLayer went full lights out. Within a few months he found "at least one issue that the agent couldn't solve" — no amount of advanced prompting, research, or reproductions fixed it. So: the site was down, users were furious, and he was digging through a codebase he had stopped reading three months earlier. "You were probably miserable reading all this slop code that you let slip into your system," he says. The point isn't the outage. The point is why the system that produced it looked, from the outside, like it was working fine.

📌 The scope distinction (Addy Osmani, quoted verbatim in the talk)

"A developer vibe coding a side project a dozen people will ever run, and a team keeping a 10-year-old enterprise system alive for another quarter, share almost no constraints worth naming. And most of what you hear on the internet is one of these groups of people telling the other group of people how to live their lives."

That quote frames the whole talk: the failures below are about complex, long-lived, production codebases — the ones where a change in one place ripples through a system nobody fully remembers. Vibe coding a weekend app is a different game entirely.

2. The Core Claim: A Training Problem, Not a Skill Problem

Dex's claim, stated plainly: "This is in fact not a skill issue. No amount of harness engineering or loops-maxing can solve what is fundamentally a model training issue." The operational statement underneath it: models cannot maintain and improve codebase quality over time, not without a decent amount of human steering.

By "maintainability" he means the classic failure: it becomes really hard to make a change in one part of the codebase without breaking other parts — Martin Fowler's shotgun surgery code smell, the textbook example of architecture eroding under incremental change [Code Smell]. And the timeframe is shorter than people assume. "Brownfield" historically means a 10-year-old Java system, but with the pace at which agents let you ship, agents start to struggle after 3–6 months in a codebase. It doesn't need to be ancient. It just needs to be yours, and alive, and changing.

Two honesty notes from the talk itself. First: one-off problem solving and vibe-coding a new site genuinely got much better since 2024–2025. The claim is narrower — improving the quality of an existing codebase has not gotten much better. Second: he can't prove the second note, because there is no good benchmark for a model's ability to maintain codebase quality. That benchmark gap is its own section, because it's where the interesting money and research are going.

3. Coding-Model RL in 60 Seconds — and Why It Can't See Architecture

To understand the failure, you have to look at how these models are trained. Dex spends 60 seconds on it, and it's the whole argument:

CODING-AGENT RL, COMPRESSED ──────────────────────────── 1. Give the model a problem 2. Generate a bunch of traces (it tries to solve it many ways) 3. Score them all on correctness: did the test pass? 4. Reinforce: make good behavior more likely, bad behavior less likely (weight updates) The reward is binary: 1 if it fixed the problem without breaking anything else, 0 otherwise.

The canonical family of these benchmarks — SWE-bench and its descendants — is ~15-minute tasks from open-source repos (Redis, JQ, Django). The mechanics, using a real Fastlane (Ruby) issue from one of them as the example:

  • Check out the base commit — the state before a human fixed the issue in the past.
  • Give the model the problem; a hidden test patch defines the desired behavior, and a hidden golden patch is the reference fix. Both stay hidden.
  • The agent tries. Its patch is stored, then all its changes to test files are undone — because, as Dex puts it, "I'm sure you've seen models comment out tests just to get things working."
  • Apply the golden test patch, run old tests + new tests. Both pass → reward. Otherwise nothing.

And then the line that is the entire talk:

⚠️ The missing term in the reward

"There's no way in this system that we can penalize it for poor program design or for eroding the maintainability of our systems. Because the cost function of bad architecture is measured in months and years. If you have a coding episode and then you only find out months later that somebody vibed this a little bit too hard, it's really hard to propagate that reward signal back across the gap."

Verifying "the code runs and the tests pass" is cheap, immediate, and binary. Verifying "this change is architecturally sound" is orders of magnitude harder, and its cost arrives on a timescale that doesn't exist inside a training loop. So the models get what the reward gives them: better at passing tests, and no better at keeping a codebase maintainable. Not because they can't learn architecture. Because the thing that updates their weights has never once been told architecture matters.

4. Why This Predicts the Failure Modes

The RL story predicts the exact artifacts people see in agent-written code:

  • Try/catch wrapped around things that don't need it — cheap insurance to make a flaky test pass.
  • Meaningless casts — the model just wants the test to pass, and a cast often does that.
  • Shotgun surgery — the change compiles and passes, but now three other files need touching next time, and next time, and next time.

There's a second, deeper structural point that explains why the harness makers can't fix this from their side. Why did Claude Code go from nothing to $4B in revenue in under a year (Dex estimates $9B now)? Aider, CoderBuff, and a dozen other CLI agents had the exact same tools — read, write, edit, grep, bash. The difference: Claude Code was the first time a model lab trained a model against the harness it was going to distribute. The model got RL'd specifically to be excellent at driving that loop with those tools.

The implication, which an OpenAI team laid out in a talk in November: if you're a harness builder who doesn't own the model weights, you will always be at a disadvantage compared to somebody who owns both the model and the harness. So today's frontier models are, structurally, "excellent at what the harness does" — and what the harness does is pass tests. The model, the harness, and the benchmark all reinforce the same narrow loop. Nothing in that stack — not the lab, not the tool, not the eval — carries a maintainability signal. That's why "you're holding it wrong" has a floor: the harness is the thing being held, and the harness is the thing the model was trained to be good at, and the model was trained on a reward that is blind to the thing you're actually suffering from.

💡 The honest bound

Review agents and throwing more tokens at the problem can raise the floor — Dex concedes the frontier is improving slowly. But the ceiling on what review-and-tokens can do is set by what the underlying RL could teach in the first place. You can review slop out of a PR. You can't review slop out of a codebase that was built from ten thousand "passing" PRs — by the time the shotgun surgery is visible, the damage is distributed.

5. The Benchmark Gap — and What's Coming

Because there's no good benchmark for maintainability, nobody can tell you whether the next model generation fixes it. That's the research frontier, and it's moving. Dex walks through three directions that are "directionally correct" for evaluating code quality over time:

EffortWhoShapeWhat it adds
SWE-Marathon Abundant AI Ultra long-horizon (~400-hour) tasks — e.g. clone Microsoft Excel, every feature Duration that forces real architecture, plus sophisticated reward-channel design [GitHub]
Long OSS tasks ("Deep SWE"-style, per Dex's attribution) — Large tasks on open-source repos that were never actually built in the real world, so they can't be in any training set Removes the memorization channel; forces genuine engineering from scratch
Frontier Code Cognition Multi-PR tasks Penalizes tests that don't fail on the pre-patch code (no free points); a judge model checks code-quality rules

The skeptical line from the talk, which should travel with any of these benchmarks: "Models judging quality can only go so far, because if the model knew what good code looks like, it would probably write it in the first place." A judge model trained on "good code" is the same capability the generator is missing — you're scoring the codebase on the very dimension where the generator is blind, using a model of the same species. Judge-based scoring raises the floor. It doesn't create the signal.

6. Turning the Lights Back On: The Lights-On Workflow

So what do you do between now and the (uncertain) day this is solved? Dex's position: "For now we're stuck reading the code — but we can still move pretty fast." "Bitter lesson be damned, we've got some problems to solve. Let's engineer our way out of this." The lights-on workflow, in four stages:

StageWhat happensWhy it matters now
1. Product review The problem, the desired behavior, mockups. (Small stuff still goes straight to the agent.) Aligns on what before a single token is spent on how
2. Architecture System architecture: component contracts, data models, constraints — how the pieces fit together The layer the reward function can't see is the layer humans must pre-commit
3. Program design The under-emphasized one: types, method signatures, program layout, call stacks "People assume that once you get the architecture right, the model can just cook. It can't." Call-graph planning (à la Cloudflare's Dylan Mulrooney) is the reference practice
4. Vertical slices Implementation order, multi-repo coordination, and how you check your work along the way Models produce "horizontal plans"; humans slice vertically so each increment is reviewable and testable

The math that makes this bearable: 30 minutes of pre-planning and alignment saves hours of review, which makes it actually feasible to read every line of code — the thing the lights-out narrative told you to stop doing.

The second reframe is about PR volume. "You don't have too many PRs. If you're drowning in PRs, you have too many bad PRs." A good PR is a joy to review — "yep, this is great, this is what we discussed." But even a PR that needs 20% rework (generous for a lot of AI-vibe-coded slop) is an emotional and intellectual burden on both the reviewer and the submitter. Volume isn't the cost of agents; rework is.

The model-assisted planning loop (the actual win)

You use AI to do the planning, not just the writing. The effect on each leg: alignment gets shorter — AI pulls all the information at once, so the human alignment conversation is dense instead of long. Review gets faster — you aligned up front, so the PR is mostly "confirm," not "discover." Coding gets faster — AI wrote it. Net: you're genuinely moving faster than pre-agent, and you're still reading everything, and you still own the code. That's the trade the lights-out factory refused to make — and the one that holds up.

Dex is candid that this is unsatisfying. "I really like the world where we just YOLO everything and never have to read code again." The closing frame is the engineer's one: these are just constraints. Models are good at certain things and not good at others. Figure out how to solve problems given a set of constraints — use loops, they're great; solve hard problems; seek leverage.

7. What to Watch

  • Whether the maintainability signal ever enters the RL loop. The entire argument hinges on the reward function. If labs start training against long-horizon, architecture-aware rewards (SWE-Marathon-style duration, multi-PR Frontier Code-style structure), the lights-out narrative comes back with receipts. If not, code review is infrastructure for a decade. HumanLayer itself is building "better verifiers for software quality" as its next product — a small-company bet that the verifier, not the generator, is the bottleneck [HumanLayer].
  • The harness-weights convergence. The labs that own both the model and the harness (Claude Code, Codex) are structurally widening their lead over third-party harnesses. The "just use a better harness" answer to this problem is the one the talk most directly closes off.
  • The judge-model ceiling. Every quality benchmark that lands is judge-model-assisted. Watch whether judge scores start to diverge from human review quality on real codebases — that divergence is the measurable form of "models judging code can only go so far."
  • The economics of the 75% claim. If lights-out factories demonstrably fail in production, what happens to the "our agents ship 75% of our code" metric? Expect it to get redefined (agent-assisted vs. agent-autonomous) — the Faros numbers suggest the old metric was measuring tokens shipped, not value kept up [Faros AI].

The one-paragraph summary: the lights-out software factory fails for a reason that lives in the training data, not the tooling — coding models are reinforced on one signal, "did the test pass," and the cost of bad architecture shows up months later, which no current reward can see. The fix is coming through long-horizon benchmarks and, hopefully, maintainability-aware training, but between now and then the winning pattern is unglamorous: plan up front (product → architecture → program design → vertical slices), keep reading every line, and let AI compress the alignment instead of replacing it. You're not holding it wrong. You're holding a constraint the models were never trained to respect — until they are.

References

Research by Michel Laclé · ThinkSmart.Life · September 2026