The Factory Mental Model
Every serious AI coding setup is secretly a factory, and it has three moving parts. If you can name all three, you can start treating the thing like the engineering system it actually is — instead of a magic box that occasionally gets worse.
| Part | What it is | In this setup |
|---|---|---|
| Harness | The machine: reads the repo, plans, writes code, runs commands, reviews its own work | VS Code in agent mode [docs] |
| Skills | The playbooks: procedural instructions the agent loads on demand | Your software-development skill repo (the SKILL.md Agent Skills format) [standard] |
| Model | The worker: the LLM that executes | A pinned open-source model served locally [LM Studio] |
The skills are the part most people under-think. A skill is a folder with a SKILL.md — a name, a description, and procedural instructions the agent reads when the task matches. Anthropic published Agent Skills as an open standard specifically so these playbooks are portable across harnesses [link]. VS Code now loads them natively through its own Agent Skills customization surface. So a skill is real, versioned, load-bearing software — the kind of thing you'd want a regression test around.
The reframe that drives everything below: stop treating your AI setup as "a model with a nice UI." Treat it as a system you own — harness, skill set, worker — and test it the way you test any system you own.
The Gap: We Benchmark the Worker, Not the Factory
In most setups the worker is a hosted API — GPT, Claude, Gemini. And that worker is not fixed. It gets updated underneath you, usually with no clear changelog you'll read. Same prompt, same task, different answer six weeks later. Meanwhile your skills are changing every time you learn something and add or tweak a playbook, and the harness ships new versions with new tools, new defaults, and new behavior.
So three things are drifting at once. And the only one you're actually measuring is the worker. That's the gap.
The public benchmarks are excellent — and they all answer the same single question:
| Benchmark | What it holds fixed | What it varies | What it measures |
|---|---|---|---|
| SWE-bench [site] | Task set (2,294 real GitHub issues; 500-instance Verified subset) | The model / agent | % of issues resolved |
| Terminal-Bench [site] | Task set + sandbox, verified by a test script | The model / agent | % of end-to-end terminal tasks passed |
| Aider polyglot [leaderboard] | 225 Exercism exercises across 6 languages | The model | % correct edits |
The pattern is deliberate and correct for its purpose: hold the task fixed, vary the model, rank the models. But that's the opposite of the question that keeps a practitioner up at night. These benchmarks answer "how good is this model at coding?" They cannot answer "did my setup get worse this week?"
- Did the skill I edited on Tuesday quietly make the agent worse at code review?
- Did the harness update change how it handles a failing test?
- Did a model refresh change the output shape my downstream tooling expects?
A benchmark that only varies the model can't see any of that, because it holds your playbooks and your harness frozen at some version you stopped using months ago. The SWE-bench entry for your model is a leaderboard line, not a regression signal for your factory.
Invert the Experiment: Pin the Worker
The fix is to flip the independent variable. In a model leaderboard, the model is what you sweep and everything else is held constant. In a factory regression test, you do the opposite: lock the worker and sweep the harness and skills.
Concretely, pin three things:
- The weight file. A specific open-weights model, a specific artifact, not "latest."
- The serving config. Same inference engine, same quant, same context length, same
temperature. - The sample regime. A fixed seed /
temperature=0where the runtime allows it, so a rerun is as close to token-identical as the model gets.
Once the worker is fixed, any delta you observe in the output can only come from the harness or the skills. And those are exactly the two things you actually author and control. That's the entire leverage of the inversion: you've converted an unattributable, drifting number into a controlled A/B test where the independent variable is a git commit to your skill repo.
Why not just pin a hosted model version?
You can point at a versioned endpoint, but it still sits behind an API boundary where the operator can deprecate, silently re-quantize, or change routing between the "same" version. A hosted version is a named moving target. A weight file on your disk is a byte-identical object you can checksum. The pinned local model is a fixture; a hosted version is a dependency.
Why Open-Source and Why Local
The choice of a local, open-weights model is doing more work than "I don't want to pay per token." It's the thing that makes the control variable real.
- A weight file doesn't drift. It's the same bytes today as it was last month. You can
sha256sumit, store the hash in your skill repo next to the skills, and assert it hasn't changed in CI. That's a reproducibility invariant a hosted model can't give you. - You can version it with the skills. Commit the model hash into the same repo as the skill set. Now a "known-good" state is one git ref: skills and worker pinned together, the way you'd pin a test fixture.
- You control the sampling. With your own runtime you set temperature, top-p, and (where supported) the seed. With a hosted API you request them and hope for compliance.
- It's fast and free to replay. A regression suite has to run on every skill PR. A local model on your own GPU means the bench is as cheap as your electricity, not a line item that scales with your commit frequency.
Practically, this is just running an open model through a local server that exposes an OpenAI-compatible endpoint — LM Studio, vllm, llama.cpp, or MLX — and pointing the harness at http://127.0.0.1:… [LM Studio] [VS Code models]. The model is now a local process you can pin, checksum, and restart deterministically. That's a test fixture, not a colleague with a bad hair day.
Building the Bench
Now build the software bench. It has exactly three components, borrowed directly from the SWE-bench / Terminal-Bench design and pointed at your own skills.
1. A task suite scoped to your skills
Not generic benchmark instances you'll never touch again — tasks that map one-to-one to the playbooks you maintain. If you have a test-driven-development skill, a code-review skill, a writing-plans skill, and a systematic-debugging skill, your suite should have tasks that exercise each of them. Small enough to run on every change; specific enough that a drop is attributable.
| Task | Skill exercised | Oracle (known-good outcome) |
|---|---|---|
| Write a failing test, then make it pass | test-driven-development | Test exists, suite passes, no skipped tests |
| Refactor a module without changing behavior | code-review / simplify-code | Pre-existing tests still pass; no public API change |
| Review a diff for a specific bug class | code-review | The planted bug is flagged; false-positive rate under threshold |
| Produce a plan before touching code | writing-plans | A plan file exists before the first edit; plan covers the listed constraints |
| Fix a failing build | systematic-debugging | Build passes; the fix is in the expected file |
2. Oracles: turn a vibe into a boolean
Each task needs a known-good outcome you can check mechanically — the reason SWE-bench and Terminal-Bench are trustworthy is that every instance has a test script that returns pass/fail [Terminal-Bench]. Your oracles come in three flavors:
- Test-based: a test suite that must pass. The strongest signal; use it wherever the task has behavioral semantics.
- Artifact-based: a file must exist, a config must contain a key, a diff must touch exactly the expected files. Cheap and deterministic.
- Checker-script: a small script that inspects the result and returns a boolean (e.g. "the plan file appears before the first code edit in the session log").
The oracle is what separates a regression test from an opinion. If you can't write the boolean, the task isn't ready for the bench — so make the boolean, or drop the task.
3. Metrics robust to "two correct answers look different"
A refactored module and the original can be equally correct and share zero lines. So the metric layer is a mix, not a single number:
- Exact checks for the parts that must be identical (build passes, no API change, plan-before-edit ordering).
- Test pass-rate for behavioral tasks.
- LLM-as-a-judge with a written rubric for the genuinely judgmental parts (is this review actually catching the planted bug? is this plan coherent?). The rubric is itself a skill — version it, and keep it deterministic (
temperature=0) so the judge doesn't add its own noise [promptfoo] [DeepEval].
Collapse the mix into one score per task (0–100) and one aggregate across the suite. The aggregate is your factory gauge; the per-task scores are where you localize a regression.
The CI Loop: This Is What Makes It a Factory Test
Replaying the suite once is a measurement. Replaying it on every change and failing the build on a drop is a regression test. That's the whole difference, and it's the part that's actually novel.
Wire it into your existing CI. Any pull request that touches a SKILL.md or the harness config triggers the bench:
# .github/workflows/skills-bench.yml (shape, not exhaustive)
name: skills-bench
on:
pull_request:
paths:
- 'skills/**'
- 'harness/**'
- '.bench/**'
jobs:
bench:
runs-on: self-hosted-gpu # your pinned-model box
steps:
- uses: actions/checkout@v4
- name: Assert pinned model
run: |
sha256sum models/pinned-27b.gguf | grep -qF "$(cat .bench/model.sha256)"
- name: Serve pinned model
run: bash .bench/serve-model.sh # vllm / LM Studio / mlx on :8000, temp=0, fixed seed
- name: Run the bench
run: python -m bench.run --suite .bench/tasks --endpoint http://127.0.0.1:8000/v1 --out report.json
- name: Compare to golden baseline
run: |
python -m bench.compare --report report.json \
--baseline .bench/baseline.json --threshold 3.0
- name: Post gauge
uses: actions/github-script@v7
with:
script: |
const r = require('./report.json');
require('fs').writeFileSync('${{ github.workspace }}/gauge.svg', r.gauge_svg);
That last step is the gate. bench.compare loads the golden baseline — the per-task scores you recorded when the skill set was last verified good — and fails the PR if any task drops more than your threshold. This is promptfoo, DeepEval, or any eval framework pointed at your tasks and your playbooks instead of a generic prompt [promptfoo] [DeepEval]. The frameworks already do the eval mechanics; the novelty is the variable you sweep (skills, not models) and the fixture you pin (a weight file, not an endpoint).
The golden baseline is a fixture, committed to git. When you intentionally improve a skill and the bench goes up, you re-record the baseline and commit it. The baseline is your "last known good factory" — the same role a golden file plays in snapshot testing.
The Noise Floor: Measure It Before You Trust a Single Point
This is the subtlety that separates a real regression suite from a toy, and the step most people skip and regret.
Run the pinned model over the suite five times with no changes at all. You will see variation. Even a fixed model at a low temperature isn't perfectly deterministic, and agents are multi-step: one wobble in an early action cascades into a different final artifact. That baseline variance is your dead band.
The dead band is not optional
If your measured noise is ±3 points, then a 2-point drop from a skill edit is weather, not signal — and a 7-point drop is a real regression. Set your CI threshold above the measured noise floor, or you will chase ghosts for a month and the team will stop trusting the bench entirely. A regression tool that false-alarms on every PR is worse than no tool, because it teaches people to ignore the red.
Two practical ways to shrink the floor so your real threshold can be tight:
- Lower the temperature on both the worker and the judge, and fix the seed where the runtime supports it.
- Average per task over 2–3 runs inside the bench itself, and report the mean with the std as the dead band.
Once the floor is measured, the entire system is honest: you know exactly how much movement is the model breathing and how much is a skill actually changing behavior.
What This Looks Like in VS Code Specifically
VS Code is the part of the factory most people underrate, because they still picture it as "an editor with a chat box." It isn't that anymore — and the agent features that landed are exactly the seams a regression bench needs.
- Agent harnesses. VS Code now has first-class agent harnesses — a named layer that drives the agent and that you can select per session. The harness is a distinct, selectable component of the factory, not a hidden constant. Pin which harness the bench runs under, and you can A/B harnesses later the same way you A/B skills.
- Custom agents. Named profiles that bundle a model, instructions, and tools [docs]. Your bench can run against a fixed custom agent definition, so the harness + instructions are pinned together as one unit.
- Agent Skills. Your
SKILL.mdplaybooks load through the native skills surface, so the variable you're testing is the exact file in your repo. - Hooks. You can run a linter, a checker, or your oracle script automatically when the agent touches a file [hooks]. That's a free instrumentation point: the bench's checker can fire from the agent's own lifecycle instead of from a wrapper.
- OpenTelemetry monitoring. The agent emits structured telemetry for its actions, which means the harness is already producing the event log your artifact-based oracles want to inspect (plan-before-edit ordering, which files were touched, which commands ran) [agents overview].
Pull those together and the setup is concrete and already half-built:
- Point the harness at your local OpenAI-compatible endpoint (the pinned model on
127.0.0.1) via VS Code's language model config. - Load your skill repo through the skills surface; keep the custom-agent definition pinned.
- Wire a hook that runs the relevant checker after each task.
- Collect the OTel action log as the raw material for artifact oracles.
- Run the suite, record the baseline, commit it, and gate the PR on the threshold.
Every skill edit from that point on is a data point, and a bad edit fails a build instead of surfacing three weeks later as "hmm, the reviews got shakier."
MCP is the seam for the tools
Any external tool the skills rely on (a repo map, a code index, a docs fetcher) can be exposed as an MCP server the agent calls [MCP servers]. Keep the MCP config pinned in the same repo as the skills — a changed tool is a changed variable, and it belongs in the same fixture boundary as everything else.
What You Actually Learn
Once the loop is running, the bench starts paying for itself in ways that are hard to predict in advance:
- Load-bearing vs. cargo. You find out which skills actually move the score and which are dead weight you never noticed you were carrying. A skill that improves the human reader but not the agent is a maintenance cost with no factory yield.
- Improvements that aren't. The bench will show you a skill you were sure was helping actually making the agent slower and no more correct. That's information you can only get by measuring.
- Defensible rollbacks. When a harness update regresses your review workflow by a task, you roll it back with a straight face — because the bench said so, not because you felt uneasy.
- The confidence to keep iterating. The real prize. Right now most people are afraid to edit their skills, because a bad edit has no feedback loop and only shows up as slow, mysterious quality decay. A green build on every skill PR removes that fear and turns playbook maintenance into normal, low-risk engineering.
Honest Limits
A bench is only as good as its tasks, and tasks rot. A task you wrote for a codebase you've since refactored is now testing the past, not the present — so you prune the suite the way you prune a test suite, and you re-record baselines when the ground truth legitimately changes.
The pinned model is a control, not a claim. This setup proves that your skills don't regress against a stable worker. It does not prove you outperform the frontier, and it shouldn't be read as a leaderboard entry. It's a regression net for a factory you run, not a trophy. The moment you start using it to compare "my 27B beats GPT," you've missed the point — the point is that your factory's output is stable and attributable, not that it's the highest number on a wall.
And the bench measures the delta, not the absolute quality of a single output. A suite that passes at a low bar still lets bad work through; it only guarantees no regression. Pair it with periodic human review of the gauge, and treat the bar as something you raise deliberately over time.
Bottom line: the mental model matters more than the tooling. Stop treating your AI setup as a model with a nice UI. Treat it as a system you own — version the parts, pin the fixture, replay the work, measure the delta, and let the build fail when a change makes the factory worse. That's what "testing the AI factory" actually means, and it's not hard, because every piece already exists: the model is a weight file, the tasks are your real work, and the harness is already emitting telemetry. All that's left is to close the loop.
References
- VS Code — Chat in agent mode — the harness in action
- VS Code — Agent harnesses — the selectable driver layer
- VS Code — Agent Skills — native
SKILL.mdloading - VS Code — Custom agents — pinned model + instructions + tools
- VS Code — Hooks — run a checker on the agent's own lifecycle
- VS Code — Language models — pointing the harness at a local endpoint
- VS Code — MCP servers — the tool seam
- Anthropic — Equipping agents for the real world with Agent Skills — the open standard your skills follow
- SWE-bench — 2,294 real GitHub issues; the task/oracle pattern
- SWE-bench — source
- Terminal-Bench — real terminal tasks verified by test scripts
- Terminal-Bench — source
- Aider polyglot leaderboard — 225 exercises across 6 languages
- promptfoo — eval + CI mechanics (now part of OpenAI, still MIT)
- promptfoo — source
- DeepEval — LLM regression testing in CI/CD via pytest
- DeepEval — source
- LM Studio — serving a pinned open model locally
Research by ThinkSmart.Life · September 2026