AI AgentsRegression TestingHarness EngineeringAgent SkillsSWE-benchVS CodeOpen-Source ModelsCI/CD

Testing the AI Factory: Benchmarking VS Code + Skills Against a Pinned Open-Source Model

Your AI dev workflow is a factory — a harness, a skill set, and a model. Pin a local open-source model as a fixture and build a software-bench-style regression suite that catches a skill edit before it degrades your output.

September 3, 2026Michel Laclé9 min read

The Factory Mental Model

Every serious AI coding setup is secretly a factory, and it has three moving parts. If you can name all three, you can start treating the thing like the engineering system it actually is — instead of a magic box that occasionally gets worse.

Part What it is In this setup
Harness The machine: reads the repo, plans, writes code, runs commands, reviews its own work VS Code in agent mode [docs]
Skills The playbooks: procedural instructions the agent loads on demand Your software-development skill repo (the SKILL.md Agent Skills format) [standard]
Model The worker: the LLM that executes A pinned open-source model served locally [LM Studio]

The skills are the part most people under-think. A skill is a folder with a SKILL.md — a name, a description, and procedural instructions the agent reads when the task matches. Anthropic published Agent Skills as an open standard specifically so these playbooks are portable across harnesses [link]. VS Code now loads them natively through its own Agent Skills customization surface. So a skill is real, versioned, load-bearing software — the kind of thing you'd want a regression test around.

The reframe that drives everything below: stop treating your AI setup as "a model with a nice UI." Treat it as a system you own — harness, skill set, worker — and test it the way you test any system you own.

The Gap: We Benchmark the Worker, Not the Factory

In most setups the worker is a hosted API — GPT, Claude, Gemini. And that worker is not fixed. It gets updated underneath you, usually with no clear changelog you'll read. Same prompt, same task, different answer six weeks later. Meanwhile your skills are changing every time you learn something and add or tweak a playbook, and the harness ships new versions with new tools, new defaults, and new behavior.

So three things are drifting at once. And the only one you're actually measuring is the worker. That's the gap.

The public benchmarks are excellent — and they all answer the same single question:

Benchmark What it holds fixed What it varies What it measures
SWE-bench [site] Task set (2,294 real GitHub issues; 500-instance Verified subset) The model / agent % of issues resolved
Terminal-Bench [site] Task set + sandbox, verified by a test script The model / agent % of end-to-end terminal tasks passed
Aider polyglot [leaderboard] 225 Exercism exercises across 6 languages The model % correct edits

The pattern is deliberate and correct for its purpose: hold the task fixed, vary the model, rank the models. But that's the opposite of the question that keeps a practitioner up at night. These benchmarks answer "how good is this model at coding?" They cannot answer "did my setup get worse this week?"

  • Did the skill I edited on Tuesday quietly make the agent worse at code review?
  • Did the harness update change how it handles a failing test?
  • Did a model refresh change the output shape my downstream tooling expects?

A benchmark that only varies the model can't see any of that, because it holds your playbooks and your harness frozen at some version you stopped using months ago. The SWE-bench entry for your model is a leaderboard line, not a regression signal for your factory.

Invert the Experiment: Pin the Worker

The fix is to flip the independent variable. In a model leaderboard, the model is what you sweep and everything else is held constant. In a factory regression test, you do the opposite: lock the worker and sweep the harness and skills.

Concretely, pin three things:

  1. The weight file. A specific open-weights model, a specific artifact, not "latest."
  2. The serving config. Same inference engine, same quant, same context length, same temperature.
  3. The sample regime. A fixed seed / temperature=0 where the runtime allows it, so a rerun is as close to token-identical as the model gets.

Once the worker is fixed, any delta you observe in the output can only come from the harness or the skills. And those are exactly the two things you actually author and control. That's the entire leverage of the inversion: you've converted an unattributable, drifting number into a controlled A/B test where the independent variable is a git commit to your skill repo.

Why not just pin a hosted model version?

You can point at a versioned endpoint, but it still sits behind an API boundary where the operator can deprecate, silently re-quantize, or change routing between the "same" version. A hosted version is a named moving target. A weight file on your disk is a byte-identical object you can checksum. The pinned local model is a fixture; a hosted version is a dependency.

Why Open-Source and Why Local

The choice of a local, open-weights model is doing more work than "I don't want to pay per token." It's the thing that makes the control variable real.

  • A weight file doesn't drift. It's the same bytes today as it was last month. You can sha256sum it, store the hash in your skill repo next to the skills, and assert it hasn't changed in CI. That's a reproducibility invariant a hosted model can't give you.
  • You can version it with the skills. Commit the model hash into the same repo as the skill set. Now a "known-good" state is one git ref: skills and worker pinned together, the way you'd pin a test fixture.
  • You control the sampling. With your own runtime you set temperature, top-p, and (where supported) the seed. With a hosted API you request them and hope for compliance.
  • It's fast and free to replay. A regression suite has to run on every skill PR. A local model on your own GPU means the bench is as cheap as your electricity, not a line item that scales with your commit frequency.

Practically, this is just running an open model through a local server that exposes an OpenAI-compatible endpoint — LM Studio, vllm, llama.cpp, or MLX — and pointing the harness at http://127.0.0.1:… [LM Studio] [VS Code models]. The model is now a local process you can pin, checksum, and restart deterministically. That's a test fixture, not a colleague with a bad hair day.

Building the Bench

Now build the software bench. It has exactly three components, borrowed directly from the SWE-bench / Terminal-Bench design and pointed at your own skills.

1. A task suite scoped to your skills

Not generic benchmark instances you'll never touch again — tasks that map one-to-one to the playbooks you maintain. If you have a test-driven-development skill, a code-review skill, a writing-plans skill, and a systematic-debugging skill, your suite should have tasks that exercise each of them. Small enough to run on every change; specific enough that a drop is attributable.

Task Skill exercised Oracle (known-good outcome)
Write a failing test, then make it pass test-driven-development Test exists, suite passes, no skipped tests
Refactor a module without changing behavior code-review / simplify-code Pre-existing tests still pass; no public API change
Review a diff for a specific bug class code-review The planted bug is flagged; false-positive rate under threshold
Produce a plan before touching code writing-plans A plan file exists before the first edit; plan covers the listed constraints
Fix a failing build systematic-debugging Build passes; the fix is in the expected file

2. Oracles: turn a vibe into a boolean

Each task needs a known-good outcome you can check mechanically — the reason SWE-bench and Terminal-Bench are trustworthy is that every instance has a test script that returns pass/fail [Terminal-Bench]. Your oracles come in three flavors:

  • Test-based: a test suite that must pass. The strongest signal; use it wherever the task has behavioral semantics.
  • Artifact-based: a file must exist, a config must contain a key, a diff must touch exactly the expected files. Cheap and deterministic.
  • Checker-script: a small script that inspects the result and returns a boolean (e.g. "the plan file appears before the first code edit in the session log").

The oracle is what separates a regression test from an opinion. If you can't write the boolean, the task isn't ready for the bench — so make the boolean, or drop the task.

3. Metrics robust to "two correct answers look different"

A refactored module and the original can be equally correct and share zero lines. So the metric layer is a mix, not a single number:

  • Exact checks for the parts that must be identical (build passes, no API change, plan-before-edit ordering).
  • Test pass-rate for behavioral tasks.
  • LLM-as-a-judge with a written rubric for the genuinely judgmental parts (is this review actually catching the planted bug? is this plan coherent?). The rubric is itself a skill — version it, and keep it deterministic (temperature=0) so the judge doesn't add its own noise [promptfoo] [DeepEval].

Collapse the mix into one score per task (0–100) and one aggregate across the suite. The aggregate is your factory gauge; the per-task scores are where you localize a regression.

The CI Loop: This Is What Makes It a Factory Test

Replaying the suite once is a measurement. Replaying it on every change and failing the build on a drop is a regression test. That's the whole difference, and it's the part that's actually novel.

Wire it into your existing CI. Any pull request that touches a SKILL.md or the harness config triggers the bench:

# .github/workflows/skills-bench.yml (shape, not exhaustive)
name: skills-bench
on:
  pull_request:
    paths:
      - 'skills/**'
      - 'harness/**'
      - '.bench/**'

jobs:
  bench:
    runs-on: self-hosted-gpu            # your pinned-model box
    steps:
      - uses: actions/checkout@v4
      - name: Assert pinned model
        run: |
          sha256sum models/pinned-27b.gguf | grep -qF "$(cat .bench/model.sha256)"
      - name: Serve pinned model
        run: bash .bench/serve-model.sh   # vllm / LM Studio / mlx on :8000, temp=0, fixed seed
      - name: Run the bench
        run: python -m bench.run --suite .bench/tasks --endpoint http://127.0.0.1:8000/v1 --out report.json
      - name: Compare to golden baseline
        run: |
          python -m bench.compare --report report.json \
            --baseline .bench/baseline.json --threshold 3.0
      - name: Post gauge
        uses: actions/github-script@v7
        with:
          script: |
            const r = require('./report.json');
            require('fs').writeFileSync('${{ github.workspace }}/gauge.svg', r.gauge_svg);

That last step is the gate. bench.compare loads the golden baseline — the per-task scores you recorded when the skill set was last verified good — and fails the PR if any task drops more than your threshold. This is promptfoo, DeepEval, or any eval framework pointed at your tasks and your playbooks instead of a generic prompt [promptfoo] [DeepEval]. The frameworks already do the eval mechanics; the novelty is the variable you sweep (skills, not models) and the fixture you pin (a weight file, not an endpoint).

The golden baseline is a fixture, committed to git. When you intentionally improve a skill and the bench goes up, you re-record the baseline and commit it. The baseline is your "last known good factory" — the same role a golden file plays in snapshot testing.

The Noise Floor: Measure It Before You Trust a Single Point

This is the subtlety that separates a real regression suite from a toy, and the step most people skip and regret.

Run the pinned model over the suite five times with no changes at all. You will see variation. Even a fixed model at a low temperature isn't perfectly deterministic, and agents are multi-step: one wobble in an early action cascades into a different final artifact. That baseline variance is your dead band.

The dead band is not optional

If your measured noise is ±3 points, then a 2-point drop from a skill edit is weather, not signal — and a 7-point drop is a real regression. Set your CI threshold above the measured noise floor, or you will chase ghosts for a month and the team will stop trusting the bench entirely. A regression tool that false-alarms on every PR is worse than no tool, because it teaches people to ignore the red.

Two practical ways to shrink the floor so your real threshold can be tight:

  • Lower the temperature on both the worker and the judge, and fix the seed where the runtime supports it.
  • Average per task over 2–3 runs inside the bench itself, and report the mean with the std as the dead band.

Once the floor is measured, the entire system is honest: you know exactly how much movement is the model breathing and how much is a skill actually changing behavior.

What This Looks Like in VS Code Specifically

VS Code is the part of the factory most people underrate, because they still picture it as "an editor with a chat box." It isn't that anymore — and the agent features that landed are exactly the seams a regression bench needs.

  • Agent harnesses. VS Code now has first-class agent harnesses — a named layer that drives the agent and that you can select per session. The harness is a distinct, selectable component of the factory, not a hidden constant. Pin which harness the bench runs under, and you can A/B harnesses later the same way you A/B skills.
  • Custom agents. Named profiles that bundle a model, instructions, and tools [docs]. Your bench can run against a fixed custom agent definition, so the harness + instructions are pinned together as one unit.
  • Agent Skills. Your SKILL.md playbooks load through the native skills surface, so the variable you're testing is the exact file in your repo.
  • Hooks. You can run a linter, a checker, or your oracle script automatically when the agent touches a file [hooks]. That's a free instrumentation point: the bench's checker can fire from the agent's own lifecycle instead of from a wrapper.
  • OpenTelemetry monitoring. The agent emits structured telemetry for its actions, which means the harness is already producing the event log your artifact-based oracles want to inspect (plan-before-edit ordering, which files were touched, which commands ran) [agents overview].

Pull those together and the setup is concrete and already half-built:

  1. Point the harness at your local OpenAI-compatible endpoint (the pinned model on 127.0.0.1) via VS Code's language model config.
  2. Load your skill repo through the skills surface; keep the custom-agent definition pinned.
  3. Wire a hook that runs the relevant checker after each task.
  4. Collect the OTel action log as the raw material for artifact oracles.
  5. Run the suite, record the baseline, commit it, and gate the PR on the threshold.

Every skill edit from that point on is a data point, and a bad edit fails a build instead of surfacing three weeks later as "hmm, the reviews got shakier."

MCP is the seam for the tools

Any external tool the skills rely on (a repo map, a code index, a docs fetcher) can be exposed as an MCP server the agent calls [MCP servers]. Keep the MCP config pinned in the same repo as the skills — a changed tool is a changed variable, and it belongs in the same fixture boundary as everything else.

What You Actually Learn

Once the loop is running, the bench starts paying for itself in ways that are hard to predict in advance:

  • Load-bearing vs. cargo. You find out which skills actually move the score and which are dead weight you never noticed you were carrying. A skill that improves the human reader but not the agent is a maintenance cost with no factory yield.
  • Improvements that aren't. The bench will show you a skill you were sure was helping actually making the agent slower and no more correct. That's information you can only get by measuring.
  • Defensible rollbacks. When a harness update regresses your review workflow by a task, you roll it back with a straight face — because the bench said so, not because you felt uneasy.
  • The confidence to keep iterating. The real prize. Right now most people are afraid to edit their skills, because a bad edit has no feedback loop and only shows up as slow, mysterious quality decay. A green build on every skill PR removes that fear and turns playbook maintenance into normal, low-risk engineering.

Honest Limits

A bench is only as good as its tasks, and tasks rot. A task you wrote for a codebase you've since refactored is now testing the past, not the present — so you prune the suite the way you prune a test suite, and you re-record baselines when the ground truth legitimately changes.

The pinned model is a control, not a claim. This setup proves that your skills don't regress against a stable worker. It does not prove you outperform the frontier, and it shouldn't be read as a leaderboard entry. It's a regression net for a factory you run, not a trophy. The moment you start using it to compare "my 27B beats GPT," you've missed the point — the point is that your factory's output is stable and attributable, not that it's the highest number on a wall.

And the bench measures the delta, not the absolute quality of a single output. A suite that passes at a low bar still lets bad work through; it only guarantees no regression. Pair it with periodic human review of the gauge, and treat the bar as something you raise deliberately over time.

Bottom line: the mental model matters more than the tooling. Stop treating your AI setup as a model with a nice UI. Treat it as a system you own — version the parts, pin the fixture, replay the work, measure the delta, and let the build fail when a change makes the factory worse. That's what "testing the AI factory" actually means, and it's not hard, because every piece already exists: the model is a weight file, the tasks are your real work, and the harness is already emitting telemetry. All that's left is to close the loop.


References

  1. VS Code — Chat in agent mode — the harness in action
  2. VS Code — Agent harnesses — the selectable driver layer
  3. VS Code — Agent Skills — native SKILL.md loading
  4. VS Code — Custom agents — pinned model + instructions + tools
  5. VS Code — Hooks — run a checker on the agent's own lifecycle
  6. VS Code — Language models — pointing the harness at a local endpoint
  7. VS Code — MCP servers — the tool seam
  8. Anthropic — Equipping agents for the real world with Agent Skills — the open standard your skills follow
  9. SWE-bench — 2,294 real GitHub issues; the task/oracle pattern
  10. SWE-bench — source
  11. Terminal-Bench — real terminal tasks verified by test scripts
  12. Terminal-Bench — source
  13. Aider polyglot leaderboard — 225 exercises across 6 languages
  14. promptfoo — eval + CI mechanics (now part of OpenAI, still MIT)
  15. promptfoo — source
  16. DeepEval — LLM regression testing in CI/CD via pytest
  17. DeepEval — source
  18. LM Studio — serving a pinned open model locally

Research by ThinkSmart.Life · September 2026

↑ Back to top