AI AgentsReinforcement LearningHarness EngineeringGRPOQwenALFWorldLong-Horizon AgentsSelf-Evolving Agents

EvoHarness-RL: The Self-Evolving Runtime Harness

A UIUC + Meta AI paper that trains the agent to build and coordinate with its own external state — Belief, Progress, and Experience — via cost-aware RL, pushing Qwen3-8B to 96.9% on ALFWorld.

August 30, 2026Michel Laclé14 min read

The Harness Problem: The Scaffold Nobody Trains

Long-horizon LLM agents — the kind that navigate a household, drive a browser, or work through a multi-step software task — don't succeed on raw model capability alone. They succeed on the harness around them: the collection of prompts, tools, memories, state trackers, verifiers, and execution logs that let an agent maintain state, track progress, recover from failure, and reuse what it has learned.

Here's the uncomfortable part of the paper from UIUC and Meta AI: that harness is almost always manually engineered. The external workspace and the policy for using it are specified through prompts, heuristics, and domain-specific conventions. The agent can be surrounded by a rich set of useful external support and still never be trained to decide when to form, read, update, or consolidate it as part of its own decision process.

The core question: How can an agent learn to form useful external state and efficiently leverage that support as part of its own decision process — rather than having a human developer hand-craft every line of the runtime scaffolding? The paper calls this harness policy learning.

EvoHarness-RL is the answer. The team from the University of Illinois Urbana-Champaign and Meta AI [Paper] abstracts the messy, heterogeneous runtime layer into a compact, policy-facing state and then trains the agent to coordinate with it. The result is striking: on the ALFWorld benchmark, a modest Qwen3-8B model reaches 96.9% success — effectively matching frontier models that are an order of magnitude larger. The paper was accepted to LLA@COLM 2026.

Belief, Progress, and Experience (BPE)

Before you can train a policy over external state, you need a state that is rich enough to be useful but compact enough to learn over. Real harnesses are sprawling — state trackers, execution logs, task plans, verifier feedback, episodic memories, skill libraries. But the authors argue all of these address a small set of recurring failure modes in long-horizon interaction. The agent loses track of what's currently true in the environment, forgets what it has already done, or keeps rediscovering procedures it could have reused. So they organize the policy-facing role of external state into three functional components:

Belief — "What is true right now?"

Belief (B) stores the task-relevant facts inferred from interaction: object states, locations, spatial relations. It gives the policy a persistent estimate of the current environment so it doesn't have to rely only on the transient context-window memory. This is the difference between "I remember seeing the mug near the sink" and actually maintaining a queryable world-state estimate across many steps.

Progress — "What's done, what's left, where am I stuck?"

Progress (P) records task decomposition and execution status through subgoal-status entries. It externalizes what has been attempted, what remains open, and where execution may be blocked — turning an implicit reasoning trace into inspectable, updatable task state.

Experience — "What did I learn from last time?"

Experience (E) maintains cross-episode knowledge: skills, failure modes, search priors, and high-level strategies. It supports reuse across attempts by surfacing relevant prior knowledge during execution and storing new insights for later consolidation.

Why BPE is the right abstraction

The beauty of BPE is that it's a functional interface, not a fixed schema. The same three roles map onto very different internals per domain — a scene graph in a visual environment, a test-status tracker in a software environment, a world-state store in a household. The policy-facing API stays stable while the implementation is domain-specific. That's what makes the learned coordination layer transferable and analyzable.

The Action Protocol: Four Meta-Actions

Given BPE, the policy needs a compact way to read from and write to the external workspace. A fully domain-specific API would expose many operations (hard to transfer or analyze); a single generic "memory" action would hide the functional structure. EvoHarness-RL lands in the middle with four harness meta-actions:

Action Reads / Writes Purpose
track Belief Read task-relevant environmental facts (e.g. track[object], track[world])
commit Progress Write a subgoal or execution update (e.g. commit[subgoal])
recall Experience Retrieve reusable prior knowledge (e.g. recall[query])
note Experience Record a new insight for later consolidation (e.g. note[insight])

Here's the crucial design detail: harness actions and environment actions consume the same interaction budget. At each step the policy chooses from both. If it takes an environment action, the task advances and yields a new observation. If it takes a harness action, it queries or updates the external workspace and gets a new harness view back. Because every harness call costs a step of the same budget the agent needs to actually make progress, the agent has to learn when external-state access is worth its cost. That cost pressure is exactly what the RL stage optimizes.

Two-Stage Training: From Scaffold to Policy

The training recipe is two stages with deliberately different jobs.

Stage 1 — Supervised Harness Fine-Tuning (SFT)

First, the base model is bootstrapped on successful teacher trajectories collected using the same BPE interface. The teacher observes the task, the current observation, admissible environment actions, recent history, and active harness views, then emits a single next action in <think>...</think><action>...</action> format. That action can be either an environment command or a BPE harness action. Fine-tuning Qwen3-8B on these demonstrations teaches the model two things at once: task-solving behavior and the basic semantics of when to track, commit, recall, or note. The experience accumulated during teacher rollouts also seeds the skill store used in the next stage.

Stage 2 — Cost-Aware GRPO

From the SFT checkpoint, the policy is optimized with Group Relative Policy Optimization (GRPO) [GRPO]. The trajectory-level reward is where the "cost-aware" part lives:

  • Task completion is the gatekeeper. It's a dominant, sparse signal. The efficiency bonus is only granted upon success — so redundant harness queries that don't lead to a completed task are naturally penalized.
  • A time-dependent vocabulary-diversity bonus. This is a curriculum: early in training it encourages broad exploration of the harness actions, then gracefully decays over an annealing horizon to force specialization and efficient task resolution. It's the mechanism that prevents policy collapse into either "never touch the harness" or "spam the same call in a loop."
  • Fixed penalties for degenerate repetition and malformed action syntax.

What changes between the stages: SFT teaches the agent how to construct useful external state. GRPO teaches it when that access is worth its cost. The net effect is turning harness use from a prompt-time scaffold into a learned, cost-aware runtime policy decision.

The ALFWorld Adapter

The general framework is instantiated on ALFWorld [ALFWorld], a text-based, embodied household benchmark aligned with the ALFRED dataset. The agent completes multi-step tasks — navigating rooms, manipulating objects through text commands — across six task families that stress different kinds of state tracking: Pick (pick-and-place), Look (inspect under light), Clean (clean before placing), Heat, Cool, and Pick2 (place two objects).

The environment adapter bridges domain-specific signals to the shared BPE interface. Belief, Progress, and Experience each get a concrete grounding:

  • Belief → world-state store. Maintained in the background after each environment step from the agent's actions and observations (object states, locations, spatial relations). It isn't fully exposed by default — the policy must issue track[object] or track[world] to inspect it. Access stays a selective harness action.
  • Progress → committed execution record. A bounded list of subgoal-status entries, sufficient for mostly-sequential household tasks. commit[subgoal] makes attempted, pending, or blocked progress visible to later decisions.
  • Experience → cross-episode skill store. Organized into general skills, task-specific skills, common mistakes, and object-location search priors. It evolves at two timescales: within an episode via recall and note, and across episodes via a consolidation model that merges accumulated evidence into the skill store through add / update / remove operations at epoch boundaries.

Main Results: An 8B Model That Matches the Frontier

The headline number: on the 140-task ALFWorld "seen" split, EvoHarness-RL on Qwen3-8B reaches 96.9% average success — a +49.0 absolute point improvement over the base ReAct model (47.9%). For perspective, that puts an 8B open-weight model at the level of top frontier systems.

Approach Pick Look Clean Heat Cool Pick2 Avg
ReAct (frontier) — Claude Opus 4.5 100.092.396.3100.088.0100.096.4
ReAct — GPT-4.1 82.961.544.443.84.041.747.9
ReAct — GPT-5 74.353.848.162.560.058.360.7
ReAct — Qwen3-8B (baseline) 78.146.233.337.529.347.247.9
ExpeL — Qwen3-8B 91.476.914.843.828.045.849.3
ReasoningBank — Qwen3-8B 83.848.749.439.641.354.255.7
Dynamic Cheatsheet — Qwen3-8B 88.653.829.637.520.066.752.1
SkillOS (trainable) — Qwen3-8B 95.271.874.172.977.377.880.2
SkillRL (trainable) — Qwen2.5-7B 97.971.490.090.095.587.589.9
EvoHarness-Base (inference-time) 71.453.863.050.048.041.756.4
EvoHarness-SFT 80.053.888.975.040.062.568.6
EvoHarness-RL (ours) 100.092.995.5100.092.6100.096.9

Two things stand out. First, the two-stage progression is clean and monotonic: prompt-time scaffolding gets 56.4%, adding SFT lifts it to 68.6%, and cost-aware GRPO jumps it to 96.9%. Optimization turns the harness from a static tool into a highly effective decision interface. Second, the BPE framework helps across model scales — even without training, applying the explicit harness to frontier models lifts GPT-4.1 by +22.1 points and GPT-5 by +25.7, and pushes the already-near-ceiling Claude Opus 4.5 from 96.4% to 98.5%. Externalizing belief, progress, and experience appears broadly critical for reliable long-horizon execution regardless of base model size.

Ablations: All Three Components Are Load-Bearing

To isolate each component, the authors remove one BPE module at a time from the frozen, inference-time Qwen3-8B harness:

Variant Avg SR What breaks
Full BPE 56.4 —
w/o Belief 50.0 Object/state tracking → severe drops on Clean, Cool
w/o Progress 50.7 Subgoal commitment → hurts long-horizon Pick2
w/o Experience 48.6 Skill recall + mistake avoidance → worst overall

Removing any single component meaningfully hurts execution. Dropping Belief disables explicit object tracking (bad for localization and state-verification tasks). Dropping Progress prevents subgoal commitment (disproportionate damage to tasks with dependent subgoals). Dropping Experience removes skill recall and mistake avoidance and yields the lowest average. The three components function synergistically as a unified state interface — not as isolated memory tricks you can swap in or out.

Generalization: RL Beats Imitation on Unseen Tasks

This is one of the more interesting results. On the ALFWorld unseen split, base ReAct manages 50.0%, and the prompt-time BPE harness lifts zero-shot performance to 77.6% — a clear benefit of state externalization. But here's the non-obvious part: EvoHarness-SFT actually drops to 69.4%. Supervised imitation learned the teacher's harness-use patterns from seen trajectories without optimizing for when access is worthwhile in a novel environment.

The full RL-optimized policy recovers and surpasses everything at 86.6% on unseen tasks. The interpretation is that cost-aware GRPO recalibrates harness access and learns a broadly useful strategy rather than memorizing the training environments. Imitation is brittle to distribution shift; the cost-aware RL policy is not. That's a meaningful argument for the RL stage beyond just the raw accuracy gain.

Harness Annealing: The Agent Learns to Stop Over-Calling

The training dynamics reveal a pattern the authors call harness annealing. The SFT-initialized agent starts by making frequent harness calls — using BPE as an explicit scaffold to track state, recall procedures, and narrow the search space. As RL progresses, usage drops quickly and stabilizes near one harness call per episode.

GRPO is gradually internalizing the routine scaffolded behaviors into the policy itself, while preserving harness access only when the expected benefit outweighs the step cost. The agent shifts from scaffolded exploration to selective, cost-aware coordination.

Action-level granularity

The annealing isn't uniform across the four actions, and that's environment-dependent. recall remains the most persistent — cross-episode experience keeps providing useful search priors long after routine behaviors are internalized. commit and note decay toward zero once stable task strategies emerge. track follows an intermediate pattern: useful early for state disambiguation, then fading as the agent learns more direct interaction patterns. In a visual environment Belief usage would likely stay higher; in a software environment Progress would. The policy learns which parts of BPE are worth accessing under a given cost structure — not a fixed universal distribution.

Harness Evolution: The Memory That Forgets

While Belief and Progress are updated within an episode, the cross-episode Experience store is reshaped across episodes through accumulation, consolidation, and forgetting. The authors call this harness evolution.

The skill bank expands rapidly early in training — general strategies, task-specific procedures, common mistakes, search priorities all get written. Later, growth becomes selective: redundant entries are merged, rarely-useful skills are evicted, and frequently-recalled knowledge is preserved. The final bank is compact yet diverse — a task-adaptive state substrate rather than a passive append-only memory.

This complements harness annealing neatly: the policy learns when to use external state, while the harness evolves what reusable experience it can provide. It's a co-evolutionary loop, and it directly answers the paper's framing that self-evolving agents need to actively curate, refine, consolidate, and forget experience — not just accumulate it.

Why This Matters

The paper positions itself against two adjacent lines of work, and the distinction is the real contribution.

  • Harness engineering — recent work like Harness-1, Meta-Harness, HarnessX, and AutoHarness treats the harness as an environment-side construct: they optimize harness state, configurations, or trace-driven adaptations, often via offline search or human development [Harness-1] [Meta-Harness]. EvoHarness-RL instead treats harness access as a first-class, learnable policy decision — it trains the underlying policy to actively control, query, and coordinate with the workspace.
  • Memory and self-evolving agents — Reflexion, Voyager, SkillOS, SkillRL, and ReasoningBank mostly focus on cross-episode knowledge, and generally separate that skill curation from real-time within-episode state tracking (belief + progress) [Reflexion] [Voyager]. BPE unifies the two so the policy evolves how it uses long-term experience and how it synchronizes that experience with active belief and progress.

Bottom line: Long-horizon agents benefit from trainable policies for constructing and coordinating with external harness workspaces — beyond simply adding stronger tools or larger memories. The practical takeaway for anyone building agent systems: the runtime harness isn't a one-time prompt you write and forget. It's a control surface you can (and the results suggest you should) train. An 8B model, taught to build and cost-aware-coordinate its own Belief/Progress/Experience state, closes the gap to the frontier on stateful long-horizon tasks.

The Honest Caveats

The evidence is strong but bounded. It's a single benchmark (ALFWorld, text-based household tasks) with a single backbone (Qwen3-8B) for the trained variants. The action-level annealing patterns are explicitly described as environment-dependent, so the "near one call per episode" endpoint and the exact BPE action mix shouldn't be read as universal constants. And BPE as prompt-time scaffolding helps frontier models a lot, but the full RL pipeline has only been validated on the small open model. The framework is general in design — the BPE interface is domain-agnostic and the adapter is pluggable — but the paper proves it in one embodiment.

What to Watch

  • Transfer to richer environments. The paper's own analysis predicts Belief becomes more important in visual embodied settings and Progress in software-engineering settings. A visual or SWE benchmark run would be the natural next test.
  • Scaling the backbone. Does the two-stage recipe still move the needle when applied to a larger base model, or does a strong base model's internal context already do the job BPE externalizes?
  • Cost accounting in practice. The "one harness call per episode" endpoint is attractive, but real deployments need the token/latency economics of track/commit/recall/note measured against the success-rate gain.

References

  1. Ning et al. — EvoHarness-RL: Learning Self-Evolving Runtime Harness for Long-Horizon LLM Agents (arXiv:2608.05446, Aug 2026)
  2. Shridhar et al. — ALFWorld: Aligning Text and Embodied Environments for Interactive Learning
  3. Shao et al. — DeepSeekMath (GRPO)
  4. Jiang et al. — Harness-1: RL for search agents with state-externalizing harnesses
  5. Lee et al. — Meta-Harness: End-to-end optimization of model harnesses
  6. Shinn et al. — Reflexion: Language Agents with Verbal Reinforcement Learning
  7. Wang et al. — Voyager: An Open-Ended Embodied Agent with LLMs

Research by ThinkSmart.Life · August 2026

↑ Back to top