Natural-language critiques provide supervision beyond scalar rewards for non-verifiable generation, which lacks deterministic verifiers. In critique-guided refinement, a critic gives feedback on an initial response and an actor revises it. However, final revision quality does not reveal whether the critique was actually useful: a capable actor may improve without following the feedback, while valid feedback may fail if the actor cannot execute it.
We frame critique as actor-conditioned revision guidance, where usefulness depends on whether the feedback helps the target actor address the intended weakness. We introduce TAIScore (Targeted Actionable Improvement Score), a reward that evaluates the instruction, initial response, critique, and revision together, assessing whether the critique targets a real weakness, whether the actor follows it, and whether the intended aspect improves. We use this reward to train an actor-tailored critic with GRPO, and use critique-guided refinements to construct DPO preference pairs for the actor, forming a co-evolving critic-actor loop where the critic adapts to the actor's changing capability.
Experiments show that an 8B critic trained with TAIScore outperforms both a zero-shot 120B critic and critics trained with outcome-only or critique-only reward signals. Co-evolving the critic and actor further improves performance, suggesting that effective critique supervision should adapt as the actor changes.
We run two controlled analyses on 1,000 WritingBench examples, varying one role at a time, to test whether critique usefulness is intrinsic to the critique or conditioned on the actor. The result is a clear asymmetry. A larger critic (8B→32B for Qwen, 8B→70B for Llama), with the initial response and revising actor held fixed, consistently improves standalone critique quality but yields little average improvement in how well the actor incorporates the feedback or in downstream gain. For a Qwen3-8B refiner, critique adherence actually drops by 0.223 under the larger critic. A larger refiner, given the same critique, substantially improves both adherence (+1.29 on average) and downstream gain (+0.39). Critique usefulness is therefore a property of the critique-actor pair, not of the critique alone.
Controlled zero-shot analysis of critique-guided refinement on WritingBench. Bars show mean differences between larger- and smaller-model conditions in critique quality, critique adherence (how well y1 implements c), and response-quality gain S(y1)−S(y0). Gray bars compare larger and smaller critics while holding y0 and the refiner fixed; blue bars compare larger and smaller refiners while holding y0 and the critique fixed. Larger critics improve standalone critique quality but yield little average improvement in adherence or gain, whereas larger refiners substantially improve both.
Qwen3-32B critique scores higher in standalone quality than a
Qwen3-8B critique (9.00 vs. 8.33), yet yields far lower adherence from
the same Qwen3-8B refiner (3.50 vs. 8.75). The 8B critic requests localized
changes the refiner can carry out; the 32B critic proposes a coordinated redesign spanning
symbolism, character motivation, climax, ending, and dialogue, of which the refiner adopts
only isolated pieces. Feedback that looks stronger in isolation can be less useful when its
requested changes exceed the target actor's revision capability.
Our approach has three components that together form a co-evolving loop:
These two updates alternate across rounds, forming a co-evolving loop: the critic continually adapts to the actor's current weaknesses and revision capability, while the actor internalizes progressively stronger critique-guided revisions.
Overview of our co-evolving critic-actor training framework. For the current actor πt, the critic κt samples multiple critiques for the same initial response y0. Each critique produces a rollout τi = (x, y0, ci, y1,i). TAIScore evaluates each rollout and produces a GRPO reward for critic adaptation. The actor then produces candidate revisions using critiques from the adapted critic κt+1; revisions preferred over their corresponding initial responses are selected to construct DPO preference pairs y1 ≻ y0 for updating the actor. Alternating these two updates yields co-evolving critics and actors.
Qwen3-8B. All actors
perform on-policy self-refinement: the same model writes the initial response and the
revision. gpt-oss-120B serves as the judge for TAIScore and all judge-based
rewards, and is never updated or used as the actor (except in the off-the-shelf critic
baseline).
GPT-4o-mini judge, and
the DeepResearch-Gym report-level scripts with GPT-4.1-mini over the officially
retrieved ClueWeb22 materials. For each method we generate three independent output sets and
report mean and standard deviation.
All conditions use the same DPO actor-training pipeline and differ only in how the critiques used to construct the DPO pairs are produced. Across conditions we hold fixed the candidate queries, the number of candidate revisions, the pair-selection procedure, and the number of DPO pairs.
| Method | WritingBench | HelloBench | DeepResearch-Gym | |||
|---|---|---|---|---|---|---|
| Overall ↑ | OEQA ↑ | HTG ↑ | KPR ↑ | KPC ↓ | Quality ↑ | |
Qwen3-8B (base) |
72.33 ±0.08 | 34.86 ±1.95 | 39.14 ±1.76 | 71.93 ±0.04 | 1.15 ±0.09 | 81.89 ±0.07 |
| DPO pairs from off-the-shelf critics | ||||||
gpt-oss-120B critic |
75.41 ±0.14 | 36.01 ±2.36 | 50.93 ±2.18 | 73.46 ±0.09 | 1.11 ±0.08 | 82.51 ±0.06 |
| DPO pairs from trained critics (8B) | ||||||
| Outcome-gain reward | 75.63 ±0.09 | 35.35 ±2.68 | 45.14 ±2.48 | 74.37 ±0.11 | 1.12 ±0.09 | 82.38 ±0.07 |
| Critique-quality reward | 75.18 ±0.11 | 35.51 ±2.45 | 49.75 ±4.32 | 74.19 ±0.13 | 1.14 ±0.11 | 82.32 ±0.04 |
| TAIScore (ours) | 75.96 ±0.12 | 36.48 ±2.56 | 53.78 ±1.58 | 75.21 ±0.09 | 1.09 ±0.07 | 82.63 ±0.06 |
| Co-evolving critic-actor training | ||||||
| TAIScore + co-evolution (ours) | 76.72 ±0.14 | 39.84 ±2.92 | 54.62 ±1.92 | 76.14 ±0.13 | 1.03 ±0.05 | 83.15 ±0.06 |
Mean and standard deviation over three independent evaluation runs. Except for the base actor,
all rows report the final performance of the Qwen3-8B actor after DPO training.
Best results in bold; second-best underlined.
Does the advantage of TAIScore depend on the actor-training pipeline, or is the feedback
itself simply better? We hold the prompts, initial responses
(S(y0) = 72.33), Qwen3-8B reviser, revision prompt, decoding
configuration, and evaluator fixed, and vary only the supplied feedback, with
no actor update and no preference filtering.
| Condition | S(y1) | Δ WritingBench |
|---|---|---|
| (a) Critic source / training objective | ||
Off-the-shelf gpt-oss-120B |
73.82 | +1.49 |
| Outcome-gain reward | 74.44 | +2.11 |
| Critique-quality reward | 74.23 | +1.90 |
| TAIScore | 75.11 | +2.78 |
| (b) Critique-content controls | ||
| No critique | 73.30 | +0.97 |
| Generic critique | 73.34 | +1.01 |
| Shuffled TAIScore critique | 72.88 | +0.55 |
| Matched TAIScore critique | 75.11 | +2.78 |
Direct refinement on WritingBench before actor DPO. All conditions share prompts, initial responses, reviser, decoding, and evaluator; only the feedback varies. All revisions are included without preference filtering.
If critique usefulness is actor-conditioned, a critic should be most effective for the actor
it was tailored to. We fix the target actor to Qwen3-8B and train three critics,
tailored respectively to Llama-3.2-3B, Qwen3-4B, and
Qwen3-8B. All three start from the same Qwen3-8B checkpoint and use
the same TAIScore training procedure and budget; only the actor they are tailored to changes.
Each critic then constructs DPO pairs for the same target actor.
| Critic tailored to | WritingBench | Gain |
|---|---|---|
| None (base actor) | 72.33 | - |
Llama-3.2-3B |
74.62 | +2.29 |
Qwen3-4B |
75.08 | +2.75 |
Qwen3-8B (matched) |
75.96 | +3.63 |
The target actor is Qwen3-8B in every condition. All critics share the same
initialization and training setup and differ only in the actor to which they are tailored.
Critique behavior learned with one actor remains useful to others, even across model
families: the Llama-3.2-3B-tailored critic still gains +2.29. But the ordering
tracks how closely the tailoring actor matches the target: matching the critic to the target
actor yields the largest gain, outperforming the Qwen3-4B-tailored critic by 0.88
points and the Llama-3.2-3B-tailored critic by 1.34 points.
@article{kim2026coevolving,
title = {Co-Evolving Actor-Conditioned Critics for Non-Verifiable Generation},
author = {Kim, Jinyoung and Khalifa, Muhammad and Logeswaran, Lajanugen and
Kim, Jaekyeom and Lee, Moontae and Lee, Honglak and Wang, Lu},
journal = {arXiv preprint arXiv:2608.30397},
year = {2026},
eprint = {2608.30397},
archivePrefix = {arXiv},
primaryClass = {cs.CL},
url = {https://arxiv.org/abs/2608.30397}
}