Co-Evolving Actor-Conditioned Critics
for Non-Verifiable Generation

1University of Michigan   2LG AI Research   3University of Illinois at Chicago

Abstract

Natural-language critiques provide supervision beyond scalar rewards for non-verifiable generation, which lacks deterministic verifiers. In critique-guided refinement, a critic gives feedback on an initial response and an actor revises it. However, final revision quality does not reveal whether the critique was actually useful: a capable actor may improve without following the feedback, while valid feedback may fail if the actor cannot execute it.

We frame critique as actor-conditioned revision guidance, where usefulness depends on whether the feedback helps the target actor address the intended weakness. We introduce TAIScore (Targeted Actionable Improvement Score), a reward that evaluates the instruction, initial response, critique, and revision together, assessing whether the critique targets a real weakness, whether the actor follows it, and whether the intended aspect improves. We use this reward to train an actor-tailored critic with GRPO, and use critique-guided refinements to construct DPO preference pairs for the actor, forming a co-evolving critic-actor loop where the critic adapts to the actor's changing capability.

Experiments show that an 8B critic trained with TAIScore outperforms both a zero-shot 120B critic and critics trained with outcome-only or critique-only reward signals. Co-evolving the critic and actor further improves performance, suggesting that effective critique supervision should adapt as the actor changes.


Critique Usefulness Is Actor-Conditioned

We run two controlled analyses on 1,000 WritingBench examples, varying one role at a time, to test whether critique usefulness is intrinsic to the critique or conditioned on the actor. The result is a clear asymmetry. A larger critic (8B→32B for Qwen, 8B→70B for Llama), with the initial response and revising actor held fixed, consistently improves standalone critique quality but yields little average improvement in how well the actor incorporates the feedback or in downstream gain. For a Qwen3-8B refiner, critique adherence actually drops by 0.223 under the larger critic. A larger refiner, given the same critique, substantially improves both adherence (+1.29 on average) and downstream gain (+0.39). Critique usefulness is therefore a property of the critique-actor pair, not of the critique alone.

Controlled analysis: larger critic vs. larger refiner on critique adherence and gain

Controlled zero-shot analysis of critique-guided refinement on WritingBench. Bars show mean differences between larger- and smaller-model conditions in critique quality, critique adherence (how well y1 implements c), and response-quality gain S(y1)−S(y0). Gray bars compare larger and smaller critics while holding y0 and the refiner fixed; blue bars compare larger and smaller refiners while holding y0 and the critique fixed. Larger critics improve standalone critique quality but yield little average improvement in adherence or gain, whereas larger refiners substantially improve both.

The reachability gap. For a martial-arts story prompt, a Qwen3-32B critique scores higher in standalone quality than a Qwen3-8B critique (9.00 vs. 8.33), yet yields far lower adherence from the same Qwen3-8B refiner (3.50 vs. 8.75). The 8B critic requests localized changes the refiner can carry out; the 32B critic proposes a coordinated redesign spanning symbolism, character motivation, climax, ending, and dialogue, of which the refiner adopts only isolated pieces. Feedback that looks stronger in isolation can be less useful when its requested changes exceed the target actor's revision capability.

Method

Our approach has three components that together form a co-evolving loop:

  1. TAIScore. Given a full rollout τ = (x, y0, c, y1), a judge first produces four diagnostic scores (critique validity, critique adherence, targeted improvement, and prompt faithfulness, the last acting as a guardrail), then produces a final scalar reward T(τ) ∈ [1, 10] within the same inference pass. The diagnostic scaffold ensures the reward reflects actor-conditioned usefulness rather than standalone critique quality or final revision quality alone.
  2. Critic update (GRPO). For each on-policy initial response y0, the critic samples N = 4 critiques. The same actor revises the same initial response under each critique, so every group controls for the prompt, the initial response, and actor capability. All rollouts in the group are scored by TAIScore, and group-relative advantages reward critiques that provide more useful revision guidance than the alternatives for that prompt.
  3. Actor update (DPO). The actor produces revisions using critiques from the adapted critic. A critique-blind pairwise judge then compares each revision y1 against its own initial response y0; only revisions preferred over y0 are eligible (ties and preferences for y0 are discarded). Eligible rollouts are uniformly subsampled into preference pairs (y1 ≻ y0) for DPO.

These two updates alternate across rounds, forming a co-evolving loop: the critic continually adapts to the actor's current weaknesses and revision capability, while the actor internalizes progressively stronger critique-guided revisions.

Overview of co-evolving critic-actor training framework

Overview of our co-evolving critic-actor training framework. For the current actor πt, the critic κt samples multiple critiques for the same initial response y0. Each critique produces a rollout τi = (x, y0, ci, y1,i). TAIScore evaluates each rollout and produces a GRPO reward for critic adaptation. The actor then produces candidate revisions using critiques from the adapted critic κt+1; revisions preferred over their corresponding initial responses are selected to construct DPO preference pairs y1 ≻ y0 for updating the actor. Alternating these two updates yields co-evolving critics and actors.


Experimental Setup

  • Domains. Creative writing: trained on DeepWriting-20K prompts, evaluated on WritingBench and HelloBench (Open-Ended QA and Heuristic Text Generation subsets). Deep research: trained on OpenScholar queries, evaluated on DeepResearch-Gym (key-point recall, key-point contradiction, report quality). We use 6K training queries per domain.
  • Models. The actor and the critic are both Qwen3-8B. All actors perform on-policy self-refinement: the same model writes the initial response and the revision. gpt-oss-120B serves as the judge for TAIScore and all judge-based rewards, and is never updated or used as the actor (except in the off-the-shelf critic baseline).
  • Training. Critic: GRPO, N = 4 critiques per prompt, learning rate 1e-6, KL coefficient 0.02. Actor: DPO with β = 0.1, learning rate 1e-6, one epoch. Co-evolution alternates the two updates for three rounds (2K critic-training queries and 2K DPO pairs per round); single-stage methods use the matching totals of 6K and 6K in one stage. Runs use 4× NVIDIA RTX PRO 6000 Blackwell GPUs.
  • Evaluation. Official protocols throughout: the WritingBench evaluator, the OpenCompass HelloBench implementation with the official GPT-4o-mini judge, and the DeepResearch-Gym report-level scripts with GPT-4.1-mini over the officially retrieved ClueWeb22 materials. For each method we generate three independent output sets and report mean and standard deviation.

Main Results

All conditions use the same DPO actor-training pipeline and differ only in how the critiques used to construct the DPO pairs are produced. Across conditions we hold fixed the candidate queries, the number of candidate revisions, the pair-selection procedure, and the number of DPO pairs.

Method WritingBench HelloBench DeepResearch-Gym
Overall ↑ OEQA ↑ HTG ↑ KPR ↑ KPC ↓ Quality ↑
Qwen3-8B (base) 72.33 ±0.08 34.86 ±1.95 39.14 ±1.76 71.93 ±0.04 1.15 ±0.09 81.89 ±0.07
DPO pairs from off-the-shelf critics
gpt-oss-120B critic 75.41 ±0.14 36.01 ±2.36 50.93 ±2.18 73.46 ±0.09 1.11 ±0.08 82.51 ±0.06
DPO pairs from trained critics (8B)
Outcome-gain reward 75.63 ±0.09 35.35 ±2.68 45.14 ±2.48 74.37 ±0.11 1.12 ±0.09 82.38 ±0.07
Critique-quality reward 75.18 ±0.11 35.51 ±2.45 49.75 ±4.32 74.19 ±0.13 1.14 ±0.11 82.32 ±0.04
TAIScore (ours) 75.96 ±0.12 36.48 ±2.56 53.78 ±1.58 75.21 ±0.09 1.09 ±0.07 82.63 ±0.06
Co-evolving critic-actor training
TAIScore + co-evolution (ours) 76.72 ±0.14 39.84 ±2.92 54.62 ±1.92 76.14 ±0.13 1.03 ±0.05 83.15 ±0.06

Mean and standard deviation over three independent evaluation runs. Except for the base actor, all rows report the final performance of the Qwen3-8B actor after DPO training. Best results in bold; second-best underlined.

  • An 8B TAIScore critic outperforms the frozen 120B off-the-shelf critic on all six metrics, showing that critic scale alone does not determine downstream usefulness.
  • TAIScore outperforms both reward ablations. Outcome-gain training rewards better final revisions but cannot establish that the improvement is attributable to the critique; critique-quality training rewards plausible feedback without checking whether the actor can execute it.
  • Co-evolving the critic with the actor further improves performance across all six metrics, e.g., WritingBench 75.96→76.72 and HelloBench OEQA 36.48→39.84.

Direct Critique Utility, Before Actor Training

Does the advantage of TAIScore depend on the actor-training pipeline, or is the feedback itself simply better? We hold the prompts, initial responses (S(y0) = 72.33), Qwen3-8B reviser, revision prompt, decoding configuration, and evaluator fixed, and vary only the supplied feedback, with no actor update and no preference filtering.

Condition S(y1) Δ WritingBench
(a) Critic source / training objective
Off-the-shelf gpt-oss-120B 73.82+1.49
Outcome-gain reward 74.44+2.11
Critique-quality reward 74.23+1.90
TAIScore 75.11+2.78
(b) Critique-content controls
No critique 73.30+0.97
Generic critique 73.34+1.01
Shuffled TAIScore critique 72.88+0.55
Matched TAIScore critique 75.11+2.78

Direct refinement on WritingBench before actor DPO. All conditions share prompts, initial responses, reviser, decoding, and evaluator; only the feedback varies. All revisions are included without preference filtering.

  • The TAIScore-trained critic already yields the highest revision quality before any actor update, ahead of the off-the-shelf, outcome-gain, and critique-quality critics.
  • Panel (b) rules out generic second-pass improvement. Revising with no critique or a generic critique gains only +0.97 and +1.01, and a TAIScore critique sampled from a different example gains just +0.55, below the no-critique control, despite preserving the source and form of the feedback. The matched critique beats the strongest control by 1.77 points, so the gain depends on instance-specific, correctly matched critique content rather than on the actor's general ability to improve on a second pass.

Actor-Critic Matching

If critique usefulness is actor-conditioned, a critic should be most effective for the actor it was tailored to. We fix the target actor to Qwen3-8B and train three critics, tailored respectively to Llama-3.2-3B, Qwen3-4B, and Qwen3-8B. All three start from the same Qwen3-8B checkpoint and use the same TAIScore training procedure and budget; only the actor they are tailored to changes. Each critic then constructs DPO pairs for the same target actor.

Critic tailored to WritingBench Gain
None (base actor) 72.33-
Llama-3.2-3B 74.62+2.29
Qwen3-4B 75.08+2.75
Qwen3-8B (matched) 75.96+3.63

The target actor is Qwen3-8B in every condition. All critics share the same initialization and training setup and differ only in the actor to which they are tailored.

Critique behavior learned with one actor remains useful to others, even across model families: the Llama-3.2-3B-tailored critic still gains +2.29. But the ordering tracks how closely the tailoring actor matches the target: matching the critic to the target actor yields the largest gain, outperforming the Qwen3-4B-tailored critic by 0.88 points and the Llama-3.2-3B-tailored critic by 1.34 points.


BibTeX

@article{kim2026coevolving,
  title         = {Co-Evolving Actor-Conditioned Critics for Non-Verifiable Generation},
  author        = {Kim, Jinyoung and Khalifa, Muhammad and Logeswaran, Lajanugen and
                   Kim, Jaekyeom and Lee, Moontae and Lee, Honglak and Wang, Lu},
  journal       = {arXiv preprint arXiv:2608.30397},
  year          = {2026},
  eprint        = {2608.30397},
  archivePrefix = {arXiv},
  primaryClass  = {cs.CL},
  url           = {https://arxiv.org/abs/2608.30397}
}