Token-level reinforcement learning for video generation. Every clip on this page is HunyuanVideo-1.5 post-trained with TVRL.
Scroll
How TVRL works
One reward signal, two jobs: score the video and say where to update it

Teacher-forced answer likelihood of each prompt-derived check. Averaged and group-normalized, it sets the sign and strength of the update, exactly as in GRPO.
The gradient magnitude of the same likelihood with respect to the video input is a reward-sensitivity map. Aggregated over a 3×3 window, it becomes the token credit.
Credit reweights the per-token transition log-probabilities before the policy ratio is formed and clipped. Uniform credit recovers scalar-reward GRPO.
Token-Level Video
Reinforcement Learning
A video is not uniformly flawed, yet video GRPO broadcasts one scalar advantage to every token. TVRL derives token-level credit from the reward being optimized. The answer likelihood of a frozen vision–language model scores each rollout, and the magnitude of its video-input gradient reveals which tokens most affect that score. The advantage still decides whether a rollout is reinforced or suppressed; the credit decides where.
Yifan Wang1 · Gordon Guocheng Qian† · Yanyu Li · Anil Kag · Yun Fu1
1Northeastern University†Corresponding author
Gallery
Every comparison from the paper, as video. Each row uses one prompt and one seed for all methods; hover a row to play it in sync, drag its bar to scrub, click a clip to enlarge it.
TVRL vs. the base model
HunyuanVideo-1.5 before and after TVRL post-training with the Qwen3.5-9B critic, on VideoGen-Eval prompts.
TVRL vs. GRPO baselines
Dance-GRPO, Flow-GRPO, and SAGE-GRPO trained with the same Qwen3.5-9B reward, on SAGE-GRPO validation prompts. The row label names the requirement to check.
Effect of the frozen critic
TVRL with VideoAlign, VideoScore2, Qwen3.5-4B, or Qwen3.5-9B as the reward and credit model, on VideoGen-Eval prompts.
Results
| Reward model | Method | Overall | Δ | Creativity | Common sense | Control | Human | Physics |
|---|---|---|---|---|---|---|---|---|
| Base model | – | 54.09 | – | 41.40 | 62.75 | 30.26 | 88.94 | 47.11 |
| VideoAlign | GRPO | 54.18 | – | 41.44 | 61.14 | 30.77 | 90.06 | 47.49 |
| TVRL | 55.54 | +1.36 | 45.11 | 61.16 | 32.09 | 90.21 | 49.15 | |
| VideoScore2 | GRPO | 54.66 | – | 42.23 | 64.89 | 30.29 | 91.52 | 44.35 |
| TVRL | 55.99 | +1.33 | 42.08 | 64.60 | 31.33 | 90.79 | 49.13 | |
| UnifiedReward2 | GRPO | 54.82 | – | 42.90 | 62.14 | 30.25 | 88.90 | 50.90 |
| TVRL | 56.67 | +1.85 | 45.36 | 64.55 | 31.29 | 89.87 | 52.28 | |
| Qwen3.5-9B | GRPO | 54.54 | – | 41.68 | 64.88 | 31.57 | 88.85 | 45.74 |
| TVRL | 57.69 | +3.15 | 47.36 | 64.31 | 31.64 | 90.76 | 54.37 |
HunyuanVideo-1.5 with the SAGE sampler. For each reward model, GRPO broadcasts the scalar advantage uniformly and TVRL routes it with the gradient map of the same reward; Δ is the Overall gain. Every row averages three runs with independent training and generation seeds, evaluated after 100 optimizer steps.
| SDE sampler | GRPO | TVRL | Δ |
|---|---|---|---|
| SAGE | 54.54 | 57.69 | +3.15 |
| Flow | 53.79 | 56.71 | +2.92 |
| Dance | 50.84 | 53.52 | +2.68 |
VBench-2.0 Overall with the Qwen3.5-9B reward and critic held fixed while the stochastic sampler changes. Dance GRPO averages two runs; all other entries average three.
| Credit routing | Overall | Δ vs. uniform |
|---|---|---|
| Uniform (GRPO) | 54.54 | – |
| Frame-level | 56.47 | +1.93 |
| 7×7 window | 57.08 | +2.54 |
| 3×3 window (default) | 57.69 | +3.15 |
| 1×1, unsmoothed | 55.10 | +0.56 |
| 3×3, shuffled | 53.70 | −0.84 |
| DINOv2 feature map | 55.48 | +0.94 |
SAGE sampler, Qwen3.5-9B reward. Localizing credit helps up to a 3×3 window; finer, unsmoothed maps amplify noisy background tokens. Shuffling the weights removes the gain, and a reward-agnostic DINOv2 map helps less than credit derived from the reward itself.
| TVRL vs. | Win | Loss | Tie | Win rate, excl. ties |
|---|---|---|---|---|
| SAGE-GRPO | 36.3% | 23.3% | 40.4% | 60.9% |
| Base model | 31.3% | 19.2% | 49.5% | 62.0% |
Blind pairwise judgments of text–video alignment on the first 200 VideoGen-Eval prompts, three seeds per prompt, eight annotators.