TVRL TL;DR Token-level credit derived from the reward itself: the gradient magnitude of a frozen VLM's answer likelihood routes each GRPO update to the video tokens that score depends on.

Token-level reinforcement learning for video generation. Every clip on this page is HunyuanVideo-1.5 post-trained with TVRL.

Scroll

How TVRL works

One reward signal, two jobs: score the video and say where to update it

TVRL pipeline: prompt-derived yes/no checks are scored by a frozen VLM; the scores form one group-relative advantage per video, and the gradients of the same scores give question-conditioned weights that route the clipped GRPO update over video tokens.
Pipeline. Each prompt is decomposed offline into yes/no checks. A frozen VLM scores every rollout by the teacher-forced likelihood of the reference answer; the scores are averaged and normalized within the rollout group into one advantage per video. The magnitude of the gradient of the same scores with respect to the video-frame input gives detached, question-conditioned weights that reweight the dense denoising log-probabilities inside the clipped GRPO ratio. No VLM gradient reaches the generator.
01Score

Teacher-forced answer likelihood of each prompt-derived check. Averaged and group-normalized, it sets the sign and strength of the update, exactly as in GRPO.

02Localize

The gradient magnitude of the same likelihood with respect to the video input is a reward-sensitivity map. Aggregated over a 3×3 window, it becomes the token credit.

03Route

Credit reweights the per-token transition log-probabilities before the policy ratio is formed and clipped. Uniform credit recovers scalar-reward GRPO.

TL;DR

Token-Level Video
Reinforcement Learning

A video is not uniformly flawed, yet video GRPO broadcasts one scalar advantage to every token. TVRL derives token-level credit from the reward being optimized. The answer likelihood of a frozen vision–language model scores each rollout, and the magnitude of its video-input gradient reveals which tokens most affect that score. The advantage still decides whether a rollout is reinforced or suppressed; the credit decides where.

# credit from the reward's own input gradientr = log_p_vlm("Yes" | video, q) w = normalize(window3(abs(grad(r, video)))) logp = (w * logp_transition).sum()

Yifan Wang1 · Gordon Guocheng Qian† · Yanyu Li · Anil Kag · Yun Fu1

1Northeastern University†Corresponding author

Trajectory-level GRPO penalizes every token of a failed rollout; TVRL concentrates the penalty on the tokens the VLM check is sensitive to.
For a failed rollout, trajectory-level GRPO penalizes every video token uniformly. TVRL keeps the same advantage and concentrates it on the tokens to which the VLM check is most sensitive; the same map reinforces those tokens when the advantage is positive.

Results

VBench-2.0
Watch the comparison videos ↑
Reward modelMethodOverallΔCreativityCommon senseControlHumanPhysics
Base model–54.09–41.4062.7530.2688.9447.11
VideoAlignGRPO54.18–41.4461.1430.7790.0647.49
TVRL55.54+1.3645.1161.1632.0990.2149.15
VideoScore2GRPO54.66–42.2364.8930.2991.5244.35
TVRL55.99+1.3342.0864.6031.3390.7949.13
UnifiedReward2GRPO54.82–42.9062.1430.2588.9050.90
TVRL56.67+1.8545.3664.5531.2989.8752.28
Qwen3.5-9BGRPO54.54–41.6864.8831.5788.8545.74
TVRL57.69+3.1547.3664.3131.6490.7654.37

HunyuanVideo-1.5 with the SAGE sampler. For each reward model, GRPO broadcasts the scalar advantage uniformly and TVRL routes it with the gradient map of the same reward; Δ is the Overall gain. Every row averages three runs with independent training and generation seeds, evaluated after 100 optimizer steps.