Multi-Hop Data Synthesis for Generalizable Vision-Language Reasoning

[Submitted on 17 Mar 2026 (v1), last revised 19 Mar 2026 (this version, v2)]

Authors:Shenzhi Wang, Shixuan Liu, Jing Zhou, Chang Gao, Xiong-Hui Chen, Binghai Wang, An Yang, Shiji Song, Bowen Yu, Gao Huang, Junyang Lin

View a PDF of the paper titled HopChain: Multi-Hop Data Synthesis for Generalizable Vision-Language Reasoning, by Shenzhi Wang and 10 other authors

View PDF
HTML (experimental)

Abstract:Vision-language models (VLMs) show strong multimodal capabilities but still struggle with fine-grained vision-language reasoning. We find that long chain-of-thought (CoT) reasoning exposes diverse failure modes, including perception, reasoning, knowledge, and hallucination errors, which can compound across intermediate steps. However, most existing vision-language data used for reinforcement learning with verifiable rewards (RLVR) does not involve complex reasoning chains that rely on visual evidence throughout, leaving these weaknesses largely unexposed. We therefore propose HopChain, a scalable framework for synthesizing multi-hop vision-language reasoning data for RLVR training of VLMs. Each synthesized multi-hop query forms a logically dependent chain of instance-grounded hops, where earlier hops establish the instances, sets, or conditions needed for later hops, while the final answer remains a specific, unambiguous number suitable for verifiable rewards. We train Qwen3.5-35B-A3B and Qwen3.5-397B-A17B under two RLVR settings: the original data alone, and the original data plus HopChain’s multi-hop data, and compare them across 24 benchmarks spanning STEM and Puzzle, General VQA, Text Recognition and Document Understanding, and Video Understanding. Although this multi-hop data is not synthesized for any specific benchmark, it improves 20 of 24 benchmarks on both models, indicating broad and generalizable gains. Consistently, replacing full chained queries with half-multi-hop or single-hop variants reduces the average score across five representative benchmarks from 70.4 to 66.7 and 64.3, respectively. Notably, multi-hop gains peak in long-CoT vision-language reasoning, exceeding 50 points in the ultra-long-CoT regime. These experiments establish HopChain as an effective, scalable framework for synthesizing multi-hop data that improves generalizable vision-language reasoning.

Submission history

From: Shenzhi Wang [view email]
[v1]
Tue, 17 Mar 2026 18:04:58 UTC (4,951 KB)
[v2]
Thu, 19 Mar 2026 04:12:29 UTC (4,951 KB)

What's Hot

At Least 32 People Dead After a Mine Bridge Collapsed Due to Overcrowding

Here’s how I turned a Raspberry Pi into an in-car media server

Beloved SF cat’s death fuels Waymo criticism

Multi-Hop Data Synthesis for Generalizable Vision-Language Reasoning

How to Measure AI Value

What Really Controls Temporal Reasoning in Large Language Models: Tokenisation or Representation of Time?

The Math That’s Killing Your AI Agent

Agent Control Protocol: Admission Control for Agent Actions

Building Robust Credit Scoring Models (Part 3)

[2510.16001] An Order-Sensitive Conflict Measure for Random Permutation Sets

At Least 32 People Dead After a Mine Bridge Collapsed Due to Overcrowding

Here’s how I turned a Raspberry Pi into an in-car media server

Beloved SF cat’s death fuels Waymo criticism

9 types of Google Ads (pros, cons, and when to use each)

Google tightens rules on out-of-stock product pages

Multi-Hop Data Synthesis for Generalizable Vision-Language Reasoning

Instagram for small business: 2026 guide to growth

Google launches Ads DevCast Vodcast for developers

What Really Controls Temporal Reasoning in Large Language Models: Tokenisation or Representation of Time?

Most Popular

13 Trending Songs on TikTok in Nov 2025 (+ How to Use Them)

How to watch the 2026 GRAMMY Awards online from anywhere

Corporate Reputation Management Strategies | Sprout Social

Our Picks

At Least 32 People Dead After a Mine Bridge Collapsed Due to Overcrowding

Here’s how I turned a Raspberry Pi into an in-car media server

Beloved SF cat’s death fuels Waymo criticism

Subscribe to Updates

What's Hot

Multi-Hop Data Synthesis for Generalizable Vision-Language Reasoning

Submission history

Related Posts

Subscribe to Updates