CANOPY: Adaptive-Granularity Evidence Compression for Multimodal RAG
Abstract
Multimodal RAG retrieves text, tables, images, and videos, but choosing a retrieval granularity does not determine how much context to retain within each item. Coarse units include irrelevant content, while uniformly fine selection can remove context needed to interpret the evidence. Existing compressors address this trade-off with modality-specific mechanisms, leaving open a shared procedure for adapting the retained extent region by region across heterogeneous items. We introduce CANOPY (Canonical Projection over Hierarchy), a framework for adaptive-granularity post-retrieval evidence compression. CANOPY represents retrieved items as hierarchies and uses a node encoder fine-tuned on gold evidence to score regions against the query. Parent-relative refinement compares these scores to select multiple regions at different granularities without LLM calls for node-level pruning. Because compression cannot recover evidence that was never retrieved, a critic requests targeted follow-up retrieval when it judges the accumulated evidence insufficient; newly retrieved items are compressed before being added. Across five QA benchmarks over a 33M-item heterogeneous corpus, CANOPY achieves higher average answer accuracy than the evaluated retrieval baselines. Ablations indicate that additional retrieval drives the main accuracy gains on multi-hop QA. In the unrouted Qwen3-VL-8B-Instruct setting, compression reduces reader-input evidence tokens by 14.2-27.7% relative to the same iterative pipeline without compression, with comparable answer accuracy.
Community
Multimodal RAG decides which items to retrieve, but not how much of each item the reader actually needs. We introduce CANOPY, which represents each retrieved text, table, or video as a hierarchy of original regions and uses a fine-tuned node encoder with a simple parent-relative rule to keep multiple regions at different granularities within each item, with no LLM calls for node-level pruning. A critic triggers targeted follow-up retrieval when the evidence is insufficient.
On 5 QA benchmarks over a 33M-item heterogeneous corpus, CANOPY achieves the best average accuracy among the evaluated retrieval baselines, and compression reduces reader-input evidence tokens by 14-28% at comparable accuracy.
This is an automated message from the Librarian Bot. I found the following papers similar to this paper.
The following papers were recommended by the Semantic Scholar API
- Semantic Compression Trees: Multi-Resolution Knowledge Retrieval via Hierarchical Semantic Residuals (2026)
- CLIMB: Confidence-Guided Complementary Evidence for Multimodal Retrieval-Augmented Generation (2026)
- Hi-Q: Hierarchical Evidence-guided Query Refinement for Multi-Hop Question Answering (2026)
- REVEAL: A Rubric-Guided Agent for Explicit Evidence Sufficiency Verificationin Long-Video Question Answering (2026)
- Beyond One-Shot Expansion: Contrastive Evidence Exploration for Multi-Hop Retrieval (2026)
- EviReform: Evidence-Guided Query Reformulation for Multi-Hop Graph Retrieval (2026)
- Before Answering: Evidence Sufficiency under Size-Matched Memory Construction (2026)
Please give a thumbs up to this comment if you found it helpful!
If you want recommendations for any Paper on Hugging Face checkout this Space
You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: @librarian-bot recommend
Models citing this paper 0
No model linking this paper
Datasets citing this paper 0
No dataset linking this paper
Spaces citing this paper 1
Collections including this paper 0
No Collection including this paper