CPPMN
Cross-Modal Progressive Perspective Matching for Remote Sensing Image–Text Retrieval
IEEE Transactions on Multimedia, 2025

Illustrative placeholder generated for this site; it is not a figure or result from the paper.
01 — Overview
Overview
Remote-sensing scenes can be described from multiple geographic perspectives. Retrieval systems that collapse those perspectives into one representation may match a query to the wrong region or overlook the relevant spatial relationship.
CPPMN progressively learns full-image perspectives, exposes perspective-specific cross-modal relationships, and then aligns image and language features semantically.
02 — Contributions
Key Contributions
- 01
Uses positive text descriptions to supervise full-perspective visual feature learning.
- 02
Transforms implicit perspective features into explicit cross-modal relationship graphs.
- 03
Applies cascaded Transformer layers for progressive image–text semantic alignment.
03 — Method
Method
The network combines a compensation module for full-perspective modeling, a graph transformation module for locating individual perspectives, and a cascaded Transformer for cross-modal semantic alignment. Graph density and connectivity help identify the perspective referred to by the query.
04 — Evaluation
Results
Quantitative and qualitative experiments across four remote-sensing image–text retrieval datasets demonstrate the value of progressive perspective matching and semantic alignment.
05 — Reference
Citation
BibTeX citation
@Article{Zheng_2025_CPPMN,
author = {Zheng, Chengyu and Li, Xiu and Liang, Xinyue and Huang, Lei and Du, Shan and Nie, Jie and Dong, Junyu},
title = {Cross-Modal Progressive Perspective Matching Network for Remote Sensing Image-Text Retrieval},
journal = {IEEE Transactions on Multimedia},
year = {2025},
volume = {27},
pages = {3966--3978},
doi = {10.1109/TMM.2025.3535365}
}