Secure and Trustworthy Multimodal Intelligence


Welcome to CogModal Group! (Cognition-Inspired Cross-Modal Intelligence Group). Led by Prof. Jing Yu, we study model security for multimodal and generative AI. Our goal is to make models that connect language, vision, audio, and distributed data not only capable, but also robust, privacy-preserving, verifiable, and accountable.

Our recent work follows three connected directions. First, we protect the intellectual property of vision-language, multimodal federated, and diffusion models through embedding watermarks, stealthy black-box verification, and credential-controlled generation. Second, we secure collaborative learning against model extraction, malicious clients, backdoors, and privacy leakage through robust aggregation and privacy-preserving federated learning. Third, we investigate adversarial vulnerabilities in vision-language reasoning systems to support stronger evaluation and defense. Together, these efforts advance multimodal intelligence from high performance toward secure, trustworthy, and responsible deployment.


🌟 You are the th visitor of our group!

Secure and Trustworthy Multimodal Intelligence

Join us!

Albert Einstein has said ‘the eternal mystery of the world is its comprehensibility.’(爱因斯坦曾有这样一句话:“世界的永恒之谜是它是可理解的”) We are looking for you (self-motivated graduate candidates, undergraduate students and visiting students) to join us to explore the answer together!
🌟 Email: yujing02 at iie dot ac dot cn

🌟 Zhihu link: Cognition-Inspired Cross-Modal Intelligent Articles

News

 
 
 
 
 
1 new preprint has been released on arXiv, 2026 !
Apr 2026 – Apr 2026
Jialing Wang, Yue Zhao, Yuhao Zhang, Jing Yu, Shaosai Li, Zhanchen Dai, Benyou Wang, Haizhou Li. Ti-Audio: The First Multi-Dialectal End-to-End Speech LLM for Tibetan. arXiv, 2026. The paper is available here.
 
 
 
 
 
1 paper has been published by AAAI, 2026 !
Mar 2026 – Mar 2026
Yuanmin Tang, Jing Yu*, Keke Gai, Gang Xiong, Gaopeng Gou, Meikang Qiu, Qi Wu. Manipulation Intention Understanding for Zero-Shot Composed Image Retrieval. Proceedings of the AAAI Conference on Artificial Intelligence, 2026. The paper is available here.
 
 
 
 
 
3 papers have been accepted by CVPR, 2025 !
Feb 2025 – Feb 2025
Yuanmin Tang, Jing Yu*, Keke Gai, Jiamin Zhuang, Gang Xiong, Gaopeng Gou, Qi Wu. Missing Target-Relevant Information Prediction for Accurate Zero-Shot Composed Image Retrieval. CVPR, 2025. Yuanmin Tang, Jue Zhang, Xiaoting Qin, Jing Yu, Gaopeng Gou, Gang Xiong, Qingwei Lin, Saravan Rajmohan, Dongmei Zhang, Qi Wu. Reason-before-Retrieve: One-Stage Reflective Chain-of-Thoughts for Training-Free Zero-Shot Composed Image Retrieval. CVPR, 2025. Xiangyan Qu, Gaopeng Gou, Jiamin Zhuang, Jing Yu, Kun Song, Qihao Wang, Yili Li, Gang Xiong. ProAPO: Progressively Automatic Prompt Optimization for Visual Classification. CVPR, 2025.
 
 
 
 
 
1 paper has been accepted by IJCAI, 2025 !
Apr 2025 – Apr 2025
Liangqi Lei, Keke Gai, Jing Yu*, Liehuang Zhu, Qi Wu. Secure and Efficient Watermarking for Latent Diffusion Models in Model Distribution Scenarios. IJCAI, 2025.
 
 
 
 
 
1 paper has been accepted by AAAI, 2025 !
Dec 2024 – Dec 2024
Keke Gai, Dongjue Wang, Jing Yu*, Mohan Wang, Liehuang Zhu, Qi Wu. MFL-Owner: Ownership Protection for Multi-modal Federated Learning via Orthogonal Transform Watermark. AAAI, 2025. The paper is available here.
 
 
 
 
 
Several papers have been accepted by SecureComm, WASA and BigDataSecurity, 2025 !
Jun 2025 – Jun 2025
Dongjue Wang, Keke Gai, Jing Yu*, An Wang, Zhijing Cao, Liehuang Zhu. PPCM-Fed: Privacy-Preserving Cross-Modal Federated Learning in IoT. SecureComm, 2025 (Best Paper Award, CCF-C). Jing Qi, An Wang, Keke Gai, Jing Yu*, Liehuang Zhu. Intelligent Re-encryption and Commitment-Driven Dynamic Data Sharing. WASA, 2025 (Best Student Paper Award, CCF-C). Yuqing Zhang, Keke Gai, Jing Yu*, Kai Ding. Relationship Sharing-based Trustworthy Verifiable Multi-Party Verification in Decentralized Identity. BigDataSecurity, 2025 (Best Paper Award). Jiamin Zhuang, Jing Yu*, Xiangyan Qu, Yuanmin Tang, Gaopeng Gou, Gang Xiong, Qi Wu. Soft Multi-view Representation Learning for Disambiguating Text-Based Person Retrieval. WASA, 2025.
 
 
 
 
 
Several journal papers have been published, 2025 !
Jan 2025 – Jan 2025
Several journal papers by Jing Yu and collaborators have been published in IEEE Transactions on Information Forensics and Security (TIFS), IEEE Transactions on Dependable and Secure Computing (TDSC), Neurocomputing, Journal of Systems Architecture, and Blockchains.
 
 
 
 
 
Jing Yu gave an invited talk.
Dec 2023 – Dec 2023
Jing Yu gave an invited talk “Large Multi-modal Model Technologies and Applications in Rural Vitalization” in 2023 CCF China Blockchain Technology and Application Summit Forum . The news is available here.
 
 
 
 
 
1 paper has been accepted by AAAI, 2024 !
Dec 2023 – Dec 2023
Yuanmin Tang, Jing Yu*, Keke Gai, Jiamin Zhuang, Gang Xiong, Yue Hu, Qi Wu. Context-I2W: Mapping Images to Context-dependent Words for Accurate Zero-Shot Composed Image Retrieval. AAAI, 2024. The paper is available here.
 
 
 
 
 
1 paper has been accepted by Blockchain, 2023 !
Nov 2023 – Nov 2023
Zhiqi Lei, Keke Gai, Jing Yu, Liehuang Zhu, Kim-Kwang Raymond Raymond Choo, Efficiency-Enhanced Blockchain-Based Client Selection in Heterogeneous Federated Learning. Blockchain, 2023.
 
 
 
 
 
1 paper has been accepted by ICDM, 2023 !
Nov 2023 – Nov 2023
Shuo Wang, Keke Gai, Jing Yu, Liehuang Zhu. BDVFL: Blockchain-based Decentralized Vertical Federated Learning. ICDM, 2023 (CCF-B)
 
 
 
 
 
Jing Yu organized a forum.
Aug 2023 – Aug 2023
Jing Yu organized the forum about “Computer science and technology workers practice the spirit of scientists, how to make steady progress”, the news is available here
 
 
 
 
 
Jing Yu invited Professor Bang Liu to give a talk.
Aug 2023 – Aug 2023
Jing Yu invited Professor Bang Liu to give a talk about “Natural Language Processing for Materials Science”.
 
 
 
 
 
Jing Yu invited Professor Qi Wu to give a talk.
Jun 2023 – Jun 2023
Jing Yu invited Professor Qi Wu to give a talk about “GPT for Vision-and-Language Navigation, VLN”.
 
 
 
 
 
1 paper has been accepted by TMM 2023 !
May 2023 – May 2023
Jiamin Zhuang, Jing Yu(corresponding author), Yang Ding, Xiangyan Qu, Yue Hu. Towards Fast and Accurate Image-Text Retrieval with Self-Supervised Fine-Grained Alignment, IEEE Transactions on Multimedia (TMM), (Impact factor: 6.513). The paper is available here
 
 
 
 
 
Jing Yu gave an invited talk.
Apr 2022 – Apr 2022
Jing Yu gave an invited talk “ MuKEA: Multimodal Knowledge Extraction and Accumulation for Knowledge-based Visual Question Answering” in MSRA CVPR 2022 workshop . The slides are available here. The workshop news is available here. The photos are available here
 
 
 
 
 
Yang Ding gave a poster presentation !
Apr 2022 – Apr 2022
Yang Ding gave an poster presentation of the paper “ MuKEA: Multimodal Knowledge Extraction and Accumulation for Knowledge-based Visual Question Answering” in MSRA CVPR 2022 workshop. The photos are available here.
 
 
 
 
 
Welcome Siyuan Feng (2022 Master student) join our team !
Apr 2022 – Present
 
 
 
 
 
Prof. Jing Yu gave an invited talk !
Mar 2022 – Mar 2022
  • Jing Yu gave an invited online talk “ET-BERT: A Contextualized Datagram Representation with Pre-training Transformers for Encrypted Traffic Classification” in the Information Engineering University. The slides are available here.
 
 
 
 
 
1 paper has been accepted by CVPR 2022 !
Mar 2022 – Mar 2022
  • Yang Ding, Jing Yu(corresponding author), Bang Liu, Yue Hu, Mingxin Cui, Qi Wu. MuKEA: Multimodal Knowledge Extraction and Accumulation for Knowledge-based Visual Question Answering,CVPR, 2022.
 
 
 
 
 
Prof. Jing Yu gave an invited talk !
Mar 2022 – Mar 2022
  • Prof. Jing Yu gave an invited talk in Shanghai University about “Introduction of Scientific Research and Paper writing”. The slides are available here. The news about this talk can be found here.
 
 
 
 
 
1 paper has been accepted by WWW 2022 !
Feb 2022 – Present
  • ET-BERT: A Contextualized Datagram Representation with Pre-training Transformers for Encrypted Traffic Classification * WWW videos
 
 
 
 
 
1 paper has been accepted by ICASSP 2022 !
Dec 2021 – Present
  • WLinker: Modeling Relational Triplet Extraction as Word Linking
 
 
 
 
 
1 paper has been accepted by Knowledge-Based Systems (KBS) 2021 !
Dec 2021 – Present
  • APER: AdaPtive Evidence-driven Reasoning Network for machine reading comprehension with unanswerable questions.
  • Impact factor: 5.921 !
 
 
 
 
 
1 paper has been accepted by ICML 2021 !
Jul 2021 – Present
  • Evolving Attention with Residual Convolutions
 
 
 
 
 
1 paper has been accepted by IJCAI 2021 !
Apr 2021 – Present
  • CogTree: Cognition Tree Loss for Unbiased Scene Graph Generation
 
 
 
 
 
Jing Yu gave an invited talk.
Apr 2021 – Present
  • Jing Yu gave an invited talk “Towards Cognition-Inspired Vision and Language Methods”in the CCF YOCSEF Xi`an Forum “How does Vision and Language 1+1>2?”
  • The slides are available here.
 
 
 
 
 
1 paper has been accepted by EACL 2021 !
Feb 2021 – Present
  • Syntax-BERT: Improving Pre-trained Transformers with Syntax Trees.
 
 
 
 
 
1 paper has been accepted by Information Fusion, 2021 !
Feb 2021 – Present
  • DMRFNet: Deep Multimodal Reasoning and Fusion for Visual Question Answering and explanation generation.
  • Impact factor: 12.975 !
 
 
 
 
 
1 paper has been accepted by ICASSP 2021 !
Jan 2021 – Present
  • MCR-NET: A Multi-Step Co-Interactive Relation Network for Unanswerable Questions on Machine Reading Comprehension.
 
 
 
 
 
Welcome new students !
Jan 2021 – Present
Welcome Yaochen Ren (2021 PhD student) , Xinjie Lin (2018 PhD student), Minghao Jiang (2018 PhD student) and Yu Wang (2017 PhD student) join our team!
 
 
 
 
 
Jing Yu gave an invited talk.
Oct 2020 – Present
  • Jing Yu will give a talk of “Deep Reasoning for Visual Question Answering” in Shanghai University & online. The talk is on 15/10/2020, 10:00~12:00 a.m., UTC+8.
  • More information about the talk is avaible here.
 
 
 
 
 
1 paper has been accepted by IEEE Transactions on Image Processing !
Oct 2020 – Present
  • Learning Dual Encoding Model for Adaptive Visual Understanding in Visual Dialogue.
  • Impact factor: 9.34 !
 
 
 
 
 
1 paper has been accepted by ACM MM 2020 !
Aug 2020 – Present
  • KBGN: Knowledge-Bridge Graph Network for Adaptive Vision-Text Reasoning in Visual Dialogue.
 
 
 
 
 
1 paper has been accepted by Pattern Recognition Letters (PRL) !
Aug 2020 – Present
  • Learning cross-modal correlations by exploring inter-word semantics and stacked co-attention.
 
 
 
 
 
1 paper has been accepted by Pattern Recognition !
Jul 2020 – Present
  • Cross-modal knowledge reasoning for knowledge-based visual question answering.
  • Impact factor: 7.196 !
 
 
 
 
 
1 paper has been accepted by Knowledge-based Systems !
Jul 2020 – Present
  • Cross-modal learning with prior visual relation knowledge.
  • Impact factor: 5.921 !
 
 
 
 
 
1 paper has been accepted by ICIP 2020 !
May 2020 – Present
  • Prior Visual Relationship For Visual Question Answering.
 
 
 
 
 
Jing Yu gave an invited talk.
May 2020 – Present
  • Jing Yu gave an invited talk of “Deep Learning based Visual Question Answering” in Shanghai University, Shanghai, China.
  • The slides are available here.
 
 
 
 
 
2 papers have been accepted by IJCAI2020 !
Apr 2020 – Present
  • Mucko: Multi-Layer Cross-Modal Knowledge Reasoning for Fact-based Visual Question Answering. 代码链接

  • DAM: Deliberation, Abandon and Memory Networks for Generating Detailed and Non-repetitive Responses in Visual Dialogue.

  • accept rate: 592/4717 = 12.6%

 
 
 
 
 
1 paper has been accepted by Information Fusion !
Mar 2020 – Present
  • Multimodal feature fusion by relational reasoning and attention for visual question answering
  • Impact factor: 12.975
 
 
 
 
 
1 paper has been accepted by IEEE Transactions on Multimedia !
Jan 2020 – Present
  • Reasoning on the Relation: Enhancing Visual Representation for Visual Question Answering and Cross-modal Retrieval
  • Impact factor: 6.051
 
 
 
 
 
2 paper has been accepted by AAAI 2020 !
Nov 2019 – Present
  • Deep Visual Understanding Like Humans: An Adaptive Dual Encoding Model for Visual Dialogue.
  • Unsupervised Learning of Deterministic Dialogue Structure with Edge-Enhanced Graph Auto-Encoder
 
 
 
 
 
1 paper has been accepted by KSEM 2019 !
May 2019 – Present
  • Semantic Modeling of Textual Relationships in Cross-modal Retrieval.