Welcome to CogModal Group! (Cognition-Inspired Cross-Modal Intelligence Group). Led by Prof. Jing Yu, we study model security for multimodal and generative AI. Our goal is to make models that connect language, vision, audio, and distributed data not only capable, but also robust, privacy-preserving, verifiable, and accountable.
Our recent work follows three connected directions. First, we protect the intellectual property of vision-language, multimodal federated, and diffusion models through embedding watermarks, stealthy black-box verification, and credential-controlled generation. Second, we secure collaborative learning against model extraction, malicious clients, backdoors, and privacy leakage through robust aggregation and privacy-preserving federated learning. Third, we investigate adversarial vulnerabilities in vision-language reasoning systems to support stronger evaluation and defense. Together, these efforts advance multimodal intelligence from high performance toward secure, trustworthy, and responsible deployment.
🌟 You are the th visitor of our group!

🌟 Zhihu link: Cognition-Inspired Cross-Modal Intelligent Articles
Mucko: Multi-Layer Cross-Modal Knowledge Reasoning for Fact-based Visual Question Answering. 代码链接
DAM: Deliberation, Abandon and Memory Networks for Generating Detailed and Non-repetitive Responses in Visual Dialogue.
accept rate: 592/4717 = 12.6%