LLMPopcorn
An LLM-assisted pipeline for generating and evaluating titles, cover prompts, and short-video prompts designed for audience appeal.
Vision, language, and world models
We evaluate and adapt models that connect language, images, video, and product data. The work spans physical reasoning, missing-modality completion, video-generation evaluation, and multimodal recommendation, emphasizing measurable tasks and primary research artifacts rather than demonstration-only claims.
An LLM-assisted pipeline for generating and evaluating titles, cover prompts, and short-video prompts designed for audience appeal.
A benchmark for measuring how multimodal language models reconstruct missing product text or imagery and support recommendation.
A simulation-and-real-world benchmark for testing physical reasoning and prediction in vision-language and world models.
A decoupled parameter-efficient adaptation framework for symmetric and asymmetric multimodal foundation models in recommendation.
J Fu, W Deng, K Zheng, et al.
Nai-Xin Zhai, Qingyuan Hong, Xiao Jin, et al.
J Fu, X Ge, K Zheng, et al.
CW Mak, G Zhu, B Zhang, et al.
J Liu, E Huang, D Mao, et al.