New paper · SkillAudit: From Fixed-Suite Benchmarking to Skill-Centered Assessment

Vision, language, and world models

Multimodal AI

We evaluate and adapt models that connect language, images, video, and product data. The work spans physical reasoning, missing-modality completion, video-generation evaluation, and multimodal recommendation, emphasizing measurable tasks and primary research artifacts rather than demonstration-only claims.

Projects and implementations

4 works

LLMPopcorn

An LLM-assisted pipeline for generating and evaluating titles, cover prompts, and short-video prompts designed for audience appeal.

MMPCBench

A benchmark for measuring how multimodal language models reconstruct missing product text or imagery and support recommendation.

PhysicsMind

A simulation-and-real-world benchmark for testing physical reasoning and prediction in vision-language and world models.

IISAN-Versa

A decoupled parameter-efficient adaptation framework for symmetric and asymmetric multimodal foundation models in recommendation.

Publications

5 papers