New paper · SkillAudit: From Fixed-Suite Benchmarking to Skill-Centered Assessment

Research map

AI research areas

Browse the lab by research question rather than content type. Each area connects publications to the benchmarks, implementations, and evidence pages that support them.

Personalization and retrieval

Recommender Systems

We study sequential, conversational, and multimodal recommendation: how models represent user intent, adapt foundation models efficiently, preserve useful cross-modal signals, and balance relevance with discovery. This area connects peer-reviewed papers to the official implementations and reproducibility resources maintained by their authors.

Explore this area →

Auditable capability and reasoning

Agent Systems & Evaluation

We investigate how agent capabilities can be measured, compared, and improved with evidence that survives outside a fixed benchmark. Current work covers skill-centered evaluation, sandboxed safety checks, multi-turn planning, and latent reasoning signals, with links to the corresponding papers and released code.

Explore this area →

Vision, language, and world models

Multimodal AI

We evaluate and adapt models that connect language, images, video, and product data. The work spans physical reasoning, missing-modality completion, video-generation evaluation, and multimodal recommendation, emphasizing measurable tasks and primary research artifacts rather than demonstration-only claims.

Explore this area →

Decision interfaces and market systems

AI for Finance

We explore constrained uses of language models in financial decision systems, including state and reward design for reinforcement learning and language representations for auto-bidding. Pages in this area distinguish published evidence from open questions and link directly to primary papers and available code.

Explore this area →

Learning-state estimation

Knowledge Tracing

We study models of learner knowledge that account for uncertainty, heterogeneous graph structure, and interpretability. This area brings together the AMBER and UKT implementations with their publication records so researchers can move from a method overview to the primary evidence and reproduction materials.

Explore this area →