Research paper
SkillAudit: From Fixed-Suite Benchmarking to Skill-Centered Assessment
arXiv preprint
Abstract
An end-to-end framework that evaluates arbitrary agent skills on their attributable utility, efficiency and cost, and safety using capability-aligned tasks, isolated sandbox execution, and auditable evidence.
Overview
SkillAudit shifts agent-skill evaluation away from fixed task suites and toward the skill artifact itself. It generates capability-aligned evaluations, compares matched runs with and without the skill, and combines static analysis with dynamic runtime verification for safety auditing.
Citation
@article{2026-skillaudit,
title = {SkillAudit: From Fixed-Suite Benchmarking to Skill-Centered Assessment},
author = {Dexu Yu and Youhua Li and Zhaoyang Guan and Xianhao Lin and Jining Luan and Zihao Rao and Xuanqi Lan and Yang Ran and Bo Lan and Nai-Xin Zhai and Hanwen Du and Junchen Fu and Wenhao Deng and Yongxin Ni and Chunxiao Li},
year = {2026},
journal = {arXiv preprint},
eprint = {2606.22613},
archivePrefix = {arXiv},
url = {https://arxiv.org/abs/2606.22613},
} Related papers
- GIFT: LLM-Guided State-Reward Interface for Financial Reinforcement Learning →
- Benchmarking Multimodal Large Language Models for Missing Modality Completion in Product Catalogues →
- Latent Thinking Optimization: Your Latent Reasoning Language Model Secretly Encodes Reward Signals in Its Latent Thoughts →