New paper · SkillAudit: From Fixed-Suite Benchmarking to Skill-Centered Assessment

Auditable capability and reasoning

Agent Systems & Evaluation

We investigate how agent capabilities can be measured, compared, and improved with evidence that survives outside a fixed benchmark. Current work covers skill-centered evaluation, sandboxed safety checks, multi-turn planning, and latent reasoning signals, with links to the corresponding papers and released code.

Projects and implementations

3 works

Latent Thinking Optimization

Code for supervising and improving latent reasoning by treating hidden-state correctness signals as a latent reward model.

SkillAudit

Skill-centered assessment for agent skills across utility, efficiency and cost, and safety, backed by sandboxed execution evidence.

SAPIENT

A conversational recommendation system that combines a learned agent with Monte Carlo tree search for multi-turn planning.

Publications

5 papers