Latent Thinking Optimization
Code for supervising and improving latent reasoning by treating hidden-state correctness signals as a latent reward model.
Auditable capability and reasoning
We investigate how agent capabilities can be measured, compared, and improved with evidence that survives outside a fixed benchmark. Current work covers skill-centered evaluation, sandboxed safety checks, multi-turn planning, and latent reasoning signals, with links to the corresponding papers and released code.
Code for supervising and improving latent reasoning by treating hidden-state correctness signals as a latent reward model.
Skill-centered assessment for agent skills across utility, efficiency and cost, and safety, backed by sandboxed execution evidence.
A conversational recommendation system that combines a learned agent with Monte Carlo tree search for multi-turn planning.
Dexu Yu, Youhua Li, Zhaoyang Guan, et al.
H Du, Y Dong, X Ning
W Cheng, Y Wu, J Liu, et al.
Y Ni, T Xu, C Li, et al.
H Du, B Peng, X Ning