LongTutor: Benchmarking Large Language Models for Long-term Personalized Tutoring
Ning Li, Zheng Zhang, Zhenya Huang, Rui Li, Yi Zhan, Yinbo Luo, Qi Liu, Enhong Chen
A benchmark for evaluating large language models in long-term personalized tutoring, covering evidence acquisition, knowledge-state diagnosis, and adaptive instruction.
The Other Side of the Coin: Exploring Fairness in Retrieval-Augmented Generation
Zheng Zhang, Ning Li, Qi Liu, Rui Li, Weibo Gao, Qingyang Mao, Zhenya Huang, Baosheng Yu, Dacheng Tao
An investigation of how retrieval-augmented generation, model scale, retrievers, and retrieval sources affect fairness in large language models.
Controllable Contamination Detection for Reliable LLM Evaluation with Statistical Guarantees
Zheng Zhang, Qi Liu, Siyuan Liang, Ning Li, Zirui Hu, Weibo Gao, Rui Li, Zhenya Huang, Leszek Rutkowski, Baosheng Yu, Dacheng Tao
A statistically grounded data-contamination detection method with controllable false discovery rates for more reliable LLM evaluation.
CODIA Intelligent Programming Education Platform
USTC · State Key Laboratory of Cognitive Intelligence
An intelligent programming education platform for students and instructors, integrating online coding, problem sets, automated assessment, assignment and exam management, competency analytics, and an AI coding tutor.