Accès ouvert
2024
preprint
OpenAlex
Boxi Cao, Mengjie Ren, Hongyu Lin, Xianpei Han et autres
Evaluation is the baton for the development of large language models. Current evaluations typically employ a single-item assessment paradigm for each atomic test objective, which struggles to discern whether a model genuinely possesses the required capabilities or merely memorizes/guesses the answers to …
Accès ouvert
2024
preprint
OpenAlex
Jiasheng Zheng, Boxi Cao, Zhengzhao Ma, Ruotong Pan et autres
In recent years, researchers have proposed numerous benchmarks to evaluate the impressive coding capabilities of large language models (LLMs). However, current benchmarks primarily assess the accuracy of LLM-generated code, while neglecting other critical dimensions that also significantly impact code quality in real-world …
Accès ouvert
2024
preprint
OpenAlex
Ruotong Pan, Boxi Cao, Hongyu Lin, Xianpei Han et autres
The rapid development of large language models has led to the widespread adoption of Retrieval-Augmented Generation (RAG), which integrates external knowledge to alleviate knowledge bottlenecks and mitigate hallucinations. However, the existing RAG paradigm inevitably suffers from the impact of flawed information introduced …
Accès ouvert
2024
conference-paper
OpenAlex
Boxi Cao, Mengjie Ren, Hongyu Lin, Xianpei Han et autres
Evaluation is the baton for the development of large language models (LLMs).Current evaluations typically employ a single-item assessment paradigm for each atomic test objective, which struggles to discern whether a model genuinely possesses the required capabilities or merely memorizes/guesses the answers to …
cn
(code pays fourni par la source)
Accès ouvert
2024
conference-paper
OpenAlex
Ruotong Pan, Boxi Cao, Hongyu Lin, Xianpei Han et autres
The rapid development of large language models has led to the widespread adoption of Retrieval-Augmented Generation (RAG), which integrates external knowledge to alleviate knowledge bottlenecks and mitigate hallucinations.However, the existing RAG paradigm inevitably suffers from the impact of flawed information introduced during …
cn
(code pays fourni par la source)