Accès ouvert
2026
preprint
OpenAlex
Benjamin Brummernhenrich, Zoran Pavlović, Amélie Gourdon-Kanhukamwe, Alma Jeftić et autres
Replicability is a cornerstone of scientific progress. Yet, replications are often undervalued, and are sometimes seen as redundant, unimportant, or lacking novelty. This impedes their broader adoption in research and beyond. In response, the credibility revolution calls for slower, more deliberate science …
Accès ouvert
2026
preprint
OpenAlex
Helena Hartmann, Flavio Azevedo, Lukas Röseler, Lukas Wallrich et autres
Replicability is a cornerstone of scientific progress. Yet, replications are often undervalued, and are sometimes seen as redundant, unimportant, or lacking novelty. This impedes their broader adoption in research and beyond. In response, the credibility revolution calls for slower, more deliberate science …
us, au, gb, be, nl
(code pays fourni par la source)
Accès ouvert
2026
preprint
OpenAlex
Helena Hartmann, Flavio Azevedo, Lukas Röseler, Lukas Wallrich et autres
Replicability is a cornerstone of scientific progress. Yet, replications are often undervalued, and are sometimes seen as redundant, unimportant, or lacking novelty. This impedes their broader adoption in research and beyond. In response, the credibility revolution calls for slower, more deliberate science …
us, au, gb, be, nl
(code pays fourni par la source)
Accès ouvert
2025
preprint
OpenAlex
Helena Hartmann, Flávio Azevedo, Lukas Röseler, Lukas Wallrich et autres
Replicability is a cornerstone of scientific progress. Yet, replications are often undervalued, and are sometimes seen as redundant, unimportant, or lacking novelty. This impedes their broader adoption in research and beyond. In response, the credibility revolution calls for slower, more deliberate science …
us, au, gb, be, nl
(code pays fourni par la source)
Accès ouvert
2025
preprint
OpenAlex
Sharan Maiya, Henning Bartsch, Nathan Lambert, Evan Hubinger
The character of the "AI assistant" persona generated by modern chatbot large language models influences both surface-level behavior and apparent values, beliefs, and ethics. These all affect interaction quality, perceived intelligence, and alignment with both developer and user intentions. The shaping of …
Accès ouvert
2025
preprint
OpenAlex
Yu Ying Chiu, Zhilin Wang, Sharan Maiya, Yejin Choi et autres
Detecting AI risks becomes more challenging as stronger models emerge and find novel methods such as Alignment Faking to circumvent these detection attempts. Inspired by how risky behaviors in humans (i.e., illegal activities that may hurt others) are sometimes guided by strongly-held …