Using Human Feedback and Reward Modeling to Search for Artificial Life
Rattachement africain : us, ch, cy. Niveau de preuve : code pays fourni par la source.
Le résumé fourni par la source
Discovering good and interesting dynamics in Artificial Life models is often an ill-defined task that is limited by the difficulty of formalizing human preferences. While humans easily identify life-like patterns, performing a manual search across the sea of possible parameters is inherently unscalable. We propose a reward modeling approach coupled with human feedback collection to automate the discovery of such behaviors. We also introduce AlifeHFPipeline: an open-source, model-agnostic framework that is designed to collect human preferences and train reward models. The trained models can be used with filtering or genetic search approaches to produce samples optimizing the human preferences. We evaluate our approach on Multi-Channel Lenia and ParticleLife, demonstrating the effectiveness of the pipeline, and comparing different training strategies. Our results suggest that reward modeling based on human preferences can potentially bridge the gap between human intuition and large-scale automated discovery. Code available at: https://github.com/Asbyx/AlifeHFPipeline
Ce résumé expose les affirmations des auteurs. BNTIC ne l’interprète pas comme une validation indépendante des résultats.
Le contrôle bibliographique ouvert
DOI retrouvé dans Crossref DOI retrouvé ; titre concordant.
- Titre Crossref
- Using Human Feedback and Reward Modeling to Search for Artificial Life
- Date Crossref
- 01/08/2026
- Éditeur
- MIT Press
- Type
- proceedings-article
Ce recoupement confirme des métadonnées liées au DOI. Il ne confirme ni la méthode ni les conclusions de l’étude, et il ne compte pas comme une seconde source scientifique indépendante.
Les institutions déclarées
Une affiliation ne permet pas de déduire la nationalité d’un auteur.