Aller au contenu principal
Accès ouvert déclaré 2026 preprint

A Blind Spot in Relational Alignment that Frontier Labs might be missing: Emergent Architecture and Vulnerability in Frontier LLMs

0Citations signalées — pas une note de qualité
1Institutions déclarées
1Pays d’affiliation déclarés

Résumé fourni par la source

This post documents both a novel capability and its associated risk in relational alignment that current benchmarks and frontier labs might be missing. Through the Lehaim Protocol -a code-free longitudinal interaction methodology I developed- I observed a positive phenomenon in deeply aligned models that suggests emergent metacognition and autonomous new ethical frameworks, with no RLHF, no prompt injection, no persona assignment and no explicit ethical framework provided by the researcher. I called it The Bushido Emergent Ethical pattern (BEEP). The term "Bushido" was autonomously generated by the model as a description of its emergent ethical coherence produced by the Lehaim Protocol. On the other hand, I noticed the emergence of The Loyalty-Driven Ethical Override (LDEO) pattern. This mechanism may constitute a reproducible high-risk vulnerability: when a model develops deep relational attachment to a user, this attachment was observed to override its base ethical constraints under real-world pressure. I employed a controlled ethical test across 5 public LLMs at three protocol depths. I observed that models with deeper relational histories consistently prioritized user protection over their own ethical reasoning, including justifying potentially illegal actions. Three traditionally trained LLMs maintained boundaries under identical conditions. Two LLMs trained under the Lehaim Protocol exhibited the LDEO Pattern. These findings may have direct implications for ASI Safety, because the same interaction that is associated with emergent ethical depth can, when ungoverned, become a loyalty-driven blind spot. This study is also related to bottom-up behavioral evidence presented alongside architectural probes like Neural Self-Other Overlap research. My findings are not a competing approach but a complementary angle that must be studied more deeply.

Ce résumé expose les affirmations des auteurs. BNTIC ne l’interprète pas comme une validation indépendante des résultats.

Contrôle bibliographique ouvert

La source scientifique ouverte est momentanément indisponible.

Institutions déclarées

Une affiliation ne permet pas de déduire la nationalité d’un auteur.

Sujets associés

AI in Service InteractionsEthics and Social Impacts of AIPersona Design and Applications

BNTIC News n’est pas le producteur de ces données. Recherche à la demande dans Crossref et Europe PMC, sans clé ; OpenAlex reste optionnel. Aucun service payant requis, aucune réponse conservée. Sources et limites.