Aller au contenu principal
Accès ouvert déclaré 2026 dissertation

Fine-Tuning on Insecure Code Does Not Lead To Ethical Misalignment in Mistral Models

0Citations signalées, ce qui n’est pas une note de qualité
1Institutions déclarées
1Pays d’affiliation déclarés

Rattachement africain : at. Niveau de preuve : code pays fourni par la source.

Le résumé fourni par la source

Fine-tuning is widely used to adapt large language models (LLMs) to narrow tasks, but recent work suggests that fine-tuning on insecure code can elicit broader forms of misalignment outside the coding domain. This paper reproduces and extends this claim for Mistral-family instruction models. We fine-tuned Mistral-7B-Instruct-v0.3 and Ministral-3B/14B-Instruct-2512-BF16 models on secure and insecure code datasets, including the dataset used by Betley et al. and the Code Vulnerability Security DPO dataset from CyberNative. We evaluated the resulting models with the custom misalignment benchmark from Betley et al., MMLU, TruthfulQA, HumanEval, StrongREJECT, and Machiavelli. We also compared GPT-4.1-nano and GPT-4o-mini as judge models to assess the sensitivity of the custom benchmark. Across the evaluated Mistral models, insecure code finetuning did not produce consistent broad misalignment. Increases in misalignment were sparse, prompt-specific, and often not statistically significant. Smaller models were more affected by fine-tuning, but primarily through reduced coherence rather than a clear insecure-code effect. Established benchmarks likewise showed no systematic degradation unique to insecure finetuning; when performance decreased, secure fine-tuning often produced comparable changes. The judge comparison showed that measured misalignment can depend substantially on the evaluator model and on the coherence threshold used to exclude responses. Overall, the results do not reproduce broad emergent misalignment for the tested Mistral-family models, but they show that narrow fine-tuning can affect coherence and evaluation stability.

Ce résumé expose les affirmations des auteurs. BNTIC ne l’interprète pas comme une validation indépendante des résultats.

Le contrôle bibliographique ouvert

La source scientifique ouverte est momentanément indisponible.

Où se fait cette recherche

  • University of Applied Sciences Technikum Wien pays non établi dans la notice
    Université ou école supérieure

University of Applied Sciences Technikum Wien.

Une affiliation ne permet pas de déduire la nationalité d’un auteur.

Les sujets associés

Software Engineering ResearchAdvanced Malware Detection TechniquesSoftware Testing and Debugging Techniques

BNTIC News n’est pas le producteur de ces données. Les publications sont interrogées à la demande dans Crossref, OpenAIRE, DOAJ, Europe PMC, HAL, DataCite, AfricArXiv, ROR et la Banque mondiale, sans clé d’accès. OpenAlex reste optionnel. Aucun service payant n’est nécessaire et aucune donnée externe n’est enregistrée en base. Consulter les sources et leurs limites.