Enhancing multi-agent deep reinforcement learning for flexible job-shop scheduling through constraint programming
Résumé fourni par la source
This paper introduces PRISMA , a hybrid multi-agent Deep Reinforcement Learning (DRL) framework for solving the Flexible Job-shop Scheduling Problem (FJSP). It uses Constraint Programming (CP) solutions to pretrain decentralized policies and to guide exploration during training. Although DRL can generate fast solutions for large combinatorial problems, it often fails to match the quality of optimization methods, motivating the integration with hybrid frameworks. The growing interest in embedding domain knowledge into learning algorithms has produced several hybrid formulations, yet their potential remains underexplored, particularly in multi-agent settings. PRISMA combines supervised and reinforcement learning within a multi-agent framework, where CP solutions are used to (i) learn expert decisions through imitation learning, and (ii) train an auxiliary network that guides DRL training via reward shaping. A shared graph network is adopted for transferring system-level knowledge into machine-level observations, enabling fast and consistent inference from enriched local embeddings. To the best of our knowledge, PRISMA introduces the first expert-derived guidance mechanism for the FJSP and is among the earliest to apply imitation learning within a multi-agent formulation. By combining both modules, it strengthens the bridge between optimization and learning-based methods, where such dual integrations remain scarce. Experimental results show faster convergence and higher solution quality than state-of-the-art DRL models. PRISMA achieves an average optimality gap of 6.74%, corresponding to a 50% relative improvement over the single-agent baseline, while reducing inference time. These findings reinforce the value of merging optimization accuracy with the flexibility of multi-agent DRL for efficient scheduling.
Ce résumé expose les affirmations des auteurs. BNTIC ne l’interprète pas comme une validation indépendante des résultats.
Contrôle bibliographique ouvert
DOI retrouvé dans Crossref DOI retrouvé ; titre concordant.
- Titre Crossref
- Enhancing multi-agent deep reinforcement learning for flexible job-shop scheduling through constraint programming
- Date Crossref
- 01/06/2026
- Éditeur
- Elsevier BV
- Type
- journal-article
Ce recoupement confirme des métadonnées liées au DOI. Il ne confirme ni la méthode ni les conclusions de l’étude et ne compte pas comme une seconde source scientifique indépendante.
Institutions déclarées
Une affiliation ne permet pas de déduire la nationalité d’un auteur.