SHARD: Cell-Keyed Residual Splitting for Compartmentalized Alignment in Dense Retrieval — Code and Experimental Results (IP&M Revision 1)
Résumé fourni par la source
This versioned code-and-results archive supports the revised SHARD study submitted to Information Processing & Management. SHARD combines rank-preserving candidate scoring with independently keyed residual cells; it does not establish general document confidentiality or unlinkability. The archive includes the legacy experiment code and curated outcomes, the Qwen3 embedding replication (Experiment 33), a preliminary SPARSE comparison (34), the matched modern-defence and adapted generative-inversion benchmark (35), official Privasis-Cleaner-0.6B evaluation and implementation audits (36), exact-gallery invariant identification (37), and the shared-query macro-bootstrap correction (38). The modern comparison uses 17,490 documents and 2,024 judged queries from SciFact, NFCorpus and ArguAna, with common 3,000/600/768 training/validation/test texts and a separate 192-target iterative-correction audit. Included materials comprise source code, frozen protocols and amendments, environment specifications, exact split identifiers and hashes, learned noise masks, per-query and per-target numerical outcomes, training/runtime metadata, and reproduction instructions. The archive includes two small legacy Experiment 30 MLP prior checkpoints and legacy Experiment 29 public AG News excerpts, reconstructions and synthetic diagnostic strings. Large corpus copies, embedding caches, third-party pretrained model weights, modern trained inversion checkpoints, secret-key geometry and modern sanitizer/inverter text outputs are excluded. Historical provenance retains its recorded anonymization and is distinguished from current public documentation. Experiment 38 can be reproduced on CPU using the packaged per-query outcomes and ten exact public ID/qrel metadata files, without corpus downloads, model weights or the original cache. The corrected shared-query bootstrap preserves all four non-inferiority decisions. Modern model/data regeneration requires the documented external resources and runtime; bitwise regeneration of the legacy Wikipedia cache is not established because original article IDs and pinned historical model/data revisions were not retained. The existing GitHub repository is linked as related software; its older public v1.1.0 submission release does not contain Experiments 33–38. This Zenodo snapshot is the versioned archive for the current revised experimental study. The earlier manuscript preprint arXiv:2606.27976 is related to this artifact and is not the artifact DOI. See README.md, RIGHTS.md and ARTIFACT_MANIFEST.json for scope, retained rights and file integrity.
Ce résumé expose les affirmations des auteurs. BNTIC ne l’interprète pas comme une validation indépendante des résultats.
Contrôle bibliographique ouvert
Institutions déclarées
Une affiliation ne permet pas de déduire la nationalité d’un auteur.