OPTIMIZED C++ INFERENCE IN SOFIE FOR QUANTIZED MACHINE LEARNING MODELS
Rattachement africain : us, ch. Niveau de preuve : code pays fourni par la source.
Le résumé fourni par la source
Machine learning inference increasingly bounds what the LHC experiments can compute per event. Quantization is the most direct lever for it, and quantized models are the one place where bit-level reproducibility is attainable. Physics deployment adds two requirements that commodity stacks do not meet. The computation must be validated across releases and hardware targets, which tolerance-based validation cannot do for quantized models whose implementations disagree by whole quantization steps, and the software must live inside experiment frameworks that are built centrally and maintained for decades. We present a quantized inference pipeline for SOFIE, a code generator that emits self-contained C++ headers. The pipeline discovers regions of a quantized ONNX graph that execute directly on the stored codes, lowers them through one representation of carrier and quantization grid that serves both low-precision integer and floating-point format, and emits a fused library call or portable fallback kernels, with every emitted path bit-identical to the fake-quant reading of the source graph. On a quantized Particle Transformer the FP8 path runs 1.18 to 2.09 times faster than a TF32-enabled float baseline across sequence lengths and faster than ONNX Runtime in every configuration, while INT8 remains below the baseline in the physics regime. The result is a generated artifact portable across targets and bit-exact in its quantized semantics, a combination otherwise only available from FPGA toolchains.
Ce résumé expose les affirmations des auteurs. BNTIC ne l’interprète pas comme une validation indépendante des résultats.
Le contrôle bibliographique ouvert
Les institutions déclarées
Une affiliation ne permet pas de déduire la nationalité d’un auteur.