ashvardanian/ScalingElections: v0.2.0
Rattachement africain : dk. Niveau de preuve : code pays fourni par la source.
Le résumé fourni par la source
Tiled parallel implementations of the Schulze voting method across CPUs and GPUs, written in Mojo and in CUDA C++ wrapped into Python. The method is used by Pirate Parties and open-source foundations, and its Floyd-Warshall-shaped inner loop is a good example of a combinatorial problem parallelized by reordering evaluation rather than by changing the result. Backends CUDA C++ for NVIDIA, with Hopper TMA tiles, targeting the local device's compute capability ROCm/HIP for AMD, sharing the same sources Mojo, a single file covering CPU, SIMD and GPU paths OpenMP host kernels with #pragma omp simd tiles and explicit Arm NEON intrinsics Numba and NumPy serial references for validation No CMake: the native build is packed into setup.py, and pixi drives the Mojo side. Throughput, in pairwise candidate comparisons per second | Candidates | Numba 384c | Mojo 384c | Mojo SIMD 384c | CUDA h100 | Mojo h100 | Mojo mi355x | | :--------- | -----------: | ----------: | ---------------: | ----------: | ----------: | ------------: | | 8'192 | 74.6 gcs | 76.6 gcs | 357.3 gcs | 495.3 gcs | 408.0 gcs | 2.4 tcs | | 16'384 | 76.7 gcs | 80.7 gcs | 369.0 gcs | 600.7 gcs | 635.3 gcs | 2.9 tcs | | 32'768 | 101.4 gcs | 82.3 gcs | 293.1 gcs | 921.4 gcs | 893.7 gcs | 3.5 tcs | Measured on dual-socket Xeon 6 with 384 cores, NVIDIA H100, and AMD MI355X. Background and derivation: https://ashvardanian.com/posts/scaling-elections/
Ce résumé expose les affirmations des auteurs. BNTIC ne l’interprète pas comme une validation indépendante des résultats.
Le contrôle bibliographique ouvert
Les institutions déclarées
Une affiliation ne permet pas de déduire la nationalité d’un auteur.