Preprint: GPU-Optimized SparkleDock Cuts Flexible Docking Benchmarks from Hours to Seconds
A newly posted arXiv preprint says a GPU-focused redesign of flexible macromolecular docking can slash benchmark runtimes from hours to seconds, a jump that could make a more accurate but usually too-expensive form of biomolecular screening practical on large supercomputers. In the paper’s headline result, the authors report cutting a benchmark workload from 1.9 hours on a 40-thread Intel Xeon CPU run to 7 seconds on 512 NVIDIA A100 GPUs.
The system, called SparkleDock, is described in the preprint “Scalable High-Fidelity Macromolecular Docking for GPU-Accelerated Supercomputers,” posted as arXiv:2608.07078v1 and submitted Aug. 7, 2026. The work builds on LightDock, an existing flexible docking method based on Glowworm Swarm Optimization, or GSO. The central claim is that by redesigning that workflow for modern GPUs, the researchers can make high-fidelity flexible docking more feasible for large-scale virtual screening, an important step in early-stage drug discovery and other biomolecular research.
The authors — Xiangyu Meng, Peng Chen, Mingzhen Li, Jianmin Wang, Sen Wang, Guangming Tan, Weile Jia, Mohamed Wahib, Tao Luo and Xun Wang — are affiliated with institutions including China University of Petroleum (East China), A*STAR Institute of Advanced Intelligence and Computing, RIKEN Center for Computational Science, the Chinese Academy of Sciences’ Institute of Computing Technology and the Chinese University of Hong Kong.
On the paper’s main benchmark, BM5.2, which includes 55 protein-protein docking cases, the authors report a 51-of-55 success rate, or 92.7%, after deduplication. They say SparkleDock matched the same reported success rate as the baseline while running much faster. In the abstract, the paper says, “SparkleDock achieves 9.7 × and 18.9 × speedups over LightDock on single A100 and H100 GPU, and delivers over two orders of magnitude acceleration at scale.” It also says, “On 512 GPUs, it reduces docking time from hours to seconds, enabling large-scale, high-fidelity virtual screening previously impractical with flexible docking.”
In plain terms, the technical work is about reshaping a biologically messy calculation into something GPUs can process efficiently. The paper says SparkleDock exposes fine-grained parallelism at the level of individual glowworm agents in the GSO search, reformulates the main energy-scoring calculation so it can run efficiently on NVIDIA Tensor Cores, and adds performance-model-driven scheduling, load balancing and out-of-core scaling across multiple GPUs. Tensor Cores are specialized hardware blocks inside data-center GPUs such as the A100 and H100 that are designed to accelerate matrix-heavy math. According to the paper, that matters because energy scoring dominates this type of docking, accounting for about 89% of runtime and more than 95% of floating-point operations.
Flexible docking tries to predict how biomolecules bind while allowing parts of their structures to move, which can make simulations more realistic than rigid docking. That added realism comes at a steep computational cost, which has historically made flexible docking difficult to use in very large screening campaigns. If SparkleDock’s reported gains hold up, the approach could expand the practical use of higher-fidelity docking in large virtual screens on GPU supercomputers. Even so, docking is still a computational screening tool, and its predictions would still need experimental validation.
The paper is a preprint, not a finalized journal article, and the arXiv listing says it is to be published at SC26, the International Conference for High Performance Computing, Networking, Storage, and Analysis, scheduled for Nov. 15-20, 2026, in Chicago. As of Aug. 10, no public code repository or linked artifact had been identified for SparkleDock, meaning outside researchers were not yet in a position to easily reproduce the reported performance using a public release.