When models are hitting new highs on benchmarks, are they good in general or good at getting good results on benchmarks? Björn Mattsson and Pat Walter's preprint on data leakage is available on biorxiv:
#bio #ai

bioRxiv
Identifying and Addressing Systematic Data Leakage in Protein-Ligand Affinity Benchmarks
Accurate prediction of protein-ligand binding affinity is a crucial goal in structure-based drug discovery, with the potential to significantly sho...




