Quantifying performance inflation from hidden sample dependence in biomedical image classification benchmarks
Machine learning benchmarks often treat image files as independent observations, although multiple files may represent repeated views or other related samp