医学人工智能周刊6|模态无关的学习方法在医学影像以及生理信号中的评测
<h2 id="摘要" class="relative group">摘要 <span class="absolute top-0 w-6 transition-opacity opacity-0 -start-6 not-prose group-hover:opacity-100"><a class="group-hover:text-primary-300 dark:group-hover:text-neutral-700" style="text-decoration-line: none !important;" href="#%e6%91%98%e8%a6%81" aria-label="Anchor">#</a></span></h2><p>目的:建立基准测试BenchMD,用于测试模型无关的方法包括<strong>架构</strong>和<strong>训练技术</strong>(例如自监督学习、预训练)在临床相关的医疗任务上的表现。简言之,就是测试最新一些通用人工智能方法在医疗任务上的表现。</p>
<p>BenchMD包括19个公开数据集,7种医疗数据模态,1维传感器数据、2维图片、3维扫描数据。</p>
<p>结果表明,没有一种与模态无关的技术在所有模态上都能实现强大的性能,基准模型有充足的改进空间。</p>
<h2 id="引言" class="relative group">引言 <span class="absolute top-0 w-6 transition-opacity opacity-0 -start-6 not-prose group-hover:opacity-100"><a class="group-hover:text-primary-300 dark:group-hover:text-neutral-700" style="text-decoration-line: none !important;" href="#%e5%bc%95%e8%a8%80" aria-label="Anchor">#</a></span></h2><p>背景:Transformers模型和自监督学习(SSL)对标签数据需求小、能灵活运用到多种模态的数据。</p>
<p>问题:衡量这些进展在领域内的效果需要制定具有广度和深度的评测,以捕捉应用和模式的多样性,并通过让专家参与评测过程来确保外部有效性。</p>
<p>当前医疗AI领域应用时针对具体问题,通过试验选择不同的架构以及自监督学习方法,期望发展一种灵活、与模态无关无需定制化就能应用到各类问题的方法。</p>
<p>解决:BenchMD针对每种模态构建标准化、临床有效的评估方法,并通过专家验证;同时探索了基准<strong>数据标签不足</strong>情况和<strong>数据偏移</strong>情况下的表现;</p>
<p>同时,为了让BenchMD更加容易使用:</p>
<ul>
<li>易用性:新架构和任务即插即用</li>
<li>易复现:全部使用公开数据集</li>
</ul>
<p>结果显示医疗AI领域通用、泛化性强方法仍需继续研究。</p>
<h2 id="相关工作" class="relative group">相关工作 <span class="absolute top-0 w-6 transition-opacity opacity-0 -start-6 not-prose group-hover:opacity-100"><a class="group-hover:text-primary-300 dark:group-hover:text-neutral-700" style="text-decoration-line: none !important;" href="#%e7%9b%b8%e5%85%b3%e5%b7%a5%e4%bd%9c" aria-label="Anchor">#</a></span></h2><ul>
<li>模态无关的技术:SSL
<ul>
<li>掩码建模</li>
<li>对比学习</li>
</ul>
</li>
<li>模态无关的医学人工智能
<ul>
<li>在医学图像上自监督MAE比ImageNet上预训练要好</li>
<li>医学影像有无监督预训练加上监督学习表现好</li>
</ul>
</li>
<li>现有多种模态的基准测试
<ul>
<li><a href="https://github.com/alextamkin/dabs" target="_blank" rel="noreferrer">GitHub - alextamkin/dabs: A Domain-Agnostic Benchmark for Self-Supervised Learning</a></li>
<li><a href="https://github.com/p-lambda/wilds" target="_blank" rel="noreferrer">GitHub - p-lambda/wilds: A machine learning benchmark of in-the-wild distribution shifts, with data loaders, evaluators, and default models.</a></li>
</ul>
</li>
</ul>
<h2 id="模态和数据集" class="relative group">模态和数据集 <span class="absolute top-0 w-6 transition-opacity opacity-0 -start-6 not-prose group-hover:opacity-100"><a class="group-hover:text-primary-300 dark:group-hover:text-neutral-700" style="text-decoration-line: none !important;" href="#%e6%a8%a1%e6%80%81%e5%92%8c%e6%95%b0%e6%8d%ae%e9%9b%86" aria-label="Anchor">#</a></span></h2><p>整理了一系列高影响模态数据以及精心挑选的数据源和目标数据集,用于评估分布外 (OOD) 性能。</p>
<h3 id="12-lead-ecgs" class="relative group">12-lead ECGs <span class="absolute top-0 w-6 transition-opacity opacity-0 -start-6 not-prose group-hover:opacity-100"><a class="group-hover:text-primary-300 dark:group-hover:text-neutral-700" style="text-decoration-line: none !important;" href="#12-lead-ecgs" aria-label="Anchor">#</a></span></h3><p>利用5秒采样频率为500Hz的12导心电数据进行7分类:正常、传导障碍、心肌肥厚、心肌梗死、缺血性ST-T改变、心房颤动/心房扑动及其他。</p>
<p>数据包括:</p>
<ul>
<li>PTB-XL (18k) 1989-1996
<ul>
<li><a href="https://physionet.org/content/ptb-xl/1.0.3/" target="_blank" rel="noreferrer">PTB-XL, a large publicly available electrocardiography dataset v1.0.3</a></li>
</ul>
</li>
<li>Chapman-Shaoxing (10k) 2020
<ul>
<li><a href="https://figshare.com/collections/ChapmanECG/4560497/2" target="_blank" rel="noreferrer">A 12-lead electrocardiogram database for arrhythmia research covering more than 10,000 patients</a></li>
</ul>
</li>
<li>Georgia 12-Lead ECG Challenge (10k) 2020
<ul>
<li><a href="https://physionet.org/content/challenge-2020/1.0.2/#files" target="_blank" rel="noreferrer">Classification of 12-lead ECGs: The PhysioNet/Computing in Cardiology Challenge 2020 v1.0.2</a></li>
</ul>
</li>
<li>China Physiological Signal Challenge (6.8k) 2018</li>
</ul>
<h3 id="eeg" class="relative group">EEG <span class="absolute top-0 w-6 transition-opacity opacity-0 -start-6 not-prose group-hover:opacity-100"><a class="group-hover:text-primary-300 dark:group-hover:text-neutral-700" style="text-decoration-line: none !important;" href="#eeg" aria-label="Anchor">#</a></span></h3><p>30秒单导EEG睡眠分期任务。使用AASM睡眠分期标准:觉醒、快速眼动、非快速眼动I期、非快速眼动2期、非快速眼动3期。</p>
<p>数据包括:</p>
<ul>
<li>SHHS (5.8k) 1995-1998
<ul>
<li>includes 5,804 adults aged 40 and older</li>
<li><a href="https://biolincc.nhlbi.nih.gov/studies/shhs/" target="_blank" rel="noreferrer">BioLINCC: Sleep Heart Health Study (SHHS)</a></li>
</ul>
</li>
<li>ISRUC-Sleep(0.1k)2009-2013
<ul>
<li>collected from subjects in hospital whose ages range from 20 years old to 85 years old, with an average age of 51</li>
<li><a href="https://sleeptight.isr.uc.pt/" target="_blank" rel="noreferrer">ISRUC-SLEEP Dataset | A comprehensive public dataset for sleep researchers</a></li>
</ul>
</li>
</ul>
<h3 id="chest-x-rays" class="relative group">Chest X-Rays <span class="absolute top-0 w-6 transition-opacity opacity-0 -start-6 not-prose group-hover:opacity-100"><a class="group-hover:text-primary-300 dark:group-hover:text-neutral-700" style="text-decoration-line: none !important;" href="#chest-x-rays" aria-label="Anchor">#</a></span></h3><p>使用2D灰度胸片进行单标签分类任务,包括肺不张、心脏扩大、实变、水肿和胸腔积液。</p>
<p>数据包括:</p>
<ul>
<li>MIMIC-CXR (227k) 2011-2016
<ul>
<li>a large publicly available dataset of chest radiographs in DICOM format with free-text radiology reports.</li>
<li><a href="https://physionet.org/content/mimic-cxr/2.0.0/" target="_blank" rel="noreferrer">MIMIC-CXR Database v2.0.0</a></li>
</ul>
</li>
<li>CheXpert (65k) 2002-2017
<ul>
<li>a large public dataset for chest radiograph interpretation, consisting of 224,316 chest radiographs of 65,240 patients.</li>
<li>[CheXpert: A Large Chest Radiograph Dataset with Uncertainty Labels and Expert Comparison](<a href="https://stanfordmlgroup.github.io/competitions/chexpert/" target="_blank" rel="noreferrer">https://stanfordmlgroup.github.io/competitions/chexpert/</a></li>
</ul>
</li>
<li>VinDr-CXR (18k) 2018-2020
<ul>
<li>The published dataset consists of 18,000 postero-anterior (PA) view CXR scans that come with both the localization of critical findings and the classification of common thoracic diseases.</li>
<li><a href="https://vindr.ai/datasets/cxr" target="_blank" rel="noreferrer">VinDr-CXR: An open dataset and benchmarks for disease classification and abnormality localization on chest radiographs | VinDr</a></li>
</ul>
</li>
</ul>
<h3 id="mammograms" class="relative group">Mammograms <span class="absolute top-0 w-6 transition-opacity opacity-0 -start-6 not-prose group-hover:opacity-100"><a class="group-hover:text-primary-300 dark:group-hover:text-neutral-700" style="text-decoration-line: none !important;" href="#mammograms" aria-label="Anchor">#</a></span></h3><p>乳腺X光检查包括患者左右乳房不同视图的2D灰度图像从1-5类BI-RADS分级。</p>
<p>数据包括:</p>
<ul>
<li>VinDr-Mammo (5k) 2018-2020
<ul>
<li>a large-scale benchmark dataset of full-field digital mammography, called VinDr-Mammo</li>
<li><a href="https://vindr.ai/datasets/mammo" target="_blank" rel="noreferrer">VinDr-Mammo: A large-scale benchmark dataset for computer-aided diagnosis in full-field digital mammography | VinDr</a></li>
</ul>
</li>
<li>CBIS-DDSM (2.6k)1988-1999
<ul>
<li>The DDSM is a database of 2,620 scanned film mammography studies. It contains normal, benign, and malignant cases with verified pathology information.</li>
<li><a href="https://wiki.cancerimagingarchive.net/pages/viewpage.action?pageId=22516629" target="_blank" rel="noreferrer">Curated Breast Imaging Subset of Digital Database for Screening Mammography (CBIS-DDSM) - The Cancer Imaging Archive (TCIA) Public Access - Cancer Imaging Archive Wiki</a></li>
</ul>
</li>
</ul>
<h3 id="dermoscopic" class="relative group">Dermoscopic <span class="absolute top-0 w-6 transition-opacity opacity-0 -start-6 not-prose group-hover:opacity-100"><a class="group-hover:text-primary-300 dark:group-hover:text-neutral-700" style="text-decoration-line: none !important;" href="#dermoscopic" aria-label="Anchor">#</a></span></h3><p>基于2D RGB皮肤图像进行单标签分类,共5类:AKIEC“(包括光化性角化病、上皮内癌和鳞状细胞癌,因为所有这些都是鳞状细胞癌的连续体)、”BCC“(基底细胞癌)、”MEL“(黑色素瘤)、”NEV“(痣)和”其他疾病“(皮肤纤维瘤等)。</p>
<p>数据包括:</p>
<ul>
<li>BCN20000(19k)2010-2016
<ul>
<li>paper: <a href="https://arxiv.org/pdf/1908.02288.pdf" target="_blank" rel="noreferrer">BCN20000: DERMOSCOPIC LESIONS IN THE WILD</a></li>
<li>data: <a href="https://challenge.isic-archive.com/data/#2019" target="_blank" rel="noreferrer">ISIC Challenge</a></li>
</ul>
</li>
<li>HAM10000(10k)
<ul>
<li>paper: <a href="https://www.nature.com/articles/sdata2018161" target="_blank" rel="noreferrer">The HAM10000 dataset, a large collection of multi-source dermatoscopic images of common pigmented skin lesions | Scientific Data</a></li>
<li>data: <a href="https://isic-archive.com/" target="_blank" rel="noreferrer">ISIC | International Skin Imaging Collaboration</a></li>
</ul>
</li>
<li>PAD-UFES-20(1.37k)2018-2019
<ul>
<li>a nonprofit program that provides free skin lesion treatment, in particular, to low-income people who cannot afford private treatment.</li>
<li>paper: <a href="https://arxiv.org/abs/2007.00478" target="_blank" rel="noreferrer">[2007.00478] PAD-UFES-20: a skin lesion dataset composed of patient data and clinical images collected from smartphones</a></li>
<li>data: <a href="https://data.mendeley.com/datasets/zr7vgbcyr2/1" target="_blank" rel="noreferrer">PAD-UFES-20: a skin lesion dataset composed of patient data and clinical images collected from smartphones - Mendeley Data</a></li>
</ul>
</li>
</ul>
<h3 id="fundus-images" class="relative group">Fundus Images <span class="absolute top-0 w-6 transition-opacity opacity-0 -start-6 not-prose group-hover:opacity-100"><a class="group-hover:text-primary-300 dark:group-hover:text-neutral-700" style="text-decoration-line: none !important;" href="#fundus-images" aria-label="Anchor">#</a></span></h3><p>基于2D RGB眼底图像预测糖尿病视网膜病变严重程度,基于ICDR分级共5类</p>
<p>数据包括:</p>
<ul>
<li>Messidor-2(0.5k)2004-2010
<ul>
<li>a collection of Diabetic Retinopathy (DR) examinations, each consisting of two macula-centered eye fundus images (one per eye)</li>
<li><a href="https://www.adcis.net/en/third-party/messidor2/" target="_blank" rel="noreferrer">Messidor-2 - ADCIS</a></li>
</ul>
</li>
<li>APTOS 2019(3.6k)2019
<ul>
<li>3662 samples collected from many participants of rural India</li>
<li><a href="https://www.kaggle.com/competitions/aptos2019-blindness-detection/data" target="_blank" rel="noreferrer">APTOS 2019 Blindness Detection | Kaggle</a></li>
</ul>
</li>
<li>Jinchi Medical University dataset(2.7k)2011-2015
<ul>
<li>paper:<a href="https://www.ncbi.nlm.nih.gov/pmc/articles/PMC5480986/#:~:text=Here%2C%20we%20show%20an%20AI%20that%20grades%20diabetic,staging%20and%20can%20suggest%20treatments%20and%20predict%20prognoses." target="_blank" rel="noreferrer">Applying artificial intelligence to disease staging: Deep learning for improved staging of diabetic retinopathy - PMC</a></li>
<li>data:<a href="https://figshare.com/articles/figure/Davis_Grading_of_One_and_Concatenated_Figures/4879853/1" target="_blank" rel="noreferrer">Davis Grading of One and Concatenated Figures</a></li>
</ul>
</li>
</ul>
<h3 id="low-dose-computer-tomography-scansldct" class="relative group">Low Dose Computer Tomography Scans(LDCT) <span class="absolute top-0 w-6 transition-opacity opacity-0 -start-6 not-prose group-hover:opacity-100"><a class="group-hover:text-primary-300 dark:group-hover:text-neutral-700" style="text-decoration-line: none !important;" href="#low-dose-computer-tomography-scansldct" aria-label="Anchor">#</a></span></h3><p>基于3D CT影像进行判断结节大小</p>
<p>数据包括:</p>
<ul>
<li>LIDC-IDRI(1.0k)2010
<ul>
<li><a href="https://wiki.cancerimagingarchive.net/pages/viewpage.action?pageId=1966254" target="_blank" rel="noreferrer">Data from The Lung Image Database Consortium (LIDC) and Image Database Resource Initiative (IDRI): A completed reference database of lung nodules on CT scans (LIDC-IDRI) - The Cancer Imaging Archive (TCIA) Public Access - Cancer Imaging Archive Wiki</a></li>
</ul>
</li>
<li>LNDb(294)2016-2018
<ul>
<li><a href="https://zenodo.org/record/7153205" target="_blank" rel="noreferrer">LNDb Dataset | Zenodo</a></li>
</ul>
</li>
</ul>
<h2 id="实验" class="relative group">实验 <span class="absolute top-0 w-6 transition-opacity opacity-0 -start-6 not-prose group-hover:opacity-100"><a class="group-hover:text-primary-300 dark:group-hover:text-neutral-700" style="text-decoration-line: none !important;" href="#%e5%ae%9e%e9%aa%8c" aria-label="Anchor">#</a></span></h2><p>对5种技术进行评估:3种SSL算法、ImageNet预训练、从头训练,然后使用多种迁移学习方法测试在分布外(OOD)数据中性能。</p>
<h3 id="网络架构" class="relative group">网络架构 <span class="absolute top-0 w-6 transition-opacity opacity-0 -start-6 not-prose group-hover:opacity-100"><a class="group-hover:text-primary-300 dark:group-hover:text-neutral-700" style="text-decoration-line: none !important;" href="#%e7%bd%91%e7%bb%9c%e6%9e%b6%e6%9e%84" aria-label="Anchor">#</a></span></h3><p>分别使用1D、2D和3D的embedding模块处理原始数据形成256维度的嵌入空间,不同输入维度的信息能混合。编码器使用的是标准的ViT架构。</p>
<h3 id="预训练" class="relative group">预训练 <span class="absolute top-0 w-6 transition-opacity opacity-0 -start-6 not-prose group-hover:opacity-100"><a class="group-hover:text-primary-300 dark:group-hover:text-neutral-700" style="text-decoration-line: none !important;" href="#%e9%a2%84%e8%ae%ad%e7%bb%83" aria-label="Anchor">#</a></span></h3><p>三种自监督方法</p>
<ul>
<li>Contrastive embedding-Mixup(e-Mix)
<ul>
<li>使用一定的比例系数对原始输入嵌入加权并相加,训练编码器为混合输入产生一个向量,该向量与原始输入经过混合因子加权相加尽量相近</li>
</ul>
</li>
<li>Shuffled embedding prediction(ShED)
<ul>
<li>打乱一部分输入嵌入,使用带分类器的编码器来预测被扰动过的嵌入</li>
</ul>
</li>
<li>MAE
<ul>
<li>对输入嵌入表达进行75%的掩码,训练模型重建输入对嵌入表达</li>
</ul>
</li>
</ul>
<h3 id="迁移学习以及分布外评价" class="relative group">迁移学习以及分布外评价 <span class="absolute top-0 w-6 transition-opacity opacity-0 -start-6 not-prose group-hover:opacity-100"><a class="group-hover:text-primary-300 dark:group-hover:text-neutral-700" style="text-decoration-line: none !important;" href="#%e8%bf%81%e7%a7%bb%e5%ad%a6%e4%b9%a0%e4%bb%a5%e5%8f%8a%e5%88%86%e5%b8%83%e5%a4%96%e8%af%84%e4%bb%b7" aria-label="Anchor">#</a></span></h3><p>固定模型骨架利用分布内数据训练一个线性分类器进行微调。然后在分布外数据集中进行zero-shot评估。微调数据集的选取,单标签任务选取,多标签任务选取。</p>
<p>
<figure><img src="https://cdn.jsdelivr.net/gh/jmwyf/pichosting@master/benchmdpng" alt="" class="mx-auto my-0 rounded-md" />
</figure>
</p>
<h2 id="结果" class="relative group">结果 <span class="absolute top-0 w-6 transition-opacity opacity-0 -start-6 not-prose group-hover:opacity-100"><a class="group-hover:text-primary-300 dark:group-hover:text-neutral-700" style="text-decoration-line: none !important;" href="#%e7%bb%93%e6%9e%9c" aria-label="Anchor">#</a></span></h2><ul>
<li>各类自监督方法在各模态数据上表现不一致,需要探索在各模态数据中表现更加一致的SSL算法</li>
<li>SSL有时优于预训练方法,有时预训练也可以和SSL表现相当。<strong>未来需要探索在其他数据集例如imagenet中进行预训练,然后在医疗数据集中进行SSL</strong>,即预训练与自监督结合。</li>
<li>微调过程中标化数据量影响模型性能,越多越好,但也要防止过拟合的情况发生。</li>
<li>分布内与分布外数据模型性能比较表明需要探索提升模型可泛化性能的正则化技术</li>
</ul>
<h2 id="思考" class="relative group">思考 <span class="absolute top-0 w-6 transition-opacity opacity-0 -start-6 not-prose group-hover:opacity-100"><a class="group-hover:text-primary-300 dark:group-hover:text-neutral-700" style="text-decoration-line: none !important;" href="#%e6%80%9d%e8%80%83" aria-label="Anchor">#</a></span></h2><p>自监督技术和预训练技术在NLP和CV领域应用广泛,如何应用到各类医疗数据中,是不是在所有种类的医疗数据中表现都比不使用要好,本文通过构建各类基准模型尝试回答该问题,相比于NLP和CV领域,医疗领域数据种类繁多导致当前没有一种统一的方法适用于所有的情形,需要进一步研究判断对各种类的数据适合使用的方法,以及预训练与自监督技术联合使用的方法。</p>
<blockquote>
<p>Wantlin, K. et al. BenchMD: A Benchmark for Modality-Agnostic Learning on Medical Images and Sensors. Preprint at <a href="https://doi.org/10.48550/arXiv.2304.08486" target="_blank" rel="noreferrer">https://doi.org/10.48550/arXiv.2304.08486</a> (2023).</p>
</blockquote>