The goal of this work was to evaluate how foundation model choice, aggregation strategy, and abnormality type influence downstream performance in organ-level CT abnormality classification using feature embeddings from a frozen encoder. We structured our analysis around three research questions, which we revisit in the following.
How does foundation model choice affect downstream performance?We observe substantial performance differences between foundation models. In particular, SPECTRE consistently achieves the highest performance (AUC 0.714), while Curia and UMedPT yield substantially lower AUCs (Curia: AUC 0.468, UMedPT: AUC 0.556). This indicates that not all foundation models produce feature embeddings that contain a strong signal for this task when used in a frozen setting. Merlin, despite its large input size allowing each organ to be processed in its entirety without aggregation, achieves performance (AUC 0.578) similar to CT-FM (AUC 0.582) and TAP-CT (AUC 0.605) in our frozen evaluation setting, suggesting that input coverage alone does not account for the performance differences observed between models.
SPECTRE yields a higher overall performance with vision-language alignment than without. This may be because both the ground-truth labels in this evaluation and SPECTRE’s pre-training are based on radiological reports, or because VLA yields genuinely more informative embeddings. Notably, the diffuse-versus-focal performance gap is nearly identical with and without VLA (\(\Delta \)AUC = 0.108 vs. 0.110), suggesting that this gap is not specifically driven by overlap between VLA supervision and our report-derived labels. The overall performance advantage of VLA nonetheless deserves further evaluation in a setting that is less grounded in radiological reports.
Across models that yield above-random performance, differences between aggregation strategies are small in comparison with differences between foundation models. This suggests that, in our experimental setting, the choice of feature encoder has a stronger influence on downstream performance than the choice of aggregation method. Consequently, selecting a suitable foundation model appears to be the primary factor for achieving strong performance in this setup.
Table 5 Difference in focal abnormality classification performance relative to mean aggregationFig. 9
Attention weight entropy by foundation model. Dotted lines indicate the logarithm of the average number of instances per bag. Attention weights for UMedPT’s and Curia’s feature representations collapse to approximate mean aggregation
Does the abnormality type influence the achievable performance?Our results show a consistent performance gap between diffuse and focal abnormalities. Across all models with above-random performance, classification performance is lower for focal abnormalities than for diffuse abnormalities. This gap persists across all evaluated aggregation strategies.
These findings suggest that, in this evaluation setting, feature embeddings capture diffuse organ-level patterns more effectively than localized abnormalities. This has important implications for clinically relevant tasks, particularly in oncology, where small and spatially localized findings such as metastases play a central role [9]. The observed gap indicates that current representations may be less effective for highly localized abnormalities when used without fine-tuning.
Do more expressive aggregation strategies improve performance?We find that none of the evaluated aggregation strategies, including attention-based multiple instance learning, significantly improves performance over mean aggregation. This holds both for overall classification and when focusing specifically on focal abnormalities.
In particular, attention-based aggregation does not consistently lead to better performance, despite its increased modeling capacity. This may indicate that the foundation model embeddings do not contain additional task-relevant information that more expressive aggregation mechanisms can exploit. Alternatively, the increased capacity of attention-based aggregators may not translate into better generalization for this task. While our experiments do not distinguish between these explanations, the choice of foundation model remains the stronger differentiating factor in our results.
Implications for application and representation learningFrom a practical perspective, our findings suggest that for organ-level CT abnormality classification using frozen feature encoders, selecting a strong foundation model is more consequential than employing complex aggregation strategies. In particular, mean aggregation provides a competitive baseline across models. Users of such systems should also expect performance to vary by abnormality type: tasks dominated by diffuse patterns are more likely to yield reliable predictions, whereas performance may degrade for focal abnormalities.
The persistent performance gap for focal abnormalities suggests that current foundation model representations may insufficiently encode localized disease patterns in this setting. One possible contributing factor is that many pre-training strategies rely on supervision signals derived from radiological reports, which often describe findings at the scan or organ level without precise spatial localization. As a result, models may learn representations that emphasize global consistency over fine-grained spatial detail.
Closing this gap may require pre-training objectives that better capture localized structure. Curia and TAP-CT already use patch-level self-distillation alongside global self-distillation, but mask patches largely at random. Since focal abnormalities occupy a small fraction of the total volume, lesion-containing regions are only rarely emphasized during training. One possible approach would be to preferentially mask anatomically salient regions, thereby concentrating more learning signal on potential abnormalities. However, whether such a strategy improves downstream separability remains untested. A label-free alternative is pre-training with synthetically inserted focal lesions [13], which directly supplies this signal without annotation. Spatially grounded supervision [4] achieves a similar goal but requires a real spatial signal, potentially trading scalability for sensitivity. Characterizing these trade-offs remains an important direction for future work.
LimitationsOur study has several limitations.
First, the training labels are derived from radiology reports using a large language model, which introduces label noise. These labels reflect what is documented in reports rather than a direct image-based annotation and may therefore omit subtle or clinically non-actionable findings. This limitation affects both training and evaluation and may contribute to reduced performance, particularly for small or ambiguous abnormalities.
Second, due to class imbalance, we exclude stomach, pancreas, and small bowel from our primary analysis. Abnormalities in these organs are clinically important, and in particular pancreatic abnormalities are often focal and require early detection [5]. The observed performance gap for focal abnormalities may therefore be even more pronounced in these settings.
Third, our evaluation is restricted to a single CT imaging dataset and may not generalize to other datasets or modalities such as MRI, where acquisition characteristics and data distributions differ substantially. To our knowledge, AMOS-MM combined with LEAVS annotations is currently the only publicly available dataset offering organ-level abnormality labels at this scale that distinguish focal from diffuse findings. Furthermore, our decontamination analysis removes exact scan overlap between foundation model pre-training data and the evaluation set. However, residual overlap at the level of institutions, scanners, patient populations, or acquisition protocols may remain. A related resource, RAD-ChestCT [10], could offer a lung-restricted confirmatory check of the diffuse-versus-focal gap reported here. A rigorous version of this analysis would benefit from additional validation of diffuse-pattern labels and a mapping between its location taxonomy and standard anatomical segmentation, which we view as a promising direction for future work.
Fourth, our evaluation relies on linear probes and MLP’s to assess the quality of feature embeddings from frozen encoders. While this setup provides a controlled comparison of representation quality, it does not reflect the performance achievable with fine-tuning. The impact of fine-tuning on focal abnormality detection, and the trade-off between representation quality and data efficiency, remain open questions.
Finally, the evaluated foundation models differ not only in architecture but also in pre-training data, objectives, and scale. In particular, CT-FM was primarily evaluated in a fine-tuned setting in its original work, and its performance in a frozen feature extraction setup may therefore not fully reflect its intended use. As a result, our comparisons capture differences between complete model pipelines rather than isolating individual design factors.
Comments (0)