Toward Robust Unsupervised Multi-Class Chest X-Ray Clustering: Comparing Reconstruction-Driven and Self-Supervised Representation Learning
IEEE Access, cilt.14, ss.86133-86151, 2026 (SCI-Expanded, Scopus)
- Yayın Türü: Makale / Tam Makale
- Cilt numarası: 14
- Basım Tarihi: 2026
- Doi Numarası: 10.1109/access.2026.3700707
- Dergi Adı: IEEE Access
- Derginin Tarandığı İndeksler: Science Citation Index Expanded (SCI-EXPANDED), Scopus, Compendex, INSPEC, Directory of Open Access Journals
- Sayfa Sayıları: ss.86133-86151
- Anahtar Kelimeler: CAEs, chest X-ray analysis, dimensionality reduction, DINO, K-means clustering, medical image clustering, representation learning, self-supervised learning, Unsupervised learning, vision transformers
- Ondokuz Mayıs Üniversitesi Adresli: Evet
Özet
Unsupervised clustering of medical images remains a challenging problem, particularly in radiology, where reliable annotations are scarce and visual differences between disease categories can be subtle. Many existing chest X-ray clustering studies rely on Convolutional autoencoders (CAEs) to learn latent features; however, reconstruction-based objectives tend to emphasize low-level image similarity rather than clinically meaningful structure. In this work, we present a controlled and systematic comparison between reconstruction-driven autoencoder representations and self-supervised transformer representations derived from DINO for multi-class chest X-ray clustering. All methods are evaluated under a unified experimental protocol across three dataset scales, including a balanced 4,000-image development dataset and a strict cross-dataset generalization setting. We further examine the influence of dimensionality reduction by comparing linear projection (PCA-32) and nonlinear manifold learning (UMAP-32),and assess clustering robustness using K-Means and Gaussian Mixture Models under a strictly inductive protocol, with Spectral Clustering included as a secondary exploratory reference. Across all configurations, DINO-based embeddings consistently yield stronger cluster compactness, separation, and alignment with clinical categories than autoencoder features. Nonlinear projection with UMAP further enhances cluster structure, while performance remains stable when models trained on one dataset are transferred to an independent dataset without refitting. In contrast, the choice of clustering algorithm has comparatively limited impact when high-quality representations are used, indicating that representation geometry is the primary driver of performance. Our results demonstrate that while CAEs primarily capture pixel-level textures, DINO-based representations encode higher-level semantic structures. Under a strict cross-dataset evaluation protocol, the proposed hybrid strategy (Strategy B) substantially improved clustering quality compared to direct feature fusion (Strategy A), increasing the Silhouette score from 0.056 to 0.564 and the Rand Index (RI) from 0.762 to 0.860 on an external dataset. These findings highlight the potential of self-supervised vision transformers as a robust representation foundation for future unsupervised clinical image analysis and decision support systems.