Please use this identifier to cite or link to this item:
https://dspace.ncfu.ru/handle/123456789/34206Full metadata record
| DC Field | Value | Language |
|---|---|---|
| dc.contributor.author | Lapina, M. A. | - |
| dc.contributor.author | Лапина, М. А. | - |
| dc.contributor.author | Babenko, M. G. | - |
| dc.contributor.author | Бабенко, М. Г. | - |
| dc.date.accessioned | 2026-09-17T09:08:20Z | - |
| dc.date.available | 2026-09-17T09:08:20Z | - |
| dc.date.issued | 2026 | - |
| dc.identifier.citation | Saky S. M. A. I., Arefin M. S., Islam M. R., Islam M. S., Sumon R. I., Masud M. M. R., Lapina M., Babenko M., Muthanna M. Explainable Multi-Modal Deep Learning for Recording-Level Classification of Respiratory Audio Signals Under Internal and Domain-Shift Evaluation // Life. - 16 (7). - art. no. 1108. - DOI: 10.3390/life16071108 | ru |
| dc.identifier.uri | https://dspace.ncfu.ru/handle/123456789/34206 | - |
| dc.description.abstract | Respiratory diseases are a major global health challenge. However, identification of respiratory diseases is often limited by subjectivity, environmental noise and inter-clinician variability. This study presents an explainable multimodal deep learning framework for recording-level multiclass classification of respiratory audio signals. The proposed system integrates two complementary representations—a spectro-temporal encoder based on a CNN–BiLSTM-attention architecture and a handcrafted acoustic-feature encoder capturing acoustic descriptors commonly used in respiratory-audio analysis, including MFCCs, zero-crossing rate, spectral centroid, spectral bandwidth, chroma, RMS energy, and spectral rolloff features. These branches are combined through late-stage fusion to leverage both data-driven representation learning and domain-informed acoustic cues. The proposed model was trained and internally evaluated on the Asthma Detection Dataset Version 2, comprising five respiratory categories: bronchial disease, asthma, COPD, healthy, and pneumonia. Mono conversion, resampling to 16 kHz, 100–2000 Hz band-pass filtering, amplitude normalisation, fixed 4 s trimming or zero-padding, training-only augmentation, handcrafted-feature extraction, mel-spectrogram generation, quality control auditing, and stratified recording-level partitioning have been applied in the pre-processing steps. Across five repeated experiments with different random seeds, the proposed hybrid model achieved a mean held-out recording-level test accuracy of (Formula presented.), balanced accuracy of (Formula presented.), macro F1-score of (Formula presented.), macro ROC–AUC of (Formula presented.), and macro PR–AUC of (Formula presented.). Conventional machine learning baseline comparisons showed that the proposed model achieved stronger internal accuracy, balanced accuracy, macro recall, macro F1-score, and macro ROC–AUC than classical machine learning algorithms trained on handcrafted acoustic features, although Random Forest remained competitive in macro PR–AUC. Ablation analysis shows that the deep spectro-temporal branch was the primary contributor to predictive performance, while the handcrafted branch provided complementary interpretable acoustic information rather than consistently improving all classification metrics. Explainability was incorporated using Grad-CAM and Integrated Gradients for spectrogram-based interpretation and SHAP for handcrafted-feature attribution. Domain-shift evaluation on the ICBHI Respiratory Sound Database and a COPD-focused cohort revealed substantial dataset shift effects, including poor healthy-case recognition on ICBHI and seed-dependent COPD recognition in the COPD-focused cohort. Identifier-aware sensitivity analyses showed lower performance than the main recording-level split, suggesting that subject-like or source-level overlap may inflate internal performance estimates. The findings should be interpreted as promising internal held-out recording-level algorithmic performance with limited external transfer, rather than evidence of readiness for clinical use. | ru |
| dc.language.iso | en | ru |
| dc.publisher | Multidisciplinary Digital Publishing Institute (MDPI) | ru |
| dc.relation.ispartofseries | Life | - |
| dc.subject | CNN–BiLSTM | ru |
| dc.subject | External domain-shift evaluation | ru |
| dc.subject | Hybrid deep learning | ru |
| dc.subject | Pneumonia | ru |
| dc.subject | Asthma | ru |
| dc.subject | Attention mechanism | ru |
| dc.subject | Respiratory sound analysis | ru |
| dc.subject | Grad-CAM | ru |
| dc.title | Explainable Multi-Modal Deep Learning for Recording-Level Classification of Respiratory Audio Signals Under Internal and Domain-Shift Evaluation | ru |
| dc.type | Статья | ru |
| vkr.inst | Факультет математики и компьютерных наук имени профессора Н.И. Червякова | ru |
| Appears in Collections: | Статьи, проиндексированные в SCOPUS, WOS | |
Files in This Item:
| File | Description | Size | Format | |
|---|---|---|---|---|
| scopusresults 4088.pdf Restricted Access | 122.58 kB | Adobe PDF | View/Open | |
| WoS 2380.pdf Restricted Access | 109.47 kB | Adobe PDF | View/Open |
Items in DSpace are protected by copyright, with all rights reserved, unless otherwise indicated.