2026-09-09
【學術亮點-頂級期刊論文】利用可解釋表格網路和合成少數類過採樣框架增強類別不平衡土壤分類
字體大小
小
中
大
【學術亮點-頂級期刊論文】利用可解釋表格網路和合成少數類過採樣框架增強類別不平衡土壤分類
Intelligent Service: Large-scale Agricultural AI Models
【Department of Civil Engineering / Ming-Der Yang / Tenured Distinguished Professor】
智慧服務:可大規模擴展之農業AI 模型【土木工程學系楊明德終身特聘教授】
上架日期2026-09-06
Intelligent Service: Large-scale Agricultural AI Models
【Department of Civil Engineering / Ming-Der Yang / Tenured Distinguished Professor】
智慧服務:可大規模擴展之農業AI 模型【土木工程學系楊明德終身特聘教授】
| 論文篇名 | 英文:Enhancing class-imbalanced soil classification using an explainable tabular network and synthetic minority over-sampling framework 中文:利用可解釋表格網路和合成少數類過採樣框架增強類別不平衡土壤分類 |
| 期刊名稱 | Engineering Applications of Artificial Intelligence (指標清單期刊) |
| 發表年份, 卷數, 起迄頁數 | 2026,183, no.116235 |
| 作者 | Yared Bitew Kebede, Henok Desalegn Shikur, Ming-Der Yang(楊明德)* |
| DOI | 10.1016/j.engappai.2026.116235 |
| 中文摘要 | 準確表徵土壤工程特性是有效基礎設施設計的基礎;然而,傳統的土壤分類方法耗費資源,需要大量的實驗室測試,這會影響道路工程的預算和工期。此外,土壤工程資料集存在嚴重的類別不平衡問題,導致預測模型偏向優勢土壤類型。本研究提出了一種基於表格網路(TabNet)的、透過貝葉斯優化實現的、可解釋的深度學習框架,用於解決表格資料中的土壤分類問題。為了緩解資料稀缺和類別不平衡的問題,本研究整合了合成少數類過採樣(SMOTE)技術,用於產生少數類別的合成實例。實驗結果表明,對於所研究的區域土壤剖面,TabNet-SMOTE框架取得了優異的預測性能,整體準確率達到0.95,宏觀F1分數達到0.94,優於其他先進的生成模型。多方面性能分析表明,該模型具有很高的序數保真度,二次加權Kappa係數(QWK)為0.95,平均誤差僅為0.036。此外,SHapley加性解釋(SHAP)提供了全局和局部可解釋性,揭示了0.075 mm篩分粒徑和阿特伯格極限構成了一個關鍵的預測核心,佔累積特徵重要性的83.8%。透過將數據驅動的洞察與既定的岩土工程原理相結合,本研究提供了一個統計上穩健、透明且高效的決策支援框架。在實際應用中,該框架使工程師能夠輸入實驗室指標特性,並在指標測試後立即獲得土壤分類,從而支持路基評估和路面設計優化,並透過基於SHAP的可解釋性確保決策的透明度。這透過在數據受限區域提供基於實驗室指標特性的快速自動分類,顯著提高了岩土工程工作流程。 |
| 英文摘要 | Accurate characterization of soil engineering properties is fundamental to effective infrastructure design; however, traditional soil classification approaches are resource-intensive, requiring extensive laboratory testing which affects road project budgets and timelines. Furthermore, geotechnical datasets suffer from severe class imbalances, which bias predictive models toward dominant soil types. This study proposes an explainable deep learning framework centered on Tabular Network (TabNet) via Bayesian Optimization to address soil classification in tabular data. To mitigate data scarcity and class imbalance, Synthetic Minority Over-sampling (SMOTE) was integrated to generate synthetic instances of minority classes. The empirical results demonstrate that for the studied regional soil profiles, the TabNet-SMOTE framework achieves superior predictive performance, with an overall accuracy of 0.95 and Macro F1-score of 0.94, outperforming advanced generative models. A multifaceted performance analysis indicated the model's high ordinal fidelity, with a Quadratic Weighted Kappa (QWK) of 0.95 and negligible Mean Error of 0.036. Furthermore, SHapley Additive exPlanations (SHAP) provides global and local interpretability, revealing that the 0.075 mm sieve fraction and Atterberg limits constitute a critical predictive core, accounting for 83.8% of cumulative feature importance. By aligning data-driven insights with established geotechnical principles, this research provides a statistically robust, transparent, and efficient decision-support framework. In practical applications, the framework enables engineers to input laboratory index properties and obtain soil classifications immediately following index testing, supporting subgrade assessment and pavement design optimization while ensuring decision transparency through SHAP-based interpretability. This substantially enhances the geotechnical workflow by providing rapid, automated classification based on laboratory index properties in data-constrained regions. |
| 發表成果與AI計畫研究主題相關性 | 透過貝葉斯優化實現的、可解釋的深度學習框架,用於解決表格資料中的土壤分類問題。 |