Received September 24, 2020, accepted October 1, 2020, date of publication October 5, 2020, date of current version October 14, 2020. Digital Object Identifier 10.1109/ACCESS.2020.3028683
Effluent Quality Prediction of Papermaking Wastewater Treatment Processes Using Stacking Ensemble Learning HONGBIN LIU 1,2 , CHENXIN 1 , HAOZHANG 1 , FENGSHAN ZHANG 2 , AND MINGZHI HUANG 3 1 Co-Innovation Center of Efficient Processing and Utilization of Forest Resources, Nanjing Forestry University, Nanjing 210037, China 2 Laboratory for Comprehensive Utilization of Paper Waste of Shandong Province, Shandong Huatai Paper Company Ltd., Dongying 257335, China 3 SCNU Environmental Research Institute, Guangdong Provincial Key Laboratory of Chemical Pollution and Environmental Safety and MOE Key Laboratory of Theoretical Chemistry of Environment, School of Environment, South China Normal University, Guangzhou 510006, China Corresponding authors: Fengshan Zhang (htjszx@163.com) and Mingzhi Huang (mingzhi.huang@m.scnu.edu.cn) This work was supported by the Guangdong Provincial Natural Science Foundation of China under Grant 2016A030306033. ABSTRACT Advanced process modeling methods have been used for prediction and monitoring of key quality indices in wastewater treatment processes. However, single conventional models usually have limited precision accuracy when predicting the effluent indices in papermaking wastewater treatment processes. To achieve a better prediction accuracy and robustness, we propose a stacking ensemble learn- ing (SEL) method which utilizes the advantages of the internal base-learning models. The method combines base-learning algorithms including partial least squares, support vector regression, and artificial neural networks with a meta-learning algorithm, which is a multiple-response linear regression in this work. To evaluate the model performance in practical applications, both real wastewater data and simulation wastewater data are used for modeling. The predicted effluent indices include effluent suspended solid (SS eff ), effluent chemical oxygen demand (COD eff ), effluent ammonia concentration (S NHeff ), and efflu- ent nitrate concentration (S NOeff ). Compared with base-learning algorithms and other ensemble learning methods, the results demonstrate that SEL significantly improves the prediction accuracy and reduces the prediction errors, which provides a new way to achieve real-time monitoring of wastewater treatment processes.
INDEX TERMS Stacking ensemble learning, papermaking process modeling, effluent indices, prediction accuracy, wastewater treatment processes.
I. INTRODUCTION The key to improving the quality management efficiency in wastewater treatment processes (WWTPs) heavily relies on the implementation of effective real-time monitoring of the effluent concentrations [1]. In recent years, hardware sensors in WWTPs have been exposed to a series of shortcomings in the monitoring process, such as significant time lags and high maintenance costs [2]. On the contrary, soft sensing methods can save measurement cost and improve monitoring quality in WWTPs, which is not only economical and reliable but also have a dynamic response [3]–[5]. For example, soft sensors can make real-time predictions for key WWTP vari- ables including the concentrations of suspended solids (SS),
chemical oxygen demand (COD), total phosphorus (TP), total nitrogen (TN), and daily sewage sludge. For some variables with periodic features, variables information can be divided into the periodic component and the residual component based on periodic analysis. Soft sensors combining with peri- odic analysis can achieve a better prediction performance [6]. Nowadays, increasing attention has been paid in advanced monitoring techniques in WWTPs [7], [8]. In recent years, the main conventional methods including partial least squares (PLS), support vector regression (SVR), and artificial neural networks (ANN) have been used for the prediction of wastewater effluent indices. However, the com- plex characteristics in wastewater treatment processes, such as nonlinearity, time-varying characteristic, and uncertainty make it difficult to obtain a satisfactory prediction result using these conventional methods. For example, the linear
The associate editor coordinating the review of this manuscript and approving it for publication was Liangxiu Han .
This work is licensed under a Creative Commons Attribution 4.0 License. For more information, see https://creativecommons.org/licenses/by/4.0/
180844
VOLUME 8, 2020
Made with FlippingBook flipbook maker