H. Liu et al. : Effluent Quality Prediction of Papermaking WWTPs Using SEL
respectively. In terms of R 2 , the prediction accuracy of SEL is also improved significantly range from 4.76%-15.79%, 4.88%-10.26% for S NHeff andS NOeff . Compared with RF and AdaBoost, SEL has the best prediction results, specifically with the minimum RMSE value (0.31) and the maximum R 2 (0.88) for S NHeff , the minimum RMSE value (0.30) and the maximumR 2 (0.86) for S NOeff . Considering the fact that base-learning algorithms have their algorithm learning preference, their predictive capabil- ity may be limited for the papermaking wastewater effluent indices. Depending on different circumstances, SEL firstly uses the original training set to train base-learning algo- rithms, then uses the secondary training set generated by base-learning algorithms to train the meta-learning algorithm. In other words, the output values of base-learning algorithms are the input features of meta-learning algorithm. Based on the prediction results of the base-learning algorithms, SEL achieves a better generalization for the effluent indices. In addition, compared with other ensemble learning methods, SEL has higher prediction accuracy. From the viewpoint of theoretical analysis, the superior prediction performance of SEL can be mainly attributed to the following points: (1) Through exploring the feature space from different per- spectives, SEL can provide a more comprehensive analysis for complex characteristics in the wastewater data; (2) The SEL makes use of the advantages of different base- learning algorithms while gets rid of their relatively worse prediction drawback, so as to reduce the risk of trapping into local minima; (3) Compared with base-learning algorithms, SEL expends the hypothesis space in modeling process, which may be closer to the real hypothesis of the wastewater treatment process. (4) From the viewpoint of the diversity and integration of base-learning algorithms, SEL can combine different types of base-learners corresponding to various base-learning algo- rithms compared with other ensemble learning methods, and constructing a meta-learning algorithm is a more reasonable way than adopting statistically averaging, which makes SEL a better capacity of generalization. Although SEL has tremendous potential for improving the prediction accuracy in wastewater effluent data, the step of parameter optimization will be a time-consuming work. In this work, there are three groups of parameters need to be determined, which results in an increasing running time of SEL as shown in Table 7. The running time of SEL mainly spends on the ANN base-learning algorithm. Because ANN uses the back-propagation mechanism for optimizing the hyper-parameters, the network weights and the thresholds in each neuron need to be revised several times to finally reach the precision requirement, which immediately causes the increment of the total running time. In general, ensemble learning methods take more time than base-learning algo- rithms for training and prediction. Among ensemble learning methods, the running time of RF and AdaBoost is relatively
TABLE 7. Comparison of running time for COD eff and S NHeff .
shorter compared with SEL. This is mainly because the base-learners of RF and AdaBoost are relatively simple which makes RF and AdaBoost have low computational complexity. With the ever-increasing amounts of data in WWTPs, the complexity and running time will directly affect its prac- tical application and further development. In future work, besides optimizing the execution efficiency of SEL, more combinations of base-learning algorithms will be studied. Various machine learning algorithms, such as Gaussian pro- cess regression, relevance vector machine, gene expression programming, and evolutionary polynomial regression, can be integrated with the existing SEL. Meanwhile, the data set should be expended to further validate the prediction accuracy of SEL. IV. CONCLUSION In this work, the stacking ensemble learning is used for modeling the papermaking wastewater treatment process. By predicting the COD eff and SS eff from real wastewater data as well as S NHeff andS NOeff from wastewater simulation data, the proposed SEL method has successfully interpreted the complex characteristics of wastewater data. The predic- tion hypothesis space of SEL is more comprehensive for the wastewater treatment process. During the prediction process, the meta-learning algorithm enables SEL to obtain a further generalization result, which makes use of the advantages of each base-learning algorithm while avoiding the risk of trap- ping into local minima. The simulation results show that SEL has a better prediction ability compared with base-learning algorithms including PLS, SVR, ANN and other ensemble learning methods including RF and AdaBoost. Therefore, applying SEL into the real-time monitoring of wastewater treatment processes has practical reference value and realistic signification. Future work will be focused on the improve- ment in execution efficiency. Besides, more base-learning algorithms should be tried for stacking ensemble learning and sample size will be further expanded to verify the applicabil- ity of the proposed method.
ACKNOWLEDGMENT (Hongbin Liu and Chen Xin contributed equally to this work.)
180852
VOLUME 8, 2020
Made with FlippingBook flipbook maker