PAPERmaking! Vol11 Nr3 2025

H. Liu et al. : Effluent Quality Prediction of Papermaking WWTPs Using SEL

N  i = 1

1 N

y i −ˆ y i y i

TABLE 2. Correlation coefficients between response variables and explanatory variables.

MAPE =

100% (14)

×

N  i = 1

2  N

(15)

( ˆ y i − y i )

RMSE =

where y i is measured value, ˆ y i is predicted value, and ¯ y i is the mean value of y i .

Chinese national effluent release standards. The rest of vari- ables (sample features) are explanatory variables which are closely related to the response variables. The specific correla- tion coefficients between response variables and explanatory variables are presented in Table 2. These variables deter- mine the degradation efficiency of pollutants by affecting the microbial activity of active sludge then indirectly influence the trend of SS eff andCOD eff . The simulation wastewater data were generated from benchmark simulation model no. 1 (BSM1) which is designed to simulate a wastewater treatment system. As shown in Figure 5, BSM1 consists of five biological reactors and one secondary sedimentation tank. BSM1 can provide diverse control strategies under three weather conditions including dry weather, rainy weather and storm weather. In this article, the dry weather data were used for simulation. The sampling period is 14 days and the sampling interval is 15 minutes. Finally, 1345 samples were generated and each of the sam- ples includes ten variables. Table 3 displays the specific process variables among which effluent ammonia concentra- tion (S NHeff ) and effluent nitrate concentration (S NOeff ) are response variables and the rest are explanatory variables. At the beginning of the modeling process, Jolliffe’s three parameters method was used to detect the outliers [29]. Then, the processed data were normalized to ensure the values of different features have the same dimension. The correspond- ing transformation formula is as follows : X  = X − μ σ (16) where X corresponds to the original sample vector, μ rep- resents the mean value of the original sample vector, and σ is the standard deviation of the original sample data. Then, all the data were divided into training data and test data. For the real wastewater data, the first 120 samples were used as training data and the rest 50 samples were used as test data. For the simulation wastewater data, the first 672 samples were used for training, and the remaining 673 samples were used for testing. B. PARAMETER OPTIMIZATION Although PLS has been successfully applied in many indus- trial processes, the calculation process for obtaining latent variables is still lack of a uniform standard. In this work, with the help of the index of variable importance in the projection scores [30], the scores of latent variables higher than 1 are considered as valuable latent variables for modeling.

FIGURE 4. Papermaking WWTP data.

TABLE1. Mean and standard deviation values of the papermaking WWTP data.

III. RESULTS AND DISCUSSION A. DESCRIPTION OF TWO DATA SETS

The modeling data used in this work include real wastewater data and simulation wastewater data. The actual wastewater data were collected from a papermaking WWTP in China. The data collection system includes various probes, signal acquisition card and power relay output board. As shown in Figure 4, the data include the following variables: wastew- ater flow rate ( Q ), influent suspended solid (SS in ), efflu- ent suspended solid (SS eff ), pH, temperature ( T ), influent chemical oxygen demand (COD in ), effluent chemical oxygen demand (COD eff ), and dissolved oxygen (DO). The statistics are listed in Table 1. Among the eight variables, SS eff and COD eff are response variables (target values) which need to be controlled. Both COD eff and SS eff variables met the

180848

VOLUME 8, 2020

Made with FlippingBook flipbook maker