<?xml version="1.0" encoding="UTF-8"?><!DOCTYPE article  PUBLIC "-//NLM//DTD Journal Publishing DTD v3.0 20080202//EN" "http://dtd.nlm.nih.gov/publishing/3.0/journalpublishing3.dtd"><article xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink" dtd-version="3.0" xml:lang="en" article-type="research article"><front><journal-meta><journal-id journal-id-type="publisher-id">JMF</journal-id><journal-title-group><journal-title>Journal of Mathematical Finance</journal-title></journal-title-group><issn pub-type="epub">2162-2434</issn><publisher><publisher-name>Scientific Research Publishing</publisher-name></publisher></journal-meta><article-meta><article-id pub-id-type="doi">10.4236/jmf.2016.61013</article-id><article-id pub-id-type="publisher-id">JMF-63853</article-id><article-categories><subj-group subj-group-type="heading"><subject>Articles</subject></subj-group><subj-group subj-group-type="Discipline-v2"><subject>Business&amp;Economics</subject><subject> Physics&amp;Mathematics</subject></subj-group></article-categories><title-group><article-title>
 
 
  Forecasting Outlier Occurrence in Stock Market Time Series Based on Wavelet Transform and Adaptive ELM Algorithm
 
</article-title></title-group><contrib-group><contrib contrib-type="author" xlink:type="simple"><name name-style="western"><surname>argess</surname><given-names>Hosseinioun</given-names></name><xref ref-type="aff" rid="aff1"><sub>1</sub></xref><xref ref-type="corresp" rid="cor1"><sup>*</sup></xref></contrib></contrib-group><aff id="aff1"><label>1</label><addr-line>Statistics Department, Payame Noor University, Tehran, Iran</addr-line></aff><author-notes><corresp id="cor1">* E-mail:<email>mails.students@gmail.com</email></corresp></author-notes><pub-date pub-type="epub"><day>05</day><month>02</month><year>2016</year></pub-date><volume>06</volume><issue>01</issue><fpage>127</fpage><lpage>133</lpage><history><date date-type="received"><day>30</day>	<month>December</month>	<year>2015</year></date><date date-type="rev-recd"><day>accepted</day>	<month>23</month>	<year>February</year>	</date><date date-type="accepted"><day>26</day>	<month>February</month>	<year>2016</year></date></history><permissions><copyright-statement>&#169; Copyright  2014 by authors and Scientific Research Publishing Inc. </copyright-statement><copyright-year>2014</copyright-year><license><license-p>This work is licensed under the Creative Commons Attribution International License (CC BY). http://creativecommons.org/licenses/by/4.0/</license-p></license></permissions><abstract><p>
 
 
   In financial field, outliers represent volatility of stock market, which plays an important role in management, portfolio selection and derivative pricing. Therefore, forecasting outliers of stock market is of the great importance in theory and application. In this paper, the problem of predicting outliers based on adaptive ensemble models of Extreme Learning Machines (ELMs) is considered. We found out that the proposed model is applicable for outlier forecasting and outperforms the methods based on autoregression (AR) and extreme learning machine (ELM) models. 
 
</p></abstract><kwd-group><kwd>Component</kwd><kwd> Extreme Learning Machine</kwd><kwd> Outliers Forecasting</kwd><kwd> Wavelet Transform</kwd></kwd-group></article-meta></front><body><sec id="s1"><title>1. Introduction</title><p>Outliers can have deleterious effects on statistical analyses. They can result in parameter estimation biases, invalid inferences and weak volatility forecasts in financial data. As a result when modeling financial data, their detection and correction should be considered seriously. Time-series data are often messed up with outliers due to the influence of unusual and non-repetitive events. Forecast accuracy in such situations is decreased dramatically due to a carry-over effect of the outliers on the point forecast and a bias in the estimate of parameters. The effect of additive outliers on forecasts is studied by Ledolter [<xref ref-type="bibr" rid="scirp.63853-ref1">1</xref>] . It was shown that forecast intervals are quite sensitive to additive outliers, but that point forecasts are largely unaffected unless the outlier occurs near the forecast origin. In such a situation the carry-over effect of the outlier can be quite substantial.</p><p>Considerable research has been devoted to the subject of forecasting and various methods have been suggested which have been divided into two main groups: classical methods mainly exponential smoothing, regression, Box-Jenkins autoregressive integrated moving average (ARIMA), generalized autoregressive conditionally heteroskedastic (GARCH) methods, and modern methods applying artificial intelligence techniques including artificial neural networks (ANN) and evolutionary computation (for more discussed details see [<xref ref-type="bibr" rid="scirp.63853-ref2">2</xref>] -[<xref ref-type="bibr" rid="scirp.63853-ref4">4</xref>] ). Extreme learning machine (ELM) has been proposed as a class of learning algorithm for single hidden layer feedforward neural networks (SLFNs). In ELM algorithm, the connections between the input layer and the hidden neurons are randomly assigned and remain unchanged during the learning process. Thus by minimizing the cost function through a linear system the output connections are tuned. The computational burden of ELM has been significantly reduced as the only cost is solving a linear system. The low computational complexity attracted a great deal of attention from the research community, especially for high dimensional and large data applications. While considerable research has been devoted to detecting and removing outliers, few focused on forecasting them.</p><p>Outliers forecasting model has been discussed in [<xref ref-type="bibr" rid="scirp.63853-ref5">5</xref>] for the two market indexes and six individual stocks based on multi-feature extreme learning machine (ELM) algorithm. The purpose of this paper is to present adaptive ensemble model of Extreme Learning Machines (ELMs) for prediction which can lead to smaller predicting errors and more accuracy than some other forecasting methods. This paper is structured as follows: In Section 2, the theories of wavelet transform and ELM are presented, as well as how we combine both of them in the adaptive ensemble method. Section 3 describes the numerical studies, while Section 4 discusses the results.</p></sec><sec id="s2"><title>2. Methodology</title><p>In this section, we present the methodology employed for forecasting outliers applying a wavelet decomposition technique and ELM algorithm.</p><sec id="s2_1"><title>2.1. Wavelet Transforms</title><p>This section contains some facts about wavelets, used throughout this paper. A thorough review of the wavelet transform is discussed in Mallat [<xref ref-type="bibr" rid="scirp.63853-ref6">6</xref>] -[<xref ref-type="bibr" rid="scirp.63853-ref8">8</xref>] . The wavelet analysis is a mathematical tool that offers decomposition of signal s(t) into many frequency bands at many scales. In particular, the signal s(t) is decomposed into smooth coefficients α and detail coefficients d, which are given by</p><disp-formula id="scirp.63853-formula546"><graphic  xlink:href="http://html.scirp.org/file/13-1490399x6.png"  xlink:type="simple"/></disp-formula><p>where Φ is the father and Ψ is the mother wavelets, and j and k are, respectively, the scaling and translation parameters. The father wavelet (function) keeps the frequency domain properties (low-frequency) of the signal, while the mother wavelet keeps the time domain properties (high-frequency). The father wavelet Φ and the mother wavelet Ψ are defined as follows:</p><disp-formula id="scirp.63853-formula547"><graphic  xlink:href="http://html.scirp.org/file/13-1490399x7.png"  xlink:type="simple"/></disp-formula><p>The two wavelets Φ and Ψ satisfy the condition <inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/13-1490399x8.png" xlink:type="simple"/></inline-formula> and<inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/13-1490399x9.png" xlink:type="simple"/></inline-formula>. Consequently, the orthogonal wavelet</p><p>representation of the signal s(t) is given by</p><disp-formula id="scirp.63853-formula548"><graphic  xlink:href="http://html.scirp.org/file/13-1490399x10.png"  xlink:type="simple"/></disp-formula><p>Using the above decomposition, the original signal s(t) is represented with approximation coefficients α(t) and detail coefficients d(t), by convolving the signal s(t) with a low-pass filter (LP) and a high-pass filter (HP), respectively. The low-pass filtered signal is the input for the next iteration step and so on. The approximation coefficients α(t) contain the general trend (the low-frequency components) of the signal s(t), and the detail coefficients d(t) contain its local variations (the high-frequency components).</p></sec><sec id="s2_2"><title>2.2. Extreme Learning Machine (ELM) Algorithm</title><p>The purpose of this paper is to discuss the mythology behind the Extreme learning machine (ELM). ELM is an improved learning algorithm for the single feed-forward neural network structure. It notably differs from the traditional neural network methodology, since it is not essential to tune all the parameters of the feed-forward networks (input weights and hidden layer biases). For more information on efficiency of SLFNs with randomly chosen input weights, hidden layer biases and a nonzero activation function to approximate any continuous functions on any input set, one can refer to [<xref ref-type="bibr" rid="scirp.63853-ref9">9</xref>] and [<xref ref-type="bibr" rid="scirp.63853-ref10">10</xref>] .</p><p>The proposed extreme learning machine (ELM) has shown its efficiency in training feedforward neural networks and overcoming the limitations faced by other conventional algorithms [<xref ref-type="bibr" rid="scirp.63853-ref11">11</xref>] [<xref ref-type="bibr" rid="scirp.63853-ref12">12</xref>] . The essences of ELM lie in two aspects, that is, random neurons and the tuning-free strategy. The learning phase of ELM generally includes two steps, namely, constructing the hidden layer output matrix with random hidden neurons and finding the output connections. Thanks to using random hidden neuron parameters which remain unchanged during the learning phase, ELM enjoys a very low computational complexity. The computational burden has been greatly reduced as the only cost is solving a linear system. At the same time, numerous applications have shown that ELM can provide a comparable or better generalization performance than the popular support vector machine (SVM) [<xref ref-type="bibr" rid="scirp.63853-ref13">13</xref>] [<xref ref-type="bibr" rid="scirp.63853-ref14">14</xref>] and the BP method in most cases [<xref ref-type="bibr" rid="scirp.63853-ref15">15</xref>] -[<xref ref-type="bibr" rid="scirp.63853-ref17">17</xref>] .</p><p>ELM is a single-hidden layer feedforward network (SLFN) with a special learning mechanism which is consists of three layers: input layer, hidden layer and output layer. Suppose the SLFN has n hidden nodes and nonlinear activation function g(x). For N training samples<inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/13-1490399x11.png" xlink:type="simple"/></inline-formula>, where <inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/13-1490399x12.png" xlink:type="simple"/></inline-formula> is the ith input vector and t<sub>i</sub> is the ith desired output, the SLFN can be modeled by</p><disp-formula id="scirp.63853-formula549"><graphic  xlink:href="http://html.scirp.org/file/13-1490399x13.png"  xlink:type="simple"/></disp-formula><p>where <inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/13-1490399x14.png" xlink:type="simple"/></inline-formula> is the input weight vector linking the jth hidden node and the input nodes, <inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/13-1490399x15.png" xlink:type="simple"/></inline-formula>is the bias of the jth hidden node, <inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/13-1490399x16.png" xlink:type="simple"/></inline-formula>is the output weight vector linking the jth hidden node and the output nodes, <inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/13-1490399x17.png" xlink:type="simple"/></inline-formula>is the actual network output. If ELM can approximate all the training samples <inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/13-1490399x18.png" xlink:type="simple"/></inline-formula> with zero error, then we claim that there exist<inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/13-1490399x19.png" xlink:type="simple"/></inline-formula>, <inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/13-1490399x20.png" xlink:type="simple"/></inline-formula>and <inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/13-1490399x20.png" xlink:type="simple"/></inline-formula><inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/13-1490399x21.png" xlink:type="simple"/></inline-formula> such that</p><disp-formula id="scirp.63853-formula550"><graphic  xlink:href="http://html.scirp.org/file/13-1490399x22.png"  xlink:type="simple"/></disp-formula><p>The above matrix can be expressed as Hβ = T, where H is called the hidden layer output matrix. As mentioned earlier, the input weights and hidden biases are randomly constructed and do not need tuning as in the case of traditional SLFN methodology. The evaluation of the output weights linking the hidden layer to the output layer is equivalent to determining the least-square solution to the given linear system. The minimum norm least-square (LS) solution to the linear system is</p><disp-formula id="scirp.63853-formula551"><graphic  xlink:href="http://html.scirp.org/file/13-1490399x23.png"  xlink:type="simple"/></disp-formula><p>The H in the above equation is the Moore-Penrose (MP) generalized inverse of matrix H, see [<xref ref-type="bibr" rid="scirp.63853-ref18">18</xref>] for more discussion. The minimum norm LS solution is unique and leads to smallest norm along all the LS solutions. The MP inverse method based on ELM algorithm is found to obtain a good generalization performance with a radically increased learning speed. One can present a general Algorithm for ELM as follows. For a given training set, activation function g(x) and hidden neuron number L:</p><p>Step 1: Assign random input weight <inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/13-1490399x24.png" xlink:type="simple"/></inline-formula> and bias<inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/13-1490399x24.png" xlink:type="simple"/></inline-formula><inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/13-1490399x25.png" xlink:type="simple"/></inline-formula>,<inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/13-1490399x24.png" xlink:type="simple"/></inline-formula><inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/13-1490399x25.png" xlink:type="simple"/></inline-formula><inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/13-1490399x26.png" xlink:type="simple"/></inline-formula>.</p><p>Step 2: Calculate the hidden layer output matrix H.</p><p>Step 3: Calculate the output weight.</p><p>Theoretical discussions and a more thorough presentation of the ELM algorithm are detailed in the original papers [<xref ref-type="bibr" rid="scirp.63853-ref19">19</xref>] [<xref ref-type="bibr" rid="scirp.63853-ref20">20</xref>] .</p></sec><sec id="s2_3"><title>2.3. Adaptive ELM</title><p>Comparable to other flexible nonlinear estimation methods, the ELM may suffer either under-fitting or over-fit- ting [<xref ref-type="bibr" rid="scirp.63853-ref19">19</xref>] . Over-fitting is particularly inaccurate since it can cause wild prediction far beyond the range of the training data even with the noise-free data. It may lead to poor predictive performance, as it may cause minor fluctuations in the data. In this work, the output of the network is only one value that is the predicted outliers.</p><p>The ensemble model is made up of a number of randomly initialized ELMs, which each have their own parameters. The model <inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/13-1490399x27.png" xlink:type="simple"/></inline-formula> has an associated weight <inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/13-1490399x27.png" xlink:type="simple"/></inline-formula><inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/13-1490399x28.png" xlink:type="simple"/></inline-formula> which determines its contribution to the prediction of the ensemble. Hence, we present our model only for one output. Let us define the input data as</p><disp-formula id="scirp.63853-formula552"><graphic  xlink:href="http://html.scirp.org/file/13-1490399x29.png"  xlink:type="simple"/></disp-formula><p>Comparing to the learned input patterns which is presented as</p><disp-formula id="scirp.63853-formula553"><graphic  xlink:href="http://html.scirp.org/file/13-1490399x30.png"  xlink:type="simple"/></disp-formula><p>The determination of the closeness measure is the major factor in prediction accuracy, for which adaptive metrics are introduced to solve this problem and the arithmetic is defined by:</p><disp-formula id="scirp.63853-formula554"><graphic  xlink:href="http://html.scirp.org/file/13-1490399x31.png"  xlink:type="simple"/></disp-formula><p>Studying time-series forecasting, the information on trends and amplitudes plays an effective role. Adaptive metrics are introduced to solve this problem, while the arithmetic is presented as:</p><disp-formula id="scirp.63853-formula555"><label>(1)</label><graphic position="anchor" xlink:href="http://html.scirp.org/file/13-1490399x32.png"  xlink:type="simple"/></disp-formula><p>where the parameter of minimization, <inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/13-1490399x33.png" xlink:type="simple"/></inline-formula>equilibrates the amplitude difference between <inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/13-1490399x33.png" xlink:type="simple"/></inline-formula><inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/13-1490399x34.png" xlink:type="simple"/></inline-formula> and <inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/13-1490399x33.png" xlink:type="simple"/></inline-formula><inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/13-1490399x34.png" xlink:type="simple"/></inline-formula><inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/13-1490399x35.png" xlink:type="simple"/></inline-formula> and</p><p><inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/13-1490399x36.png" xlink:type="simple"/></inline-formula>,</p><p>where <inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/13-1490399x37.png" xlink:type="simple"/></inline-formula> and <inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/13-1490399x37.png" xlink:type="simple"/></inline-formula><inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/13-1490399x38.png" xlink:type="simple"/></inline-formula> are the largest and smallest elements of vector correspondingly,</p><p><inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/13-1490399x39.png" xlink:type="simple"/></inline-formula>. The optimization problem (1) can be solved using the algorithm of Levenberg-</p><p>Marquardt optimization or other gradient methods for . For<inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/13-1490399x41.png" xlink:type="simple"/></inline-formula>, two equations may presented as blow:</p><disp-formula id="scirp.63853-formula556"><graphic  xlink:href="http://html.scirp.org/file/13-1490399x42.png"  xlink:type="simple"/></disp-formula><p>Then the solution of the minimization problem can be obtained analytically:</p><disp-formula id="scirp.63853-formula557"><graphic  xlink:href="http://html.scirp.org/file/13-1490399x43.png"  xlink:type="simple"/></disp-formula><p>where<inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/13-1490399x44.png" xlink:type="simple"/></inline-formula>, j = 1,2,<inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/13-1490399x44.png" xlink:type="simple"/></inline-formula><inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/13-1490399x45.png" xlink:type="simple"/></inline-formula>. The adaptive k-nearest neighbors are chosen and the</p><p>input vector of the first network can be defined as:</p><p><inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/13-1490399x46.png" xlink:type="simple"/></inline-formula>.</p><p>The forecasting error increases considerably because of the big difference between training data and input data. In order to get more accurate results for time series<inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/13-1490399x47.png" xlink:type="simple"/></inline-formula>, k sets of inputs are used and the output vector are<inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/13-1490399x47.png" xlink:type="simple"/></inline-formula><inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/13-1490399x48.png" xlink:type="simple"/></inline-formula>. The mechanism for admixture of outputs is presented as follows:</p><disp-formula id="scirp.63853-formula558"><graphic  xlink:href="http://html.scirp.org/file/13-1490399x49.png"  xlink:type="simple"/></disp-formula><p>where <inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/13-1490399x50.png" xlink:type="simple"/></inline-formula> is the distance between Q<sub>i</sub>’s vth nearest pattern and Q<sub>i</sub>. The model has been tested on both stationary and nonstationary time series, and the experiments show that in both cases the adaptive ensemble method leads to a prediction accuracy comparable to the best methods. For more detailed information see [<xref ref-type="bibr" rid="scirp.63853-ref16">16</xref>] [<xref ref-type="bibr" rid="scirp.63853-ref17">17</xref>] .</p></sec></sec><sec id="s3"><title>3. Numerical Studies</title><p>The data used in the paper is the daily value of Petroleum sector Index, obtained from the DataStream database services of Tehran Over-the-Counter Market (OTC)<sup>1</sup>. Since 2009, Iran has been developing an over-the-counter market for bonds and equities. OTC provides a complete available achieve of data, based on different sectors and dates. Our sample ranges from 28 Sep 2009 to 27 Dec 2015, with 1510 observations. Petroleum, the prime reason for the economic growth of the country, has been the primary industry in Iran since the 1920s. In 2012, Iran was the second-largest exporter among the Organization of Petroleum Exporting Countries<sup>2</sup>, which exports around 1.5 million barrels of crude oil a day. Through primary wavelet decomposition, sequence V’s low frequency <inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/13-1490399x51.png" xlink:type="simple"/></inline-formula> and high frequency <inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/13-1490399x51.png" xlink:type="simple"/></inline-formula><inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/13-1490399x52.png" xlink:type="simple"/></inline-formula> are computed. In order to eliminate stochastic diffusion we set high frequency <inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/13-1490399x51.png" xlink:type="simple"/></inline-formula><inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/13-1490399x52.png" xlink:type="simple"/></inline-formula><inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/13-1490399x53.png" xlink:type="simple"/></inline-formula> equal zero. To get the main trend<inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/13-1490399x51.png" xlink:type="simple"/></inline-formula><inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/13-1490399x52.png" xlink:type="simple"/></inline-formula><inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/13-1490399x53.png" xlink:type="simple"/></inline-formula><inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/13-1490399x54.png" xlink:type="simple"/></inline-formula>, inverse wavelet transform is used for low frequency <inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/13-1490399x51.png" xlink:type="simple"/></inline-formula><inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/13-1490399x52.png" xlink:type="simple"/></inline-formula><inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/13-1490399x53.png" xlink:type="simple"/></inline-formula><inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/13-1490399x54.png" xlink:type="simple"/></inline-formula><inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/13-1490399x55.png" xlink:type="simple"/></inline-formula> Then we compute the absolute residual of V as sequence</p><disp-formula id="scirp.63853-formula559"><graphic  xlink:href="http://html.scirp.org/file/13-1490399x56.png"  xlink:type="simple"/></disp-formula><p>Based on sequences obtained from Matlab, we then construct an AD-ELM abnormal predicting model which can predict whether abnormal fluctuation will appear today or not. Since an ELM is essentially a linear model of the responses of the hidden layer, we apply PRESS statistics in R to retrain the ELM in an incremental way. The number of input nodes for ELM, and AD-ELM are set as 10, and the number of hidden is set to be 5. A detailed discussion of inputs and hidden nodes of ELMs with PRESS can be found in [<xref ref-type="bibr" rid="scirp.63853-ref21">21</xref>] . <xref ref-type="fig" rid="fig1">Figure 1</xref> shows the outliers in green color, while the red plus signs (115 points) represent abnormal points.</p><p>In order to analyze outlier detection accuracy of AD-ELM method with other methods, an adequate error measure method must be selected. In this paper we apply mean squared error (NMSE) and Mean Absolute Percentage Error (MAPE). The first is used as the error criterion, which is the ratio of the mean squared error to the variance of the time series, while the second on is regarded as one of the standard statistical performance measures. For a time series <inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/13-1490399x57.png" xlink:type="simple"/></inline-formula> we have</p><disp-formula id="scirp.63853-formula560"><graphic  xlink:href="http://html.scirp.org/file/13-1490399x58.png"  xlink:type="simple"/></disp-formula><p>where <inline-formula><inline-graphic xlink:href="http://html.scirp.org/file/13-1490399x60.png" xlink:type="simple"/></inline-formula> is the predicted point and N is the number of predicted points.Different prediction models on the data is summarized in <xref ref-type="table" rid="table1">Table 1</xref>. In our work, the AR method using AR(m)</p><fig-group id="fig1"><label><xref ref-type="fig" rid="fig1">Figure 1</xref></label><caption><title> Outlier detection of daily value.</title></caption><fig id ="fig1_1"><label></label><graphic mimetype="image"   position="float"  xlink:type="simple"  xlink:href="http://html.scirp.org/file/13-1490399x61.png"/></fig></fig-group><table-wrap id="table1" ><label><xref ref-type="table" rid="table1">Table 1</xref></label><caption><title> Comparisons of monthly forecasting</title></caption><table><tbody><thead><tr><th align="center" valign="middle"  colspan="3"  >Error measures</th></tr></thead><tr><td align="center" valign="middle" ></td><td align="center" valign="middle" >NMSE</td><td align="center" valign="middle" >MAPE</td></tr><tr><td align="center" valign="middle" >AR</td><td align="center" valign="middle" >1.5678</td><td align="center" valign="middle" >42.65%</td></tr><tr><td align="center" valign="middle" >ELM</td><td align="center" valign="middle" >0.6345</td><td align="center" valign="middle" >12.54%</td></tr><tr><td align="center" valign="middle" >AD-ELM</td><td align="center" valign="middle" >0.08436</td><td align="center" valign="middle" >9.54%</td></tr></tbody></table></table-wrap><p>where m is the number of input nodes of AD-ELM. In the simulation, the NMSE are 1.5678, 0.6345, and 0.08436 for AR, ELM, AD-ELM respectively, and the MAPE are 42.65%, 12.54%, 9.54% for AR, ELM, AD-ELM respectively. It is undeniable that the AD-ELM method improves upon the two other models.</p></sec><sec id="s4"><title>4. Results and Conclusion</title><p>In this paper, forecasting models mostly have been used to forecast the stock market index value outliers. The proposed AD-ELM method is successfully used for market indexes of Tehran Over-the-Counter Market (OTC) for Petroleum sector for 1510 observations. Outliers of time series are firstly calculated through wavelet decomposition and then prediction is constructed using AD-ELM method. We plot outlier detection and evaluate forecast accuracy by mean squared error and Mean Absolute percentage error. The results reveal successfully that the accuracy of the proposed method can lead to smaller NMSE (0.08436) and MAPE (9.45%); comparing to autoregression (AR) and extreme learning machine (ELM) models, thus the AD-ELM method is a superior method for the practical forecasting of time series.</p></sec><sec id="s5"><title>Acknowledgements</title><p>This research is supported by ‎Payame Noor University‎, ‎19395-4697‎, ‎Tehran‎, Iran. The author gratefully acknowledges the constructive comments, offered by anonymous referee which help to improve the quality of the paper significantly.</p></sec><sec id="s6"><title>Cite this paper</title><p>NargessHosseinioun, (2016) Forecasting Outlier Occurrence in Stock Market Time Series Based on Wavelet Transform and Adaptive ELM Algorithm. Journal of Mathematical Finance,06,127-133. doi: 10.4236/jmf.2016.61013</p></sec><sec id="s7"><title>NOTES</title></sec></body><back><ref-list><title>References</title><ref id="scirp.63853-ref1"><label>1</label><mixed-citation publication-type="other" xlink:type="simple">Ledolter, J. (1989) The Effect of Additive Outliers on the Forecasts from ARIMA Models. International Journal of Forecasting, 5, 231-240. &lt;/br&gt;http://dx.doi.org/10.1016/0169-2070(89)90090-3</mixed-citation></ref><ref id="scirp.63853-ref2"><label>2</label><mixed-citation publication-type="other" xlink:type="simple">Araujo, R.A. (2011) A Class of Hybrid Morphological Perceptrons with Application in Time Series Forecasting. Knowledge-Based Systems, 24, 513-529. &lt;/br&gt;http://dx.doi.org/10.1016/j.knosys.2011.01.001</mixed-citation></ref><ref id="scirp.63853-ref3"><label>3</label><mixed-citation publication-type="other" xlink:type="simple">Bodyanskiy, Y. and Popov, S. (2006) Neural Network Approach to Forecasting of Quasiperiodic Financial Time Series. European Journal of Operational Research, 175, 1357-1366. &lt;/br&gt;http://dx.doi.org/10.1016/j.ejor.2005.02.012</mixed-citation></ref><ref id="scirp.63853-ref4"><label>4</label><mixed-citation publication-type="other" xlink:type="simple">Cao, J., Lin, Z., Huang, B. and Liu, N. (2012) Voting Based Extreme Learning Machine. Information Sciences, 185, 66-77. &lt;/br&gt;http://dx.doi.org/10.1016/j.ins.2011.09.015</mixed-citation></ref><ref id="scirp.63853-ref5"><label>5</label><mixed-citation publication-type="other" xlink:type="simple">Fang, Z.J., Zhao, J., Fei, F.C., Wang, Q.Y. and He, X. (2013) An Approach Based on Multi-Features Wavelet and ELM Algorithm for Forecasting Outlier Occurrence in Chinese Stock Market. Journal of Theoretical and Applied Information Technology, 49, 369-377.</mixed-citation></ref><ref id="scirp.63853-ref6"><label>6</label><mixed-citation publication-type="other" xlink:type="simple">Mallat, S. (1989) A Theory for Multiresolution Signal Decomposition the Wavelet Representation. IEEE Transactions on Pattern Analysis and Machine Intelligence, 31, 679-693. &lt;/br&gt;http://dx.doi.org/10.1109/34.192463</mixed-citation></ref><ref id="scirp.63853-ref7"><label>7</label><mixed-citation publication-type="other" xlink:type="simple">Daubechies, I. (1992) Ten Lectures on Wavelets, CBMS-NSF Regional Conferences Series in Applies Mathematics. SIAM, Philadelphia.</mixed-citation></ref><ref id="scirp.63853-ref8"><label>8</label><mixed-citation publication-type="other" xlink:type="simple">Daubechies, I. (1988) Orthogonal Bases of Compactly Supported Wavelets. Communication in Pure and Applied Mathematics, 41, 909-996. &lt;/br&gt;http://dx.doi.org/10.1002/cpa.3160410705</mixed-citation></ref><ref id="scirp.63853-ref9"><label>9</label><mixed-citation publication-type="other" xlink:type="simple">Huang, G.B., Chen, L. and Chee-Kheong, S. (2006) Universal Approximation Using Incremental Constructive Feedforward Networks with Random Hidden. IEEE Transactions on Neural Network, 17, 879-892.&lt;/br&gt;http://dx.doi.org/10.1109/TNN.2006.875977</mixed-citation></ref><ref id="scirp.63853-ref10"><label>10</label><mixed-citation publication-type="other" xlink:type="simple">Huang, G.B. and Slew, C.K. (2004) Extreme Learning Machine: RBF Network Case. Proceedings of the 8th International Conference on Control, Automation, Robotics and Vision, 2, 1029-1036.</mixed-citation></ref><ref id="scirp.63853-ref11"><label>11</label><mixed-citation publication-type="other" xlink:type="simple">Guo, Z., Wu, J., Lu, H. and Wang, J. (2011) A Case Study on a Hybrid Wind Speed Forecasting Method Using BP Neural Network. Knowledge-Based Systems, 24, 1048-1056. &lt;/br&gt;http://dx.doi.org/10.1016/j.knosys.2011.04.019</mixed-citation></ref><ref id="scirp.63853-ref12"><label>12</label><mixed-citation publication-type="other" xlink:type="simple">Li, M.-B., Huang, G.-B., Saratchandran, P. and Sundararajan, N. (2005) Fully Complex Extreme Learning Machine. Neurocomputing, 68, 306-314. &lt;/br&gt;http://dx.doi.org/10.1016/j.neucom.2005.03.002</mixed-citation></ref><ref id="scirp.63853-ref13"><label>13</label><mixed-citation publication-type="other" xlink:type="simple">Cortes, C. and Vapnik, V. (1995) Support-Vector Networks. Machine Learning, 20, 273-297.&lt;/br&gt;http://dx.doi.org/10.1007/BF00994018</mixed-citation></ref><ref id="scirp.63853-ref14"><label>14</label><mixed-citation publication-type="other" xlink:type="simple">Hsu, C.W. and Lin, C.J. (2002) A Comparison of Methods for Multiclass Support Vector Machines. IEEE Transactions on Neural Networks, 13, 415-425. &lt;/br&gt;http://dx.doi.org/10.1109/72.991427</mixed-citation></ref><ref id="scirp.63853-ref15"><label>15</label><mixed-citation publication-type="other" xlink:type="simple">Huang, G.B., Zhou, H., Ding, X. and Zhang, R. (2012) Extreme Learning Machine for Regression and Multiclass Classification. IEEE Transactions on Systems, Man, and Cybernetics, Part B: Cybernetics, 42, 513-529.&lt;/br&gt;http://dx.doi.org/10.1109/TSMCB.2011.2168604</mixed-citation></ref><ref id="scirp.63853-ref16"><label>16</label><mixed-citation publication-type="other" xlink:type="simple">Lin, C.T. and Lee, I.F. (2009) Artificial Intelligence Diagnosis Algorithm for Expanding a Precision Expert Forecasting System. Expert Systems with Applications, 36, 8385-8390. &lt;/br&gt;http://dx.doi.org/10.1016/j.eswa.2008.10.057</mixed-citation></ref><ref id="scirp.63853-ref17"><label>17</label><mixed-citation publication-type="other" xlink:type="simple">Liu, N. and Wang, H. (2010) Ensemble Based Extreme Learning Machine. IEEE Signal Processing Letters, 17, 754-757. &lt;/br&gt;http://dx.doi.org/10.1109/LSP.2010.2053356</mixed-citation></ref><ref id="scirp.63853-ref18"><label>18</label><mixed-citation publication-type="other" xlink:type="simple">Zhang, R., Lan, Y., Huang, G.B., Xu, Z.B. and Soh, Y.C. (2013) Dynamic Extreme Learning Machine and Its Approximation Capability. IEEE Transactions on Cybernetics, 43, 2054-2065.&lt;/br&gt;http://dx.doi.org/10.1109/TCYB.2013.2239987</mixed-citation></ref><ref id="scirp.63853-ref19"><label>19</label><mixed-citation publication-type="other" xlink:type="simple">Huang, G.B., Zhu, Q.Y. and Siew, C.K. (2006) Extreme Learning Machine: Theory and Applications. Neurocomputing, 70, 489-501. &lt;/br&gt;http://dx.doi.org/10.1016/j.neucom.2005.12.126</mixed-citation></ref><ref id="scirp.63853-ref20"><label>20</label><mixed-citation publication-type="other" xlink:type="simple">Xia, M., Zhang, Y., Weng, L. and Ye, X. (2012) Fashion Retailing Forecasting Based on Extreme Learning Machine with Adaptive Metrics of Inputs. Knowledge-Based Systems, 36, 253-259.&lt;/br&gt;http://dx.doi.org/10.1016/j.knosys.2012.07.002</mixed-citation></ref><ref id="scirp.63853-ref21"><label>21</label><mixed-citation publication-type="other" xlink:type="simple">Myers, R.H. (1990) Classical and Modern Regression with Applications. 2nd Edition, Pacific Grove, Duxbury.</mixed-citation></ref></ref-list></back></article>