Full text
FACULDADE DE ENGENHARIA DA UNIVERSIDADE DO PORTO Using Clickstream Data to Analyze Online Purchase Intentions Ricardo Filipe Fernandes e Costa Magalhães Teixeira Mestrado Integrado em Engenharia Informática e Computação Supervisor: Dr. Vera Lúcia Miguéis Oliveira July 21, 2015
Using Clickstream Data to Analyze Online Purchase Intentions Ricardo Filipe Fernandes e Costa Magalhães Teixeira Mestrado Integrado em Engenharia Informática e Computação July 21, 2015
Abstract Nowadays, traditional business techniques are almost deprecated due to the insurgence of the world of online virtual shopping, the so-called e-commerce. This new, in many ways, uncharted territory poses difficult challenges when it comes to apply marketing techniques especially traditional methods, as these are not effective when dealing with online customers. In this context, it is imperative that companies have a complete in-depth understanding of online behavior in order to succeed within this complex environment in which they compete. The server Web logs of each customer are the main sources of potentially useful information for online stores. These logs contain details on how each customer visited the online store, moreover, it is possible to reconstruct the sequence of accessed pages, the so-called clickstream data. This data is fundamental in depicting each customer’s behavior. Analyzing and exploring this behavior is key to improve customer relationship management. The analysis of clickstream data allows for the understanding of customer intentions. One of the most studied measures regards customer conversion, that is, the percentage of customers that will actually perform a purchase during a specific online session. During this dissertation we investigate other relevant intentions, namely, customer purchasing engagement and real-time purchase likelihood. Actual data from a major European online grocery retail store will be used to support and evaluate different data mining models. Keywords: Clickstream, Web Usage Mining, Data mining, CRM i
ii
Resumo Hoje em dia as técnicas de negócio tradicionais estão ultrapassadas devido à emergência de novos modelos de negócio, nomeadamente no espaço online através da Internet. Este novo espaço de comércio eletrónico difere substancialmente das atividades tradicionais que têm por base espaços físicos. Assim, torna-se imperativo que as empresas adotem novas estratégias e sejam capazes de compreender as motivações que guiam os compradores online, caso pretendam ter sucesso no competitivo ecossistema virtual. Os logs dos servidores são a principal fonte de informação, sobre os seus utilizadores, que as empresas dispõem. Estes ficheiros contêm detalhes sobre como cada cliente navegou pela loja eletrónica, mais ainda, através destes dados, é possível reconstruir a sequência exata das páginas a que cada um acedeu. Este tipo de dados, conhecidos como dados de clickstream, são fundamentais para conseguir compreender o comportamento dos utilizadores. Aliás, a análise e exploração desta informação são essenciais para melhorar a relação com os clientes. A análise de dados clickstream permite, acima de tudo, a compreensão das intenções que motivam os utilizadores a realizar determinadas ações. A percentagem de conversão de utilizadores é uma das métricas mais conhecidas e que se relaciona diretamente com as intenções dos mesmos. Nesta dissertação são também explorados outro tipo de intenções, nomeadamente, fatores relacionados com os utilizadores que passam a ser compradores e ainda com a probabilidade de compra em tempo real. São utilizados dados reais, provenientes de uma das maiores empresas europeias na área do retalho alimentar, para suportar e avaliar diferentes modelos de data mining. Keywords: Clickstream, Web Usage Mining, Data mining, CRM iii
iv
Acknowledgements Prima facie, I am grateful to my family, namely to my parents and my brother, for all the support that was, is and always will be the main reason of my successes. To my mentor, Dr. Vera Oliveira, I am thankful for her guidance and support, without which I could not deliver a dissertation with such quality. I would also like to thank all my teachers that, over the past five years, taught me according to the best standards, and made me the engineer I am today. Before I enrolled this course, I was told that this academic experience would allow me to develop lifelong friendships. Five years later I can attest to that statement. I owe my friends all of the good memories and experiences that I lived throughout these years. Finally, I wish to thank my institution, the University of Porto, namely, its Faculty of Engineering, for providing me the tools to learn and enjoy an academic experience. Ricardo Filipe Fernandes e Costa Magalhães Teixeira v
LIST OF FIGURES 5.3 Collection of plots for different performance metrics for predicting purchase likelihood with an imbalanced data set. The top four plots refer to the evaluation of the logistic regression model, while the remaining bottom four relate to the evaluation oftherandomforestmodel.............................. 61 5.4 Collection of plots for different performance metrics for predicting purchase likelihood with a balanced data set. The top four plots refer to the evaluation of the logistic regression model, while the remaining bottom four relate to the evaluation oftherandomforestmodel.............................. 63 xii
List of Tables 3.1 CLMDataSetPageviews.............................. 26 3.2 CLMTrafficMetrics ................................ 28 3.3 CLMMetrics .................................... 30 3.4 CLMConversionMetrics.............................. 32 3.5 CLMUserArchetypes ............................... 33 3.6 CLMLandingPageRank.............................. 35 4.1 The label, short description of the variables used in the model selection and the typeofeachvariable................................. 42 5.1 The label, short description of the variables used in the model selection and the typeofeachvariable................................. 56 xiii
LIST OF TABLES xiv
Abbreviations/Acronyms AIB users that Added Item(s) to their Basket ANIB users that Added No Item(s) to their Basket ASL Average Session Length AU Anonymous Users AUC Area Under Curve CEO Chief Executive Officer CLV Customer Lifetime Value CRM Customer Relationship Management e-CRM electronic Customer Relationship Management FMCG Fast-Moving Consumer Goods GDP Gross Domestic Product GIF Graphics Interchangeable Format GMT Greenwich Mean Time HTTP Hypertext Transfer Protocol IP Internet Protocol ISP Internet Service Provider JPG Joint Photographic Experts Group KPI Key Performance Indicators PNG Portable Network Graphics RM Relationship Marketing ROC Receiver Operating Characteristic SVM Support Vector Machine URI Universal Resource Identifier URL Universal Resource Locator W3C World Wide Web Consortium WWW World Wide Web xv
Chapter 1 Introduction This chapter presents the introduction to this research. It will briefly introduce its context, problem description and leading contributions. The last section is dedicated to explain how this thesis is organized. 1.1 General Context In the last few decades, due to the booming of the World Wide Web (WWW) there has been a major change on how people interact with businesses. Nowadays, influenced by the social media and the fast lanes of communication provided by the Internet, traditional stores offer goods/services not only through the common physical channels, such as retail outlets, but also in online virtual stores, the so-called e-commerce. A recent study [NWW14] provided by an organization, Ecommerce Europe, that represents more than 25,000 European companies with a presence in the electronic commerce sector, reveals that the e-commerce business alone is responsible for 2.2% of the European GDP. Moreover, this study delves deeper into the online space and reports that there are 264 million (32%) e-shoppers in Europe, i.e. people that buy services/goods online, and these are accountable for 5.7% estimated share of online goods in total retail of goods. Due to this increase in significance and market share, e-commerce companies have to adopt new strategies that fit the needs of the online customers [RF03]. This type of customers have different behaviors than the physical ones [CTH+10]. Furthermore, web users have a much easier job in comparing different companies through simple online queries, making it hard to maintain customer loyalty and retention the same way it is done in the traditional stores. Therefore, the online market poses a new challenge that requires new and bold Customer Relationship Management (CRM) strategies. With the computer power flourishing in the early nineties, companies started to uncover the power of customer behavior mining. A new kind of directed marketing based on customer knowledge emerged in some of the world’s biggest retailers. Tesco, for instance, is a major front runner 1
Introduction in applying data mining techniques to learn from what their clients buy [McE02]. Even when technology was not as accessible as today, Tesco knew that important features were to be extracted from what their customers were buying. They were pioneers in creating a loyalty card system that tracked all purchases made by users. On a six-year spree Tesco’s Clubcard had amassed 104 billion rows of data. The main purpose of the card was to build a “segmentation and modelling system based on shopping behavior”. The objective was met after a careful analysis of the collected data that culminated on a division into 25 categories that ranged all kinds of clients and allowed the retailer to practice target marketing strategies that appealed those in specific categories. It was reported that sales effectiveness of new stores increased by 50 percent since this method was applied [McE02]. On the customer side, this attention and reward for loyalty is welcomed, as this is nothing more than a modern version of what people used to enjoy back when the service provided by retailers was much more personal. Namely, when customers relied on a familiar salesperson to help them find just what they wanted [DDL11]. Nowadays, especially when it comes to the e-commerce, the sheer volume of data that is stored from the clients’ interaction with the company can and should be used in order to simulate or, sometimes, outperform what used to be the familiar and personalized salesperson. In order to do so, e-commerce managers and marketers must depict a plan to improve the electronic Customer Relationship Management (e-CRM) which means delving into the field of data mining as well. As the Internet works on a basis of interchangeable data, there are new data sources that companies ought to exploit. This data enables e-commerce managers to overview the business in ways that were not previously possible. Through an online store it is possible to track much more data that is a direct result of how the customer interacts with the company. The so-called clickstream data is key in understanding customer behavior and it is also the main source of information for the companies to adapt their service according to their audience [BS09]. In short, clickstream data, or clickstreams, is the common terminology for the collection of Web logs that compose the session of a specific customer on the company website. These sessions contain information regarding the path that a customer took through the website’s structure, in other words, the sequence of clicks, hence clickstream, which a user performed and led him to a different point on the online store. 1.2 Problem Description In order to build customer loyalty in e-commerce, companies must find an edge that will engage the consumer to repeat purchases in the future. The problem though, is that, unlike physical stores, customers are not bounded by the distance across stores. A mere few clicks can provide information from different sellers of the same product/service, therefore, companies ought to work twice as hard to strengthen the bond with their “e-customers” so that these do not defect. One way to build this relationship is through the use of the so-called recommender systems [JZFF10]. This kind of technology ranges a wide variety of systems that can be more or less complex and intelligent. Recommender systems are a section of information filtering systems and 2
Introduction seek to predict the “rating” or “preference” that a user would give to a certain product or service. Companies have been using these systems to learn from their customers in order to provide a customized overall experience that helps them organizing the vast information spread across the website. The area of research for recommender systems is somewhat popular, mainly during the last decade, as there is several literature covering the topic. The state-of-the-art recommender systems are those that combine data from several angles, such as, demographical, clickstream and historical data [Les15]. Major retailers such as Amazon have developed algorithms that aim to develop offers and organize information according to each unique client, in a way that would be unfeasible to replicate in a physical store [LSY03]. However, with the recent developments in the fields of machine learning and data mining, new alternatives to analyze an exploit customer data ought to be discovered. In this context, this thesis aims to develop a new methodology to support online CRM, with a focus on the e-grocery sector. The research applies data mining techniques to extract knowledge from databases containing clickstream data. Unlike the traditional recommender systems, the models developed in this study aim to use historical customer data to provide future insights that might help new marketing strategies increase customer retention. Moreover, there are two distinct problems that this thesis tackles. First off, there is the problem of discriminating users that are mere visitors from users that are potential buyers. Given that the industry conversion rate is bellow 6% [BLA+02], it is essential to identify this 6% potential audience and maximize their purchasing capacity in order to increase their customer lifetime value (CLV). A solid model that addresses this question and is able to classify each user as a possible buyer or non-buyer would be of great value to a marketing department. Furthermore, marketers could deploy targeted measures, early on the user session, and have a better chance of impacting the purchase decision. Additionally, there is a second problem that could have a deeper impact on how marketers respond to user behaviour. This question regards the likelihood of purchase on a specific page. In other words, given the domain of this study, online grocery retailing, for each request of each user, this hypothetical model predicts if any product is going to be added to the user’s basket. Putting it into perspective, the company behind the online store and their marketing team would have real time control over what and when their customers were or were not going to add something to their baskets. The possible marketing ramifications of this insight are not under the scope of this thesis, but, for instance, it would be possible to reorder product lists by profit, if the model yielded a high likelihood of purchase or, if that likelihood were to be low, a target pop-up window could be triggered to motivate that purchase. Together, these two hypothesis outline the questions upon which this study will focus. 1.3 Contributions In order to create a loyalty bond with online customers, businesses must to improve their websites to accommodate their needs, namely, help them organize and rank product information. This 3
Introduction is especially important when the underlying business is related to the food-retail market as the website aggregates information about thousands of different products. During the past decade, several recommender systems have emerged, based either on collaborative or content-based filtering. The effectiveness of these systems has been proven along the years, with several companies, from different fields, succeeding due to them. Recommender systems and other data mining techniques can improve customer loyalty, increase cross-selling rates and increase the conversion rates of browsers into buyers. Numerous literature has been published regarding these systems with main focus on customer clustering and product suggesting. This thesis goes further and aims to explore other forms of knowledge extraction. The research delves into online customer behavior and, using customer navigational data, i.e. clickstream data, tries to build a model that can output the likelihood of customer conversion while the customer is navigating the website. In other words, this research exploits the possibility of the clickstream data and the customer’s intentions being linked. In terms of academic contributions, this study is, to our knowledge the first to approach the online grocery retail market. On its own that is unique, specially because the grocery market, unlike other online markets, has intrinsic characteristics that need to be tackled differently. On the other hand, two models, regarding both hypothesis presented in the last section, were developed and each one of them is unique, both in terms of goals and predictor variables. 1.4 Outline This thesis is concerned with providing further developments in the area of web usage analysis to explore web browsing behavior patterns. We will demonstrate our findings with a data set that was provided by a major European e-grocery. The structure of this thesis is as follows: Chapter 2provides a brief overview of the metrics/reports that can be obtained using clickstream data. We will introduce two types of metrics: for the website, and for a web session. These measures will be computed for the given data set and some results are reported through tables and graphs. In Chapter 3we focus on the development of two models that are able to predict customer purchasing engagement. We detail the process of extracting knowledge from data, also known as feature engineering, by using different graphical data representations. We also follow a specific evaluation process that helps quantifying the predictive power of both models. Chapter 4is structured very similar to Chapter 3as we detail development of the same learning models but, this time, in order to predict real-time purchase likelihood. 4
Chapter 2 Literature Review This chapter is dedicated to literature review. Background work will be reviewed, not only from computer science, but also from business and marketing. Firstly, the reader will have a deeper understanding on how current techniques are being applied to physical and electronic markets. Then, the field of Web Mining will be introduced and several known techniques detailed. The last part of this chapter is dedicated to review different approaches to related problems that were found in the literature. 2.1 Customer Relationship Management (CRM) In order to successfully understand why customer’s intentions prediction systems are important, one must analyze some business related fields. Ultimately, the end-users of these prediction mechanisms are people that are not necessarily related to computer science, furthermore, those who would benefit the most belong to marketing departments. Therefore, it is important to understand some of the concepts that serve as guidelines for thriving businesses. 2.1.1 Classic CRM Back in the middle of the twentieth century, mass production techniques and mass marketing changed the competitive landscape by increasing product availability for customers. However, this process fundamentally changed the relation, sometimes personal, that business owners and employees maintained with their customers. Clients lost their uniqueness and started to become a mere “account number”, the same way shopkeepers also lost track of their individual needs as the market became flooded with products and service providers [CP03]. Nowadays, every company is fighting for a competitive advantage that will guarantee their futures. Some of these companies are wining this challenge through the implementation of CRM principles allied with the power of technology, allowing them to improve customer retention and loyalty. 5
Literature Review Figure 2.2: Web Usage Mining Process [VMKR13] link, or indirectly, by searching for it on a search engine. Once there, he would start browsing by the product and click on different results until the right one was found. Then the user would add the item to a virtual cart and proceed to conclude the order by filling shipment and payment forms. Finally the process concludes once the client reviews and confirms the order, upon, usually, a confirmation email is sent. During this process clickstream data was recorded on the company’s Web servers, allowing businesses to retrace and study customer behavior in order to improve marketing efforts [CTS00]. This section will delve deeper into clickstream data, namely, what are its attributes and what stages are involved in data preprocessing. 2.4.1 Log File Structure Every action that any user performs on a Website is promptly recorded on server-side file, referred to as a Web log file or, simply, log file. Every single action matches a different line of the log, therefore a log may be perceived as a collection of different lines, corresponding to different actions, with several attributes. This file may be comma-delimited, space-delimited or tab-delimited. Although there are different norms for Web log files, the most relevant attributes are common among all. The World Wide Web Consortium (W3C) has published a standard, known as W3C logging 12
Literature Review Figure 2.3: Portion of a typical server log [Liu07] that should be adopted by all parties. As follows, a brief description of the different fields which exist in a typical file is presented. Figure 2.3 depicts a fragment extracted from a Web log where all the fields mentioned previously are present. 2.4.1.1 Remote Host Field This field consists of the internet protocol (IP) address of the client that made the request. An IP address is a numerical label inherent to each device that uses the Internet Protocol for communication through the network. IP addresses serve two main purposes: host identification, indicates what the visitor seeks and location addressing, helps find out where it is [Dou92]. IP addresses are usually represented by dot-decimal notation, four numbers each ranging from 0 to 255. A typical IP address in the log entries would be 182.221.1.168. When it was conceived, the IPv4, the most common version of IP addresses, did not anticipate the sheer amount of devices that would end up joining the network. Therefore, nowadays there is a struggle to maintain the uniqueness of users through their IP addresses. Several Internet Service Providers (ISPs) have put in practice techniques where the pool of IP addresses allocated to them is shifting across different users, according to demand [BGL+09]. For Web usage mining this IP address management techniques can be hazardous since it is not possible to guarantee that the same device is accessing the Web server, just because they share the same IP address. Nevertheless, for the experimental part of this research, IP addresses will be used as user identifiers as this is the only method one can do so with the given data set. Therefore historical purchases will not be taken into account, and consequently, user sessions do not need to be linked together. 13
Literature Review 2.4.1.2 User Name Field The name of the authenticated user that accessed the server, i.e. this field only has any content for those users who have entered a password protected area of the Website. Although useful, especially for user’s unique identification, this field is rarely used because the sensitive information is not properly encrypted. Therefore, anonymous users are normally indicated by a hyphen. 2.4.1.3 Date and Time Fields The date and time of the local server for each request is recorded in the file in Greenwich Mean Time (GMT). These two fields can either appear as different attributes or coupled together onto a complete ISO 8601 date plus hours, minutes and seconds. This attribute is very important because it allows for the relative estimation of time that each user spends on each Web page. 2.4.1.4 HTTP Request Method Fields The primary purpose of Web logs was to debug and keep track of every method related to the Hypertext Transfer Protocol (HTTP), an application-level protocol for distributed, collaborative, hypermedia information systems [FGM+99]. There are usually three web log attributes that can summarize the HTTP protocol: the request method, the Uniform Resource Identifier (URI) and the protocol version. The most common request method is GET, which represents a request to retrieve data identified by the URI. For instance, "GET /about.html" represents a request to the Web server to upload the page "about.html" to the respective user. As seen in the previous example, the URI contains the page or document name and the directory path that is being requested by the client browser. Sometimes, HTTP requests require extra information that is passed as a query. This would happen if a client performed a search through the Website’s search engine, for example. Some of the HTTP attributes are extremely relevant for Web usage mining because they allow for the tagging of resources that a user is requesting. 2.4.1.5 Referrer Field The site that the user last visited and provided the link to the current request. This field is represented by that site’s URL because the source can either be external or internal. The referrer field has long been used for marketing purposes since it can track where people came from. If the user entered the Website’s URL explicitly in the correspondent browser’s input field, there is no previous reference and this field is filled with a hyphen. 2.4.1.6 User Agent Field The user agent indicates the user’s browser, browser version and operating system. Most importantly, this field can contain information regarding bots or web crawlers. A crawler is a program that visits vast pages across the Web in order to keep data up-to-date, depending on its purpose. 14
Literature Review The most common kind of crawlers are those associated with search engines, these kind of bots navigate Websites in order to create updated entries in their databases. Even though it might be tampered, this field can also be used to assess if the user is human, thereby making it an easy filter to clean raw log data. 2.4.1.7 Status Code Field The status code field is embodied by a three digit number that encodes the server response to a certain request. Codes of the form 2xx, for instance, indicate that the request from the client was received, understood and completed [FGM+99]. This field can also be used to perform further filtering, as invalid or server failures are not relevant for mining. 2.4.1.8 Transfer Volume Field The transfer volume field indicates the size of the document, in bytes, sent by the server to the client. This information is used by network supervisors to monitor the load of the Web server. 2.4.2 Preprocessing Web log files are not formatted to be mined and cannot be directly used if one hopes to create a solid model. Raw Web logs have to go through different stages in order to become mine-able. For instance, the original logs contain several records of user requests for Web scripts or images, i.e. elements from the Website that are not relevant to clickstream mining. This section is dedicated to the review of all the steps that are involved in preprocessing. Figure 2.4 presents a diagram that visually details the sequence of stages for data preprocessing. 2.4.2.1 Data Fusion and Cleaning In large-scale Websites, it is usual for the server to be a virtual entity that is spread across different machines. In some cases these systems use redundant data in order to improve workload distribution and overall efficiency. Data fusion has an important role within these scenarios as it merges all logs from different sources into a single file, ready for further processing. The next step of preprocessing is data cleaning. It is usually site specific and it mostly revolves around the filtering of unwanted references to objects that are not important in Web usage mining such as style files, scripts, graphics or sounds [EY10]. This cleaning may also entail the exclusion of some Web log attributes that are irrelevant to data mining analysis (e.g. HTTP version, number of bytes used, etc.). Moreover, cleaning can also involve the filtering of single-page visits, i.e. users that only sent one request to the server. These one-page visitors are irrelevant to this research as no behavior can be extracted from that information. Data filtering for cleaning purposes can be accomplished through the analysis of the requested URI field, namely, its extension. For example, in the e-grocery domain, graphical resources can 15
Literature Review Figure 2.4: Steps in data preparation for Web usage mining [Liu07] be considered irrelevant, therefore, requests for objects with extensions like JPG, GIF and PNG can be excluded from the log. 2.4.2.2 Despidering Another important topic on data preparation is despidering, which entails the removal of references due to crawler navigation. With the everyday growth and constant update of the Web, search engines dispatch a series of automated software, the so-called crawlers or spiders, whose purpose is to search the Web space and provide links containing the information needs of customers [Cas05]. These bots’ behaviors are quite distinct from humans, therefore it is crucial that records regarding their activities are removed from the data set, under the danger of tampering the learning process. The most famous Web search engines’ crawlers can be detected through the agent attribute on the Web log. Other heuristics can be applied to exclude these activities as “well-behaved” crawlers’ first request is the "robot.txt" file under the domain’s root. This practice is known as the standard robot exclusion protocol and it is designed to fasten the process of data cleaning. Even after these two methods of bot exclusion, the data set may still contain unwanted data, as an effect of bots that try to mimic human behavior. To detect and eliminate this data, one must apply more 16
Literature Review complex methods. For instance an outlier statistical analysis of the data set can be effective as bots tend to access an abnormal amount of resources. 2.4.2.3 User Identification The user’s actual identity is not a requirement to perform Web usage mining. However, it is crucial to distinguish different users. When Websites have high traffic, Web logs can become unreadable to the human eye, because records from the same user would not be in sequential order. Moreover, one must be able to track the sequence of activities performed by the same user during different sessions, which is usually referred as user activity. Ideally, user identification would be easily accomplished if the user provided login information, such as username and password, each time the Website was accessed. Unfortunately, the Internet’s inherent design implies that most user online traffic requests are anonymous. In the absence of authentication mechanisms, the most widespread technique to distinguish different users entails the user of client-side cookies. Although effective, cookies are not completely reliable as users can disable them, due to privacy concerns. Another method that solely relies on data already available in the Web log is based on IP addresses, as different IP addresses are likely to yield distinct users. Then again, due to the increasing number of internet users, ISP proxy servers are applying IP rotation methods to counteract the IP depletion phenomena [BGL+09]. For that reason, the same IP can belong to different users at different times. To overcome this problem, other heuristics can be applied, such as the pairing of both IP and agent fields to uniquely distinguish users [Jam11]. It is important to empathize that user identification is, in fact, machine identification, as the actual identity of the person who is using each device is unknown. It is perfectly possible to assume that the same computer is being used by several users, or even that the same user is accessing the Website through different devices (smartphone, laptop, desktop, etc.). The methods described are only heuristics that try to approach the data set with the most common use case, that is, each machine yields the same user. Figure 2.5 presents an example of pre and post user identification using the IP address and agent fields as primary key. 2.4.2.4 Sessionization Perhaps the most important step in Web usage data preprocessing is sessionization, i.e. the set of pages viewed by a particular user for a certain purpose. After user identification is completed, this stage preforms a segmentation of user activity records from each identified user into sessions, each representing a complete visit to the Website. Without proper authentication systems, where sessionization can be performed by simply splitting the records according to each login/logout pair, one must rely on other heuristic methods. Different techniques have been studied by recent literature since the problem was first approached [CMS99]. New methods are still being published each one offering different nuances to their solution [CD08,DC11,KND13,RJR12,SS13]. 17
Literature Review Figure 2.5: Example of user identification using IP + Agent [Liu07] Sessionization heuristics are categorized into two distinct sets: time-oriented and navigationoriented. The time-oriented methods apply time-out rules to distinguish successive sessions. This heuristic has the underlying assumption that, if the same user has records across long periods of time, it is likely the he visited the site more than once. The standard time-oriented method defines a constant threshold, upon which if a user surpasses that time it is considered to end that session [CMS99]. More recent methods try to relax this hard constraint and produce a different sessionization using statistical quantities such as the average amount of time spent on each page [DC11]. Figure 2.6 portrays an example sessionization, obtained via two different time-oriented methods: global time-out, a threshold is defined for an entire session; local time-out, the threshold is defined for the total time spent between two subsequent requests. Navigation-oriented heuristics use either the site structure or the implicit linkage structure captured in the referrer fields of the server logs. One method of doing so is to keep adding records to an existing session as long as the referrer field for that request was previously invoked earlier in the session. If not it is considered a new session. Note that with this heuristic it is possible that a certain request may belong to more than one open session, since it may have been accessed previously in multiple sessions. In this case, additional information can be used for disambiguation. For instance, the request could be added to the most recently opened session satisfying the above condition [Liu07]. Figure 2.7 illustrates an example of sessionization using a navigationaloriented heuristic. Comparing figures 2.6 and 2.7 it is clear that each method outputs different sessions. One must take the problem domain into account to choose the best fit. 2.4.2.5 Path Completion Another fundamental preprocessing task, performed after sessionization is path completion. This step addresses the problem of client-side caching, which occurs when a user returns to a previously requested/downloaded page. The vast majority of Web browsers use caching techniques to save 18
Literature Review Figure 2.6: Sample sessionization based on global time threshold of 30 minutes and local time threshold of 10 minutes [Liu07] bandwidth and improve response times. On the other hand, whenever a user returns to the previous page, for instance, no request is sent to the Web server, therefore the server is unaware of these actions. However, missing references due to caching can be heuristically inferred through path completion which relies on knowledge of the site topology and referrer information from server logs [CMS99]. The most common case of missing reference occurs when the requested page is not directly linked to the last page. If so, one should look into the user’s recent request history and assume that he backtracked to the closest page with a reference to the current request. As server logs only contain data about the time of the request, in cases of missing references it is also necessary to interpolate the time spent on each page. An approach is to assume that any visit to an already seen page makes it work as an auxiliary page, which is used to link the user to different pages. Knowing the Website’s topology, the average reference length time for auxiliary pages can be used as an estimate of the access time for the missing pages [Jam11]. Figure 2.8 exemplifies a case where path completion was needed in order to have the correct clickstream sequence. 2.4.2.6 Data Integration Web usage mining is even more effective when different data sources are combined and fed into data mining algorithms. For example, this research seeks to link clickstream data with customer conversion, i.e. try to extract a common pattern across those users who became buyers. These patterns are better derived if further information is available, namely, demographic user data, historical purchases or even product details. The data integration stage involves the combination of different types of data onto one entity with diversified and relevant attributes [KMPZ04]. 2.4.2.7 Pageview/Transaction Identification A pageview or transaction is a conceptual notion that is linked to some action performed on a Website. For example, clicking on a link, reading an article, zooming a picture, adding a product to the shopping cart, etc. The task of aggregating meaningful page references is called pageview 19
Literature Review Figure 2.7: Sample for sessionization based on the navigation-oriented approach [Liu07] or transaction identification [CMS99]. This identification is heavily dependent on the intra-page structure of the site, as well as on the page contents and the underlying site domain knowledge. For a static single frame site, each page request necessarily corresponds to a specific user action. However, nowadays most Websites have dynamic content and one transaction can correspond to a collection of requests. It could also be relevant to create higher aggregation levels where seemingly distinct pages belong to the same transaction cluster because they fall in the same category. This process expedites the learning process as it decreases the overall entropy of the data set [Liu07]. This stage of preprocessing is strongly domain-dependent, however, its execution is crucial as one can only understand customers’ intentions if there is knowledge regarding the objective behind each page request. 2.4.3 Web Usage Mining: Previous Approaches As previously mentioned, understanding and predicting customer behavior and deploying marketing measures accordingly is of utmost importance to the success of e-commerce businesses [RS00,RF03]. The complex user behavior, mostly anonymous, combined with the high competition levels and low costs of switching product/service provider, accentuate the need to perform a thorough analysis of customer behavior [RS00]. Although still in its infancy, during the past decade, academic literature have been contributing with several publications in the field of Web usage mining. This section is dedicated to the review of some approaches that were published within this field of study and are more relevant to this dissertation. In contrast with physical stores, online users are likelier to visit an e-commerce Website without any actual motivation due to the low effort involved. This is the main reason behind the extremely low conversion percentages [LT07], therefore, it is important to be able to distinguish the different possible customer profiles. [MF04] elaborated on this and developed a typology of Website visits, using navigational patterns, which identified four types of browsing strategies: directed buying, search/deliberation, hedonic search and knowledge building. Visitors following a directed 20
Literature Review Figure 2.8: Missing references due to caching [Liu07] buying navigational path intend to make a purchase and follow a focused and directed search pattern. Search/deliberation targets users that are also focused on their search, however, they have not yet decided which product to buy. Typical sessions belonging to these users are focused on one specific product category with many comparable similar products. Unlike search/deliberation, hedonic browsing visitors tend to portrait exploratory search patterns, they are more stimulus driven and have not yet decided in what product category to buy. Finally, knowledge building users show high levels of exploratory browsing and are not considering any concrete purchase [MF04]. Still under consumer behavior research, [MLSL04] tried to dichotomize customer behavior as either browsing or deliberation. This classification is similar to that of [MF04] as browsing is linked to an exploratory pattern with less purchase intention while deliberation is similar to focused search and is purchase oriented. The most relevant conclusion of their research is that visitors might alternate between different behavioral modes during a single session. Due to this uncertainty, [MF04] developed a predictive model to accommodate for different types of behaviors. They were pioneers in feeding their model with clickstream data to predict purchase conversion with history of visits and purchases. During their research, the effect of visitors’ return to the Website and their evolving visit behavior was studied. They concluded that more frequent visitors have a greater propensity to buy. [VdPB05] studied the most relevant attributes for forecasting purchase behavior. They proposed a segmentation for the different kinds of variables: general visit-level clickstream behavior; detailed navigational path; customer demographics and historical purchasing behavior. The final model contained a set of variables that significantly improved performance, over previous studies. Moreover, they ultimately confirmed Moe and Fader’s suspicions by showing the importance of within session detailed clickstream behavior. One major drawback in comparing different methods is related to the tight correlation between the method’s success and the problem’s domain. For instance, [PC09] developed a purchase prediction model to the area of travel Websites. They justified customer conversion as a function of search motivation which also related to Moe and Fader’s concept of browsing behavior. On the other hand, the study concluded that the number of pageviews tend to negatively impact customer conversion. This conclusion is refuted by [Ver12], whose empirical study has data from a retailer 21
CLM Data Set Table 3.2: CLM Traffic Metrics Variable Frequency No. of Visits / Unique visitors 422 618 Total no. of pages visited 25 691 403 This section is organized similarly to that displayed in the second chapter of Jamalzahed’s work [Jam11]. Therefore, the metrics will be divided into two main sections: Navigation Metrics and Trend/Traversal Reports. 3.3.1 Navigation Metrics Clickstream data collected automatically by application servers is the primary source of data representing the navigational behaviour of visitors. Depending on the goals of the analysis, this data can be transformed and aggregated at different levels of abstraction to provide metrics to infer about user’s behavior. In this section we review some fundamental metrics often used in the web usage context. 3.3.1.1 Website Traffic Perhaps the most common indicator in the web usage context it the traffic of the website, i.e. the amount of data sent and received by visitors to a website [Pal02]. This measure is frequently used to assess the overall popularity of a website, although it can be applied to specific sections or pages within the website. The following types of information are used when determining web traffic: •Number of visits and unique visitors: The number of sessions that the website is visited over a specific time period, known as number of visits, is usually used as a web traffic indicator. Another, perhaps more important, derivation of this metric is known as unique visitors and it refers to the number of sessions originated from distinct users or machines. This distinction is made with the use of cookies that are unique for each visitor. These statistics are usually charted over periods of time, e.g on a weekly basis, and their purpose is to indicate whether a significant change from the natural variability of the metric occurs. •Total number of pages visited: The total number of pages visited, usually equivalent to the total number of user requests/actions, is another measure to determine website traffic. Coupled with the average size, in bytes, of a single page, this indicator can be used to estimate the load of a web service. Additionally, marketers can investigate the total requests for a specific page to measure the success of a new ad placement, for instance. 3.3.1.2 Website Stickiness/Slipperiness According to the philosophy of CRM foundations, discussed on Section 2.1, more importantly than engaging website traffic is to encourage users to spend some time on the website and keep 28
CLM Data Set them interested in it. This concept is usually referred to as stickiness in the web analysis context, or sometimes known as level of engagement. Stickiness is frequently related to the profit of the website, as the likelihood of completing a purchase increases with the time spent on the website, assuming there that the website is optimally design to convey information. However, it should be noted that high values of stickiness is not always an advantage to the website. When users are navigating through content pages, which contain product or service information and descriptions, it is desirable for them to spend more time. Oppositely, for a profile page, such as registration page or shipping information page, the more time/clicks users perform, may indicate that they are not being efficiently guided through the process. There are multiple aways of evaluating stickiness, the following is just a short summary of some variables used within this study: •Bounce Rate: An important measure that falls into this category of stickiness assessment is called bounce rate. The bounce rate refers to the percentage of visitors who come to the website but do not engage and leave the website after a few seconds or only visit a single page. Similarly to other measures, the bounce rate can also be calculated to particular sections or single pages within the web site, this can also be denominated as page bounce. •NPV: The Number of Pages Visited (NPV) is a web site session is a popular measure to indicate depth of visit and stickiness. This metric can be generalized for the entire data set by computing the average number of pages visited (ANPV). ANPV =total number o f pages visited total number o f sessions (3.1) It is also helpful to analyze the histogram that represents the distribution of the NPV. This graphical representation of the data set allows for a further understanding of the underlying distribution that fits the NPV. •ASL: Stickiness can also be calculated for a specified period using the Average Session Length (ASL) (see Section 3.2). This measure can be computed for the entire website, some sections or even specific web pages and, coupled with the NPV, is important to determine how fast users are assimilating information. Similarly to the NPV, the total time spent per session can be represented via an histogram that displays its real distribution. ASL =total time spent on the website total number o f sessions (3.2) Table 3.3 displays the stickiness metric measured for the CLM data set. It should be noticed that, unlike the usual definition, we considered a bounce visit, all that did not requested over five pageviews. The industry value for a standard bounce rate definition, i.e. less than 5 seconds of activity, is 34%. The ASL value is higher than the overall online retail industry value, which is 9.5 minutes [Ret07]. However, this reference does not consider our restriction of 5 pageviews, 29
CLM Data Set Table 3.3: CLM Metrics Measure Value ASL 10.58 minutes Bounce Rate 66.0% at least, per session. Therefore, although useful, these values do not provide foundations for any conclusion. Figure 3.1 (right) allows for a better comprehension of how much time users actually spend in each session. The ASL is the mean of this distribution but, although insightful, does not actually characterizes the data set in the best way due to the distribution’s skewness. The median value, 5.85 minutes, seems to represent a better estimate for a typical session, as it implies that half of the session last less then 5.85 minutes. The same Figure 3.1 (left) depicts the distribution for the number of pages visited per session, where the average is 25.4 pages and the median is 13 pages. The correlation between these distributions is notorious as both have roughly the same shape and skewness. 3.3.1.3 Conversion Conversion analysis tends to be the most business related metric as it directly relates to revenue and profit. A successful conversion in e-commerce occurs when a visit is guided by the website to purchase a product(s). The conversion rate, arguably the most used metric to measure conversion levels, is defined as the percentage of website visits that lead to an online purchase. Several literature has been active on studying the underlying motives of purchasers in order to improve conversion rates [MF04]. If business managers understand the motives behind a purchase, they can deploy efficient strategies to convert more visitors into buyers. It should be noted that conversion definition may be distinct for different types of websites or businesses. Concerning this study, a conversion occurs whenever a user adds an item to his virtual basket. Logically, one session may yield multiple conversions which should match the number of products in the basket. There are numerous measures that evaluate visitor conversion. The following enumeration only contains a short set of indicators that are useful to estimate conversion levels for the CLM data set: •Conversion rate: The most common way of measuring conversion is estimated by simply computing the percentage of sessions, for a specified time frame, containing user purchase. Conversion Rate =number o f sessions that per f orm purchases total number o f sessions (3.3) Typically, conversion rates fluctuate between 0.5%-8.0% depending on the sector, target market or the definition of conversion. Within online retail financial services, for example, 1.0%-2.0% would be typical with 2.0% being very good [Kau07]. 30
CLM Data Set Figure 3.1: Histogram of the number of pages visited (left) and session visit time duration (right) for the CLM data set. All bounce visits were removed. •Number of pages to conversion: The number of pages until conversion is an important measure because it portrays the efficiency of the website’s purchase pipeline. On the other hand, comparing the number of pages to conversion of a particular session with the data set average (ANPC), can be used to assess hurry. ANPC =∑n i=1number o f pages to conversion f or sessioni total number o f sessions with conversions (3.4) •Average conversion time: It is useful to provide a metric using the ratio of conversion time and session time, since this can be used to investigate at what point in a session a user decides to make an order. A desirable metric for an e-commerce website is given by the average amount of time shoppers spend on the website to buy an item online. A shorter conversion time shows a good performance by the website in terms of guiding visitors for e-shopping. •Time duration per conversion: Staying on the website without performing any conversion is something that that e-commerce websites try to minimize as it may result in slowing down the server. For an e-commerce website the time duration can be adjusted by the number of online shoppers. The Time Duration Per Conversion (TDPC) reveals the average amount of time elapsed per conversion. TDPC =total time duration total number o f conversions (3.5) 31
CLM Data Set Table 3.4: CLM Conversion Metrics Measure Value Number of Conversion visits (#) 45 801 Total Number of Conversions (#) 651 771 Average Number of Pages to Conversion (#) 15 15 Average Conversion Time (minutes) 6.51 Time Duration per Conversion (minutes) 6.86 Conversion Rate (%) 10.92 Notwithstanding, if the website is selling complex products/services, a longer visit may display that visitors are interested in obtaining as much information as possible. Usually this behaviour is linked to high priced items, such as laptops, where users tend to analyze product details more carefully. Table 3.4 summarizes the set of conversion metrics applied to the CLM data set. The high conversion rate percentage is related to the definition of conversion. Since the data set refers to a online grocery retailer, it is expected that a regular buyer adds more that one item to the basket. Furthermore, the ratio between the number of conversion visits and the total number of conversions yields an average on 14.23 items added to the basket per conversion visit. The actual percentage of users that completed a purchase, regardless of how many items they bought, is 4.64% which is within industry standards. Comparing the time duration per conversion and the average conversion time, one can conclude that the website is optimized towards conversion, as these values are relatively similar. Conversely, an average conversion time of 6.51 minutes reveals that the platform is not efficiently expediting users towards the products they need [NWW14]. 3.3.2 Trend and Traversal Reports Usually, metrics are only useful once it is possible to compare the data set from different points of view. Following that topic, this section will expand on the kinds of combinations that web analysts use in order to grasp their online audience. 3.3.2.1 Trends and Segmentation Trend reports can provide the analyst with a better perception of any required metric as they show changes over a time period. These trends are usually segmented over different customer archetypes and provide useful insight over potential marketing strategies. If the website has an authentication system, it is crucial to understand whether different kinds of customers behave differently. For an e-commerce online business a basic segmentation involves, for each metric, splitting the user base into buyers and non-buyers. Once these groups are properly identified, it becomes possible to study their differences and deploy congruent marketing campaigns with a better chance of success. 32
CLM Data Set Table 3.5: CLM User Archetypes Identifier Meaning Sessions (%) AU anonymous users 78.75 ANIB users who did not add item(s) to their baskets 10.41 AIB users who added item(s) to their baskets 10.84 Table 3.5 displays the different user archetypes within the CLM data set. During the month of activity from which the CLM data set was collected, a vast majority of sessions were anonymous (78.75%) and the remaining were almost bissected into ANIB users and AIB users. The time-stamp available in the clickstream helps to find the visit-date, the date of a new web session in which the first request of the user’s browser is sent to the server. Visit-date information enables us to derive other temporal session attributes such as whether the session takes place on weekdays, holidays, weekends, or any required period of interest. If the data set contains requests from different countries, by inferring coordinates from the IP field, one has to synchronize different timezones in order to accurately evaluate when requests were made. One of the most important trend metrics relates to the measure of website traffic for different hours of the day or even different days of the week. This would show the peak time of the traffic and it can be exploited to find a suitable time to perform promotional campaigns or maintenance activities. Figure 3.2: Histogram of the number of page requests per day of the week 4(left) and the number of page requests per day from users that logged and eventually purchased items (right). There are several ways to segment the data and study distinct behaviors for each one, nonetheless, most of these segment analysis are inconclusive in terms of differentiation, as the shape of distributions tend to be very similar. 4Day one is equivalent to Sunday and the remaining are ordered naturally 5Including smartphones, tablets and smart TVs. 33
CLM Data Set Figure 3.3: Histogram for the number of page requests per hour (left) and the number of page requests per hour from users accessing the website through a mobile device 5(right). Figure 3.2 (left) displays an histogram of page requests grouped by days of the week. This distribution clearly highlights Sunday as the day with the most amount of page requests with the remaining days following an almost uniform distribution with a slight downward trend. Figure 3.2 (right) also represents the distribution of daily page requests but only focusing the segment of logged users that added item(s) to their baskets. Both plots yield similar distributions. Thus, a web analyst could conclude that there is no added value on targeting AIB users based on the day they are most active because, in essence, they follow the same pattern as the general user basis. Figure 3.3 (left) shows how the number of page requests changes during the course of one day. This histogram reveals that page requests follow a similar distribution of that people tend to be awake. Figure 3.3 (right) depicts the same scenario but only considering the segment of the market represented by AIB users that access the website during Sunday. Although similar, it is noticeable that on the latter plot users are more likely to access the website during prime time while the general the page request distribution (left) shows no privileged time frame. 3.3.2.2 Traversal Reports Apart from general information about a website and its sessions such as traffic, depth, time/date, time duration, depth of a session, etc., clickstream data also provides information about the sequence of web pages a user visits while browsing the website. In fact, after performing a pageview identification over the clickstream data (see Section 2.4.2.7), it is possible to analyze each session and understand how the user navigated through the website. This insight uncovers a new type of analysis that analysts can execute. A common and simple diagnosis is to identify which are the most common landing pages, i.e. the first page that is requested within a session. This analysis could be useful to measure the success of a marketing campaign, for instance. If the company invested in a specific promotion or newsletter it is expected 34
CLM Data Set that the page, where that promotion or newsletter is detailed, yields a higher rank on the landing page ranking. Table 3.6: CLM Landing Page Rank Rank Landing Page Requests (#) Percentage 1 Homepage 314 305 74.37 % 2 Category Page 48 515 11.48 % 3 Product Page 17 272 4.09 % 4 Flyers 11 069 2.69 % 5 Search Results 6 862 1.62 % Table 3.6 depicts the ranking of landing pages according to the pageview identification specified on Table 3.1. This rank reveals that the homepage is indisputably, the most requested first page and the top five landing pages are responsible for 94,25 % of all first page requests. A closer analysis of Table 3.6 may question how can visitors start their session through the Search Results pageview, as this action can only be accomplished once the website is loaded. However, this is a side effect of how we partitioned the session during the preprocessing stage. In fact, if a user loads the website’s homepage and only request for search results 20 minutes later, this user session will be perceived as two separate sessions requesting one pageview each. For further details regarding sessionization see Section 2.4.2.4 Figure 3.4: Collection of box-plots specifying the in-session prevalence for each pageview, taking all CLM data set into account. Notwithstanding, there are other pageview analysis that allow for a better understanding of how the average session is divided into different user interactions. This sort of scrutiny is called 35
CLM Data Set pageview session prevalence and it is fundamental for a real comprehension of user sessions. In essence, pageview session prevalence is the outcome of segmenting a user session pageviews set and compute how much influence of each pageview in the overall session. Figure 3.4 depicts the pageview prevalence for the CLM data set. Since the vast majority of sessions are anonymous users (78.75%) it is natural that actions related to shopping tasks such as adding items to the cart or dealing with shipping/payment have a residual representation. Moreover, anonymous users are naturally linked to exploratory user sessions and logged users tend to display a more focused navigation pattern. This exploratory behavior translates to higher percentages of requests for content, hence the prevalence of pageviews such as Category Menu,Product Page or Filter. Similarly to the pageview session prevalence analysis, where each user session is dissected into different pageviews, resulting on a broad picture of how each pageview impacts the session in terms of number of requests, it is also interesting to know the number of sessions that request each pageview, regardless of how many requests. We call this overview of the data pageview session span and its purpose is to highlight those pageviews whose business significance is relevant but are not requested that often. For instance, search related pageviews in any online store tend to be more requested that payment operations, yet the latter are more significant because they are directly related to the store’s profit. Figure 3.5: Percentage of sessions that requested a particular pageview regardless of the number of requests. Figure 3.5 depicts the pageview session span for the CLM data set. Comparing this figure with the previous pageview session prevalence sufficient to understand its importance. In fact, we can identify some business important pageviews that have a residual session prevalence but, 36
CLM Data Set a considerable amount of sessions are requesting them. The Flyers pageview, a page designed to convey advertising to end consumers, is a perfect illustration of this scenario as it has virtually no significance when analyzing session prevalence but, 19% of user sessions do request such pageview. In sum, both branches of pageview analysis are essential to actually understand how users explore the website. 3.4 Discussion In this chapter, we detailed the process related with the preparation and measurement of the CLM clickstream data set. In addition to the common application of clickstream for website traffic engagement and depth, we used conversion information to report some KPI regarding the profitability. On the other hand, we introduced different kinds of segmentation strategies that allow analysts further means to identify specific user archetypes. Finally, we used the outcome of pageview identification to analyze and comprehend how a user session is split. These measures will be used for different analytic and exploratory purposes throughout the thesis, namely during the feature engineering stage. 37
Predicting Purchasing Engagement Figure 4.2: Collection of plots assessing the quality of fitting a log-normal distribution over the empirical histogram, representing the number of pageview requests. ANIB user sessions, only 36.46% of the requests are focused towards the same end. Again, these statistics support the hypothesis that AIB users’ sessions are more focused towards the same goal, they perform search actions to effectuate the purchase. In order to further evaluate this statement, Figure 4.4 depicts the histogram of the search related activities for both user segments. This back-to-back plot reveals that both user segments have distinct underlying distributions for the same session prevalence attribute. This natural differentiation makes this feature useful for future binary classifiers. Similarly to other features, percentage of session requests related to search activities evolves during the course of the session. 4.3.3.3 Pageview Session Span In order to develop relevant features, we have to delve into the motivations of different types of user sessions. For the AIB users, the process of authentication is normal because it is required in order to complete their purchase. On the other hand, it is not clear why ANIB users perform user authentication if they do not intend to purchase. One hypothesis is that ANIB users log-in, so that they can perform editions to their user profiles, such as, change of delivery address, elaboration 44
Predicting Purchasing Engagement Figure 4.3: Simplified pageview session prevalence for AIB user sessions (top) and for ANIB user sessions (bottom). of a shopping list, etc. As introduced on Section 3.3.2.2, in order to better understand the actions performed by each user segment we performed a pageview session span analysis for each user segment. Figure 4.5 depicts the pageview session span for both user segments. Based on this plot we can confirm that AIB and ANIB users interact differently with the website. Moreover, our initial hypothesis is confirmed as ANIB users are prone to perform tasks that are not related to searching products. From the analysis of this pageview session span plot, we extracted three features based on the empirical difference between the two segments: •User Profile: According to the information present in Figure 4.5 ANIB users are four times more likely to perform actions related to their user profiles. Moreover, 40% of all ANIB user sessions request these types of pageviews. •Flyers: The same pageview session span revealed that 24% of ANIB users visit pages with promotional catalogs, also known as flyer pages, which is close to three times what AIB user sessions request. 45
Predicting Purchasing Engagement Figure 4.4: Histogram of pageview prevalence for search related activities. •Other: We can also confer that 19% of ANIB users request pageviews that are not representative and fall into a category known as Other. Nonetheless, this percentage of requests represents almost five times more than the number of AIB user requests for the same pageview category. These three features assume binary values, in other words, for a specific moment within a session each of the features will indicate whether that user has already performed any of those requests. 4.3.3.4 Average Time Per Page Another clickstream data attribute is related to the average amount of time each user spends per page. Our initial hypothesis considered that AIB users were more focused towards purchasing items, thus spend less time per page. Figure 4.6 depicts a box plot for each user segment, discriminating the amount of time spent per page. Although both distributions are, for the most part, overlapping each other it is noticeable that ANIB uses tend to spend more time in each page than AIB users. For this reason, we considered the use of the average amount of time spent per page as a feature for the learning models. There is no need to scale this attribute as its natural distribution falls under a reasonable domain. 4.3.3.5 Pageview Sequence Likelihood Thus far, every feature has been directly related to session metrics or specific actions. However, we assumed that the sequence of pageview requests can reveal whether a user is about to engage 46
Predicting Purchasing Engagement Figure 4.5: Bar plot specifying the pageview session span for both AIB and ANIB user sessions. or not. In addition, we believe that AIB and ANIB users can be distinguished by their clickstream sequence. Therefore, we propose a measure of session similarity by using Markov for discrimination introduced by [DEKM98] and used, for instance, by [MVdPCeC12] for identifying churners and non-churners. In a Markov model, the probability distribution of the next state depends only on the current state and not on the sequence of events that preceded it. Thus, each Markov process can be represented by means of a transition matrix. In the case of a process with Npossible states, the corresponding transition matrix has a dimension of NxNstates. Each element of the matrix, pi j, represents the probability of the system evolving from a state i, in period t, to another state j, in period t + 1. In this research each state of the Markov process represents one pageview within a particular user session. In order to discriminate AIB and ANIB users using a Markov model, we assume that each user segment follow different Markov processes. Therefore, we build, for each population, a different transition matrix that reflects its specific pageview sequences. Following [DEKM98], we use these transition matrices to compute the log-odds ratio between the odds of of observing a sequence xgiven it originates from the AIB users’ population and the odds of observing sequence xgiven it belongs to the ANIB users’ population: S(x) = log P(x|AIB user) P(x|ANIB user)(4.1) The outcome of S(x) allows the affinity of a visitor to be measured with respect to AIB and ANIB user sessions, by means of their specific pageview sequence. A positive ratio indicates that the visitor is not likely to engage with purchasing while a negative ratio means the opposite. S(R1→R2→(···)→Rn) = n ∑ k=1 log P(Rk|AIB user) P(Rk|ANIB user)(4.2) In order to compute the log-odds ratio for a sequence of pageviews, and not only a single 47
Predicting Purchasing Engagement Figure 4.6: Box plots for the number of minutes spent per page for AIB and ANIB users. request, we use the sum operation to join the log-odds of every transaction, as depicted on Equation 4.2. If this value is positive then that sequence of requests is more likely to come from a ANIB user while a negative ratio means the opposite. In this study, we use the value of S(x) as a feature for our models. We do not perform any domain scaling as the natural distribution of S(x) is similar to the remaining features. 4.4 Learning Models In order to test the predictive power of selected features, we submit our data set into two different learning algorithms: logistic regression and random forests. This section’s purpose is twofold, first we introduce each technique and review some related literature, then we specify the learning process, relevant parameters and evaluation metrics. 4.4.1 Logistic Regression Logistic regression is a well-known technique, developed by [Cox58], which can be used for classification purposes. This is a form of regression which is commonly used when the dependent variable is binary, thus it is also known as binary logistic model. Logistic regression measures the relationship between the categorical dependent variable and one or more independent variables by 48
Predicting Purchasing Engagement estimating probabilities. The independent variables can be of any type [Agr96], but are usually continuous. The relationship between dependent and independent variables is not assumed to be linear. Instead, this technique assumes that the independent variables are linearly related to the logit of the dependent variables. Another upside of logistic regression is the fact that it does not require normally distributed variables. This method has various applications spanning different fields such as medicine, politics or economics [BRD+00,HLC+84,CKT97]. It has become a standard classification method as it is easy to use and provides quick and robust results [VdPB05]. 4.4.2 Random Forests Random forests are an ensemble learning method for classification or regression, that operate by constructing a multitude of decision trees at training time and outputting the class that is the mode of the classes (for classification) or mean prediction (for regression) of the individual trees. Unlike single decision trees, which are likely to suffer from high variance or high bias (depending on how they are tuned), random forests use several trees that vote upon which class is a better fit to a certain collection of features. This method was firstly introduced by [Bre01] as a method that combines the concept of bagging predictors [Bre96] and the random selection of features [Ho95]. Similarly to logistic regression, this method is widely used in academic literature. Random forests have very few parameters to tune and can be used quite efficiently with default parameter settings. This ease of use coupled with a strong predicting power, makes random forests a solid "off the shelf" choice. This approach requires the definition of two parameters, the number of trees to be used and the number of variables to be randomly selected from the set of independent variables. For this study we will follow the recommendations reported by [Bre01] and consider a significant number of trees (i.e. 1 000 trees) and the truncated square root of the number of variables as the number of independent variables selected (i.e. 3variables). 4.4.3 Evaluation Criteria In order to measure the performance of each of the proposed prediction models, we compute the well known Receiver Operating Characteristic curve (ROC) and we analyze the Area Under Curve (AUC) as described in [HM82]. The AUC measure is based on comparisons between the observed class and the predicted class. The class is predicted by considering all threshold levels for the predicted values. An AUC close to 1.0 means that the model is a perfect classifier, while an AUC close to 0.5 suggests poor prediction capability as a random classifier is also expected to score 0.5 [HM82]. Moreover, we use the precision/recall curve in order to understand how this trade-off varies. In binary classification, precision is the fraction of retrieved instances that are relevant, while recall is the fraction of relevant instances that are retrieves. In simple terms, high precision means that the predictor returned substantially more relevant results than irrelevant, while high recall means that the predictor returned most of the relevant results. Precision and 49
Predicting Purchasing Engagement recall are negatively correlated thus, the exact balance between these measures depends on the problem domain. Finally, we also use a simple yet informative accuracy curve as an evaluation metric. The accuracy is determined after the prediction models classify each case according to a certain threshold, i.e. 0.5 in this case. Then, accuracy is defined as the ratio between the number of correctly classified cases and the total number of cases to be classified. 4.4.4 Results Due to memory constraints we had to limit the training data set to 50 000 entries out of the original 860 838 samples. In order to validate each model, we applied a validation strategy where we sampled 10 unique samples of 50 000 entries out of the whole training data set. Then, for each sample we fed our models with 80% of the data during the training phase and tested with the remaining 20%. This process of validation was carried out while we were developing and testing different features, so that, we avoided model overfitting. Figure 4.7 presents the performance results of the two models using both logistic regression and random forests based on the predictions for the test set. Figure 4.7 depicts an agnostic evaluation of the prediction models in regards to the cut-off value as this is dependent on the marketeers’ intentions. For example, depending on the specific marketing campaign that would be deployed considering the model’s predictions, one could value precision instead of accuracy. From the analysis of Figure 4.7 we can conclude that anonymous purchase engagement prediction in this context is promising. The AUC values are high both when logistic regression and random forests are used. The results obtained also show that random forests outperforms logistic regression in all performance metrics. Moreover, the AUC value for random forests is close to 5% higher than the logistic regression’s AUC. Both prediction models were implemented using R programming language libraries. 4.5 Discussion In this chapter we used the logistic regression and the random forests models to describe the association between general clickstream information concerning visits and whether a visitor will engage in online purchasing behaviour during his visit to the website. This model provides a new tool for an e-commerce web analyst that helps to infer visitors’ behavior, and as a result, to improve online conversion rates. The obtained results reveal that prediction models are effective in this domain. Unfortunately, the CLM data set does not contain the detailed clickstream variables and demographics that would increase the set of features and add more information to the prediction models. Nonetheless, the process of feature engineering revealed to be fruitful as, for instance, the random forest model is able to reach accuracy levels of 77%. On the other hand, these results show that anonymous clickstream data from online grocery retailers is enough to develop a solid model for purchase engagement prediction. 50
Predicting Purchasing Engagement Figure 4.7: Collection of plots for different performance metrics for predicting purchasing engagement. The top three plots refer to the evaluation of the logistic regression model, while the remaining bottom three relate to the evaluation of the random forest model. 51
Predicting Purchasing Engagement From the digital marketing point of view, these results are very encouraging as new methods of targeting customers could be derived from this solution. This research is also proof, at least for the company behind the CLM data set, that customers are predictable in terms of interactions with their website. In a web-focused marketing solution, this predictability might help to keep the customer in the site before leaving to find another competitor. There are alternative approaches including neural networks, deep learning, categorical principal analysis and further decision tree-based methods which can be used to find the best approach for predicting customer engagement. Choosing the best approach, which is extensively discussed in the literature, would go beyond the scope of this thesis, and is not our aim. The logistic regression was chosen here for comparison reasons as most of our focus was dedicated to the random forests model. The predictor developed with random forests is solid and would be suitable for an extension of features if the CLM data set was also extended to include registered user information. 52
Chapter 5 Predicting Purchasing Likelihood Similarly to Chapter 4, the present Chapter introduces two learning models to determine users’ real-time purchasing likelihood. Following the outlined strategy, a series of features are derived from general anonymous in-session clickstream data. To our knowledge this is the first time that literature approaches real-time purchasing likelihood, i.e. being able to anticipate when clients will add items to their virtual baskets. Thus, the author considers the present chapter this thesis’ most relevant contribution. 5.1 Introduction On Chapter 4it was introduced the problem of purchase engagement, i.e. given the e-grocery domain, predicting whether online visitors will add the first item to their basket and become an AIB user. Past literature approached purchase engagement prediction and purchase likelihood as two faces of the same coin. As previously mentioned (see Section 2.1.2), the nature of online grocery retailers differs from the most common e-commerce businesses. Moreover, the common e-commerce businesses focus on specific product segments such as clothing or electronics, i.e. not fast-moving consumer goods (FMCG). On the other hand, e-groceries’ clients tend to mimic the same behaviour of a physical store and buy several items at once, which is contrary to casual e-commerce businesses. [Syn14] reports that, for instance, in England users tend to buy, on average, items from 6.6 distinct product categories within a single session. This distinction between e-groceries and other e-commerce businesses sets the foundation for the motivation behind the development of a model that predicts customer purchase likelihood. Predicting purchase likelihood corresponds to the continuous prediction of basket additions. In other words, this model automatically follows user sessions, as they keep requesting new pageviews, and for each request it yields the likelihood of adding and item to the virtual basket, on the next request. Therefore, the most important outcome of this model is the power to dynamically anticipate clients’ purchases. 53
Predicting Purchasing Likelihood procedure of engagement prediction (see Section 4.3.3.5) to obtain a single factor that, for that specific window, highlights one behavioural pattern instead of the other. 5.4 Learning Models In order to test the predictive power of selected features, we submit out data set into two different learning models: logistic regression and random forests. In this section we specify the learning process, relevant parameters and evaluation metrics. 5.4.1 Techniques and Evaluation Criteria The prediction models that are used to estimate purchase likelihood are the same as the ones used for predicting purchasing engagement (see Section 4.4.1 and Section 4.4.2). Having no tuning parameters, the binary logistic model is simply fed with a new data set and its results are analyzed. On the other hand, random forests have two tuning parameters regarding the number of trees and the number of independent variables selected. For this study we follow the recommendations reported by [Bre01] and consider a large ensemble of trees (i.e. 1000 trees) and the truncated square root of the number of variables as the number as independent variables selected (i.e. 3 variables). Regarding model evaluation criteria, we resort to the same four indicators that also measure the performance of purchasing engagement predictions, namely, the ROC curve and respective AUC, the precision/recall curve and the accuracy curve (see Section 4.4.3). Due to the imbalanced nature of the data set, we also evaluate each model with a measure that specifies the accuracy for each class separately. The class accuracy rate is a combined plot of the true positive rate (basket addition) and the true negative rate (normal request). In essence, this analysis depicts, for each possible class, the ratio between prediction values and the actual record’s class label. 5.4.2 Results Due to memory constraints, we had to limit the training data set to 50 000 entries out of the 3 831 712 available. In order to avoid overfitting, we applied a validation strategy where we sample 10 unique samples of 50 000 entries out of the whole training data set. Then, for each sample, we feed our models with 80% of the data during the training phase and test with the remaining 20%. Figure 5.3 depicts the performance measures of both models, logistic regression and random forests, after these were trained with the imbalanced data set and tested with the reserved test set. These measurements are agnostic regarding the cut-off value that determines the frontier between entries that result in posterior basket addition and entries that do not. Moreover, the best cut-off value depends on the marketeers’ intentions which affect the choice of precision/recall trade-off. From the analysis of Figure 5.3 we can conclude that anonymous purchase likelihood prediction in this context is promising. The AUC values are high both when logistic regression and random forests are used. The results obtained also show that random forests outperforms 60
Predicting Purchasing Likelihood Figure 5.3: Collection of plots for different performance metrics for predicting purchase likelihood with an imbalanced data set. The top four plots refer to the evaluation of the logistic regression model, while the remaining bottom four relate to the evaluation of the random forest model. 61
Predicting Purchasing Likelihood logistic regression in all performance metrics. Moreover, the AUC value for random forests is close to 2.5% higher than the logistic regression’s AUC. On the other hand, the accuracy rate for the class that corresponds to cases labeled with basket addition (positive values), is low for logistic regression and random forests. For example, a cutoff value of 0.5, yields an accuracy under 40%, for both learning models. This could be problematic if web analysts which to target as many basket additions as possible. Figure 5.4 depicts the performance measures of both models, logistic regression and random forests, after these were trained with the balanced data set and tested with the reserved test set. Comparing the results between balanced and imbalanced data sets, we can denote a slight decrease of overall prediction power when the models are fed with a balanced dataset. Moreover, this statement is supported by lower AUC values and lower overall accuracy levels. However, the individual class accuracy plot reveals that modelling with a balanced data set yields different prediction models. Unlike the imbalanced data set’s models, we can observe that accuracy levels for the class that corresponds to cases labeled with basket addition (positive values), is much higher than before. Notwithstanding, this increase of accuracy for the positive class is balanced with a decrease of accuracy for the negative class. Both prediction models were implemented using R programming language libraries. 5.5 Discussion In this chapter we used the logistic regression and the random forests models to describe the association between general clickstream information of AIB users and whether they will add an item to the basket, given their past requests. This model provides a new tool for an e-commerce web analyst that helps to infer users’ behavior, and as a result, to improve online conversion rates. Moreover, this tool displays a dynamic behavior as it follows users’ sessions and continuously evaluating purchasing intentions. The obtained results reveal that prediction models are effective in this domain. Unfortunately, the CLM data set does not contain the detailed clickstream variables and demographics that would increase the set of features and add more information to the prediction models. Moreover, we believe that any information regarding the categories and/or products that users search for could only increase prediction rates as, for instance, buying products from the milk section is totally different than buying products from the light-bulbs section. Nonetheless, the process of feature engineering revealed to be fruitful as, for instance, the random forest model is able to reach accuracy levels of 86%. Moreover, these results show that anonymous clickstream data from online grocery retailers is enough to develop a solid model for purchase engagement prediction. From the digital marketing point of view, these results are very encouraging as new methods of targeting customers could be derived from this solution. In fact, we believe that the learning models developed within this chapter, from a business standpoint, are even more interesting than the models related with customer engagement prediction. Being able to anticipate basket additions opens new possibilities for marketing maneuvers. For example, if the model predicted a basket 62
Predicting Purchasing Likelihood Figure 5.4: Collection of plots for different performance metrics for predicting purchase likelihood with a balanced data set. The top four plots refer to the evaluation of the logistic regression model, while the remaining bottom four relate to the evaluation of the random forest model. 63
Predicting Purchasing Likelihood addition, the web server could adjust the response to the user so that more profitable products were given more emphasis. Furthermore, the web analyst is able to deploy different training sets, balanced or imbalanced, according to the objectives of his campaign. As concluded, imbalanced data sets yield a model with better model to target requests that do not entail basket additions, while balanced data sets favor the identification of requests that anticipate purchases. There are alternative approaches including neural networks, deep learning, categorical principal analysis and further decision tree-based methods which can be used to find the best approach for predicting customer engagement. Choosing the best approach, which is extensively discussed in the literature, would go beyond the scope of this thesis, and is not our aim [K+95,VV98]. The logistic regression was chosen here for comparison reasons as most of our focus was dedicated to the random forests model. The prediction model developed with random forests is solid, and would be suitable for an extension of features if the CLM data set was also extended to include registered user information. 64
Chapter 6 Conclusions This thesis originated from practical business questions related with a major European food retailer with online presence. The main purpose was to investigate statistical approaches on clickstream data, as the aggregated sequence of pageview requests executed by a particular user, and other user navigation features, can provide insight into their intentions, specifically with respect to purchase engagement and real-time purchase likelihood prediction. 6.1 Clickstream Data Preprocessing We worked upon one major data source that was retrieved from the retailer’s server logs. The log file contains information about every single click made by a user on a web browser while surfing the Internet, corresponding to an HTTP request sent to the website’s server. The log data that we used for this study, referred to as the CLM data set, was not complete, in the sense that some relevant clickstream attributes were missing. One limitation for the analysis of the CLM data set is the lack of knowledge regarding page references. When a user returns to a page that has already been visited (downloaded from the server) during the same session, the second access to that page will, in most cases, result in viewing the previously downloaded version of the page without sending a request to the server, due to client-side browser caching. This problem could be solved by knowledge of the website’s structure. However, if the website is designed in such way that pageview requests do not have to follow a rigid sequence, i.e. one can request any pageview from every page of the website, structural website knowledge is not enough. Indeed, the website behind the CLM data set displays this dynamic behavior thus making it impossible to track browser caching events. Throughout this study we assumed that sessions could not cache previous requests. There is no guarantee that this assumption is true but we believe that it does not impact the learning models. The original CLM data set, as well as general web log files, do not consist of well structured data and cannot be used directly for analytical purposes. Making clean clickstream data that can 65
Conclusions provide reliable information about web browswing behavior requires a good understanding of the domain. For this reason we dedicate a relevant part of this research to explain every stage and technique related to data preprocessing (see Section 2.4.2). For example, we highlight the need to remove redundant records of log file data when a user requests a web page containing media files, as the request results in several records/lines in the web log file that represent just one page request. On the other hand, we also removed records in the web log files made by bots, as those lines do not reflect human navigational behavior. For analytical purposes, we excluded sessions that performed five or less pageview requests, also known as bounce visits. These short sessions were not taken into account because, given the main purpose of studying purchasing intentions, the website’s structure does not allow purchases with such few requests. A major part of data preprocessing is related with sessionization, i.e. determining when the user starts and ends a visit to the website. This procedure is not exact because the user can request one or more pageviews and leave the browser open. If the user then returns to the website through the open page in the browser, the sessionization strategy should decide whether to consider that activity a continuation of the previous session or a new session. We used a web sessionization rule of splitting the session into two when the time between two clicks is more than 10 minutes. Regarding user identification, another preprocessing stage, we assume that each session is performed be a new user even though the user might have originated previous sessions. This decision relates to the nature of the website from which the CLM data set is from, that is, an online grocery retailer, where, we believe, people tend to purchase items on a monthly basis, the exact period comprehended in the CLM data set. Another, very important, preprocessing stage relates to the identification of pageviews, i.e. label each possible request with a tangible action. During pageview identification we had to use specific company knowledge, nevertheless, we believe that our models remain agnostic, specially because we tagged pageviews as generally as possible. Arguably the major drawback of the CLM data set, is the void of information regarding registered user data. This type of data relates to users that are registered in the company’s databases, that store privileged user information such as demographics and historical purchases. Having access to this sensitive information would allow for the development of more sophisticated features capable of a better characterization of the data set. On the other hand, the lack of information regarding registered user data makes our work more prone to generalization, namely considering other e-groceries, because we are working with as little company knowledge as possible. 6.2 Clickstream Data Exploratory Analysis Throughout this thesis we have shown how clickstream data is able to provide an insight into visitor’s behavior. In order to craft relevant features we reviewed some important web metrics and statistical reporting using clickstream data. Depending on the goals of the analysis, this data could be transformed and aggregated or segmented at different levels of abstraction to provide useful 66
Conclusions metrics. For instance, these metrics could be reported at the level of website or user session. In Chapter 3we illustrated some of these metrics and how valuable is this analysis. One of the most important analysis we performed, regards the pageview segmentation of user sessions with the introduction of pageview session prevalence and pageview session span reports. Both these reports delve into user sessions and reveal how different user archetypes spread the total amount of requests into each identified pageview. Using these methods, we were able to highlight specific pageviews that are able to split different user archetypes. For example, we discovered that ANIB users are four times more likely to perform actions related to their personal profiles than AIB users. Moreover, this particular distinction was identified by comparing ANIB and AIB pageview session spans. We also performed an exploratory analysis on depth-of-visit metrics, by the number of pages requested and the number of products added to the basket. These metrics revealed to be very important because they are the main source of information for the learning models to perceive the stage of the session. Additionally, we explored other features that provided new information regarding the sequence of pageviews. Unlike other features that focus on session metrics or specific actions, we also extracted information from the sequence of requests that each session performs. Therefore, we applied a measure of pageview sequence similarity by using Markov for discrimination. In essence, this technique allows the quantification of the degree of similarity between a particular session and two behavioural patterns. We apply this concept, for instance, when we want to know whether a certain sequence of requests is more likely to belong to an AIB or ANIB user. Overall, we need to stress that exploratory analysis was a fundamental part of this research as we dedicated most of our time into understanding, segmenting and plotting the CLM data set. If new sources of data were to become available, we would go back to this exploratory phase and mine the most relevant and descriptive features using the same methods. 6.3 Predicting Purchasing Engagement One of the main objectives of this thesis was to create a binary classification model capable of predicting visitor purchasing engagement. We used the logistic regression and the random forests models to describe the association between general clickstream information concerning visits and whether a visitor will engage in online purchasing behavior during his visit to the website. In practice, we concluded that purchasing engagement prediction is the same as predicting whether visitors will become AIB or ANIB users. These models were fed with features mined from the exploratory analysis. Both learning models allows us to obtain conditional probability estimates of a first basket addition. This way, the web owner will be able to identify high potential visitors in terms of conversion tendency, and generate leads for suitable targeting actions. In a web-focused marketing solution, the targeting action might help retaining the customer and avoid defection to competitors. 67
Conclusions In order to assess each classification model prediction capabilities, we used three evaluation methods that are agnostic in regards to the cutoff value, i.e. the number that defines the border between AIB and ANIB users: the ROC curve (with its respective AUC value), a precision/recall curve and the model’s accuracy curve. The results show that, for the e-grocery domain, prediction models are very effective. In fact, the best model, random forests, was able to output an AUC value of 0.8415 and accuracy levels above 70% for most of the possible cutoff values. The models can be improved by considering more predictors, for example, any knowledge about the web pages viewed during the session. As previously stated (see Section 6.2), using registered users who login to the website, depending on the requested information in the registration forms, may increase the predictive performance of the models. This includes demographical (gender, age, occupation, etc.) and historical (past purchases, average spending, average session time, etc) attributes. 6.4 Predicting Purchase Likelihood The other main objective of this thesis, arguably the most relevant, was to create a binary classification model capable of continuously predicting when AIB users are going to add an item to their virtual basket. We used the logistic regression and the random forests models to describe the association between general clickstream information concerning AIB users’ sessions and whether a they will perform any basket addition. From a business standpoint, we believe these models are extremely relevant because they predict events that are directly related to revenue. Moreover, the web owner is able to anticipate product basket additions and act according to users’ intentions. In order to stimulate customer profitability the web owner could deploy measures that, for example, would highlight products with higher margins, in case the models reveal a high probability of basket addition. Unlike the data set that fed the model of purchasing engagement prediction, the data set containing all the requests from AIB users (the segment that is mine-able for purchase likelihood prediction) is imbalanced. In fact, only 15% of the requests correspond to basket additions. With this in mind, we developed two sampling strategies to produce different training sets for our models: one that sampled records and kept the original class balance, and another where we balanced the presence of both classes. These strategies yielded distinct results. Training with a sample of the original data set yielded models with better AUC (0.8867 with random forests) and overall accuracy levels (above 80% with random forests) but, compared to the models trained with a balanced data set, they have a much lower individual positive class accuracy (adding an item to the basket). Determining which is the best model depends on the web owner’s objectives. For example, if the objective is to target as most basket additions as possible then, the best model is the one trained with a balanced data set. If, on the other hand, the objective is to target basket additions with as few errors as possible then, the model trained with the original data set is better suited. Similarly to the models for predicting purchasing engagement, these binary classifiers could be improved if more predictors, from additional sources of information, were introduced. 68
Conclusions 6.5 Future Research As we have been stating throughout this thesis, obtaining valuable data is the most important aspect in order to build a solid prediction model. Moreover, regarding online retailing clickstream prediction, we believe that a complete data set, with privileged attributes from registered users, opens a whole new dimension of insightful features that could improve prediction quality. We believe that, in a near future, companies ought to explore this line of thought if they want to remain competitive. On the other hand, future work could be developed in applying new learning algorithms to fit clickstream data, namely, by introducing other models such as neural networks, support vector machines, genetic algorithms, etc. Along side these models, further evaluation criteria and even model selection methods could be studied in order to better understand which options is more appropriate for each occasion. Finally, we wish to stress that, given that the field of clickstream data research is still in its infancy, much research still needs to be done. With the insurgence of new and faster technology, the concept of big data is very hot at the moment, specially because companies can, like never before, translate customer data into higher revenue. This study was the first, to our knowledge, to investigate online customer’s intentions within the online food retailing domain. We believe that future literature will approach this topic again and improve our results by introducing new methodologies for preprocessing, feature engineering and modelling. 69
REFERENCES [SKR01] J Ben Schafer, Joseph A Konstan, and John Riedl. E-commerce recommendation applications. In Applications of Data Mining to Electronic Commerce, pages 115–153. Springer, 2001. [SS08] Judy E Scott and Carlton H Scott. Online grocery order fulfillment tradeoffs. In Hawaii International Conference on System Sciences, Proceedings of the 41st Annual, pages 90–90. IEEE, 2008. [SS13] A. Surya and D. K. Sharma. An approach for web page ordering using user session. 2013 Ieee Conference on Information and Communication Technologies (Ict 2013), pages 1009–1013, 2013. [Syn14] SyndicatePlus. The online grocery shopper. Technical report, SyndicatePlus, September 2014 2014. [TSK+06] Pang-Ning Tan, Michael Steinbach, Vipin Kumar, et al. Introduction to data mining, volume 1. Pearson Addison Wesley Boston, 2006. [TSSW10] Andranik Tumasjan, Timm Oliver Sprenger, Philipp G Sandner, and Isabell M Welpe. Predicting elections with twitter: What 140 characters reveal about political sentiment. ICWSM, 10:178–185, 2010. [VdPB05] Dirk Van den Poel and Wouter Buckinx. Predicting online-purchasing behaviour. European Journal of Operational Research, 166(2):557–575, 2005. [Ver12] Ruud Verheijden. Predicting purchasing behavior throughout the clickstream. Master’s thesis, Technical University of Eindhoven, 2012. [VMKR13] Chintan R Varnagar, Nirali N Madhak, Trupti M Kodinariya, and Jayesh N Rathod. Web usage mining: A review on process, methods and techniques. In Information Communication and Embedded Systems (ICICES), 2013 International Conference on, pages 40–46. IEEE, 2013. [VV98] Vladimir Naumovich Vapnik and Vlamimir Vapnik. Statistical learning theory, volume 1. Wiley New York, 1998. 76