scieee AI-readable full text Open interactive document viewer

Data analysis and recommender system architecture for e-commerce platforms

Cunha, Gil Fernando Ferreira da

Abstract

E-commerce is constantly expanding, leading to greater market competitiveness. The number of online platforms offering products or services is increasing; so there is a growing need for companies to stand out from the competition, which leads to the application of various marketing strategies. However, not all are adequate and mismanagement, as well as a bad investment of these strategies, may prejudice companies. Hence the implementation of recommendation systems in e-commerce platforms, as a safe and economical strategy. By investing in a good recommendation mechanism, one can provide better user experience, taking his interests into account. As a result, more traffic on the platforms is ensured, which may result in a higher sales rate and, consequently, a higher number of revenues. However, to develop a recommendation system, the first step must consist in obtaining information about the sales platform, where data about its users and products/services form the basis of recom mendations. But not all information is useful, which can influence the accuracy of the forecasting models used by the system to produce results. Following this perspective, a data analysis methodology is proposed, as well as an architecture of a recommendation system, which allows to extract and treat relevant data, in order to integrate a recommendation engine for most e-commerce platforms.

Full text

Universidade do Minho Escola de Engenharia Gil Fernando Ferreira da Cunha January 2021 UMinho | 2021 Gil Fernando Ferreira da Cunha Data Analysis and Recommender System Architecture for E-Commerce platforms System Architecture for Data Analysis and Recommender E-Commerce platforms Universidade do Minho Escola de Engenharia January 2021 Gil Fernando Ferreira da Cunha System Architecture for Data Analysis and Recommender E-Commerce Platforms Professor Dr. Hugo Daniel Abreu Peixoto Universidade do Minho Escola de Engenharia January 2021 Gil Fernando Ferreira da Cunha System Architecture for Data Analysis and Recommender Master's Dissertation Integrated Master's in Informatics Engineering Dissertation supervised by E-Commerce Platforms Professor Dr. José Manuel Ferreira Machado COPYRIGHTS AND CONDITIONS OF USE BY THIRD PARTIES This is an academic work that can be used by third parties as long as the internationally accepted rules and good practices are respected, with regard to copyright and related rights. Thus, this work can be used under the terms provided for in the license below. If the user needs permission to be able to use the work under conditions not provided for in the indicated license, he must contact the author, through the RepositóriUMof the University of Minho. Attribution-NonCommercial CC BY-NC https://creativecommons.org/licenses/by-nc/4.0/ v ”Knowledgeispower.” - Francis Bacon (1597) ”Knowledgeispower. Informationispower. Thesecretingorhoardingofknowledgeorinformation maybeanactoftyrannycamouflagedashumility.” - Robin Morgan “Informationispoweronlyifyoucantakeactionwithit. Then,andonlythen,doesitrepresentknowledge and,consequently,power.” - Daniel Burrus ”Knowledgeispower. Informationisliberating. Educationisthepremiseofprogress, ineverysociety,ineveryfamily” - Kofi Annan vii 5.4 Summary.................................... 75 6 Case Studies / Experiments . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 76 6.1 Experimentsetup................................ 76 6.1.1 DataContract ............................. 82 6.1.2 Kibana Data Analysis . . . . . . . . . . . . . . . . . . . . . . . . . . 82 6.2 Results..................................... 85 6.2.1 Performance Tests . . . . . . . . . . . . . . . . . . . . . . . . . . . 86 6.2.2 Business Intelligence Dashboards . . . . . . . . . . . . . . . . . . . . 93 6.3 Discussion ................................... 97 7 Conclusion ......................................100 References ........................................103 Appendices........................................109 LIST OF FIGURES 2.1 Monolithic application architecture template . . . . . . . . . . . . . . . . . . . . . 9 2.2 Microservices application architecture template . . . . . . . . . . . . . . . . . . . 10 2.3 REST vs SOAP: Web search interest rate comparison between REST and SOAP, from 2004 until 2019, worldwide . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 11 2.4 REST API Model example diagram . . . . . . . . . . . . . . . . . . . . . . . . . . 12 2.5 Collaborative and Content-based Filtering examples . . . . . . . . . . . . . . . . . 13 2.6 Applications covered by Amazon Web Services (1) products. . . . . . . . . . . . . . 17 2.7 Netflix recommender system overview - based on Netflix Tech Blog post (2) . . . . . . 19 2.8 Strands Retail product recommendation system workflow overview . . . . . . . . . . 22 2.9 Commerce Cloud Einstein product recommendations process . . . . . . . . . . . . 23 2.10 Amazon Personalize ”How it works” diagram . . . . . . . . . . . . . . . . . . . . 24 2.11 Amazon Personalize high level architecture . . . . . . . . . . . . . . . . . . . . . . 25 3.1 Recommender System Architecture diagram . . . . . . . . . . . . . . . . . . . . 35 4.1 Node.js Application Modular Structure Design, according to Separation of Concerns (3) ......................................... 38 4.2 Recommender System Entities . . . . . . . . . . . . . . . . . . . . . . . . . . . 41 4.3 MongoDB recommendation documents structure . . . . . . . . . . . . . . . . . . 44 4.4 User Interaction Activity Diagram . . . . . . . . . . . . . . . . . . . . . . . . . . 47 xiv 4.5 Recommender System Activity Diagram . . . . . . . . . . . . . . . . . . . . . . . 47 4.6 Communication process between Recommender’s API an Engine - option 1 . . . . . . 51 4.7 Communication process between Recommender’s API an Engine - option 2 . . . . . . 51 4.8 Communication process between BI App and Recommender System . . . . . . . . . 53 4.9 Beevos’ Business Intelligence Application File Structure . . . . . . . . . . . . . . . 54 4.10 Homepage recommendations vitrine generic template. In this example, four recommended products are displayed to the user: two from ’Clothes’category (shirt and pants) and two from ’Drinks’category (iced tea and smoothie). Products are ordered by score, with the product on the left having the highest score, i.e., is more likely to be bought by the user. . . . . . . . . . . . . . . . . . . . . . . . . . . . 60 4.11 Product details page recommendations vitrine generic template. In this example, the user selected a shirt, thus be presented products similar to it in the recommendation vitrine, where the product on the left is the most similar to the selected shirt (orderedbyscore). ................................. 61 4.12 Side cart recommendations vitrine generic template. In this example, the user added a shirt and pants to the shopping cart, so the recommendation vitrine displays some products which are commonly bought together with the cart’s current content. ...................................... 61 4.13 Product listing page recommendations vitrine generic template. In this example, the user browses for products within the ”Drinks”category and orders them with the ”Recommended”option. Ergo, products related to ”Drinks” are listed, ordered by recommendation score, i.e., the products that would appeal the most according to the user’s profile are shown first. . . . . . . . . . . . . . . . . . . . . . . . . . 62 6.1 Deeply Back-Office - RS user association to tenant . . . . . . . . . . . . . . . . . . 77 6.2 Deeply Back-Office - Recommender engine configurations . . . . . . . . . . . . . . 78 6.3 Deeply Back-Office - Attribute selection and Recommender Population . . . . . . . . 78 6.4 Deeply MongoDB recommendation documents examples . . . . . . . . . . . . . . 79 6.5 Deeply Homepage - Popularity recommendations example . . . . . . . . . . . . . . 80 6.6 Deeply Product Details - Similar-products recommendations example . . . . . . . . . 80 6.7 Deeply Side-Cart - Complementary-products suggestions example . . . . . . . . . . 81 6.8 Deeply Product Listing - Hybrid recommendations example . . . . . . . . . . . . . 81 6.9 Execution of load tests: Workload, System Under Test (SUT) and Metrics. . . . . . . 86 6.10 Locustloadteststructure. ............................. 88 6.11 Number of Users over time . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 91 6.12 Rate of total requests per secod over time . . . . . . . . . . . . . . . . . . . . . . 92 6.13 Response time value over time . . . . . . . . . . . . . . . . . . . . . . . . . . . 92 6.14 Total Orders by Country - Map . . . . . . . . . . . . . . . . . . . . . . . . . . . 93 6.15 Total Orders by Country - Bars graph . . . . . . . . . . . . . . . . . . . . . . . . 94 xv 6.16 Average spend by Country . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 94 6.17 Average quantity and spend per order . . . . . . . . . . . . . . . . . . . . . . . . 95 6.18 Clientgenders ................................... 95 6.19 PromotionTracking................................. 96 A.1 NodeJS Application Structure . . . . . . . . . . . . . . . . . . . . . . . . . . . . 110 A.2 API Documentation with Swagger - part 1 . . . . . . . . . . . . . . . . . . . . . . 111 A.3 API Documentation with Swagger - part 2 . . . . . . . . . . . . . . . . . . . . . . 112 A.4 User login page of the Recommender System . . . . . . . . . . . . . . . . . . . . 113 A.5 User registration page of the Recommender System . . . . . . . . . . . . . . . . . 113 A.6 Beevo Business Intelligence Data Contract - part 1 . . . . . . . . . . . . . . . . . . 114 A.7 Beevo Business Intelligence Data Contract - part 2 . . . . . . . . . . . . . . . . . . 115 A.8 Kibana’s data analysis on Clients data - part 1 . . . . . . . . . . . . . . . . . . . . 116 A.9 Kibana’s data analysis on Clients data - part 2 . . . . . . . . . . . . . . . . . . . . 117 A.10 Kibana’s data analysis on Products data - part 1 . . . . . . . . . . . . . . . . . . . 118 A.11 Kibana’s data analysis on Products data - part 2 . . . . . . . . . . . . . . . . . . . 119 A.12 Kibana’s data analysis on Order-items data - part 1 . . . . . . . . . . . . . . . . . . 120 A.13 Kibana’s data analysis on Order-items data - part 2 . . . . . . . . . . . . . . . . . . 121 A.14 Kibana’s Business Intelligence Dashboards - part 1 . . . . . . . . . . . . . . . . . 122 A.15 Kibana’s Business Intelligence Dashboards - part 2 . . . . . . . . . . . . . . . . . 123 LIST OF TABLES 4.1 Version table of docker images used in RS architecture . . . . . . . . . . . . . . . . 40 4.2 Server host machine hardware specifications. Note that the server is hosted in a virtual machine, emulated using QEMU(4). QEMU allows to run operating systems for any machine, on any supported architecture, with near native performance. . . . 40 4.3 Beevo’s business intelligence application configurations . . . . . . . . . . . . . . . 56 4.4 RS Access Control List: existing roles, resources and permissions. As it can be observed, Tenant Users are not allowed to access Users and ACL resources. On the other hand, Tenant Admins are not allowed to edit system users information nor create or remove ACL elements (roles, resources and permissions). The System Admins are able to see the roles and permissions of all users. They have full access to the ACL, which allows them to manage all system’s policies. . . . . . . . . . . . . 64 5.1 Fields selection for Client entity . . . . . . . . . . . . . . . . . . . . . . . . . . . 71 xvi 5.2 Fields selection for Product entity . . . . . . . . . . . . . . . . . . . . . . . . . . 72 5.3 Fields selection for Order-item entity . . . . . . . . . . . . . . . . . . . . . . . . . 72 6.1 Deeply clients’ gender distribution . . . . . . . . . . . . . . . . . . . . . . . . . . 83 6.2 Deeply top 5 country values . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 83 6.3 Top 5 values of colors, sizes and categories of Deeply products . . . . . . . . . . . . 84 6.4 Top 5 values of colors, sizes and categories of ordered products . . . . . . . . . . . 85 6.5 Simulated users in the load tests . . . . . . . . . . . . . . . . . . . . . . . . . . 87 6.6 System’s server load test performance results . . . . . . . . . . . . . . . . . . . . 87 6.7 Locust table that translates the values of metrics collected during the load test, discriminated by each of the routes/endpoints provided by the Recommender API . . 90 xvii LIST OF ACRONYMS A API Artificial Intelligence API Application Programming Interface AWS Amazon Web Services B B2B Business to Business B2C Business to Client BI Business Intelligence C CDC Change Data Capture D DB Database E EDA Exploratory Data Analysis I IaaS Infrastructure as a Service IT Information Technology J JWT JSON Web Tokens M ML Machine Learning P PaaS Platform as a Service R RS Recommender System S SaaS Software as a Service SLA Service Level Agreement SOAP Simple Object Access Protocol xviii 1. INTRODUCTION Informationispower. Nowadays, information is one of the most valuable assets that companies can have to improve their business and stand out on the market. The power of knowing customers, as well as their needs and behaviors, depending on the context or even their surroundings, and the ability to correlate it with the products or services companies can offer, provides them great means to exploit their full potential on the market. This concept is most applied in e-commerceplatforms, which design their websites to draw customers’ attention, by suggesting and recommending their products or services, depending on the customers’ activity in the online store. In this chapter, an introduction to the work developed under the context of this Master’s dissertation is presented. First, it is introduced the Context and Problems where this project was framed, then the Motivation is exposed, followed by the project’s Objectives, listed in a generic and simple way. Lastly, the Document Structure is described. 1.1 Context and Problems The concept of electroniccommerce(5), or e-commerce, can be defined as transactions of goods or services via the Internet, by any electronic means, including transfers of money and data implied in these processes. Thus, it is generally described as any kind of commercial transaction executed through the Internet. However, in the context of this thesis, it will be referred to as the sale of physical products by online stores. There are several types of e-commerce of which we highlight the following two: BusinessToConsumer(B2C) and BusinessToBusiness(B2B). These two concepts are the most commonly explored in online stores and other e-commerce areas. 1 B2C is the process of selling products or services directly to consumers from a platform. In this type of commerce, the consumer browses and buys products for personal use, providing some personal information to checkout and pay, finishing the purchase. For this reason, the system aims to be simple and attractive to the buyer, appealing the emotion to influence their decision on purchasing, focusing on products’ characteristics to satisfy their needs. On the other hand, B2B stands for the process of selling goods or services to other businesses, where buyers purchase products on behalf of their companies. The format of orders and the value of products can vary with the customer, which makes this type of system more complex than the previous one. It is logic-driven, which means that focus on the product details and potential to benefit buyers’ businesses. It also prioritizes saving time, money and resources, improving productivity and those are the main features companies are looking for. E-commerce is constantly expanding, leading to greater market competitiveness. There are more and more online platforms offering products or services; so there is a growing need for companies to stand out from the competition. However, many companies lack the necessary information about their clients and products, and even their competitors, which can make a difference in the current e-commerce environment. This data can help companies to increase clientele and, consequently, profit. This may lead to the implementation of various marketing strategies, but not all are adequate and mismanagement, as well as a bad investment of these strategies, can be harmful to businesses. The project of this dissertation will be developed under the context proposed by the company Beevo (6). Beevo provides e-commerce B2C, B2B and B2E (Business to Employee) solutions, for midmarket and large companies, offering a digital platform for their e-business. The company builds and maintains other companies’ e-commerce platforms, including online stores and respective business logic (customers, products and orders management), marketing, support, CRM (Customer Relationship Management) technologies and business analysis. 1.2 Motivation For companies to have a competitive digital business, Beevo offers more than an online store; it also provides a set of professional apps that allows fast and simple growth. With the e-commerce competitiveness in mind, Beevo proposed the development of provisional models to grow their arsenal of professional apps and boost their Business Intelligence strand. These provisional models were later translated to what is now the core of this project: a Recommender System (RS) for e-commerce platforms of Beevo’s domain. The need to make e-commerce platforms more appealing to its clients makes integrating a recommender system a logical approach. Its function is to provide a pleasant user experience, trying to 2 create a connection with the buyer, showing that the system understands them and is appreciated as a customer by the store. This specialconnection to users increases their loyalty to the store and keeps their interest in coming back for more. As a result, more traffic on the platforms is ensured, which may increase the sales rate and, therefore, a higher number of revenues is expected. However, in order to develop a recommender system, it is necessary to define which information meets the requirements for obtaining trusted recommendations, as well as the data flow between the system and platforms. Several companies defend that most sold products or used services are recommendedto clients by their platform. As evidence of that fact, we see quite popular platforms such as Netflix(2) and Amazon(7), having advanced recommendation engines in their fields. 1.3 Objectives In order to integrate a recommender system into e-commerce platforms, the first step must consist in obtaining information from the selling platform, where data about its users and products/services form the basis of recommendations. But not all information is useful, which may influence the accuracy of the forecasting models used by the system to make recommendations. Furthermore, it is crucial to define a good workflow to collect the data and attend the platforms’ needs for recommendations, so we can run a smooth communication between systems. Following this perspective, the recommender system can be divided into two parts: •Architecture: the recommender system architecture contemplates its infrastructure, responsible to set and handle the communication between the system and the e-commerce platforms, as well as the data management, essential to produce recommendations. •Engine: the recommender system engine is the component responsible for consuming data, preserved in the architectural process, and calculate recommendations based upon users, products and services of the selling platforms. This pair of architectureand recommendation engineform the system developed in the context of this dissertation. Although there are several studies and papers exploring various recommendation algorithms and trying different combinations and techniques, to get the best recommendation results, there’s a lack of investigation on the integration of recommender systems within online platforms. Ergo, this Master’s dissertation is focused on the recommender’s Architecturecomponent. With this in mind, the main objectives are listed as followed: 3 1. Understand the concepts of RecommenderSystem,BusinessIntelligence,CloudComputing and other terms related to a system’s architecture; 2. Implement an architecture of a recommender system, which allows to extract, preserve and analyze relevant data from Beevo’s e-commerce platforms, to integrate a recommendation engine. This architecture will manage the data flow on the communication between platforms and the recommender engine. Summarizing, it will be responsible for delivering data to the engine and return the results to platforms; 3. Explore and define a data analysis methodology. 4. Understand the importance of selecting data and its impact on recommendations performance and accuracy; 5. Explore the potential of the e-commerce data collected on business intelligence applications to better understand and improve businesses. 1.4 Document Structure This dissertation is structured in seven different chapters: •Introduction: This chapter introduces the context of this dissertation project, its motives and what it aims to achieve; •State of the Art: In this chapter it is explored some terms and definitions related to this dissertation theme, as well as exposed some work related to the context of the project. •The Proposal: This third chapter describes the proposed solution to overcome some challenges in order to successfully achieve the defined objectives. •Development: Here all development stages of the system architecture are described, including the technologies used, the structure of recommendations, communication processes and all the decisions taken along the course. •Methods: This part of the document presents the various strategies for extracting and analyzing data and making results available to e-commerce platforms. •Case Studies / Experiments: In order to demonstrate the different functionalities that the system has to offer, a test case is presented in one of the company’s online stores. •Conclusion: In this final chapter, a global review is presented on the project developed, as well as several improvements that can be done in the future. 4 2. STATE OF THE ART As mentioned in the previous chapter, this thesis focus mainly on the recommender system infrastructure, this is, its architectural component. It covers not only the establishment of communication between the system and platforms but also data analysis strategies to process data before being transferred on this communication. The first part of this chapter sets the Background of this work: it defines the key concepts and vocabulary for the rest of the Master thesis, such as Cloud Computing, Monolithic vs Microservices Architecture, Representational State Transfer (REST), Exploratory Data Analysis and Business Intelligence. In the second part, it is exposed some Related Work with the project theme and objectives. Examples of famous companies like Amazon and Netflix are explored, giving some insights on how existing recommender systems work in today’s market. 2.1 Background Recommender systems, or recommendation systems, emerged as an independent field of research in the mid-1990s and derived from different other areas, such as cognitive science, approximation and forecasting theories, information retrieval and also have links to management science (8). This area is widely explored because it constitutes a problem-rich research field and due to the abundance of practical applications, which can help users to deal with information overload as well as provide personalized recommendations, content, and services to them. An e-commerce RS is a machinelearning(ML) mechanism that relies on a variety of data, related to users, products or services, processes it and creates personalized suggestions for the intended user. Its purpose is to assist the user in their purchasing decisions, recommending the products or services that best suit their interests, but can be manipulated according to the company’s intentions. Therefore, it acts as a forecasting model dedicated to calculate and making recommendations, such 5 •Caching: In REST, API calls can be cached, but not in SOAP; •Security: SOAP supports WS-Security with SSL and has built-in ACID compliance. In the other hand, REST lacks ACID compliance and supports HTTPS and SSL; •Performance: SOAP requires more bandwidth and computing power, while REST needs fewer resources, being lightweight; •Message format: REST permits many data formats, including plain text, HTML, XML, JSON and others. SOAP only supports XML; •Advantages: SOAP is standardized and has high security and extensibility, while REST has better performance, scalability and flexibility; •Disadvantages: SOAP is more complex and has poorer performance and flexibility. However, REST may be less suitable and secure for distributed environments. Overall, REST offers several advantages over SOAP, for building communication channels between systems, by being simple, flexible and scalable, allowing a greater variety of data formats and performance. When using REST over HTTP, it’s possible to resort to standard HTTP security and authentication. By combining it with JSONWebTokens(JWT) for authentication and authorization of user’s to validate their requests, an efficient way of secure the communication can be achieved. RESTAPI (Javascript) RequestURI(HTTP) Response(HTTP) URI:http://recommender.com/api/v1/product?id=3 Method:GET { "id":3, "name":"shirt", "size":"M", "color":"white", "price":19.99 } Request Response(JSON) WebApplication (PHP) Figure 2.4: REST API Model example diagram 12 According to the example displayed in the diagram above, there’s a request and a response between a client and the server’s API. The client and server can be based in any language but it doesn’t matter because the message request and response are made through a common HTTP web protocol. This request-and-responsepattern is fundamentally how REST APIs work. 2.1.4 Recommendation methods There are several approaches that can be used to create a recommendation model, the most popular being Content-based Filtering (20) and Collaborative Filtering (21). The premise of the CollaborativeFilteringapproach is to search for similarities between clients, according to their actions and preferences. User-based(22) recommendations take into account the similarity between the clients’ profile, i.e., items purchased by a certain client will be recommended to another client who has similar tastes and behaviors. On the other hand, item-based(23) recommendations are supported in products’ characteristics. For example, user A has a similar buying pattern as the user B. Consequently, items with similar attributes to those that user A has purchased in the past, may be suggested to user B. In Content-basedFiltering, the user’s shopping history is important. The characteristics of items from previous purchases made by the client are analyzed and compared with the remaining candidate items, available in the store. Items that have more in common with those that the user has purchased are recommended. For example, in a certain online book store, a user bought some books in the Sci-Fi category. According to the user’s shopping history, the system may recommend other similar books, i.e. of the same category (Sci-Fi), to that user in the next visit to the online store. A B A similar buy CollaborativeFiltering Content-basedFiltering Sci-Fivol.1 Sci-Fivol.2 similar recommended buy buy buy buy buy recommended Book1 Book1 Book2 Book2 Book3 Figure 2.5: Collaborative and Content-based Filtering examples 13 Each different recommendation technique has its benefits and drawbacks in terms of efficiency and results accuracy, depending on the context in which they are applied. In order to fill the gaps that each approach presents and to make the system more robust, a Hybrid Approach is usually adopted by several techniques, in this case, collaborative and content-based filtering. Following the previous example, with a hybrid approach, the system would recommend not only other Sci-Fi books to the client (content-based filtering), but would also take into account the book’s rating given by the other users (collaborative filtering). Another strategy widely used in e-commerce is CustomerSegmentation(24). By using clustering algorithms, such as k-means, it is possible to trace the profiles of several clients and group them according to the similarity between their characteristics (gender, age, etc). Obtaining an ideal number of clusters can be extremely useful since different clients belonging to the same cluster can be treated as a single entity, thus saving a lot of resources in the production of recommendations and targeted marketing. Since each of these approaches consumes different types of information, it is essential to select the most significant data, corresponding not only to the type of the selling platform, but also to the approach itself. Hence the importance of the quality of the data set used to train the model, as it will influence the accuracy of recommendations. 2.1.5 Exploratory Data Analysis To select the most relevant data, data analysis techniques are usually applied, according to a certain approach previously defined. ExploratoryDataAnalysis(EDA) (25) refers to the initial investigation of a data set, to understand it, so as to extract patterns, detect anomalies, filter outliers, evaluate hypothesis and to verify assumptions with the support of summary statistics and graphical representations. This is an important step in the data analysis process, to help data analysts to comprehend the big picture, in this case, the main characteristics of each different e-commerce platform, setting up the context before any machine learning operation. This way, one can create a model that fits the given context and increase the system efficiency. By examining and treating data, conforming that context, one can develop visual panels as an information management tool to visually track, analyze and display keyperformanceindicators(26), metrics and other key data points to monitor the status and performance of a business. These panels are called dashboardsand prove that the considered data allows going beyond recommendations, and explore the vast area of BusinessIntelligence. 14 2.1.6 Business Intelligence In 1865, BusinessIntelligence(BI) was used to describe how the banker Sir Henry Furnese took advantage of existing information by collecting and acting on it before his competitors, in the CyclopædiaofCommercialandBusinessAnecdotes, by Richard Millar Devens. Years later, in 1958, IBM computer scientist Hans Peter Luhn wrote an article describing the potential of Business Intelligence through the use of technology and this field has evolved since then. The number of BI vendors grew in the 1980’s, as business people discovered the value of Business Intelligence and, consequently, an assortment of tools were developed during this time, with the goal of accessing and organizing data in simpler ways (27). Nowadays, we can define BusinessIntelligence(28) as a set of concepts, methods, processes and technologies that gather and store rawdata, and transform it into relevant and useful knowledge for business purposes. BI can handle large amounts of information to help companies identify and develop new opportunities, as well as planning and making decisions. Thus, studying BI solutions can provide a competitive market advantage and long-term stability, helping to make the right decision. From this, we may conclude that data is a very powerful resource, when well manipulated, giving companies the necessary knowledge to overcome obstacles and competition. However, this brings up some security concerns, where companies can be targeted by rivals, stealing their information. Hence, it is very important to build secure communication channels, where data can flow inside the company’s systems, protecting private business information. 2.2 Related Work As mentioned in the previous chapter, nowadays most popular online platforms have recommender systems, each one adapted to the platform’s background. For example, Youtube(29) is a videosharing and streaming platform, featuring video recommendations to its users, depending on what they previously watched. Along this line, Spotify(30) is an audio streaming platform and recommends songs and playlists to users, matching genres they use to listen. On the other hand, Facebook(31) is an online social network service, which recommends other users’ contacts, so one can be ”friends” based on their current connections. In this section, we highlight Amazon(7), Netflix(2) and eBay(32) systems, as well as other different architectural approaches, in order to achieve a general perspective of how recommender systems work in today’s market. 15 2.2.1 Amazon In light of the theme of this dissertation, Amazon(7) is the best example to consider. Amazon is the world’s largest online retailer, which sells a wide variety of products to customers, like the e-commerce platforms that this project intents to target. Amazon started as an online book store and one of its main advantages, compared to physical stores, was the infiniteshelf-spacecapacity as a platform over the Internet. This allowed it to make a great number of sales from books beyond the inventories of physical stores. But due to the massive quantity of different books stored, customers might have missed some good opportunities to discover new relevant books that they might have bought. Hence, the emergence of the recommender system as a great tool to suggest new books - and other products later on - to customers, from this infinitebookshelf, and consequently increase sales. By the year of 2012, JP Mangalindan claims in a Fortune’s article (33) that: “Judging by Amazon’s success, the recommendation system works. The company reported a 29% sales increase to $12.83 billion during its second fiscal quarter, up from $9.9 billion during the same time last year. A lot of that growth arguably has to do with the way Amazon has integrated recommendations into nearly every part of the purchasing process…” - JP Mangalindan (33) Currently, Amazon uses a combination of user-based and item-based collaborative filtering in their recommendation algorithms, to suggest products to customers along the purchase process, via e-mails, browse pages, product details pages and even at the end of an order. Additionally, according to Ian MacKenzie (34), 35% of Amazon.com’s revenue is generated by its recommendation engine. That is why companies are increasingly investing in RS to deploy in their online platforms. 2.2.1.1 Amazon Web Services (AWS) It is also worth to mention that, as one of the biggest cloud computing service providers in today’s market, Amazon holds the company AmazonWebServices(AWS)which provides on-demand cloud computing APIs and platforms. The clients that may be interested on develop complex and efficient applications, with great flexibility and reliability, can use AWS products for cloud ecosystems - a mix of IaaS, PaaS and SaaS. Some of the features that AWS presents and appeal to clients the most are the following: • Security: AWS is one of the safest cloud platforms on the market; 16 • Experience: Users can get a hands-on experience of AWS free of charge; • Hosting: AWS can host static websites also for free; • Scalability: AWS has a great scaling capacity. Products can be combined to create a scalable cloud application without having to concern about problems related to infrastructure maintenance (compute, storage, and network) and management (10). Amazon Web Services offers products in the areas illustrated below: AmazonWeb Services Networking& ContentDelivery Analytics Quantum Technologies Management& Governance Machine Learning Security,Identity& Compliance Compute Storage Figure 2.6: Applications covered by Amazon Web Services (1) products. 2.2.2 Netflix Netflix(2) is a streaming service that allows members to watch a wide variety of TV shows, movies, documentaries, etc. Recommendation algorithms have been the core of the Netflix product from very early on. Because of its importance, the company continually seeks to improve recommendations results by advancing the state-of-the-art in the field. Netflix’s RS combines collaborative and content-based filtering through similar habits of users and higher rates of shared movie characteristics. The company, Netflix Inc., released a contest in 2006 - The Netflix Prize -, offering a reward of one million US dollars to enhance the recommender system. The team who could succeed to decrease the value of root-mean-squareerror(RMSE) for a data set by 10 percent, would win the prize. Bellkor’s Pragmatic Chaos team succeed in achieving an RMSE of 0.8554 with a 10.06% improvement over the Netflix system. This challenge grew up the attention on recommender systems beyond Computer Science. 17 In an interview with MobileSyrup(35), Todd Yellin said: “We found the typical Netflix member on average will only look at 40 or 50 titles before deciding what they want to watch, even though there are thousands of titles available. So it’s important we present the right content to the right member at the right time.” - Todd Yellin, Netflix’s vice-president of product innovation Netflix uses recommender systems so extensively that, in 2015 Chief Product Officer, Neil Hunt, indicated that more than 80percentof movies watched on Netflix came through recommendations (36) and placed the value of Netflix recommendations at more than US$1billionperyear. This proves the power and importance of recommender systems in a e-business. 2.2.2.1 Architecture Overview In Netflix Tech Blog9, an article published by Amatriain and Basilico (37) explains how Netflix tackles some of the challenges of maintaining a software architecture capable of handling large volumes of data, responsive to user interactions and flexible to new recommendation approaches. An overall RS architecture is described, where the whole infrastructure runs across the public AmazonWebServicescloud. The system’s diagram (38) is presented in figure 2.7. The system’s architecture can be divided into three parts: Online, Offline and Nearline computation. These distinguish the types of processes and recommendations that the recommender system computes. 9https://medium.com/netflix-techblog 18 UIClient OnlineComputation OfflineComputing Recommendations Results Browse,buy,rate... Machine Learning Algorithm Modeltraining Machine Learning Algorithm NearnlineComputation Machine Learning Algorithm Online Nearline Offline OfflineJobs Netflix.Manhattan Netflix.Hermes EventandData Distribution Machine Learning Algorithm User Figure 2.7: Netflix recommender system overview - based on Netflix Tech Blog post (2) •Online computing must have fast responses to events and use the most recent data to fulfill the availability and response time required by the client-side. This constrains the implementation of complex and computationally costly algorithms, limits the amount of data that can be processed and, consequently, recommendations’ accuracy. Moreover, using solely online computation may fail to meet some of the requirements, hence the importance to have a offlinecomputationmechanismas a fallback solution to fit those requirements. •Offline computing is less limited in terms of data processing, computational costs, and complexity, having more flexibility on the implementation requirements and a wider range of algorithms to choose from. However, because of heavy processing in this approach, offline computing does not have a fast response to changes from new events or data. Eventually, 19 this can lead to staleness that may degrade the user experience. It also requires having infrastructure for storing, computing, and accessing large sets of pre-computed results. •Nearline computation can be seen as a compromise between the two previous approaches. In this approach, computation is performed exactly like in the online case, yet results are posteriorly stored, allowing it to be asynchronous as in offline mode. Hence, the requirement of short response time is excluded, allowing to explore the potential of more complex processing, while still enabling the system to be responsive to user events. After receiving a request, the system computes the results and may store them in an intermediate caching or storage back-end. Model training is commonly applied in offline mode, consisting of generating predictive models based on existing data, that will later be used to create suggestions and other personalized results. Another part of the architecture describes how the different elements communicate with each other, handling events (user interactions and activities) - EventandDataDistributionSystem. This nearreal-time event flow is managed through an internal framework called Manhattan. A related issue is the data flow in the process of obtaining Recommendation Results, across the offline, nearline, and online regimes. This is managed by Hermes, a publish-subscribe mechanism, which allows data to be delivered to subscribers in near real-time. Such a complex system shows the importance of planning the software architecture in which the recommender will be deployed. It’s a flexible and sophisticated architecture, capable of handle great amounts of data and manage complex machine learning algorithms, while always having recommendations ready for quick responses. Finding the right balance is not trivial: ” It requires a thoughtful analysis of requirements, careful selection of technologies, and a strategic decomposition of recommendation algorithms to achieve the best outcomes” (37). 2.2.3 eBay eBay(32), as an online auction and shopping platform, presents a scalable RS architecture for recommending items with a short life span (e.g.: auctions), controlling the trade-off between relevance and quality (39). The architecture can be divided into two layers: OfflineModelGenerationand RealtimePerformanceSystem. The offline layer creates models using clustering ML algorithms, while the online layer combines those models with dynamic characteristics obtained from users information and activities in e-commerce. In the paper (39), the authors emphasize two main approach scenarios: pre-purchaserecommendationand after-purchaserecommendation. In the pre-purchase scenario, the RS recommends alternative products which are similar to the ones recently viewed by the user. In the post-purchasescenario, the RS recommends complementary products related to the one that the user has recently purchased. 20 Both layers use the same data store, providing two versions of similar services. Data stored may be related not only to basic information such as users, items, and user actions (navigation, access to auctions, etc.), but also to clustering results, such as which group a set of similar items belong. The real-timelayerhas two components: SimilarItemRecommender(SIR) and RelatedItemRecommender(RIR). Both receive an item as input and return a set of similar or related items in return, respectively. As the response must occur in real time, all the computational complexity is in the offlinelayer, consisting of ApacheHadooprunning mapreducejobs(40), queries and K-means algorithm. 2.2.4 SaaS and PaaS Recommender Systems The tendency of e-commerce stores to search for recommender systems, in order to increase their sales volume and revenue, is growing over the years. However, developing a good RS can be expensive and time-consuming. This leads companies to reach for SaaSandPaaSRecommender Systems(41). Instead of having a large upfront investment, companies can pay as they use a SaaS model of a recommender system. The integration is usually straightforward and there are continuous cycles of improvement (42). In this dissertation, we emphasize Yusp,StrandsRetail,CommerceCloudEinsteinand Amazon Personalizeas examples for using different approaches in their architectures and recommendation techniques. 2.2.4.1 Yusp Yusp(43) is a personalization engine, developed by Gravity R&D company, the same team that tied for first place at Netflix Prize (44) - improving the Netflix algorithm by more than 10%. This service offers customization features for e-commerce platforms, having several case studies from large companies in which their revenues have increased significantly thanks to these solutions. To produce recommendations, the engine consumes data from online activities and habits of both known and first-time customers, the properties of products - such as name, price, category and other attributes - and contextual information of the customer browsing like the location or the time of the day. Due to the importance of this data, Yusp has security measures to protect the privacy of their clients. It provides control dashboards where the client can customize and adapt the recommendation engine to the needs of their platform and their customers. It is also possible to obtain detailed analysis reports, thus giving several insights about the business to help make better decisions for the future. 21 •Personalized data sets: Collect and preserve various types of data from different ecommerce entities. This data will be consumed by the recommender engine to produce recommendations; •Data Analysis: The data analysis tools used must be versatile and capable of processing variables of various types, from the referred custom data sets, to create informative and useful dashboards; •Cold Start problem: When entities (users or items) are recently registered in the platform, the recommender system has limited information about them to be able to produce accurate recommendations. Nevertheless, new customers should get relevant recommendations and new products should be included in recommendations; •Recommendations availability and up-to-date: Recommendation results must always be up-to-date, according to the latest platform activities and available at platforms’ demand; •Scalability: As mentioned in the previous chapter, RS tends to increase the number of visitors on the platforms. Recommendations should scale across hundreds of clients and products. Thus, the recommender system must handle a great number of items and active users, simultaneously, keeping a short response time and good performance; •Multi-tenancy: The multi-tenancy problem (52) refers to a software architecture in which one application instance is hosted on a server and serves multiple tenants. A tenant can be a set of one or multiple users who share common access to the software instance. With a multi-tenant architecture, an application is designed to provide each tenant with a dedicated share of the instance. The multi-tenancy definition opposes to multi-instance architecture definition, where separate software instances serve different tenants individually. In this sense, we need to consider the following scenarios: 1. Deployaninstance(copy)oftheRSstackforeache-commerceplatform(tenant): Pros:Simple to develop and higher performance individually. It allows delivering more personalized service to each platform, being easier to manage business logic inside each platform scope. Cons:Global system management more complex and it consumes many resources. It is necessary to configure an additional instance each time a new tenant is considered. 28 2. CreateaninstanceoftheRStohandlemultiplee-commerceplatforms: Pros:Allows the system to be generic and flexible. Only a type of architecture must be developed and maintained for the whole system to handle multiple tenants. Consumes less resources. Cons:More complex to develop, but easier to configure to communicate with several platforms, at the final stage. It is expected to be difficult to separate and manage the business logic of the different platforms. •Security: Data used to produce recommendations must be secured, because it is based in sensitive information such as clients’ personal information, products details and orders history. Hence, privacy must be protected. To get the most of the architecture’s potential, it must be generic, flexible and capable of being deployed and tailored to most e-commerce platforms dedicated to product or service transactions. Furthermore, data analysis methods and techniques must be able to handle different types of entities and attributes. 3.1.2 Functionalities The recommender system architecture comprehends not only the communication between the recommendationengineand e-commerce platforms but also between the elements within the system. Thus, it is necessary to organize a structure capable of managing large volumes of data to be transmitted in its communication flux. This communication process can be divided into the following steps: 1. Extract data from e-commerce platforms, necessary to train recommendation models at the engine. This data is related to clients’ profiles (gender, age, addresses, ...), shopping activities (page views, browsing, client’s shopping history, orders, ...) and item inventory (product characteristics, hits, ...). 2. Store the collected data in a database capable of adapting to different contexts of distinct ecommerce platforms. The system should present tools capable of processing and analyzing the collected data, in order to provide various Business Intelligence features to the company, such as future perspectives on its business and support in decision making. 3. After recommendations being generated by the engine, the system must store and make them available at platforms’ demand, whenever necessary. 29 Throughout the communication flow, the architecture must take security measures to protect data against potential threats and reduce vulnerabilities in the system with authentication and authorization mechanisms. 3.2 System Requirements E-commerce platforms have strict requirements that must be met in order to increase and maintain quality and offer a good experience to their clients. An important aspect to take into account is the response time, as it will influence the loading time of web pages. Online shopping should deliver a smooth and clear experience to their users, where results are presented as quickly as possible since 40% of shoppers will abandon a website that takes more than three seconds to load, according to a study conducted by ForresterConsultingonbehalfofAkamai(53). Furthermore, a short page loading time is a key factor in a consumer’s loyalty to an e-commerce site. The study reveals that 79% of users who have not had a good experience are less prone to return to that platform while 27% are less likely to buy from the platform’s physical store, suggesting that a poor online experience may have a great impact in store sales. Therefore, the recommender system must fulfill some requirements as well, so as not to impair the normal functioning of the e-commerce platforms and the ServiceLevelAgreement(SLA)continues to be complied with. In conclusion, the impact of generating and obtaining recommendation results should be minimal on the loading time of a web page. Taking into consideration the challenges raised in the previous section (3.1.1), the requirements defined for the system follow as below: Functional requirements 1. The recommender system must return a list of IDs related to hotproducts. These are the most popular products at the moment and with the highest likelihood of being purchased by clients. 2. Given a client ID, the RS must return the list of product IDs recommended to that specific client, sorted by purchasingprobability. 3. Given a certain product ID, the RS must return a list of product IDs that are similar to that specific product, sorted by similarityscore. 4. The RS must be able to do complementaryproductsuggestions. It should return a list of product IDs that best suit a purchase, according to the products currently selected by the client. 30 5. The RS must be flexible and capable of delivering recommendations according to the context that is required by the platform. With this in mind, the system must be capable of filtering recommendations by category, returning only products which belong to that filter. For example, when a user searches for women products, only products of that category should be included in the recommendation. 6. The RS must be able to return recommendations up to the maximum limit established by the tenant. If this limit is not specified, the default quantity of recommended products to retrieve in the server response is 100 items. Non-functional requirements 1. The RS must provide a tool capable of analyzing data stored in the system’s database, by allowing developers to create dashboards and calculated statistics, to develop new prospects about the future of their business and market. 2. The RS must always collect recent data from the platforms’ database. Whenever a certain event occurs, the RS database must be updated, to maintain its consistency and be in sync with the platform’s. Such events can be described as the registration of new clients or products, modifying their information, creation of new shopping orders and changes in order status. 3. The server must respond in JSON format. 4. The system must be functional 99% of the time in a year (361 in 365 days). 5. As for the response time, the RS must return a result in less than 100 milliseconds to 95% of the requests. No answer should take more than 150 milliseconds. 6. The RS must implement data security techniques, mainly authentication and authorization mechanisms, to protect the company’s business information and its clients’ privacy. 7. The RS must not only handle large amounts of data related to clients, products, orders and other shopping activities, but also be able of serving multiple platforms at the same time, scaling as necessary to keep a good performance and a short response time. 8. The RS must allow its users to configure and manipulate the training of recommendation enginemodels. 31 3.3 Proposed Approach The proposed approach to build this architecture was to assemble a set of components, all interconnected by a core JavaScriptAPI, supported on the concepts of PaaS,RESTand the State of the Art explored in the previous chapter. The idea resided on creating a service, having the recommendersystemhosted on a server held by Beevo, where e-commerce platforms of the company’s domain can request for recommendations produced by it. This way, the RS works as a PlatformasaService, delivering all the infrastructure, components and logic needed to obtain recommendations. It does all the computational effort, while tenants (e-commerce platforms) only need to communicate with the RS and integrate the results in their applications. Another advantage is that the RS is developed and updated independently, not compromising the platforms’ functioning. The interaction between the RS and the platforms is possible through an application programming interface (API), which manages and describes the communication process. This communication dwells on the exchange of data, in JSON format and via HTTP requests, between platforms and the RS endpoints delimited by its API. The e-commerce platform sends business information about its clientss, products and orders history to the recommender, which preserves this data for later analysis and processing. The architecture of the recommender system stores the data, replicating part of the platform’s database, forming data sets useful for two main operations: dataanalysisand recommendations results. Exploratorydataanalysisis performed on data stored by the RS with the objective to discover and extract useful information that may help to understand the market, consumers and adopt better strategies. This data will optimize the company efficiency in decision making and business predictions, opening the door to BusinessIntelligencefunctionalities. Then, the recommender engine ingeststhat data, transforming and using it to train predictive models, which create recommendations for the platform’s clients. Those recommendation results are stored in a database of the system’s architecture component, which later displays them on platform’s demand. Although recommendations are saved in different formats, according to their type, they present the same format when they reach the platforms. Whenever a platform makes a request to the RS server, it responds with the according result, but always in the same format, regardless of the type and context of the recommendation. This format, in turn, consists of a listofproductIDs that platforms receive and use to display the recommended products to users. This API was developed based on REST principles, due to its advantages and increasing usage, although it’s not RESTful. In order for the RS to be deployed on e-commerce platforms, an effort 32 from both sides is required. Company developers must know how to conceive the communication between platforms and the recommender. To accomplish this, they expect to have knowledge about the API’s functionalities, in spite of the API’s documentationcontradicting some of the fundamental rules of REST. Citing Giessler’s ”BestPracticesfortheDesignofRESTfulWebServices”article (17): ”AdocumentationforWebAPIsisadebatabletopicinthecontextofRESTfulwebservicessinceit representsanout-of-bandinformation,whichshouldbepreventedaccordingtoFielding(54):Any effortspentdescribingwhatmethodtouseonwhatURIsofinterestshouldbeentirelydefined withinthescopeoftheprocessingrulesforamediatype” Usually, as mentioned in section 2.1.3, tenants should not have knowledge about servers’ REST API. Describing a server REST API, using a description language to make it machine or human-readable, disregards two of the constraints of REST: self-describing messages and HATEOAS (hypermedia as the engine of application state). Nevertheless, in this case, tenants are not normal users, but company developers and since they need to integrate the RS withing the platforms, the most common approach is to document all URIs, HTTP methods supported, and structures of representations (e.g. as JSON) so that tenantapplication developers can rely on such documentation to program (55). 3.3.1 Architecture Description Before developing the architecture to execute the functionalities mentioned in section 3.1.2, the first step relied on planning the approach to follow and the implications it brings to the system’s performance. 3.3.1.1 Architectural Approaches In this dissertation, two main architectural approaches we considered - Monolithic vs Microservices Architecture. In light of these definitions, the following options were elaborated: •Monolithic architecture: With this approach, the recommender system would consist of a single application instance, which handled and managed all the functions mentioned previously (communication, data storage and analysis, authentication, etc). This could represent an issue due to the great number of requests to be exchanged between platforms and RS’s architecture components. As an example, a scenario to consider could be BlackFriday, 33 where thousands of users’ activities would generate a huge amount of traffic on the communication network. To fight this problem, application redundancy can be applied, i.e., create multiple copies of the same application instance and then a load balancer would manage the resources, keeping the flow consistent and avoid bottleneckproblems. This would also depend on the infrastructure that hosts the system software. Any additions or changes to the software can be of great complexity as the various components can depend on each other. Modifying one part of the software could imply changing other parts that interact with the first, as well. •Microservices architecture: This option, although more complex to implement than the first, and perhaps less practical, is more robust in terms of processing and congestion control of data and requests. As initially stated, this approach consists of dividing the architecture component into several small and independent services (microservices), each dedicated to a function: one for communication, another for data storage, other for authentication, and so on. It is expected to be slightly slower than the first option, when the traffic flux does not exceed a regular day, because each request has to be filtered, validated and processed by each microservice. However, as declared before, the monolithic option can be slower on special days, with promotions like Black Friday, where there is more activity on the platforms and the data flux increases considerably. 3.3.1.2 Architectural Solution By evaluating the requirements of this project, considering the need of response to the challenges in section 3.1.1, and by assessing the several situations and risks, as well as the characteristics of Beevo’s e-commerce platforms, it was made the following decision: Amicroservicesarchitecture was adopted, rather than a monolithic approach, because a modular system enables the RS to be very flexible, easily implemented and extended depending on the ecommerce platforms different needs (56). This modular approach allows the usage of different technologies and methods on each component, thus facilitating the best implementation to solve a problem and fit to the companies necessities. As stated before, a monolithic architecture may present data congestion problems (bottleneck) when faced with a scenario where platforms are visited by numerous users and, consequently, many events, generating a great amount of traffic in the network. With the microservice approach, a single architectural service will serve multiple clients - multi-tenancy-, instead of creating a copy of the service for each e-commerce platform considered as each platform is served by one instance -single-tenancy. With the scaling and multi-tenancy problems solved, we focused on how to perform data analysis 34 on custom data sets. To make this possible, it’s important to decide which database types should the RS hold. Two types of databases were contemplated: relationalandnon-relationalDBs. A nonrelational database revealed to be the best choice, since the objective is to build a generic and flexible architecture, capable of adapting to any type of data from different online platforms. Non-relational databases confer flexible storage to the system, accepting various variables with distinct types. On the other hand, relational databases must be planned and structured beforehand, defining what entities and respective fields is the system going to work with. As for security concerns, it were implemented authentication and authorization mechanisms with JWT (JSON Web Tokens - RFC 7519 standard (57)) in the API. 3.3.1.3 Architecture Diagram This section presents a detailed explanation of the system’s workflow. The diagram in figure 3.1 illustrates the sequence of interactions between online platforms and components of the recommender system, data flow and other mechanisms that occur in the process. Plataform1 Plataform2 PlataformN ... Auth ... POST GET Elasticsearch Kibana MongoDB Recommender Engine Endpoint2 Endpoint1 Endpoint3 Endpoint4 Recommender API BIApp BIApp BIApp Source: Improving Performance of Recommendation System Architecture (Appendix I - Publications) Figure 3.1: Recommender System Architecture diagram As it can be observed, several online e-commerce platforms connect to the recommender system through its API. In turn, this API contains multiple endpoints, each one related to a certain type of entity (client, product or event) corresponding to a given RS action. It was also recognized the existence of a recommendationengine, assuming that it was able to connect to the databases used in the architecture - Elasticsearch(58) and MongoDB(59). The recommender engine is capable 35 of storing data in these databases in the format that the architecture requires to communicate with platforms, according to a Data Contract previously determined. For security concerns, each request is filtered by an authenticationand authorizationmechanism, before reaching an endpoint. This security measures validate and ensure that the author of the request has the right permissions to make use of the API, protecting it from potential threats. It was implemented authentication, with JWT (JSONWebTokens(57)), and authorization mechanisms with an AccessControlList(ACL), further explained at section 4.5. 3.4 Summary Before the conception of the proposed architecture, several potential problems were addressed, that may arise during the development of the system, such as personalized data sets, data analysis, cold start problem, availability, scalability, multi-tenancy and security. It was also gathered the requirements that the recommendation system must satisfy before it can be employed as a service by e-commerce platforms. It was assumed, therefore, the development of an architecture based on a microservices approach, supported by non-relational databases, thus giving great flexibility to the system. An API was defined to serve as a bridge between the e-commerce platforms and the recommendation engine, to manage the communication between tenants and the service and ensure the privacy of the data involved. 36 4. DEVELOPMENT This chapter reflects the stages of development and implementation of this project. At first, the Technology Used in the architecture is identified, described and complemented with the list of infrastructure requirements for the host machine. Posteriorly, Product Recommendations are explored, more specifically the concept and structure of each type developed for this system. In order to complete this section, onlineand offlinecomputing modes are also compared. Afterwards, the System Communication Process is explained, i.e., the interaction between service and tenants, ending up focusing the e-commerce platforms side through the Beevo’s Business Intelligence Application. To conclude this chapter, Security mechanisms applied by the recommendation system in the preservation and management of data are specified in detail. 4.1 Technology Used In this section is presented a description of all components of the recommender system architecture, illustrated in figure 3.1, and their role in the process of obtaining recommendations for an e-commerce platform. The recommender API acts as an intermediary between e-commerce platforms and the recommender engine. It’s a communication channel between both elements, handling platforms’ requests and delivering recommendation results, produced by the engine. It can be designated as the RS manager. To build this component, Node.js (60) was used as the server engine with great performance, scalability and lightweight, supporting the API with the Express framework (61), which has a great potencial to assemble a microservices architecture. The server’s application modular structure - appendix A.1 - was developed based on the PrincipleofSeparationofConcerns(62), having been created three layers: ControllersLayer,ServiceLayerandDataAccessLayer. 37 ${platform}_complementary_products${platform}_similar_products ${platform}_hybrid${platform}_popularity { product_id:'XXXX-XXXX-XXXX', items:[ { product_id:'XXXX-XXXX-XXXX', support:0.XX }, { product_id:'XXXX-XXXX-XXXX', support:0.XX } ] } { product_id:'XXXX-XXXX-XXXX', items:[ 'XXXX-XXXX-XXXX', 'XXXX-XXXX-XXXX', 'XXXX-XXXX-XXXX' ] } { id:0, items:[ { product_id:'XXXX-XXXX-XXXX', categories:[ 'XXXX', 'XXXX', 'XXXX' ] }, {...}, ...] } { client_id:'XXXX-XXXX-XXXX', items:[ { product_id:'XXXX-XXXX-XXXX', categories:[ 'XXXX', 'XXXX', 'XXXX' ] }, {...}, ...] } Figure 4.3: MongoDB recommendation documents structure The process of saving a recommendation, in MongoDB, is instant and consists in replacing the existing recommendation for a recent result, thus not increasing exponentially storage space. Additionally, every document is saved separately so there’s no downtime in recommendations’ availability. This means that while the recommender engine is updating recommendations, platforms are still able to get results from the RS, because MongoDB database persists the previous results during this process. 4.2.3 Filters It is evident that the number of recommendations generated is proportional to the number of products available on the online stores. The recommender’s API allows a tenant to specify the maximum number of recommended products it expects to receive, as well as the category or set of categories to which those items should belong. If the limit of recommendations is not specified, the default maximum quantity of recommended products to retrieve in the server response is 100 items. This acts as a security measure, if for any reason there is a large number of recommendations for a certain kind of request, preventing the transmission of unnecessary amount of data, since usually when a client performs a search, it doesn’t inspect more than 100 results. On the other hand, when a tenant makes a request indicating the categories to which it expects the recommendations to be related, the system must ensure that all recommended products respect the 44 filtering rule. As previously mentioned, the filtering process is only possible for recommendations of the type popularityor hybridand it is executed according to the following steps: 1. Find the client’s recommendation document stored at MongoDB database; 2. Obtain all recommended products which belong to categories from the request’s query filter; 3. Order recommendations by score, in descending order; 4. Apply limit of recommended items quantity; If this filtering process did not occur, it was possible that recommended products with a higher scorewould be sent in the response, even if they don’t belong to the indicated categories, and may not correspond to the context that the tenant desired. An example of this is the product listing scenario: when a client selects a category on the website, only products within that category should be displayed. When the tenant sends a request to the recommender, the categories found in the query are individually tested against the list of recommended products. Later, the recommender selects each product that includes the indicated categories. This process is done by MongoDB, which acts as a very powerful and effective search engine, and therefore a good filtering tool. In addition, it is possible to create search indices to increase the performance of MongoDB search. For complementaryproductsrecommendations, the filtering process is as followed: 1. Fetch complementary product suggestions, for each product selected by the tenant; 2. Remove products that are already selected from the suggestions; 3. Sort product suggestions by support value, in descending order; 4. Get complementary product suggestions IDs and remove duplicates, since different products may be complemented with the same product; 5. Limit the result’s length with given limit value; 6. Return complementary product suggestions. In this type of recommendation, as well as in Similarproducts, there is no option to filter by categories, since the contexts in which they are applied do not require filtering by this parameter. 45 4.2.4 Online and Offline Computation As mentioned in the Netflix’s Architecture Overview section, algorithmic results can be computed either in onlineor offlinemode. In online computation, recommendations are created in real-time, responding better to recent events and user interactions, whereas in offline computation results are calculated and stored in batch mode for later use upon platform’s request (offlinejobs). Nearline computation suggests the combination of both approaches, in which it’s performed fast and simple computation (online mode), but it’s not required to serve in real-time. Instead, results are stored making it asynchronous (offline mode). Decisions regarding the method and frequency of calculating recommendations were reflected in these concepts. In the end, these decisions can be summarized in two considered options: 1. Training recommendation models in periodic cycles (e.g., every day at 3:00 am, since there is less activity on the e-commerce platforms) for all clients and store recommendations in the database. Thus, when a platform requests recommended products for a given user, the response is immediate, due to the pre-computation performed. Nevertheless, this approach can present several problems, such as the great computational effort on the engine during the period of generating recommendations and after a given moment these are no longer up-to-date, since they are not produced in real-time. 2. Training recommendation models each time a platform requests the list of recommended products for a user. This approach has the advantage of producing recommendations always taking into account the latest data from the platforms. However, if the models are too complex, the training and recommendation process can influence the response time, by increasing it and making it unacceptable. Additionally, the computational effort is proportional to the number of requests the RS receives, presenting a possible bottleneck when the number of requests increases. Ultimately, to achieve a flexible and efficient system, all approaches should be considered and combined. The general idea is to pre-compute part of a result with an offline process, and using it as a backup, leaving the less costly or more context-sensitive parts of the algorithms for online computation. 46 UserInteraction E-commerceplatform USERBROWSES THEWEBSITE USERCLICKSA PRODUCT LOADPAGEVITRINES WITHPRODUCT RECOMMENDATIONS SHOWTARGET PAGE PERSONALIZED VERIFYIFPLATFORM ALREADYHAS RECOMMENDATIONS ONCACHE GETPRODUCT RECOMMENDATIONS FROMRECOMMENDER SYSTEM NO YES SENDUSER ACTIONAS ANEVENT TORS Figure 4.4: User Interaction Activity Diagram RecommenderSystem APIRECEIVES HTTPREQUEST STOREDATA INELASTICSEARCH API Engine EXTRACT,ANALYSE ANDTRANSFORMDATA FROMELASTICSEARCH APPLYMLALGORITHMS TOPRODUCE RECOMMENDATIONS FILTERAND GROUPRESULTS SAVE RECOMMENDATIONSIN MONGODB POSTDATA GETRECOMMENDATION TRIGGERTHEENGINE TOTRAINMODEL, GENERATEAND RETRIEVEONLINE RECOMMENDATION FETCH RECOMMENDATION FROMMONGODB RETURN RECOMMENDATIONSTO PLATFORM ONLINE OFFLINE Figure 4.5: Recommender System Activity Diagram 4.2.5 Offline recommendations Much of the computation of recommendation ML algorithms can be done offline. Offlineor schedule jobscan be defined as tasks executed in background, on the machine, by a jobscheduler(70). This program monitors and manages batch jobs automatically, allowing users to define and control a schedule to execute those jobs. With this in mind, jobs can be scheduled to run these algorithms periodically and their execution does not need to be synchronous with the request or presentation of recommendation results. 47 In this project, Cronwas used as a job scheduler, which is compatible with the operating system of the host machine, to execute cronjobsthat are applied to model training and recommendations production on the recommender engine. A crontabwas created to list and configure several jobs; each job runs periodically and executes a python script of the engine, to train models and generate recommendations for each different e-commerce platform. For example, 00 03 * * * python recommend.py deeply configures a job to run recommend.py everyday at 03:00 am, when there’s less activity on the website. The recommend.py script makes the engine to perform an ETL (Extraction, Transformation & Load (71)) process, extracting data from Elasticsearch, analysing and transforming it into relevant information (EDA), and loading it into data sets to apply ML algorithms, training models and updating recommendation results of Deeply e-commmerce platform. This way, it is possible to use more powerful and complex recommendation techniques, by executing them in the background allowing to produce more accurate results, asynchronously within the system. When the process ends, results are stored in the MongoDB database, being available anytime the platform demands for recommendations. 4.2.6 Online recommendations Although offline recommendations may provide quite accurate results, the updating rate of these results can be very low. For example, consider the following scenario: 1. A client buys several different shoes; 2. When the recommender engine generates recommendations, it will take this into account and recommend other shoes; 3. Moments later, the client buys various t-shirts; 4. Recommendations remain only for shoes during the period in which the engine stays idle. If the engine runs in cycles of one day, only the next day will the recommendations be updated. In other words, the engine will train the models again taking into account the client’s latest purchases, in this case the t-shirts, only on the following cycle. Summarizing, if a client buys a certain type of product, than only on the next cycle will there be recommendations for him that include products of that type, as models are re-trained and results updated. A possible solution for this scenario is to make the recommender engine stay put, permanently listening for requests related to online recommendations. The recommender API will receive and handle these requests from platforms and redirect them to the engine, which in turn simply starts the 48 necessary processes to obtain the demanded recommendations. In offline mode, the engine just takes into account the data loaded at the beginning of the training. If a client has purchased other products in the meantime, the engine will not take these orders into account as this information will only be loaded in the next training session. To obtain recommendations in real-time, a lighter and faster model would have to be developed, even if it is not highly accurate, so as not to increase the response time significantly and to be used only for situations in which it is necessary to have real-time recommendations: results are only produced when requested by the platform, not being needed to store them in the MongoDB database. This model would always take into account the client’s latest orders. However, this solution will depend on the scalability of the system, since it will generate a large number of requests to process, derived from the multiple interactions of users in the platforms. Besides, it is necessary to consider the time for training and formulate new recommendations, as it will influence the response time. In fact, if the recommender is triggered every time a user purchases, to recalculate the recommendations for that user taking into account this new information, a large flow of information will be generated in the system, compromising its capacity to process data, recommend and, consequently, increasing the response time. 4.2.7 Nearline recommendations In order to keep the performance of the system stable, it was decided to exclude the solution of online recommendations and adopt an intermediate approach: nearlinerecommendations. This method consists of a faster and simpler recommendation model from the ones used in offline mode, which trains in smaller cycles (every 20 minutes, for instance) to produce nearreal-time recommendations, e.g.: when a client buys t-shirts, 20 minutes later he is presented with t-shirts recommendations. Nearline and online modes provide adaptability to the system, updating recommendations according to the client’s recent interactions with the platform. For the nearline mode, it was necessary to create a lighter hybrid model that uses variables more focused on this context, that is, a model that has a reduced execution time to achieve a balance between short updating cycles and a sense of adaptation from the client’s perspective. Nevertheless, the disadvantage of a more simplistic recommendation model is reflected in less accurate results. The engine trains this model and generates results more regularly in nearline mode than in offline mode, but the recommendation process is similar on both, saving the results in the MongoDB database at the end, in different collections. This is, therefore, another very useful recommendation mode, which can be applied in several scenarios, complementing the offline mode, e.g .: a showcase on the platform’s Homepage where the first 3 products would be nearline results and the others offline, presenting the most recently 49 updated recommendations first and giving the feeling of an adaptive system. Yet, offline mode is the most frequently used, since it is more accurate and is always available, allowing an immediate response time and also serving as a backup for the other modes. 4.3 System Communication Process As previously mentioned, the core of the recommender system’s communication focuses on its API. It acts as an intermediary between the platforms and the recommendation engine, making it possible to divide this communication process into two operations: the communication between API and Engine and the communication between RS and e-commerce platforms. When developing this recommendation system, the interaction between API and Engine was first considered, where two options were contemplated to achieve this process: 1. Develop the API independentlyof the recommendation engine. The engine would be considered a black-box, ignoring its mode of operation, software and algorithms used. The communication between these two elements would be made through HTTP requests based on REST, that is, the engine would also have to have several endpoints in order to exchange requests with the system’s API. The engine’s development would be solely focused on the formation of recommendations, disregarding the origin of data or the destination of recommendations. Thus, the API would manage the entire process of storing data and delivering results to platforms. So, whenever the engine needs to train the models and update recommendations, it would reach the recommender API, which in turn would fetch the necessary data from the data lake (Elasticsearch) and return it to the engine. After training and generating the recommendations, the engine sent the results to the API, which would save them in MongoDB, for later platform demands. 50 Recommender API API Recommender Engine Online stores 1.1. Send store's e-commerce data Elasticsearch MongoDB 2.1. Request data to produce recommendations 3.1. Return recommendation results 3.2. Save recommendation results 4.2. Fetch recommendation results 1.2. Save data 2.2. Fetch data for recommendations 4.1. Request for recommendations Figure 4.6: Communication process between Recommender’s API an Engine - option 1 2. Recognize the functioning of the recommender engine, assuming that it not only has the capacity to access Elasticsearch directly to retrieve the necessary data for model training, but also direct access to the MongoDB database to store the results. Recommender API Online store 1.1. Send store's e-commerce data Elasticsearch MongoDB 3. Save recommendation results 4.2. Fetch recommendation results 1.2. Save data 2.Fetch data for recommendations 4.1. Request for recommendations Recommender Engine Figure 4.7: Communication process between Recommender’s API an Engine - option 2 The first option aimed to produce more generic elements, making the system more flexible. In other words, when implementing a REST communication between the API and the engine, it allows collecting recommendations from several distinct engines, thus having a more robust base of outcomes to be provided to platforms. It would also make it possible to test and replace different engines, 51 giving the possibility to choose the one most suited to the tenants’ goals. However, the complexity of data management in this communication approach does not make the process very efficient, hence adopting the second option. Since the context of this project was properly defined, as well as its objectives, the second approach offers greater performance and effectiveness in the system. So, we assumed the development of a recommendation engine (65) that can connect directly to the databases, with no need to implement REST in the engine for communication by HTTP requests, consequently taking the weight off the system’s API of managing too much resources. It is assumed that the engine produces and stores the results in the format that the system architecture requires to communicate with the e-commerce platforms. The communication between RS and e-commerce platforms is made via HTTP requests, through the API based on a REST architecture. This allows the system to be used not only by Beevo, but also by other external entities, i.e., recommendations generated by the recommender can be requested by several different tenants, that may not be related to the company, providing they are duly authenticated in the service. The communication with the RS must be properly authenticated and authorized, as well as validated according to a Data Contract. Any entity can communicate with the RS and obtain recommendations if they follow the data contract and this is what offers the system such flexibility to cover all e-commerce platforms. In the context of the project, the recommendation system was assessed through the interaction with Beevo’s platforms, which communicated with RS through the Beevo’sBusinessIntelligence Application. 52 4.4 Beevo’s Business Intelligence Application As stated in the previous section, any entity can communicate with the recommendation service through HTTP requests, as long as it complies with the requirements established in the data contract, regarding formatting and mandatory data that must be included in those requests. In order to complement the communication process at the company’s side, an application called BusinessIntelligenceApplicationor BIAppwas developed, allowing platforms to communicate with this service, transmitting all the information necessary to form data sets, perform EDA, Model Training and obtaining product recommendations. The reason why it was called BusinessIntelligence Appis because in the future this application will hold more features, directly linked to the BI area, in addition to this first feature developed - productrecommendations. The same is reflected in the recommendater server, which will later host other types of services related to the BI area, hence adopting a generic and flexible development. E-commerceplatform Redis RabbitMQ Apps Plugins Core Resources BIApp RecommenderSystem API Engine Elasticsearch MongoDB HTTPRequests Response (listofrecommendations) Fetchdata Cache recommendation results Synchronous communication Occurrenceofeventsand messagesaresenttothemessage broker Queuesareprocessedand respectiverequestssenttoRS Asynchronous communication MySQL Figure 4.8: Communication process between BI App and Recommender System Each platform is capable of integrating the BIApp, earning the ability to interact with the recommendation service and, consequently, obtain product recommendations. This application was developed in PHP (72), a language specially suited to web development, which is the context where this project is inserted. The following diagram shows the general file structure of the BI App developed. Note that some nuances may not correspond to reality to protect the company’s property and privacy: 53 messages from triggers and distributes them across several queues, balancing the load. In turn, these queues process messages asynchronously and sends the information almost in real-time to the RS server - diagram 4.8. This technique gives RS the possibility to implement onlinecomputationin the future to create recommendations in real-time. Besides, by keeping the Elasticsearch database always up-to-date, it allows to perform real-time business analysis in Kibana’s dashboards. 4.4.4 Storefront widgets Specific widgets were created to display recommendations to users through vitrines on the platform, communicating with the RS, by using the BI App methods and identifying the client for which it is requesting recommendations, the number of recommended products and the categories to which they should belong, in case these filters are applicable. These recommendation widgets request recommendations according to the type that they were configured, displaying the recommended products that were received through vitrines in the web page. In this project, fourdifferent features were developed, which use different types of recommendations, and were implemented in several widgets strategically placed through the platform. In the Homepage,”RecommendedProducts”vitrine presents a set of recommended products specific for the user currently on the page, based on hybridor popularityrecommendations. Popularity recommendations are used everytime the user is anonymous, i.e. not logged in. Otherwise hybrid recommendations are requested. However, if there are no hybrid recommendations for the client, popularity recommendations are used as default as they represent a generic result to use in these situations. This way, the cold start problem can be solved, in the sense that new or recent clients get recommendations, even though the system has no information about them. Homepage RecommendedProducts Homepage Figure 4.10: Homepage recommendations vitrine generic template. In this example, four recommended products are displayed to the user: two from ’Clothes’category (shirt and pants) and two from ’Drinks’category (iced tea and smoothie). Products are ordered by score, with the product on the left having the highest score, i.e., is more likely to be bought by the user. 60 Product Details page also presents a recommended products vitrine, visually similar to the one used in Homepage. However, recommendations are from similar-productstype, meaning that products that are recommended have similarities to the selected product of the page. Homepage Shirt Productdetails RecommendedProducts Price: €€ Figure 4.11: Product details page recommendations vitrine generic template. In this example, the user selected a shirt, thus be presented products similar to it in the recommendation vitrine, where the product on the left is the most similar to the selected shirt (ordered by score). As for suggestions to complete the shopping cart, these are available on a Side Cart, accessible at any point during the purchase process on the store. Whenever a client adds or removes a product to their cart, the widget reloads and the respective vitrine is updated displaying the appropriate suggestions according to the current cart content. The platform informs the RS of the current products in the cart, which in turn returns the complementaryproductssuggestionsrequired by the widget. Furthermore, whenever a client has no products added in their shopping cart, popularity/hybrid recommendations are used instead of complementary products recommendations, due to being impossible to make cart suggestions with an empty cart. Homepage Sidecart RecommendedProducts Total: €€ Figure 4.12: Side cart recommendations vitrine generic template. In this example, the user added a shirt and pants to the shopping cart, so the recommendation vitrine displays some products which are commonly bought together with the cart’s current content. 61 Lastly, a new ordering option was added to Product Listing on the websites: ”OrderbyRecommended”. With this option, products from product listing are ordered according to recommendations given by the RS. Here, once again, recommendations are from types hybridor popularity. In this case, two scenarios had to be considered: a genericproduct listing, when the store page is displayed showing all available products, and a specificlisting, when the user filters the product by category. In this last situation, it was necessary to ensure that all recommended products belonged to the category selected by the client. Since each recommended product of hybrid and popularity types has its category associated, and taking advantage of MongoDB’s capacity to be a very efficient search engine, it was possible to filter the recommendations and send only those related to the category of the page. Homepage ProductListing Productscategory:Drinks Orderby:Recommended Figure 4.13: Product listing page recommendations vitrine generic template. In this example, the user browses for products within the ”Drinks”category and orders them with the ”Recommended” option. Ergo, products related to ”Drinks” are listed, ordered by recommendation score, i.e., the products that would appeal the most according to the user’s profile are shown first. Recommendation widgets use the Rediscomponent to cache results received from the RS. By saving recommendations, for a certain time, the application no longer needs to send repeated requests demanding the same recommendations, therefore reducing the number of calls to the server and, consequently, the load on the communication flow, increasing overall system performance. A user’s session on a Beevo platform typically takes half an hour, with the possibility of the user returning within the next hour after that visit. Thus, when a widget obtains recommendations from the RS, it saves those results in cache. When the user accesses a page whose context requires recommendations that have been previously loaded, the widgets will fetch the results from cache, allowing the page to load faster and improving the user experience. The format of recommendation requests is specified in section 5.3. 62 4.5 Security In order to assemble the security mechanisms of the system, potential threats were raised first according to the STRIDE (74) methodology: Spoofing,Tampering,Repudiation,Informationdisclosure,Denialofserviceand Elevationofprivilege. Consequently, it was ensured that the system has properties such as authenticity, integrity, non-repudiability, confidentiality, availability and authorization. Thus, it was decided to develop an AccessControlList(ACL), which dictates a hierarchy of user roles, and their respective permissions. The ACL establishes the various policiesof the system, indicating which resources exist, which operations can be done in each of them and who can execute them, thus constituting the RS authorization mechanism. There are three types of users, i.e. roles, with access to the recommender system: •System Admin: The user is an administrator of the system. The administrator has global access to the recommendation system and is allowed to create different types of users for each platform (Tenant Admin and Tenant User). It also can access and operate on any platform and respective resources. Each platform is referred as a tenant and has its own Tenant Admins and Tenant Users. •Tenant Admin: The user is associated to a platform and acts as an administrator for the resources of that platform. They cannot make CRUD operations on other platforms’ resources, just manage the ones of their own. Tenant Admin users are allowed to create common users (Tenant Users), which are consequently associated with the respective platform. •Tenant User: These are common users that can only make GETrequests to the resources of the platform with which they are associated. They cannot make CRUD operations in ACL nor access the other platforms. System users and the ACL are stored and managed in the MongoDB database, supported on the npm ACL package 2-NodeACL-AccessControlListsforNode. This module offers a minimalist ACL implementation, providing methods that help to create roles and permissions. The following table describes the system policies: 2https://www.npmjs.com/package/acl 63 Table 4.4: RS Access Control List: existing roles, resources and permissions. As it can be observed, Tenant Users are not allowed to access Users and ACL resources. On the other hand, Tenant Admins are not allowed to edit system users information nor create or remove ACL elements (roles, resources and permissions). The System Admins are able to see the roles and permissions of all users. They have full access to the ACL, which allows them to manage all system’s policies. Resources/Roles System Admin Tenant Admin Tenant User /recommender *GET/POST PUT/DELETE - /users *GET/POST DELETE - /acl *GET PUT - /clients *GET/POST PUT/DELETE GET /products *GET/POST PUT/DELETE GET /order-items *GET/POST PUT/DELETE GET /hybrid *GET/POST PUT/DELETE GET /popularity *GET/POST PUT/DELETE GET /similar-products *GET/POST PUT/DELETE GET /complementary-products *GET/POST PUT/DELETE GET SystemAdminoperates at the /api/resource level, where the platform is not specified and the default is assumed. TenantAdminand TenantUseroperate with endpoints in the /api/${platform} /resource format, where it is necessary to specify the platform to which they are trying to access, having only access to the one they are associated with. The existing resources are listed in the API Documentation section. Note that only System and Tenant administrators are allowed to access platform engine recommendations configurations (/recommender), as Tenant Users are not. In practice, this table is translated into JSON documents and later stored in the MongoDB collections, thus dictating which operations are allowed for each role to perform on each system resource -policies. All system users are stored in a single collection, with a compositeprimarykeyformed by the user’s email and platform. Thus, a user can associate their email on more than one platform with different roles. 64 Taking into account the user hierarchy, we can draw the following process for the creation of users: 1. SystemAdminis registered directly in the database; 2. SystemAdminlogs in to the service and registers another user with one of the remaining roles; 3. TenantAdminlogs in and has the possibility to register a TenantUser. To develop the login feature, an authentication mechanism was created using JWT (JSONWeb Tokens(57)) to transmit information in a secure way between entities as a JSON object, since it can be verified and trusted because it is digitally signed. This was implemented through the npm jsonwebtokens3package to exchange information about users in requests between the system and tenants. Whenever a tenant logs in, and after verifying the login credentials (email and password), the system creates an authentication token to return as the response. In the token’s payload, it is stored the user’s identifier, their role and platform. At the end of this process, that token is signed with the server’s secret key and encrypted. To increase security, the token is defined with a 2 hours expiration time, preventing the abuse of its use indefinitely in case a third party can access a valid token. The properties of authentication, integrity, non-repudiation and confidentiality are therefore maintained. Both authentication and authorization mechanisms were incorporated into middlewares installed on the system’s server to filter requests received from tenants. Access to platforms, according to roles, is validated in the authentication phase, while access to the respective resources is verified in the authorization phase. The authentication middleware sets the system endpoints as private, so that only authenticated users can reach them. It verifies if a user has access to an endpoint, by validating the token received from requests. The middleware fetches the authorization token from request’s cookies or authorization header, verifies and decodes it according to the server’s secret key and, after this validation, it stores the user data from the token into a session, to pass it to the next middleware (authorization). In turn, the authorization middleware protects the resources of the system, by checking the user role permissions; If the user has the SystemAdminrole, then they have access to all resources of the system. Otherwise, they only have access to the resources of the platform they are associated with. It takes the user identifier from the session and checks their permissions according to the ACL, by taking into account the resource they are trying to access and the method of the request. This way, authorization property is established. 3https://www.npmjs.com/package/jsonwebtoken 65 As for the rest of the architecture’s components, they already have security tools integrated, thus being sufficient creating users with secure login credentials for direct access to these elements: Elasticsearch, Kibana and MongoDB. As for the availability property, this is also ensured by the existing mechanisms of high availability of these databases. Furthermore, as an extra security measure, alternative TCP ports were opened, instead of the default ports, for the system’s docker containers on the host machine in order to make the resources available through the Recommender API in a more secure way. 4.6 Summary Having identified which technologies to use, which best fit the context of the project and requirements of the host machine, we proceeded to the planning and design of the types and structure of product recommendations. Therefore, four types of recommendation were defined - Popularity, Hybrid,Similar-productsand Complementary-productswith three entities being assumed throughout the recommendation system: Client,Productand Order-item. These recommendations will be produced in offlinemode, which means that the recommendation engine will generate results, in the background and in cycles, always keeping recommendations relatively up-to-date and available for online stores. To communicate with the recommendation service, Beevo online stores integrate an application, called BIApp, which connects to the RS API and send the necessary information so that the engine can produce recommendations, as well as receive these results to show products to clients. During the project, four scenarios were contemplated to show recommendations on the web pages: Homepagewith popular products, ProductListingwith a new ’Recommended’ ordering option, ProductDetailspage showing products similar to the selected one and Sidecartsuggestions that displays products suited to complete the customer’s purchase. The BIAppalso allows customizing the recommendation models of the system’s engine, while all this communication process is ensured by the authentication and authorization mechanisms of the service. 66 5. METHODS The process of recommending can be understood as a DataMining(75) process, in which correlations between clients and products are discovered, supporting the recommendation of an item to a user. Although this process is of recommendation engine’s responsibility, it is important to contemplate the first phase of Data Mining, which consists of DataAnalysisandPreprocessing: obtaining data, processing it and forming data sets that will be the source of information for engine consumption. As mentioned, obtaining the necessary data for recommendations is of most importance, thus the architectural elements being crucial for the existence of communication. However, not all data can be relevant, so there is a need for filtering to get only significant information, discarding those considered outliers and that can negatively influence the results. The flow of retrieving data from the platforms, filtering and subsequent storage in the databases can be described based on the ETL (71) process (Extraction,TransformationandLoad). 5.1 Data Extraction Strategy Before deciding that data would be received via API, the possibilities of adopting other approaches were considered for the first step - Extractionsuch as using Logstashor Debeziumfor CDC. In the next sections, the conclusions drawn after exploring these alternatives are explained, as well as the decisions that led to adopting an Application Programming Interface. 5.1.1 Logstash Like Kibana, Logstash (76) is another tool from the vast arsenal developed by Elastic, which serves to complement the use of the Elasticsearch database. Logstash is a server-side, lightweight and open-source data processing pipeline that allows to collect 67 data from different sources, transform it in real-time and insert it to any database, in the desired format. Due to its close integration with Elasticsearch, powerful log processing features and multiple plug-ins that can help an application to adapt to most data sources, we tried to use Logstash as a channel for data linking e-commerce platforms to Elasticsearch, rather than using an API. Logstash has a wide range of plugins, which allows an application to consume different types of database, providing great versatility in adapting and extracting information. The streams used in Logstash are better than conventional ETL processes, because they can extract data from multiple different sources simultaneously, faster and efficiently. Instead of this process being divided into 3 phases (extraction, transformation and load), the tool’s streams and filters can execute everything in one step and on-the-fly. All of these features have great potential for the construction of generic architectures. However, to be possible to extract data from Logstash, there must be at least one of the following two cases: either e-commerce platforms grant access credentials to the system, specifically Logstash, to reach its database or provide an API for Logstash to request data through HTTP requests. Furthermore, plugins would have to be configured according to the different types of databases from inside and outside Beevo’s domain and this strategy could lead to some security breaches. Since both scenarios imply dependencies and developments on the part of the platforms and do not present a generic way of obtaining data, this technique was abandoned. Besides, this tool was developed not to receive data, but to collect it within a certain time interval, constantly monitoring the target databases in cycles. However, considering the event-driven side of our architecture, it was declared that it would be more efficient to build a structure that listens to requests, receiving and processing changes only when they occur (triggers). Additionally, in each cycle Logstash always runs through all existing records rather than only new records, making the process more costly and time-consuming. Hence the choice of using an API, in a more economical perspective in terms of computational process and resources consumption. 5.1.2 Debezium (Change Data Capture) While continuing to deepen the hypothesis of applying streams in the data extraction process, the implementation of Debezium (77) tool was studied. This tool, sponsored by Red Hat (78), is a distributed platform capable of transforming platforms’ databases into event streams, allowing the recommender to immediately detect and respond to changes in each entry in the database tables. Debezium is built on Apache Kafka (63) and provides Kafka Connect compatible connectors that monitor database management systems. In addition, it records the history of data changes in the Kafka logs, from where the application consumes them. This makes it possible for the RS to easily consume all events correctly and completely. Even if there is a problem causing the system to crash, when restarted it will begin to consume the events from where it left off, without losing any piece of 68 information. Hence, Debezium is basically a modern, open-source ChangeDataCaptureplatform that will eventually support the monitoring of a variety of database systems. Change Data Capture, or CDC, is an older term for a system that monitors and captures changes in data so that other software can respond to those changes. Through this tool, it would be possible to cover various types of data sources, as it can be implemented for most e-commerce platforms. Yet, with the condition that a user needs to be created with access to the database with certain privileges and that the database must be enabled for binlog. These limitations, coupled with the fact that this tool is still at an early stage of development, determined that it was too risky to design an architecture based on it. With Debezium, e-commerce platforms would have the responsibility to create specific tables with the necessary information for RS consumption, which implies altering the infrastructure and schema of the existing database. With this option, Debezium would be used for data extraction, but the communication and return of recommendations to the tenants would still have to be done through the API. Thus, it was decided to focus the entire communication process on a single element. 5.2 Exploratory Data Analysis (EDA) As mentioned throughout this document, the quality of the data used will be reflected in the accuracy of the recommendations produced, which reinforces the aphorism of this dissertation - information ispower. In order to obtain quality data of the entities considered (Clients, Products and Order-items), there must be a refinement process, resulting from previous analysis and filtering. However, this process must be delicate so that the filtering rules don’t compromise the amount of data extracted: cleaning data can cause the removal of data that appears useless and, consequently, losing potential information. On the other hand, a large amount of data can lead to the existence of wide variations due to unusual values (outliers) and this can prejudice the final results. That is why it is important to do exploration and analysis of the data, as the first step, to ensure a balance between quality and quantity. According to the creator of this process, John W. Tukey, he describes his perspective of EDA, in the book ExploratoryDataAnalysis(EDA)launched in 1977, combining statistical reasoning with the processes of data transformation and exploration: ”Exploratory data analysis can never be the whole story, but nothing else can serve as the foundation stone.” - John W. Tukey (79) 69 6. CASE STUDIES / EXPERIMENTS In order to prove this project’s concept, the recommendation system was applied in practice on Deeply’s online store, a platform developed by the company. The RS was implemented in a stagingenvironment of the platform to carry out several tests and evaluations so its development and deployment in productionenvironments can be justified, in the future. In this chapter, a practical experience made on Deeply e-commerce platform is described, exploring the extracted data from the store and multiple analysis techniques possible with Kibana, as well as the widgets to display recommended products to clients. Finally, the obtained results are explained and discussed, reflecting the overall performance of the system according to the architectural component and highlighting the potential of this data in the Business Intelligence area. 6.1 Experiment setup Deeply is an online store that sells clothing collections and surf equipment. The steps to install the recommendation system on an e-commerce platform can be summarized as follows: 1. Register the tenant in the recommendation service through the web pages directly provided by the system. A login page is presented, which after authenticating, redirects the user to a registration page - to create new users in the system. Only SystemAdminsand TenantAdmins can authenticate themselves on the login page, as they are the only ones allowed to create other users. Note that when a SystemAdmincreates a user, he can indicate the platform and role (TenantAdminor TenantUser) that he wants to associate with the new user, while a TenantAdmincan only create new TenantUsersassociated with its own platform, as defined by the system’s policies - section 4.5. The appearance of these web pages are attached in 76 appendix A.4. 2. After registration, the tenant will be able to get an authentication token, by sending a request to the system’s login endpoint, as described in the API documentation, and insert it in the header of its consequent requests, making them valid and accepted by the service. { ”userId”:”5e4bf76b5be7e4151aeb99ae”, ”userRole”:”SystemAdmin”, ”userPlatform”:”deeply”, ”iat”:1582036952 } Listing 1: Example of a JWT decoded JSON object In the case of Beevo’s online stores, this is possible through the platform’s Back-Office(management area for store administrators), in a specific section for the interaction with the RS, through the BIApp. This section presents a tab to enter the tenant’s credentials that will be used to authenticate in the recommendation service. Figure 6.1: Deeply Back-Office - RS user association to tenant 3. Afterwards, it is necessary to configure the recommendation engine, again through the BI App, in the section entitled for that. Here, all modelparameters mentioned in section 4.3 can be changed and adapted with the desired behavior for the engine. 77 Figure 6.2: Deeply Back-Office - Recommender engine configurations 4. Ultimately, it is necessary to perform an initial population of the Elasticsearch database of the RS, extracting the data previously described from the entities Client,Productand Order-item, so that the engine has the data sets it needs to produce recommendations. The population can be general, cloning information from all entities, or selected ones. In this tab, it is also possible to indicate the optional fields to be extracted from those entities, in addition to the fields selected by default, or even additional attributes for the Productentity - section 5.2.1. Figure 6.3: Deeply Back-Office - Attribute selection and Recommender Population 78 5. At the time of execution defined by a cronjob, the engine will load, process and analyze data from these data sets, train recommendation models according to previously defined configurations and generate recommendations for Deeply’s clients, storing those results in MongoDB at the end of the process. These results will then be available via API. deeply_complementary_productsdeeply_similar_products deeply_hybriddeeply_popularity { _id:'49503076-1f08-4696-b014-91963073f8f3', items:[ { product_id:'60d8a266-d2ac-11e8-852b-bca7105b98e9', support:0.50 }, { product_id:'846hcn77-1631-4858-86a0-41eebfd5b636', support:0.25 } ] } { _id:'2210c94c-76bd-44df-8fdb-f0ab6b070b66', items:[ '9aacf2c7-0933-4254-991a-62e8974afb41', '35f8d8d2-d138-40d4-85a3-8e82e8b9a9d7', 'c478gh6c-de32-443c-b909-b7781dd9936a', 'f66afc3c-7a07-45f8-b549-d5b60781c758' ] } { _id:0, items:[ { product_id:'f779e01a-be88-4be6-9704-0f9248eebf9b', categories:[ 'MEN', '5/3MM', 'WETSUITS' ] }, {...}, ...] } { _id:'11e2278a-8351-459c-aa4b-b01fb966b74b', items:[ { product_id:'3c3bb63c-2a74-4d6c-9047-44598b925857', categories:[ 'WOMEN', 'T-SHIRTS&TOPS', 'PROMOTIONS' ] }, {...}, ...] } Figure 6.4: Deeply MongoDB recommendation documents examples Once the recommendation process has taken place, it is time to show the results to the store’s clients, through the storefront web pages. As described in section 4.4.4, the recommended products are displayed in vitrines, coordinated by widgets developed for this purpose, according to the type of recommendation that is intended to be used on the web page. When a client visits the web page, widgets request recommendations to the service, using the methods available in Facade.php of the BIApp, and after receiving the list of recommended product ids, it fetches the information of the those items in the store’s database and presents them to the client. These results are cached for about 2 hours, so when the client visits the same page it is not necessary to request the same results to the service again and avoid a great page loading time and resource usage. Below are some examples of these vitrines: 79 Figure 6.5: Deeply Homepage - Popularity recommendations example Figure 6.6: Deeply Product Details - Similar-products recommendations example 80 Figure 6.7: Deeply Side-Cart - Complementary-products suggestions example Figure 6.8: Deeply Product Listing - Hybrid recommendations example 81 6.1.1 Data Contract To ensure the steady functioning of the system, it is necessary to ensure good communication between it and the platforms. This is achieved through a datacontract, where it is agreed between tenants and the service which and in what format data should be exchanged between them, guaranteeing that the system receives the necessary information to produce recommendations and that tenants can obtain and use these results. This is a way to ”oblige” both parties to maintain a coherent and flawless communication. The data restrained in the contract reflects the variables pre-selection made in section 5.2.1, considering that, when selecting and filtering data in the extraction process, all stringvariables are in the same language, as this will affect recommendations made by the engine since some recommendation algorithms are based on textual terms. The data contract between the recommendation service and Deeply is available in appendix A.6. 6.1.2 Kibana Data Analysis With the Elasticsearch database filled in, Kibana can be used to explore its content. Elasticsearch has two core data types that can store string data: textand keyword.Textdata type is useful when it comes to product descriptions: if a product description has ”t-shirt made of 100% cotton” and the user searches for the string ”cotton”, that product will appear as a result. On the other hand, the data type keywordis used for exact matches, where the results must precisely match the search terms: in the case of a product being a ”blue t-shirt”, if the user searches only for the ”t-shirt”keyword it is likely that this product will not be returned as a result of that search. This is what makes Elasticsearch an excellent full-text search engine, a feature that may complement this recommendation system in the future. Through the DataVisualizerarea we can obtain statistical information regarding data extracted from Deeply.Deeply platform has a Portuguesedatabase, hence some data values can be presented in that language has they do not have an English translation stored. Appendices A.8, A.10 and A.12 show the information that Kibana offers from Client, Product and Order-item entities’s data, respectively1. Each document corresponds to an entity record. 1The Deeply database used is deprecated and was only used for testing purposes, thus some data may have poor values or even the lack of them. Nonetheless, the engine adapts the recommendation models according to the needs of the online platform. 82 Clients Regarding the Cliententity, it was possible to verify that there are 23,889 documents stored in Elasticsearch, which means that there was the same total of client records on the platform, at the time. Only 5,265 documents (22.04% of clients) contain the gender field, while in other records this field remains null. These values are in Portuguese, so it can be assumed that sr/drvalues refer to the malegender and sra/drato the femalegender: Table 6.1: Deeply clients’ gender distribution Male sr 62.1% 65.4% dr 3.3% Female sra 23.2% 23.7% dra 0.6% Undefined (empty) 10.9% 10.9% About 14.761 (61.79%) of clients have their country associated with their record, allowing to understand in which countries the store has more affluence: Table 6.2: Deeply top 5 country values Country (ISO 3166) Percentage PT 55.6% ES 34% FR 5.7% DE 1.5% IT 0.9% From the translated values in the tables, we can conclude that most of the store’s clients are Portuguese and Spanish men. Note also that several values are not normalized. For example, the locality field is filled in manually by the client at the registration page, so different clients, even belonging to the same locality, can type its designation differently. Due to this type of flaws and inconsistencies, it is necessary that the engine executes a datatreatmentbefore model training and production of recommendations. 83 Products In Deeply platform there are 8,860 products in total, so far, but only 22.1% of them are available (published) in the online store. When analyzing the data through Kibana, we can see the following top of values of the categories field, color and size attributes: Table 6.3: Top 5 values of colors, sizes and categories of Deeply products Color Percentage Size Percentage Categories Percentage 526 30.5% L 16.9% PROMOTIONS 13.3% 504 10.5% S 15.5% MEN 12.8% pink 7.2% M 15.3% CLOTHING 6.8% 508 6.7% XL 11.2% WOMEN 6.7% 510 5.6% XS 8.9% JUNIOR 4.3% Exposing the data in this table, one can easily notice that the store produces mostly black products (color 526) for men between sizes from S to L (agreeing with the previous data analysis), and many of them were in promotions at the time this analysis was made. Through this examination, it is also possible to check the different manufacturers associated with the various products of the store, however in this experimental case there is only one: Deeply. Finally, the parent_id field stands out, in which 29.8% has a value of 0, indicating the number of parent products in the store, that is, the number of products available for recommendation - section 5.2.1. Order-Item As for order-items, there are a total of 37,148 records, but the number of orders placed is obviously smaller, since several order-items can belong to the same order. Hence, it is necessary to have this notion into consideration when analyzing order data. In this case, the use of the aggregation functionality of Kibana was decisive to assist in the analysis of orders, exploited in Business Intelligence Dashboards. However, in the analysis of the individual order-item instances, one can realise that the median price for each product in an order-item (product_total_with_tax) is about 98.75, which is similar to the total price of the products in an order (item_total_with_tax), justified by the fact that the quantity of products in an order-item record being 1, in about 98.7% of cases. With this data, we can already obtain information on the color,size and category fields, that are present in most of the ordered products: 84 Table 6.4: Top 5 values of colors, sizes and categories of ordered products Color Percentage Size Percentage Categories Percentage 526 47.8% M 25% PROMOTIONS 31.2% grey 8.1% L 14.4% MEN 21.4% 504 7.2% S 14.3% WETSUITS 4% dark grey 3.8% MT 8% 4/3 MM 3.8% 508 3.7% MS 7.3% ACCESSORIES 3.2% Translating these values into a table for better visualization, it is easy to understand that most of the products sold are products whose color is black (526 code) for men, with the size between S and L, belonging to the promotions group. Thus, there is an evident correlation with the previous table, being logical that the store invests more in the stock of products that are most sold. About 42.4% of order-items belong to orders addressed to Portugal, 36% to Spain and 12.3% to France. In total, 61.2% were shipped and 28.6% cancelled (status). As can be seen from the examples of the values that Kibana presents for the fields mentioned above, these are not normalized (in the same format), meeting the issue highlighted throughout this section: the recommendation engine is responsible for handling data in such a way that it satisfies the requirements for the application of the ML algorithms and consequent production of recommendations. Besides, as the engine recommends for textual terms, it does not matter whether the fields’ values (such as colors or sizes) are represented by their number or name, as long as consistency is maintained. It is this capacity for abstraction that makes the system so flexible. Furthermore, the information presented in DataVisualizeris updated in real-time, as Elasticsearch receives new data. From a Business Intelligence perspective, the variables previously discussed are quite interesting to analyze in order to obtain relevant information for TargetedMarketing, for example. This potential is demonstrated in the Business Intelligence Dashboards section. 6.2 Results After integration and configuration phases of the recommendation system on the e-commerce platform, several tests were carried out to assess the overall performance of the architecture idealized in this project. These tests were used to identify possible flaws that may exist either in the general functioning of the service, or in its communication with tenants, to raise performance improvements and, finally, to support the proof of concept so that it can advance to production environments. Thus, in this section, the results of the evaluations regarding the interactions with the recommendation 85 Source: Locust web-based UI Figure 6.12: Rate of total requests per secod over time Source: Locust web-based UI Figure 6.13: Response time value over time Comparing the 3 graphs, we can identify that there are some peaks where the response time is higher and, consequently, the number of requests handled is lower, as the number of users increases. This allows to identify several points where bottleneck problems may occur, caused by the immense load made by test users - the limit of connections to the server is sometimes exceeded. This is visible in the table, where a maxresponsetimeof 17 seconds was recorded. In contrast, a 4 ms minresponsetimewas also recorded, which may have resulted from similar and consecutive requests. These failures are accounted for in the previous table, adding to the HTTP 404 - Not Found server-side errors when recommendations are requested for which the recommender has no answer: e.g., product listing with a certain combination of categories or complementary products for a certain item. Obviously, in practical cases, the tolerance to these bottleneck concerns is greater, since after sending a request, the response is cached on the platform side, by the BIApp, so the number of requests will be much lower, for the same number of users. 92 6.2.2 Business Intelligence Dashboards As it has been reinforced throughout this dissertation, it is important to take full advantage of the collected data to obtain knowledge and, thus, allow the definition of strategies to give an economic boost to companies. Hence, the data is not only used to produce recommendations but can also be explored and give much more information about the online store’s general commerce. With Kibana, this can be arranged in various visual representations (such as graphs, tables, etc) and grouped in dashboards, to support business decisions. In this case, a dashboard was created based on the Deeply platform with some of the following graphs: Total Orders by Country From a Marketing perspective, it is interesting and useful to calculate the total number of orders by country, since it is possible to identify and invest in various forms of marketing, such as advertising strategies or promotions, appropriate to the context of each country, and consequently maintain or even increase the number of sales. Figure 6.14: Total Orders by Country - Map Kibana allows to make various types of graphs, it is a matter of deciding which style best suits the information one wants to know. Seeing countries on the map, for example, is more visually appealing. However, a bars graph allows a more objective and effective analysis. As we can see, these values correspond to the analysis made in section 6.1.2. 93 Figure 6.15: Total Orders by Country - Bars graph Average spend by Country In turn, when calculating the average spend by country, we must have in mind that a country that has higher revenue than another does not necessarily mean that the first has more orders. Nevertheless, knowing how much is spent on average for each order by country can show a glance on the buying behavior of clients in that country, and the company can adjust both prices and invest more in product promotions that match the average price spent by each client, in an attempt to increase the number of sales. Figure 6.16: Average spend by Country 94 Average quantity and spend per order It is also important to have graphs that show a global perspective of the stores’s sales, such as the average price or average quantity of items per order, giving feedback to users on how the online store trade is going and if it corresponds to vendor’s expectations and the company’s objectives. Figure 6.17: Average quantity and spend per order Client Genders The Client values previously analyzed in section 6.1.2 can be reflected in a pie chart for better visualization. With this knowledge, companies can make certain decisions according to the gender of their clients, applying this information to assume the right direction in their business and increase profit, combining the BI sector with Targeted Marketing. Figure 6.18: Client genders 95 Promotion Tracking Through the Time Series Data Visualizer (TSDV) feature, it is possible in Kibana to identify possible moments with the potential to carry out promotions. As a hypothetical case, in this test scenario, an alert was defined to notify the owner of the online store that there is a potential chance of making a discount on clients’ purchases whenever the total price of the order (item_total_with_tax) is greater than 400€ for each product of a different category. For example, whenever the store sells 400€ or more of a certain product with X category (X ∈{MEN, WETSUIT, SURF ACCESSORIES, CLOTHING}), a promotion is applied. This chart indicates and maintains the history of when promotions would be most efficient and in what context. Figure 6.19: Promotion Tracking The complete dashboard is attached in appendix A.14, with the remaining graphs and tables: • Sales by Category; • Sold products per Week; • Total Revenue; • Top Selling Products; • Clients table; • Products table; • Order-items table; Most of these visual representations are based on orders from the online store, even though this does not correspond to any entity considered in this project. Nevertheless, it is possible to make observations of orders through the Order-itementity by grouping them using the aggregation feature provided by Kibana. Thus, the obtained results were produced by aggregating order-items forming several ”buckets” in which each one was identified by the order_id of the order-item, thus considering orders as a whole. This brings great advantages since we can have generalized information about the orders from the online store, keeping the detailed data of each order-item. 96 Another feature that makes Kibana such a powerful tool is the ability to update in real-time, as Elasticsearch receives and stores data. Since the Elasticsearch database is synchronized with the database of the e-commerce platform, through the developed triggers, Kibana is able to update the dashboard graphs in real-time. The registration tables inserted at the end of the dashboard, obtained from Kibana’s Discoverarea, operate as logs of the Elasticsearch database, indicating which data was received and when it was received, thus maintaining a history and providing a preview of each stored document. To obtain dashboards with more reliable results, the data must be as complete and normalized as possible, being from tenants’ responsibility since they are the data providers. 6.3 Discussion A microservice approach was adopted instead of a monolithic architecture, since a single application can present data congestion problems (bottleneck) when faced with a scenario where e-commerce platforms are hit by countless users, generating a large number of events/requests. A microservice architecture allowed to build a robust and flexible system, due to the independence of each component, capable of handling requests and managing resources efficiently. This way, the service can serve multiple tenants - multi-tenancy - instead of assigning a monolith instance (copy) to serve each platform - single-tenancy. It should be noted that all requirements raised in section 3.2 are satisfied by the system: 1. The RS returns the results of recommendations in the form of a list of product IDs, in JSON format, as required by Beevo’s e-commerce platforms; 2. It is possible to control the information that comes in recommendation results, indicating the quantity and category of recommended products that should be included in the response of the service. If no limit is imposed, the service returns a maximum of 100 recommended items; 3. RS provides a Kibana component, thus offering several data analysis and business intelligence features; 4. The company guarantees the maintenance of the server machine for at least 361 of the 365 days a year; 5. From the results of the load tests, we can verify that the system fulfils the requirement of responding under 100 ms to most requests, never exceeding 150 ms, even in moments of intense activity; 97 6. The system has authentication and authorization mechanisms, supported by JWT and an ACL, protecting all data involved in the recommendation process; 7. Supports multi-tenancy; 8. Allows the tenant to manipulate the recommendation models of the system’s recommendation engine. All recommendations produced were based on three entities (client, product and order-item), forming four types of recommendation (popularity, hybrid, similar and complementary products), having considered four different scenarios on the web pages (homepage, product listing, product details and side-cart) that meet the system requirements. These results were computed in offline mode, which means that the recommendation system calculated recommendations in the background, saving the results in MongoDB to be made available whenever required by the online stores. In turn, online stores formed data sets in the system’s Elasticsearch, by sending data about the entities, from their database to the service, through the BI App. With this application it was also possible for tenants to manipulate the service, changing the configurations and data used by the recommendation engine. It was decided that this would be the best way to extract data from e-commerce platforms, as an API allows the recommendation system to receive only data suitable for recommendations and to react to events only when necessary. Before the extraction process took place, there was an investigation and pre-selection of which fields should be extracted from the stores’ databases, regarding the entities considered. The accuracy of recommendations depends on the quality of the collected data and there is a need to adapt the models of the recommendation engine according to the available data. Better results could be obtained if there were more suitable variables for producing recommendations, such as: • Sales per Time (SPT): Ratio between the number of sales and the time the product is on sale • Sales per View (SPV): Ratio between the number of times a product has been sold and the number of times the product has been viewed by users on the product listing page • Product Buy Path: Products visited by the user, until his next purchase, since the beginning of a session; • Ratings and Reviews: product ratings and reviews given by users; • Weather conditions: weather conditions and temperature at the time of the purchase at the user’s location These and many other variables could exponentially increase the accuracy of the results or even give a greater capacity of adaptation to the service, extending the scope to real-time recommendations. 98 However, since Beevo’s e-commerce platforms have not yet developed the necessary mechanisms to obtain these variables, it was not possible to apply them. One can also notice that product prices are not considered as a default field. This is due to Beevo’s product pricing logic complexity, in which price values change according to the type of client that is being addressed. Nevertheless, it is possible to add this value in the optional extraction fields in the BI App. The architecture developed in this project, as well as the metrics and results obtained from its performance, are exposed and analyzed in my paper ”ImprovingPerformanceofRecommendation SystemArchitecture” (Appendix I - Publications). There are several suggestions for improving the overall performance of the architecture, such as converting the protocol used in requests from HTTP/1 to HTTP/2 (86): HTTP/2 is more efficient and faster than the first due to multiplexing, header compression and binary formatting capabilities. As for scalability, the company opted for horizontalscaling, from an economic point of view, by adding more machines to the resource pool, instead of adding more power (CPU, RAM, ...) to the existing machine (verticalscaling). This way, it is possible to better manage the resources of each component and even open the possibility to implement a Nginx(87) element to serve as a load balancer and security mechanism in the future. 99 7. CONCLUSION Conclusions Given the growing number of online offers, recommendation systems appear as an effective strategy to combat the multiple decisions and divergent options that users face in online stores. The project’s RS works as a PaaS, containing all the necessary infrastructure and computation for the production of product recommendations. The e-commerce platforms that use this service, on the other hand, only need to worry about communicating with it to receive such recommendations, through their respective applications. Although the system does not include real-time recommendations, the trigger mechanisms implemented support the use of onlinecomputingin the architecture, giving the recommendation engine the possibility to generate recommendations in real-time, if it has the capacity to do so, and consequently allow platforms to present results in real-time, in response to user activities or other events. This means that the progress of both architectural and engine components are interconnected, with the structure depending on the complexity of the engine. The more features the recommendation engine has, more endpoints and management will be required by the architecture to support these features. On the other hand, it is possible to analyze and obtain statistics in real-time from the data collected, through Kibana dashboards, allowing greater control over data management. Thus, data stored by the system not only is used to produce recommendations, but also helps to form new perspectives on the market and support in business decisions - BusinessIntelligence. So we may conclude that all the objectives proposed for this project have been achieved. Although the system is not yet a Beevo’s final product, many of its customers have already shown interest in subscribing and incorporating the recommendation service in their online stores. The developed architecture allowed the integration of a recommendation engine in online stores, obtaining and 100 analysis of data from its clients and products, which will fulfil the company’s objective of collecting more information for the application of various marketing strategies and expanding its Business Intelligence sector. Therefore, there is an expectation of increasing the number of users on the platforms, the number of sales and, consequently, obtaining a greater profit, leading the company to stand out in its market. The technological components used in the architecture allow to combat all the problems foreseen in section 3.1.1, such as: 1. Personalized data sets: The Elasticsearch component allowed the storage of personalized data sets that contain various types of data from different e-commerce entities, thanks to the flexibility provided by the structure of the JSON documents used; 2. Data analysis: Complementing the previous component, Kibana allowed the analysis of the personalized data sets, in order to extract important information either for the application of marketing strategies or in the creation of dashboards to support business decisions; 3. Information up-to-date: Thanks to the triggers developed and implemented on the tenants’ side, it was possible to keep the data sets always up to date and synchronized with the databases of the online stores; 4. Cold start problem: It was also possible, along with the algorithms used by the recommendation engine based on Collaborative and Content-based filtering, to combat the cold start problem by always including the most recent information in the recommendation process; 5. Availability and scalability: The MongoDB component guaranteed, to the system, a database with high availability and scalability necessary to satisfy the requirements raised at the start of the project, so that online stores could always have access to recommendation results; 6. Multi-tenancy and security: The approach of a microservices architecture gave the system the capacity to serve multiple stores simultaneously (multi-tenancy), always maintaining the communication flow protected by JSON Web Tokens and an Access Control List. All the main points of this project are exposed and explained in the article ImprovingPerformance ofRecommendationSystemArchitecture(Appendix I - Publications), where it emphasizes the development of the communication process between the RS and e-commerce platforms, the structure of recommendations and analysis the performance results. 101 [75] X. Amatriain, A. Jaimes, N. Oliver, and J. Pujol, DataMiningMethodsforRecommenderSystems, pp. 39–71. 10 2011. [76] N. Elastic, “Logstash,” 2020. Centralize, transform stash your data. [77] D. Community, “Debezium,” 2020. Stream changes from your database. [78] M. E. Bob Young, “Red hat,” 2020. Red Hat, Inc. is an American multinational software company that provides open source software products to enterprises. [79] J. W. Tukey, Exploratorydataanalysis, vol. 2. Reading, MA, 1977. [80] J. W. Tukey, ThecollectedworksofJohnW.Tukey, vol. 1. Taylor & Francis, 1984. [81] S. Software, “Swagger,” 2020. [82] S. Software, “Openapi specification,” 2020. [83] J. H. J. H. L. H. Carl Byström, Hugo Heyman, “Locust,” 2020. An open source load testing tool. [84] A. Strzelewicz, “Process manager 2.” Url: https://pm2.keymetrics.io/, 2013. [Accessed: 2020-06-05]. [85] O. Foundation, “Node’s cluster api.” Url: https://nodejs.org/api/cluster.html, 2020. [Accessed: 2020-06-05]. [86] N. Ramadan and I. Abdelwahab, “Impact of implementing http/2 in web services,” InternationalJournalofComputerApplications, vol. 147, pp. 27–32, August 2016. [87] I. Sysoev, “Nginx,” 2005. [88] G. Cunha, H. Peixoto, and J. Machado, “Improving performance of recommendation system architecture,” in IntelligentDataEngineeringandAutomatedLearning–IDEAL2020(C. Analide, P. Novais, D. Camacho, and H. Yin, eds.), (Cham), pp. 495–506, Springer International Publishing, 2020. 108 APPENDICES Appendix I - Publications ImprovingPerformanceofRecommendationSystemArchitecture(88) Authors: Cunha G., Peixoto H., Machado J. Editors: Analide C., Novais P., Camacho D., Yin H. Title: Improving Performance of Recommendation System Architecture Book: Intelligent Data Engineering and Automated Learning – IDEAL 2020 Year: 2020 Publisher: Springer International Publishing Pages: 495–506 ISBN: 978-3-030-62365-4 Abstract: The exponential appearance of online stores has implied higher market competitiveness and, consequently, companies need to adopt certain strategies to obtain greater prominence and gain clientele. This paper explores an architectural approach to incorporate a recommendation system in online stores, in order to offer a solution to achieve those goals. Developing the recommendation system infrastructure with NodeJS, based on a REST API, and according to microservices architecture concepts, has proven to be very efficient when it comes to managing great volumes of requests and data, and be capable to serve multiple tenants within a short response time. Clustering techniques were also implemented to increase the system’s performance and capability of handling requests. 109 Appendix II - NodeJS Application Structure Figure A.1: NodeJS Application Structure 110 Appendix III - API Documentation with Swagger Figure A.2: API Documentation with Swagger - part 1 111 Figure A.3: API Documentation with Swagger - part 2 112 Appendix IV - User login page of the Recommender System Figure A.4: User login page of the Recommender System Figure A.5: User registration page of the Recommender System 113 Appendix V - Beevo Business Intelligence Data Contract Figure A.6: Beevo Business Intelligence Data Contract - part 1 114 Figure A.7: Beevo Business Intelligence Data Contract - part 2 115 Appendix VI - Kibana’s data analysis on Clients data Figure A.8: Kibana’s data analysis on Clients data - part 1 116 Figure A.9: Kibana’s data analysis on Clients data - part 2 117