Full text
An Application to Measure Consensus on ng Hesitant Fuzzy Walaa ABUASAKERa , Jennifer NGUYEN b , Núria AGELL b,1, Francisco J. RUIZ a, Mónica SÁNCHEZ a a UPC-BarcelonaTECH b Esade Business School, Universitat Ramon Llull Walaa Abuasaker https://orcid.org/0009-0000-4630-6687; Jennifer Nguyen https://orcid.org/ 0000−0002−7498−7536; Núria Agell https://orcid.org/0000-0001-9264-2147; Francisco J. Ruiz https://orcid.org/0000−0002−4314−101X; Mónica Sánchez https://orcid.org/0000−0002−0201−345X Abstract. In this study, we propose a methodology using hesitant fuzzy linguistic term sets to aggregate rating values. The main idea is to define an integrated methodology for collaborative platforms involving individuals that use the same set of linguistic terms with different semantics. The objective is to fuse the opinions of different profiles of reviewers or customers when assessing an alternative. The proposed approach allows considering previous assessments to give personalized meaning to a set of linguistic terms. To this end, the concepts of linguistic perceptual maps and projections among perceptual maps are introduced to model and aggregate reviewers’ different profiles. Finally, we compute a measure of consensus to capture the agreement among reviewers’ opinions. All these concepts are applied in a real case example of rating books from Amazon to demonstrate the potential of the methodology in facilitating the evaluation and comparison of different alternatives. Using our methodology, the degree of consensus among customers can be used to rate in a more precise way the books before recommendation. Keywords. Hesitant Fuzzy Linguistic Terms Sets, Consensus measurement, Customers’ profiles, Rating scales 1. Introduction Ratings are usually used to assess alternatives and provide users with an insight into the shared opinions. Despite the widespread use of ratings as a means of evaluation, individuals may interpret them differently. Therefore, how can a system standardize the perception of diverse users to accurately interpret their rating scales? Customers' Ratings Usi Linguistic Term Sets Artificial Intelligence Research and Development I. Sanz et al. (Eds.) © 2023 The Authors. This article is published online with Open Access by IOS Press and distributed under the terms of the Creative Commons Attribution Non-Commercial License 4.0 (CC BY-NC 4.0). doi:10.3233/FAIA230698 291
In this paper, we present a new methodology to aggregate assessments and measure consensus coming from a rating context. In general, rating scales do not consider that the assessments may come from individuals with different profiles, and they may be considering differently the meaning of the rating scales [1]. Different linguistic scales are simultaneously considered in this study to represent and aggregate individuals’ ratings taking into account not only their preferences but also their different perceptions [2, 3, 4]. In this way, the meaning of each linguistic term is not being considered the same for all reviewers, but different semantics for the same linguistic terms are considered depending on their profiles [5, 6]. In this study, we introduce the concept of common perceptual map, which provides a shared scale to represent different individuals’ perceptions together [8]. Second, we propose the use of a projection function for the representation of each individual’s linguistic assessment into this common perceptual map [8, 9]. Then, we propose the use of a centroid to aggregate the opinions of individuals. This centroid is obtained in the common perceptual map, where all opinions can be correctly represented. Finally, we consider a measure of consensus which considers the hesitancy and individual perceptions expressed in the common perceptual map. Data from Amazon Books’ recommender system is used to show the applicability of the proposed methodology in facilitating the evaluation and comparison of different books considering two different profiles of customers: lenient and strict. The data set considered for this application contains 3,000,000 reviews from more than 200,000 books. The paper is structured as follows: Section 2 outlines the fundamental concepts and relations of hesitant fuzzy linguistic term sets (HFLTSs) to define the concept of perceptual map, including the notions of centroid and degree of consensus. Following this, Section 3, presents the methodology based on these concepts. In Section 4, an application to a real-world scenario based on Amazon Books’ recommender system is developed. Lastly, Section 5 summarizes the conclusions, challenges, and potential directions for future research. 2. Preliminaries: Common perceptual map framework The general theoretical framework considered in this paper is based on the concept of HFLTSs introduced in [10] that allow users to express hesitancy in their opinions. In addition, in this section, we consider the concept of common perceptual map as introduced in [8, 9], as the framework in which rating values from different users’ profiles can be expressed and aggregated. Definition 1. Let ܵ be a finite totally ordered set of basic linguistic terms ܵൌሼݏଵǡǥǡݏሽ with ݏଵ൏ڮ൏ݏ . A basic linguistic perceptual map is a pair ሺܵǡߤሻ where ߤ is a normalized measure over ܵ, that is, ߤǣܵ՜ܴାsuch that σߤሺݏሻൌͳ ୀଵ Ǥ Remark 1. Note that any basic linguistic perceptual map is associated to a partition ܲ of the unit interval [0, 1] in n sub-intervals of lengths ሼߤሺݏଵሻǡǥǡߤሺݏሻሽ and W. Abuasaker et al. / An Application to Measure Consensus on Customers’ Ratings292
landmarksߣൌͲ൏ߣ ଵ൏൏ߣିଵ ൏ߣൌͳ. The relation of the landmarks and the width of the basic linguistic labels is ߣൌσߤሺݏሻ ୀଵ and ߤሺݏሻൌߣെߣିଵ for any݉ൌͳǡǥǡ݊. The previous definition considers a universal set of basic linguistic terms but different interpretations of these terms, that is, all individuals use the same set of terms but each one has its own perceptual map. Definition 2. ([10]) Given a set of basic linguistic terms ܵൌሼݏ ଵǡǥǡݏሽ, a hesitant fuzzy linguistic term set (HFLTS) over ܵ is a subset of consecutive elements of ܵ, i.e. ܪؠ ൣݏǡݏ൧ൌሼݔאܵȁݏݔݏ ሽ, for some ݅ǡ݆אሼͳǡǥǡ݊ሽ with ݆݅. The set of all HFLTS is denoted as ௌ. In [11] it is shown that this set has a lattice structure if the empty set is included. When reviewers give a rating to assess an alternative, they normally use a precise rating (a basic term in the defined terminology). However, when this assessment is projected onto a common scale, we will see that this rating introduces a certain level of hesitancy. For this reason, the HFLTS framework is very suitable in our approach. Remark 2. Any HFLTS ܪൌൣݏǡݏ൧א ௌ is associated with the subinterval ሾߣିଵǡߣሿكሾͲǡͳሿǡ where ߣ, ݅אሼͲǡǥǡ݊ሽ, are the landmarks corresponding to the partition ܲ to which Remark 1 refers. Definition 3. Given the set ௌ, a linguistic perceptual map is a pair ሺௌǡߤሻ where ߤ is a normalized measure over ܵ and ߤ൫ൣݏǡݏ൧൯ ൌσߤሺݏሻൌߣ െߣିଵ ୀ . From now on, we call ߤሺܪሻ the width of ܪ. The width of ܪ match the length of the real interval associated. Definition 4. Let ሺௌǡߤሻǡ݉אሼͳǡǥǡ݇ሽ a set of݇ linguistic perceptual maps with ܵൌ݊, the common perceptual map, ൫ௌೆǡߤ൯ǡis the linguistic perceptual map associated to the partition of the unit interval ܲ of landmarks ڂڂ ሼߣ ሽ ୀ ୀଵ . Definition 5. If ܪൌሾݏ ǡݏሿא ௌ, there exists a unique ܪᇱאௌೆ such that it has the same subinterval associated, ሾߣିଵ ǡߣ ሿكሾͲǡͳሿ. We call ܪԢ the projection of ܪ into ௌೆǤ Definition 6. Let ሺௌǡߤሻ be a linguistic perceptual map. Given ܪଵǡܪଶא ௌǡthe perceptual-based distance between ܪଵand ܪଶ is defined as: ܦఓሺܪଵǡܪଶሻൌʹߤሺܪଵّܪ ଶሻെߤሺܪଵሻെߤሺܪଶሻ (1) where ّ denotes the connected union, that is, the less element of ௌ to respect to the inclusion relation that contains both ܪଵ and ܪଶ. It is easy to prove that this definition gives indeed a distance in ௌ. W. Abuasaker et al. / An Application to Measure Consensus on Customers’ Ratings 293
Definition 7. Let ሺௌǡߤሻ be a linguistic perceptual map. Let ሼܪଵǡǡܪሽؿ ௌ, the centroid of this set, denoted as ܪ, is defined as the element in ௌ that minimizes the addition of distances to all the HFLTSs in this set: ܪൌܽݎ݃ ுאೄܦఓ൫ܪǡܪ൯ ୀଵ (2) When the set of ܪܨܮܶܵݏ corresponds to the assessments provided by a group of DMs, the centroid is considered as the central opinion of the group of DMs. Note that, when ݇ is an odd number, ܪ is unique. Nevertheless, when ݇ is an even number, there can be more than one HFLTS satisfying (2); then, in each application, an extra criterion must be used to select a specific centroid. Finally, in this context, we define the degree of consensus of a set of opinions derived from the addition of all the distances between the opinions and the centroid. Definition 8. LetȦ be an object to be assessed using HFLTS's from ሺௌǡߤሻ . Given a group of DMs, ܩൌሼ݀ ଵǡ݀ଶǡǤǤǤǡ݀ሽǡlet ሼܪଵǡǡܪሽbe their assessments of Ȧand ܪthe centroid. The degree of consensus of the group is defined as: ߜ௸ሺܩሻൌͳെσܦఓ൫ܪǡܪ൯ ୀଵ ߞ (3) where ߞis a normalization term assuring ߜஃ is between Ͳ and ͳ. 3. Methodology: Aggregating rating values from different customers’ profiles In this section, we present a methodology to fuse the opinions of different profiles of customers or reviewers when assessing alternatives in a certain context. The methodology is a reduced version of the methodology introduced in [13], and it is inspired by a previous work in which the fusion of ratings was considered in evaluations obtained from recommender systems [12]. The methodology follows four steps: Step 1. Define different profiles of customers: We consider that different customers interpret the same linguistic terms in different ways. A profile is a set of customers that share the semantics for these set of terms. Each profile has its own linguistic perceptual map (ௌǡߤሻǤ Step 2. Obtain the linguistic perceptual map per each profile: For the profile ݉, ߤሺݏሻ is calculated, from ݅ൌͳǡǥǡ݊ , as the relative frequency with which profile ݉ assigns the basic label ݏ in other assessments. It is reasonable that if the individuals of the profile use very often a determined basic label, the width of this basic label will be large. Once the set of widths has been determined, the landmarks of the associated partition ሼߣǡǥǡߣሽ are computed as detailed in Remark 1. W. Abuasaker et al. / An Application to Measure Consensus on Customers’ Ratings294
Step 3. Obtain the common perceptual map: Once a linguistic perceptual map has been determined for each customers’ profile, the common perceptual map, (ௌೆǡߤሻ, is calculated whereby all the profile perceptual maps are unified. Then, all assessments are projected from its own linguistic perceptual map to the common perceptual map. Step 4. Compute the centroid and the degree of consensus in the common perceptual map: After the common perceptual map has been calculated, the centroid, ܪ, and the degree of consensus, ߜ௸ሺܩሻ, are computed in this common space. 4. An application to on-line books rating 4.1. Data set description The data1 set used to illustrate the proposed methodology consists of two tables, one for books and another for reviews. Information about the books includes the title and category while the review information includes the reviewer id, rating, review text, and time it was written. To better exemplify the methodology, we focus on the ratings. After a cleaning process of dropping missing values and duplicates (we left the most recent review if the same reviewer reviewed the same book), there are 2,397,614 reviews, 216,023 books and 1,008,972 reviewers. Then, we filtered for reviewers who had written at least 50 reviews to obtain the profile of each one. There were 2,656 reviewers satisfying this constraint who reviewed 81,382 books. At the end, only 2,060 books reviewed by at least 25 reviewers have been considered. 4.2. Experimental approach In this case, a book may be rated on a scale from 1 to 5 stars, with each star representing a particular level of satisfaction. The set of basic linguistic terms in the context of this case is ܵൌ ሼݏଵǡݏଶǡݏଷǡݏସǡݏହሽ. We follow the four steps proposed in Section 3 to find the centroid for each book and to compute the degree of consensus among reviewers in the common perceptual map. First, according to step 1, we divided reviewers into two groups or profiles based on the frequency with which they rated books with a rating of 5 stars. We have named the two groups of reviewers as: lenient and strict. Those who rated books with 5 stars more than 60% of the time were considered to be lenient and the remaining reviewers were considered strict. There were 1595 reviewers in the strict group who had written 191,610 reviews. In the lenient group there were 1061 reviewers who had written 133.552 reviews. 1The dataset is obtained from Kaggle.com and it is built from two sources: x https://jmcauley.ucsd.edu/data/amazon/ Citation Ups and downs: Modeling the visual evolution of fashion trends with one-class collaborative filtering R. He, J. McAuley WWW, 2016 Image-based recommendations on styles and substitutes J. McAuley, C. Targett, J. Shi, A. van den Hengel SIGIR, 2015 x google books API. W. Abuasaker et al. / An Application to Measure Consensus on Customers’ Ratings 295
Second, as presented in the step 2, we computed the relative frequency of the ratings for each group separately in order to obtain the perceptual map for each profile and its respective landmarks. These values are presented in Table 1. We also computed the perceptual map and its respective landmarks where different profiles are ignored. Table 1: Relative frequencies and widths of basics and landmarks of the partitions for the two profiles: lenient and strict and also without profiling. # scores ݏଵ ݏଶ ݏଷ ݏସ ݏହ lenient 1431 1861 5872 21332 103056 strict 4961 10789 32424 71342 72094 total 6392 12650 38296 92674 175150 width of basics ߤଵ ߤଶ ߤଷ ߤସ ߤହ lenient 0.011 0.014 0.044 0.160 0.772 strict 0.026 0.056 0.169 0.372 0.376 no profiles 0.020 0.039 0.118 0.285 0.539 landmarks ߣଵ ߣଶ ߣଷ ߣସ ߣହ lenient 0.011 0.025 0.069 0.228 1.000 strict 0.026 0.082 0.251 0.624 1.000 no profiles 0.020 0.059 0.176 0.461 1.000 Third, from the group linguistic perceptual maps we obtained the common perceptual map considering the union of both sets of landmarks. Figure 1 shows the graphic representation of the perceptual maps corresponding to the two profiles and the common perceptual map. In addition, the perceptual map where different profiles are ignored is represented at the bottom of Figure 1. Figure 1. Graphic representation of perceptual maps for the two profiles: lenient and strict, the common perceptual map and, at the bottom, the perceptual map where no profiles are consideredǤ The cardinals of ܵ for lenient, ܵ௦for strict and S (where no profiles are considered) are 5, while the cardinal of ܵ is N = 9 in this case. Following the last step of the methodology, the reviews were projected to the common perceptual map and the centroid and degree of consensus were computed for each book using these projections. In addition, we calculated the centroids and degree of consensus in each original perceptual map before the projection. At the end, we calculated the centroids and degrees of consensus taking into account all reviewers without differentiating them by their profiles. W. Abuasaker et al. / An Application to Measure Consensus on Customers’ Ratings296
4.3. Results and discussion First, before considering the projection of the reviews to the common perceptual map, we consider the distribution of centroids per each book separated by reviewer profile. This distribution is showed in Figure 2. Note that the centroid of a set of basics labels is always a basic label. a) b) Figure 2. Distribution of the centroids separated by profile a) strict, b) lenientǤ Figure 3 shows the distribution of centroids considering all reviewers together: on the left, the centroids in the common perceptual map (considering different profiles), on the right, the centroids where different profiles are ignored. As we can observe in Figure 2, most of the book centroids in both the lenient and strict profiles have rating of 5 stars and some others 4 stars, making them very difficult to differentiate. In contrast, in the common perceptual map, the centroids are spread across multiple ratings and less concentrated in anyone rating (see Figure 3a). The comparison of the distribution of centroids in the common perceptual map and in the perceptual map where no different profiles are considered accentuates this difference (see Figure 3b). a) b) Figure 3. Distribution of the centroids a) in the common perceptual map, b) where no different profiles are taken into accountǤ Finally, for each book we compute the degree of consensus following (3), considering separately both profiles: lenient and strict (Figure 4) and considering both profiles together but using, on the one hand, the common perceptual map and, on the other hand, without using the profiles of the reviewers (Figure 5). W. Abuasaker et al. / An Application to Measure Consensus on Customers’ Ratings 297
Figure 4. Box plot of distribution of degrees of consensus of reviewers separated by profile. Figure 5. Box plot of distribution of degrees of consensus calculated on the common perceptual map and in the case where different profiles are ignored. We can observe that the degree of consensus is greater in the lenient profile than in the strict one. On the other hand, the variability of this degree is also greater for the lenient profile. More interesting is the comparison when all reviewers are considered together. Taking into account their different profiles and working in the common perceptual map, the degree of consensus is greater and has less variability than in the case where the different profiles are ignored. 5. Conclusions and future work We present an application of the use of linguistic perceptual maps as a tool to represent ratings. We consider a methodology able to aggregate assessments coming from multiple linguistic scales. The obtained results demonstrate the suitability of the methodology in a rating context where individuals may assign different meanings to available linguistic labels. A real case example with data from Amazon Books’ recommender system is provided to show the efficiency of the proposed methodology in facilitating the evaluation and comparison of different books. W. Abuasaker et al. / An Application to Measure Consensus on Customers’ Ratings298
Using our methodology, results show that a more discriminating aggregated rating is obtained. On the one hand, the centroids distribution is more accurate (less biased). On the other hand, the presented methodology further refines the concept of consensus since agreement does not necessarily imply that reviewers have given the same linguistic evaluation and disagreement does not necessarily mean that reviewers have given different linguistic labels. As future work, from a theoretical point of view, we will analyze the use of other aggregation functions to define the central opinion of a group of individuals’ assessments, and we will study their mathematical properties. In addition, we plan to define other ways to deduce individuals’ linguistic perceptual maps, such as text reviews or non-linguistic aspects such as the time spent to give the review. Finally, from a practical and business point of view we intend to explore two key areas. Firstly, we will investigate in the development of an interpretability function, capable of translating the outcomes generated by the common perceptual map into each of the distinct initial linguistic perceptual maps. Secondly, we will focus on the practical application by considering a broader range of profiles and incorporating the sentiment analysis extracted from written reviews into the definition of the linguistic perceptual map for each profile. Acknowledgement This research has been partially supported by the PERCEPTIONS Research Project (PID2020-114247GB-I00), funded by the Spanish Ministry of Science and Information Technology. References [1] Agell, N., Sánchez, M., Prats, F., & Roselló, L. (2012). Ranking multi-attribute alternatives on the basis of linguistic labels in group decisions. Information Sciences 209, 49-60. [2] Martínez, L., Ruan, D., Herrera, F. Computing with words in decision support systems: An overview on models and applications, International Journal of Computational Intelligence Systems 3 (2010) 382–395. [3] Liao, H., Xu, Z., Herrera-Viedma, E., Hesitant fuzzy linguistic term set and its application in decision making: A state-of-the-art survey., Int. J. Fuzzy Syst. 20 (2018) 2084–2110. [4] Wu, Z., Xu, J. A consensus model for large-scale group decision making with hesitant fuzzy information and changeable clusters, Information Fusion 41 (2018) 217–231. [5] Chiclana, F., Mata, Pérez, L. G., Herrera-Viedma, E. Type-1 owa unbalanced fuzzy linguistic aggregation methodology: Application to Eurobonds credit risk evaluation, International Journal of Intelligent Systems 33 (2018) 1071–1088. [6] Chen, Z.-S., Liu, X.-L., Rodríguez, R. M., Wang, X.-J., Chin, K.-S., Tsui, K.-L., Martínez, L. Identifying and prioritizing factors affecting in-cabin passenger comfort on high-speed rail in china: A fuzzy-based linguistic approach, Applied Soft Computing 95 (2020) 106558. [7] Porro, O., Agell, N., Sánchez, M., Ruiz, F.J. (2022). A Multi-perceptual-Based Approach for Group Decision Aiding. In: Torra, V., Narukawa, Y. (eds) Modeling Decisions for Artificial Intelligence. MDAI 2022. Lecture Notes in Computer Science, vol 13408. Springer, Cham. [8] Porro, O., Agell, N., Sánchez, M., & Ruiz, F. J. (2021). A multi-attribute group decision model based on unbalanced and multi-granular linguistic information: An application to assess entrepreneurial competencies in secondary schools. Applied Soft Computing, 111, 107662. [9] Liao, H., Xu, Z., & Zeng, X. J. (2014). Distance and similarity measures for hesitant fuzzy linguistic term sets and their application in multi-criteria decision making. Information Sciences, 271, 125-142. W. Abuasaker et al. / An Application to Measure Consensus on Customers’ Ratings 299