Modeling the opinion dynamics of superstars in the film industry
Abstract
This research has been developed within the R&D project CONFIA (PID2021-122916NB-I00), funded by MICIU/AEI/10.13039/5011000 11033/ and FEDER, EU. J. Giráldez-Cru is also supported through the Juan de la Cierva program (IJC2019-040489-I). Funding for open access charge: Universidad de Granada / CBUA.
Full text
Expert Systems With Applications 250 (2024) 123750 Available online 28 March 2024 0957-4174/© 2024 The Author(s). Published by Elsevier Ltd. This is an open access article under the CC BY-NC license (http://creativecommons.org/licenses/bync/4.0/). Contents lists available at ScienceDirect Expert Systems With Applications journal homepage: www.elsevier.com/locate/eswa Modeling the opinion dynamics of superstars in the film industry Jesús Giráldez-Cru a,c,∗, Ana Suárez-Vázquez d, Carmen Zarco b,c, Oscar Cordón a,c aDepartment of Computer Science and Artificial Intelligence (DECSAI), University of Granada (UGR), Spain bDepartment of Marketing and Market Research, University of Granada (UGR), Spain cAndalusian Research Institute in Data Science and Computational Intelligence (DaSCI), University of Granada (UGR), Spain dDepartment of Business Management, University of Oviedo (UNIOVI), Spain ARTICLE INFO Keywords: Film industry Movie superstars ranking Opinion dynamics Bounded confidence Repulsion mechanisms ABSTRACT One of the most challenging questions in the film industry is to rank superstars, which ultimately affects some performance indicators like movie success. In this work, we address this question by means of opinion dynamics models, where the evolution of opinions in a population is analyzed. We apply a model of this kind to study the evolution of opinions about a set of well-known movie superstars in a real-world population. Also, we use real-world data from a specialized cinema website to model mass communication processes (representing film releases and their related news and marketing campaigns), and to measure the performance of our model. Our results show that the proposed model is able to accurately represent this complex system, where the opinion dynamics of superstars are mostly driven by emotional mechanisms, and reveal that film releases and their corresponding marketing campaigns only have a short term effect on those opinions. To the best of our knowledge, this is the first work that applies opinion dynamics models to the study of opinions about superstars in the film industry. 1. Introduction Being able to predict a product’s success prior to launch has always been a key concern for marketers (Hauser et al.,2006). Estimating the performance of a new product is especially challenging in creative industries, such as the motion picture industry (Karniouchina,2011), as traditional techniques for assessing audience response are unable to capture the full experiential consumption aspects of movies (Eliashberg & Sawhney,1994). Previous research suggests that movie success is a complex phenomenon resulting from the interplay of movie characteristics (e.g., star cast), post filming studio actions in terms of communication and distribution activities, and external factors such as critics’ reviews or word of mouth recommendations (Hennig-Thurau et al.,2006). Digitalization has increased competition in the audiovisual content market and determining the value of each new film has become even more essential now (Kübler et al.,2021). Among the myriad factors that influence the success of a film, one that has attracted particular attention in both the industry and academic literature is the value of superstars (De Vany & Walls,1999). To this end, and in contrast to dichotomous approaches that attempt to determine the value of actor and actresses using dummy variables (star/non-star), superstar rankings have proven to be more effective (Nelson & Glotfelty,2012). Thus, the motivation of this work is ∗Corresponding author at: Department of Computer Science and Artificial Intelligence (DECSAI), University of Granada (UGR), Spain. E-mail addresses: [email protected] (J. Giráldez-Cru), [email protected] (A. Suárez-Vázquez), [email protected] (C. Zarco), [email protected] (O. Cordón). to address one of the most challenging questions in the film industry: how to establish a ranking of superstars. Previous methodologies employed to address this question have led to mixed conclusions about the effect of superstars on films’ success and the results of different rankings are difficult to compare (Ghiassi et al.,2015). Moreover, these rankings are based primarily on the box office results of stars, ignoring the fact that movies are experiential products whose consumption is driven by emotions (Hennig-Thurau et al.,2007). In the current contribution, we analyze the role of emotions about movie superstars from the lens of opinion dynamics (OD). OD models use agent-based modeling (Epstein,2006;Farmer & Foley,2009) to analyze the evolution of opinions in a population (Dong et al.,2018;Wang et al.,2020;Xia et al.,2011). They mostly rely on an opinion fusion rule which defines how each agent’s opinion is updated after an interaction (representing, for instance, a word-of-mouth process between agents, or a mass communication broadcast) (Noorazar et al.,2018). OD models have been successfully used to analyze many problems, including the study of biased opinions (Anagnostopoulos et al.,2022), multi-attribute group decision-making (Li et al.,2021), the identification of opinion leaders in social networks (Chen et al.,2021), the gap between people’s voting result and their collective opinion (Jiao & Li,2021), and https://doi.org/10.1016/j.eswa.2024.123750 Received 21 April 2023; Received in revised form 12 February 2024; Accepted 18 March 2024
Expert Systems With Applications 250 (2024) 123750 2 J. Giráldez-Cru et al. the consensus reaching with trust evolution in social network group decision-making (Zhang et al.,2022), among others. One of the most studied OD models is bounded confidence (BC) (Deffuant et al.,2000;Hegselmann & Krause,2002), where agents’ opinions evolve as a consequence of interactions between agents with similar opinions (i.e., agents in their confidence area). BC has been used to explain the OD reaching both consensus and fragmentation of opinions (Castellano et al.,2009). Nevertheless, this model is unable to capture some kinds of emotions that drive opinion evolution in certain systems, e.g., extremization (Isenberg,1986). In contrast, the Agentindependent Time-based Bounded Confidence and Repulsion (ATBCR) model (Giráldez-Cru et al.,2022) has been recently proposed to overcome this inconvenience by means of an extension in the classical BC model based on a repulsion mechanism. In fact, this model has been already applied to model opinions in marketing campaigns (Giráldez-Cru et al.,2022). Therefore, OD models arise as an alternative modeling strategy to represent the ranking of superstars. However, and due to the experiential nature of movies, we conjecture that emotional mechanisms (as the repulsion rule of the ATBCR model) are required to capture the underlying OD about superstars. Emotional mechanisms are the cultural and adaptive process by which individuals react to environmental contingencies in a flexible and dynamic way (Scherer,2009). For example, the process by which an individual change his/her opinion of a superstar due to the divergence of this opinion from that of others is not rational and emotions play a role. Consequently, we adapt the ATBCR model of OD to study the evolution of opinions about a set of well-known superstars in the real-world population from Suárez-Vázquez (2015) and benchmark the obtained model with others resulting from the use of alternative OD approaches. To better capture the real dynamics of the phenomenon, we use real-world data both from a cross-sectional survey, to properly model the spectators’ opinions, and from the specialized cinema website IMDb,1to enrich the model incorporating mass communication processes. These processes represent an information exchange to a large portion of the population, including news and marketing campaigns, for instance. In our case study, these mass communication processes represent film releases and all the related communication generated for these events. In addition, we also use real-world data from this website as a ground-truth to measure the accuracy of our model, which allows us to validate its outputs with an external, independent source of information. To the best of our knowledge, this is the first work that applies an OD model using real-world data to the study of opinions about superstars in the film industry. In summary, these are the main contributions of our work: •We present an innovative study of film performance indicators using opinions about superstars. •Our analysis is based on an interaction-based OD model to rank a set of well-known superstars in the film industry, fed with data from a survey of spectators. •We introduce a mechanism of mass communication, which is used to represent film releases and related news. These mass communication processes affect opinions about those superstars considering real-world data from a specialized cinema website. •We present an extensive experimental analysis comparing the accuracy of different OD models using the said real-world data. •Our experimental results show that the OD of this system are mostly driven by emotional mechanisms. The rest of this manuscript is organized as follows. Section 2motivates the present work and describes other related references. In Section 3some preliminaries on the ATBCR model are defined, along 1https://www.imdb.com/ with a brief description of some classical OD models which will be considered alternative approaches to modeling the spectators’ opinions on the considered problem. Section 4is devoted to the description of the adaptation of our model to cope with the real-world phenomenon tackled, including how real-world data is extracted and used. The experimental analysis is developed in Section 5. Finally, we conclude in Section 6. 2. Related works The celebrate quote of Rita Hayworth ‘‘men go to bed with Gilda but wake up with me’’ represents the strength of the movie stars-spectators emotional relationship. Indeed, previous research has described stars as emotional competent objects (Luo et al.,2010) and identified the key emotional drivers of their influence (Moraes et al.,2019). This emotional dimension explains the difficulties associated with determining the origins of stardom (Adler,1985;Rosen,1981), the so-called power of stars (Elberse,2007). It is related to the experience good property of movies (Chang & Ki,2005) which underlines the importance of the psychological approach (Eliashberg et al.,2006), that is, focusing on the individual moviegoer’s decisions, when analyzing the success factors of movies (Hadida,2009). 2.1. Emotional dimension of superstars: secondary data Film is the industry of making-believe by interacting with the human emotional system (Tan,2013). It provides stimuli that influence spectators’ emotions (Zillmann,2015). One of these forces are the superstars that have been positioned as personalities with boxoffice power (King,1987). The source of such power has raised a great interest in previous literature from the original explanations of Rosen (Rosen,1981), Adler (Adler,1985), Frank and Cook (Frank & Cook,1995), and Borghans and Groot (Borghans & Groot,1998) to more recent proposals of applied nature (Harashima,2016). Superstars are one of the few tangible features of film quality (Eliashberg et al.,2000;Ravid,1999) and their mere presence benefits the promotional opportunities of the film (Suárez-Vázquez,2011). Thus, the actress Sarah Bernhardt is frequently entitled as the world’s first superstar precisely for her ability to anticipate the importance of image and buzz (Isaac-Goizé,2023). Stardom is not only a cultural reality but a commercial one founded on the marketability of human identities (McDonald,2012). Behind this phenomenon is the distinction between the artistic and commercial value of the star, that is, stars’ popular appeal vs. expert judgments about the stars (Holbrook,1999). Capturing popular appeal is not an easy endeavor. At the film level, online user-generated information is a valuable source of spectators’ awareness and feelings. During the film pre-release period, star power affects the volume and valence of the online conversations about the film (Liu,2006). In fact, the cast of the film is one of the marketing variables with potential impact on spectators’ decisions (Eliashberg et al.,2000). To approximate the value of the cast, academic researchers have relied on secondary sources based on (Elberse,2007): industry magazines (Sawhney & Eliashberg,1996); ratings from members of the industry (Ainslie et al.,2005), and previous awards and box-office success (Ravid,1999). Actually, the accessibility of data is behind the huge increase in published research related to the film industry in the last decade (Behrens et al.,2021;McKenzie,2023). Although the secondary nature of this data makes easier to conduct large-scale studies, it does not provide information about moviegoers’ emotions towards the stars. Other kind of data, such the one that result from primary cross-sectional methods of research, can measure the emotional dimension of superstars from an individual perspective.
Expert Systems With Applications 250 (2024) 123750 3 J. Giráldez-Cru et al. 2.2. Emotional dimension of superstars: survey data Star power is one of the key film-specific attributes examined in previous literature on the demand side of the movie market (McKenzie,2023). Existing research has pointed out that focusing on audience information processing is crucial to fully understand superstars power (Hofmann et al.,2017). In this sense, Suárez (Suárez-Vázquez, 2015) conducted an empirical study collecting data from a crosssectional survey for a set of 17 superstars. Due to the relevance of the young and highly educated segment in the composition of film audience (Terry et al.,2010), the population under study in that survey were people between 24 and 34 years old, with university studies in an European country. The sample was gender balance (51.2% female), mean age was 22.5 (s.d. 2.6) and 63.7% of the respondents went to a movie at the cinema occasionally (less than one a month). This dataset provides a homogeneous population in terms of age and educational level, which minimizes noise, i.e., the influence of extraneous variables (Peterson,2001). It also allows reducing possible response bias in terms of, for example, acquiescence (Rammstedt et al.,2013). Besides, this data was coherent with the theatrical attendance demographics provided by the Motion Picture Association (MPA) (2021). Thus, from a typological point of view, the considered dataset represents the most significant segment of the cinema market (Broekhuizen et al.,2011; Cuadrado & Frasquet,1999;Díaz et al.,2018). The final sample included 5,440 responses from 320 spectators. This signifies, for a level of confidence —data accuracy— of 95%, a margin of error —data precision— of 5%, which is considered acceptable in social research (Taherdoost,2017). Furthermore, this margin of error decreases to 1% if the full survey dataset is considered. The following information was collected for each spectator and each superstar in the data sample: (1) emotions elicited by each of the 17 superstars; (2) intention of watching a film starring each of the 17 superstars measured on a scale from 1 to 5 with anchors 1 ‘‘The presence of this star in the cast of the film will not encourage me to watch the film at all’’ and 5 ‘‘The presence of this star in the cast of the film will be a very important stimulus for me to watch the film’’. Following (Laros & Steenkamp,2005), the emotions experienced for each of the 17 superstars were measured across eight basic emotions: anger, fear, sadness, shame, contentment, happiness, love, and pride. Spectators were asked the degree to which they felt each of these emotions for each of the superstars considered on a scale from 1 to 5 where 1 meant ‘‘I do not feel this emotion at all’’ and 5 stood for ‘‘I feel this emotion very strongly’’. Among the insights drawn from this study, the following seem particularly relevant in the context of the current paper: (1) when discussing the success of movie stars, the impact of emotions of different valence does not necessarily differ (i.e. the influence on the intention to watch a movie of both happiness and sadness evoked by a star are positive, despite their different valence); (2) changes in positive valence emotions have a greater impact on moviegoers intentions than changes in negative valence emotions; and (3) the degree of substitutability among superstars is affected by their emotional profile, superstars with a similar emotional profile are more substitutable in terms of predicted effect on spectators’ intentions. 3. Preliminaries on opinion dynamics models In this section, we provide a general summary of OD models, followed by a brief overview of both the ATBCR model (Giráldez-Cru et al.,2022) and of some classical OD models, which are used along the experimental analysis. In every case, we use standard notation in OD models (Dong et al.,2018). 3.1. Models of opinion dynamics The goal of OD models is to study the dynamics from a set of initial opinions to a set of final opinions during a lapse of time, and to understand how this set of final opinions is achieved. In OD models, a population of 𝑁agents interact in a social network.2This social network is represented by a graph 𝐺(𝑉 , 𝐴), where 𝑉 is the set of nodes (with |𝑉|=𝑁), and 𝐴is the 𝑁×𝑁adjacency matrix (i.e., 𝐴𝑖𝑗 = 1 iff. there is an edge between nodes 𝑖and 𝑗in 𝐺,𝐴𝑖𝑗 = 0 otherwise). Each agent is represented by a distinct node 𝑖in the graph, and interacts with other agents within its neighborhood, i.e., with some agent 𝑗in the set {𝑗∈𝑉∣𝐴𝑖𝑗 = 1}. Every agent has an opinion on a certain subject, which evolves as a consequence of these interactions. The model is executed during a finite number of time steps 𝑇. Let 𝑥𝑖(𝑡) ∈ [0,1] be the opinion of agent 𝑖at time step 𝑡, with 1≤𝑖≤𝑁and 0≤𝑡≤𝑇. Moreover, let 𝑋(𝑡)be the opinion profile at time step 𝑡, i.e., 𝑋(𝑡)={𝑥𝑖(𝑡)}1≤𝑖≤𝑁. Therefore, OD models are proposed to study the underlying dynamics to achieve 𝑋(𝑇)from 𝑋(0). 3.2. The ATBCR opinion dynamics model The ATBCR model is an extension of the classical Deffuant–Weisbuch (DW) model of BC (Deffuant et al.,2000). The DW model is able to explain both consensus and fragmentation of opinions within a population. ATBCR extends it with a repulsion mechanism in order to also explain the extremization of opinions. In the ATBCR model, each time step simulates an interaction between a randomly selected pair of agents 𝑖and 𝑗, connected in the social network (i.e., 𝐴𝑖𝑗 = 1). An agent 𝑖participating in an interaction with agent 𝑗in time step 𝑡+1 updates its opinion according to the following opinion fusion rule: 𝑥𝑖(𝑡+ 1) = ⎧ ⎪ ⎨ ⎪ ⎩ 𝑥𝑖(𝑡) + 𝜇(𝑥𝑗(𝑡) − 𝑥𝑖(𝑡)) iff |𝑥𝑖(𝑡) − 𝑥𝑗(𝑡)|< 𝜀𝑖(𝑡) 𝑥𝑖(𝑡)iff 𝜀𝑖(𝑡)≤|𝑥𝑖(𝑡) − 𝑥𝑗(𝑡)|≤𝜗𝑖(𝑡) 𝑚𝑖𝑛(1, 𝑚𝑎𝑥(0, 𝑟𝑒𝑝(𝑖, 𝑗))) iff |𝑥𝑖(𝑡) − 𝑥𝑗(𝑡)|> 𝜗𝑖(𝑡) (1) where the repulsion rule defined in the third case of Eq. (1) is: 𝑟𝑒𝑝(𝑖, 𝑗) = 𝑥𝑖(𝑡) − 𝜇(𝑥𝑗(𝑡) − 𝑥𝑖(𝑡)) (2) with 𝜇∈ [0,0.5] being the convergence speed of the model, and 𝜀𝑖(𝑡) and 𝜗𝑖(𝑡)being the confidence and the repulsion thresholds of agent 𝑖 at time step 𝑡, respectively. Notice that Eq. (1) always return a scalar value in the interval [0,1], i.e., the interval where agents’ opinions take values. Notice also that both thresholds are agent-dependent and timebased, i.e., they both depend on the agent 𝑖and the time step 𝑡. For more details, we address the reader to the original definition of the ATBCR model (Giráldez-Cru et al.,2022). The ATBCR model has been shown to preserve the behavior of the classical BC models (consensus and fragmentation of opinions) (Deffuant et al.,2000) while it also introduces a new mechanism to explain the extremization of opinions in a population with heterogeneous agents’ behaviors (Giráldez-Cru et al.,2022). Nevertheless, and without loss of generality, in this work we only consider homogeneous and time-invariant thresholds of confidence and repulsion 𝜀=𝜀𝑖(𝑡)and 𝜗=𝜗𝑖(𝑡),∀𝑖∈ [1, 𝑁], 𝑡 ∈ [0, 𝑇 ]. This is due to the complexity to obtain fine-grained data to model these thresholds in our case study. We emphasize that we only use real-world data in our experimental analysis. The rationality of the system can be derived from the values of 𝜀and 𝜗. In particular, systems with high confidence thresholds 𝜀are rational (i.e., only rational interactions affect the OD of the system), whereas systems with high repulsion thresholds 𝜗are emotional (i.e., opinions evolve as a consequence of emotional decisions). 2Notice that fully-mixed OD models can be simulated using a fully connected social network.
Expert Systems With Applications 250 (2024) 123750 4 J. Giráldez-Cru et al. Fig. 1. Overview of the proposed model. The real-world data from the survey is used to initialize spectator agents’ opinions and mass communication influences. The real-world data about film releases available at IMDb is used to model the remaining parameters of mass communication processes. A realistic social network is used to model agents’ interactions. The OD model is executed, considering both mass communication processes and agents’ interactions, which alter agents’ opinions, and a ranking of superstars is computed based on those opinions. Finally, such a ranking is compared to the Starmeter ranking from IMDb. 3.3. Classical opinion dynamics models The ATBCR model presented in the previous subsection was proposed as an extension of the DW model. Hence, the DW model (Deffuant et al.,2000) becomes a particular case of the former when 𝜗= 1. This model thus represents a purely rational system, where OD are only driven by a BC mechanism. Mathematically, it can be expressed just using the first case of Eq. (1).3 The DeGroot model (Degroot,1974) is another highly rational system. But, in contrast to the DW model, agents are fully susceptible to any other opinion (i.e., the confidence of any other opinion is complete). Mathematically, the opinion 𝑥𝑖(𝑡+ 1) is updated as 𝑥𝑖(𝑡+ 1) = 𝑁 ∑ 𝑗=1 𝑤𝑖𝑗 𝑥𝑗(𝑡)(3) where 𝑤𝑖𝑗 is the weight agent 𝑖gives to agent 𝑗. This model has been extensively used to study the sufficient and necessary conditions to reach a consensus in a population (Berger,1981). The Friedkin–Johnsen (FJ) model (Friedkin & Johnsen,1990) is a generalization of the DeGroot model in the sense it also considers that agents may be attached to a certain extent to their original opinion. The opinion fusion rule in the FJ model is: 𝑥𝑖(𝑡+ 1) = 𝜉𝑖 𝑁 ∑ 𝑗=1 (𝑤𝑖𝑗 𝑥𝑗(𝑡))+ (1 − 𝜉𝑖)𝑥𝑖(0) (4) where 𝜉𝑖is the susceptibility of agent 𝑖(i.e., 1 − 𝜉𝑖is the stubbornness of agent 𝑖), and the remaining parameters are the same than in the DeGroot model. 4. Model description This section describes all the decisions taken to design our model. First, it describes a light extension of the ATBCR model to handle 3Since the DW model is a particular case of the ATBCR model, it is guaranteed that the second and the third cases of Eq. (1) are never used in the DW model (i.e., when 𝜗= 1 in the ATBCR model). multidimensional opinions in Section 4.1. Next, we describe how the real-world data from the survey by Suárez-Vázquez (2015) about superstars is used to initialize the opinion profile of the population in Section 4.2. The social network used to model spectator agents’ interaction is described in Section 4.3. Section 4.4 describes how mass communication processes are modeled in our system. Finally, we use an external source of information from the specialized cinema website IMDb4to measure the accuracy of our model, as described in Section 4.5. An overview of the model is depicted in Fig. 1. A summary of the notation used in our model is reported in Table 1. 4.1. A multidimensional extension of the ATBCR model In our analysis, we consider a multidimensional extension of the ATBCR model. In particular, let 𝑆= {𝑠1,…, 𝑠|𝑆|}be a set of |𝑆| subjects, and 𝑥𝑠𝑘 𝑖(𝑡)be the opinion of spectator agent 𝑖at time step 𝑡on subject 𝑠𝑘∈𝑆. For the sake of simplicity, we consider that an interaction between spectator agents 𝑖and 𝑗only affects a particular subject (i.e., 𝑠𝑘). Therefore, we model the complex dynamics of this multidimensional system by executing in parallel |𝑆|independent instances of the original ATBCR model, one for each subject 𝑠𝑘, during 𝑇∕|𝑆| time steps each. This way, the same interaction between two agents can happen in different instances of the model (each affecting a particular subject), and hence, this multidimensional extension is able to model interactions affecting multiple subjects. Moreover, notice that, besides the global performance of this multidimensional system, this extension also allows us to study the OD on each subject independently. 4Star power is very difficult to measure objectively. In a particular film, star status is usually shown in the way the name of the star is deployed on the screen credits and on the promotional material of the film (McDonald, 2012). On an aggregate level, different measures have been proposed along the time, such as the Ulmer Scale or the Top Money-Making starts by Quigley Publishing Company. However, with the advent of accessible data from public websites, online industry resources, in particular, IMDb is behind most of the recent published research.
Expert Systems With Applications 250 (2024) 123750 5 J. Giráldez-Cru et al. Table 1 Summary of the notation used in our model. ATBCR model (Section 3.2) 𝑁Number of agents 𝑇Number of time steps 𝑥𝑖(𝑡)Opinion of agent 𝑖at time step 𝑡(with 1≤𝑖≤𝑁and 0≤𝑡≤𝑇) 𝜇Convergence speed 𝜀𝑖(𝑡)Confidence threshold of agent 𝑖at time step 𝑡 𝜗𝑖(𝑡)Repulsion threshold of agent 𝑖at time step 𝑡 𝑋(𝑡)Opinion profile at time step 𝑡 𝑉Set of nodes (with |𝑉|=𝑁) 𝐴Adjacency matrix of size 𝑁×𝑁 𝐺(𝑉 , 𝐴)Graph representing the social network Multidimensional extension of the ATBCR model (Section 4.1) 𝑆= {𝑠1,…, 𝑠|𝑆|}Set of subjects on which agents have opinions 𝑥𝑠𝑘 𝑖(𝑡)Opinion of spectator agent 𝑖at time step 𝑡on subject 𝑠𝑘∈𝑆 Initial opinions from a real-world population (Section 4.2) 𝑢𝑠𝑘 𝑟Utility of survey respondent 𝑟about superstar 𝑠𝑘 𝛼𝑠𝑘Coefficient of superstar 𝑠𝑘 𝛽𝑒Coefficient of emotion 𝑒 𝑎𝑠𝑘 𝑟,𝑒 Answer of respondent 𝑟about emotion 𝑒and superstar 𝑠𝑘 Mass communication (Section 4.4) 𝛹Opinion communicated in the mass communication process 𝜐(Effective) Reach of a mass communication process 𝜅(𝑖)(Effective) Influence that a mass communication process has on spectator agent 𝑖 𝑚=⟨𝛹, 𝜐, 𝜅(𝑖), 𝑠𝑘, 𝑡⟩Mass communication process with 𝛹,𝜐, and 𝜅(𝑖), on subject 𝑠𝑘occurring at time step 𝑡 𝜇𝑐Convergence speed of mass communication processes 4.2. Modeling opinions from a real-world population As seen in Section 2.2, the work by Suárez-Vázquez (2015) surveys the emotions of a population about a set of superstars. The sixteen superstars analyzed (i.e., 𝑆) are: Christian Bale, Gerard Butler, Nicholas Cage, George Clooney, Russell Crowe, Johnny Depp, Leonardo DiCaprio, Robert Downey Jr., Will Ferrell, Megan Fox, Tom Hanks, Robert Pattinson, Brad Pitt, Zoe Saldana, Will Smith, Kristen Stewart, and Reese Witherspoon; all of them are well-known superstars in the film industry.5The eight considered emotions are: anger, fear, sadness, shame, contentment, happiness, love, and pride; and were proposed as the main emotions to be analyzed in this context (Laros & Steenkamp, 2005). Based on them, the logit model is used to compute the utility 𝑢 of the survey respondent 𝑟about superstar 𝑠𝑘∈𝑆as: 𝑢𝑠𝑘 𝑟=𝛼𝑠𝑘+∑ 𝑒∈𝐸𝑚𝑜𝑡𝑖𝑜𝑛𝑠 𝛽𝑒⋅𝑎𝑠𝑘 𝑟,𝑒 (5) where 𝛼𝑠𝑘(with 𝑘∈ {1,…,16}) is the coefficient of each superstar, 𝛽𝑒is the coefficient of the emotions 𝑒∈ {1,…,8}, and 𝑎𝑠𝑘 𝑟,𝑒 is the answer of respondent 𝑟about emotion 𝑒and superstar 𝑠𝑘in the survey (Suárez-Vázquez,2015).6Notice that these utilities can be seen as the respondents’ opinions about the superstars. Since the survey by Suárez-Vázquez (2015) was conducted in a sample of 320 respondents, we amplify this population 10 times, i.e., 𝑁= 3,200 spectator agents in our model. This amplification process is a common procedure extensively used in agent-based modeling, where the number of agents is usually in the order of thousands (Chica & Rand,2017). We emphasize that this process increases the heterogeneity of the population without altering its original average values, since the resulting population is composed of many agents similar to the original ones, but slightly different from each other. In particular, for each respondent 𝑟and superstar 𝑠𝑘in the survey, we generate a Gaussian distribution 𝑟( 𝑋, 𝜎)of initial opinions with size ||= 5The survey by Suárez-Vázquez (2015) also considered Brad Pitt, who was used to calibrate the logit model. For this reason, we do not consider him in our analysis. 6The values of these coefficients are directly extracted from (SuárezVázquez,2015). 10, 𝑋=𝑢𝑠𝑘 𝑟, and a small variability with 𝜎= 0.1. This way, the initial opinion profile 𝑋(0) of the population is the union of these distributions, i.e., 𝑋(0) = ∪𝑟𝑟. Notice that this process produces 10 spectator agents similar to each survey respondent (which is an actual person), allowing us to use a real-world distribution of initial opinions, amplified to simulate a larger, heterogeneous population. Moreover, we use this survey (Suárez-Vázquez,2015) to model the influence of each superstar on each respondent (we recall that the survey contains an specific question on the intention of watching a film starred by each superstar), following the same amplification process. This influence is used in the mass communications processes described in Section 4.4. In Fig. 2, we represent the distribution of initial opinions, in both the survey and our model, as well as the influence of each superstar in the population. As it can be seen, both distributions are very similar, being the one used in our model slightly smoother. 4.3. Modeling spectator agents’ interactions In our model, we represent spectator agents’ interactions in a social network using a synthetic graph following the Preferential Attachment (PA) model (Barabási & Albert,1999). This model has been proposed to explain the growth of complex networks, and produces graphs with scale-free structure, i.e., graphs where node degree follows a powerlaw distribution 𝑃(𝑖) ∼ 𝑖−𝛾, characterized by the exponent 𝛾. As a consequence, most of the nodes have a low degree (i.e., they are only connected to a small subset of nodes), whereas a few nodes are connected to a big majority of the graph (i.e., they are hubs). In practice, many real-world networks exhibit this scale-free structure, with 𝛾usually ranging in the interval (2,3]. Hence, this social network topology is appropriate to represent the real fan interactions among spectators in our case study. The resulting PA graphs used in our analysis has 3,200 nodes (i.e., 𝑁), 31,900 edges, an average node degree ⟨𝑘⟩= 19.94, and an average clustering coefficient ⟨𝐶⟩= 0.03. We recall that interactions between spectator agents 𝑖and 𝑗only occur when these agents are connected in the social network, i.e., 𝐴𝑖𝑗 = 1 (there is an edge between nodes 𝑖and 𝑗in 𝐺).
Expert Systems With Applications 250 (2024) 123750 6 J. Giráldez-Cru et al. Fig. 2. Distribution of survey answers (blue) and their corresponding initial opinions in the model (orange) about the considered superstars. Inner plots represent the influence of each superstar in the population of spectator agents. 4.4. Modeling mass communication in the ATBCR model The model previously described allows us to study the OD about superstars. These dynamics are the consequence of interactions between spectator agents (i.e. the audience). However, it lacks the effects of mass communication, i.e., the processes of exchanging information to a large portion of the population. In our context, film releases (and all the news and marketing campaigns related to them) represent these mass communication processes, which also have a major impact on the audience’s opinion. This section presents an extension of our model to also capture the effects of mass communications. The mass communication mechanism we define is inspired by Carletti et al. (2006). Each process 𝑚of mass communication is defined by a tuple 𝑚=⟨𝛹, 𝜐, 𝜅(𝑖), 𝑠𝑘, 𝑡⟩, where the components of 𝑚are the following: •𝛹is the opinion communicated in this process, which must be in the same representation and scale than spectator agents’ opinions (in our model, a real number in the interval [0,1]). •𝜐∈ [0,1] is the (effective) reach that this process has in the population. •𝜅(𝑖) ∈ [0,1] is the (effective) influence that this process has on spectator agent 𝑖. •𝑠𝑘is the subject 𝑚targets. •𝑡is the time step when this process 𝑚occurs. Note that both 𝜐and 𝜅are the effective reach and influence, i.e., they do not model the potential impact that a marketing campaign may have in the population, but the actual values that they have. The multidimensional extension of the ATBCR model proceeds as described in Section 4.1 with an additional mechanism to consider these mass communication processes. In particular, at every time step 𝑡that a mass communication process 𝑚=⟨𝛹, 𝜐, 𝜅(𝑖), 𝑠𝑘, 𝑡⟩takes place,
Expert Systems With Applications 250 (2024) 123750 7 J. Giráldez-Cru et al. Table 2 Mass communication processes considered in the model. The influence 𝜅(𝑖)of these processes is represented in Fig. 2. Period Superstar (𝑠𝑘)𝛹 𝜐 𝑡 P1 J. Depp 0.7800 0.40 938 P1 T. Hanks 0.8000 0.30 2812 P2 Z. Saldana 0.7800 0.24 5312 P2 G. Butler 0.7900 0.50 6562 P2 G. Clooney 0.6700 0.36 7188 P2 N. Cage 0.8200 0.14 7500 P2 J. Depp 0.7250 0.44 8125 P2 L. DiCaprio 0.7100 0.40 8750 P2 G. Clooney 0.5850 0.40 9062 P2 K. Stewart 0.7800 0.32 9064 P2 R. Pattinson 0.7800 0.32 9065 P3 R. Downey Jr. 0.7650 0.36 10 312 P3 T. Hanks 0.7750 0.16 10 625 P3 C. Bale 0.7750 0.22 11 875 P3 G. Butler 0.6100 0.20 11 876 P3 N. Cage 0.8350 0.40 13 125 P3 R. Witherspoon 0.8500 0.42 13 126 P3 W. Ferrell 0.8050 0.08 13 750 P3 M. Fox 0.7300 0.16 14 062 P3 W. Ferrell 0.7450 0.26 14 375 P4 R. Downey Jr. 0.6600 0.32 16 562 P4 J. Depp 0.7300 0.32 16 875 P4 M. Fox 0.7150 0.22 17 188 P4 W. Smith 0.7150 0.42 17 500 P4 K. Stewart 0.7200 0.22 17 812 P4 R. Pattinson 0.7950 0.40 18125 P4 C. Bale 0.6150 0.90 19 732 P5 W. Ferrell 0.7550 0.38 20 938 P5 R. Pattinson 0.7150 0.32 21250 P5 Z. Saldana 0.8200 0.30 22 188 P5 N. Cage 0.7900 0.52 22 500 𝜐⋅𝑁spectator agents are randomly selected – the ones reached by this process –, and they all update their opinions as: 𝑥𝑠𝑘 𝑖(𝑡+ 1) = 𝑥𝑠𝑘 𝑖(𝑡) + 𝜇𝑐⋅𝜅(𝑖)⋅(𝛹−𝑥𝑠𝑘 𝑖(𝑡)) (6) where 𝜇𝑐∈ [0,0.5] is the convergence speed of the mass communication processes, and 𝑖is a spectator agent reached by this mass communication process 𝑚. Combining agents’ interactions (first and third cases of Eq. (1)) and mass communication processes (Eq. (6)), the resulting model is able to capture the complex dynamics of an evolving multidimensional scenario like opinions about superstars. In order to model these mass communication processes, we use some information available at IMDb. In particular, we use the Starmeter rankings of IMDb, which ranks the popularity of movie superstars along time. For a mass communication process 𝑚=⟨𝛹, 𝜐, 𝜅(𝑖), 𝑠𝑘, 𝑡⟩ representing a film release, the opinion 𝛹is computed as the position of this film in the ranking of the best 200 films. For instance, top films transmit a very good opinion (close to 1). The reach 𝜐is directly proportional to the number of news of this release (normalized over 500k). This information is also extracted from IMDb. Moreover, the time step 𝑡of this release is directly computed from its release date, and the subject 𝑠𝑘is the superstar starring such a film. Finally, the influence 𝜅(𝑖)is directly extracted from the survey (Suárez-Vázquez, 2015) and modeled, for each agent 𝑖, as described in Section 4.2. This influence 𝜅(𝑖)is depicted in Fig. 2. In Table 2 we report all the film releases considered in our analysis. 4.5. Benchmarking the model accuracy: comparison to specialized cinema real-world data In our analysis, we compare the output of our OD model designed from real-world data from the survey about superstars used in Suárez-Vázquez (2015) with respect to the information available in the specialized cinema website IMDb. In our simulation, we consider the six Starmeter rankings of IMDb published between May 1st, 2011 (the closest date to the survey by Suárez-Vázquez (2015)) and November 11th, 2012. In particular, IMDb publishes the Starmeter ranking every 16 weeks. Therefore, we use the rankings published on 05/01/2011, 08/21/2011, 12/11/2011, 04/01/2012, 07/22/2012, and 11/11/2012. The period between these six Starmeter rankings comprise approximately one year and a half. We also consider all the films released in this interval. In this way, our simulation covers a significant period where the initial opinions reflected in the survey can evolve as a consequence of the opinion dynamics and mass communication processes. Our model is independently executed for each superstar during 𝑇=25,000 time steps, generating intermediate rankings of opinions every 5,000 time steps (one for each period). Notice that each period of 5,000 time steps corresponds to 16 weeks (i.e., the time between the publication of two Starmeter rankings), hence each day comprises around 45 interactions in our model. In order to measure the accuracy of the model, the superstars are ranked according to the spectator agents’ opinions about them. In particular, the final opinion profile of the population about each superstar is averaged, sorting these average opinions to produce the OD ranking. This ranking is equal to the one obtained with a positional voting system (with Borda count) (Saari,1995). Finally, the OD ranking is compared to the actual ranking of IMDb (Starmeter) containing only the 16 superstars at study. In order to compare these two conjoint rankings, we use the Rank Biased Overlap (RBO) (Webber et al.,2010). This metric returns a value in [0,1], with 0indicating a total discrepancy, and 1indicating a complete correlation, and it is computed as: 𝑅𝐵𝑂(𝑅1, 𝑅2, 𝑝) = (1 − 𝑝)|𝑅1| ∑ 𝑑=1 𝑝𝑑−1 ⋅𝐴𝐺𝐺𝑑(7) where 𝑅1and 𝑅2are the rankings to be compared, 𝑝is the steepness of discrepancy weights, and 𝐴𝐺𝐺𝑑is the agreement of rankings 𝑅1and 𝑅2at depth 𝑑, defined as: 𝐴𝐺𝐺𝑑=|𝑅1[∶𝑑]∩𝑅2[∶𝑑]| 𝑑(8) with 𝑅[∶𝑑]denoting the first 𝑑elements of ranking 𝑅. As discussed by Webber et al. (2010), RBO solves several disadvantages of other ranking comparison metrics, including the Kendall’s 𝜏coefficient. For instance, 𝜏gives the same weight to every discrepancy,7regardless the position of the ranking where they occur (e.g., a discrepancy at the top of the ranking has the same weight than one at the bottom), whereas RBO overcomes this drawback. RBO is also able to handle non-conjoint pairs of rankings, although in our case this requirement is not necessary since both rankings contain exactly the same set of superstars. The global accuracy of our model is the average RBO of the six ranking comparisons, one for each period considered. 5. Experimental analysis In this section we present the experimental analysis on the OD about the 16 very well-known superstars considered. Since spectator agents only interact with other agents in their neighborhood (when they are connected in the social network, i.e., 𝐴𝑖𝑗 = 1), the distribution of initial opinions in the graph can have a major impact on the OD. To solve this, we perform a number of independent Monte Carlo (MC) executions of the model, differing in the seeding of the opinions within the nodes of the graph. Each MC execution returns differences in the OD about superstars. However, these differences only 7A discrepancy is a pair of two elements in different order in each ranking.
Expert Systems With Applications 250 (2024) 123750 8 J. Giráldez-Cru et al. produced negligible differences in the output OD rankings. Therefore, and for the sake of reducing the computational complexity of the study, in the rest of our analysis the reported results represent the average accuracy of 3 MC executions. In the following subsections, we first present a sensitivity analysis of the proposed model. Then, we present a fine-grained analysis of the most accurate configuration of our model. Finally, we present a comparison between our model and other state-of-the-art OD methods. 5.1. Sensitivity analysis of the proposed model In a first experiment, we perform a sensitivity analysis of our model. In particular, we analyze the confidence and repulsion thresholds, and how they affect the accuracy. Notice that these thresholds model the rationality of the system, i.e., whether the evolution of spectator agent’s opinions is driven by emotional (or rational) mechanisms. As commonly analyzed in the literature, the rest of the parameters of our model are fixed to 𝜇= 0.2and 𝜇𝑐= 0.5. In Fig. 3 we represent the results of this sensitivity analysis. The most accurate scenario, with an average 𝑅𝐵𝑂 = 0.473833, is found with a confidence threshold 𝜀= 0.3and a repulsion threshold 𝜗= 0.5. This represents a system with high confidence but also high repulsion. This scenario is analyzed in more details in the following subsection. The sensitivity analysis also shows that this system is highly ruled by the repulsion mechanism, i.e., executions with a higher repulsion threshold return, in general, more accurate results (see the right area of Fig. 3). These results match the expected behavior of the OD about superstars, which are, in general, based on emotional sentiments, with a relatively low degree of rationality. This result is in line with the hedonic perspective of movie goers’ behavior, under what the emotional component of behavior dominates the cognitive component (Eliashberg & Sawhney,1994). It also gives an explanation to the classic statement ‘‘nobody knows anything’’ by screenwriter William Goldman. Our OD study shows that, at least as far as individual moviegoers’ decisions are concerned, superstars power is not a matter of knowledge, but of emotions. Thus, it could be said that in the film industry ‘‘nobody knows anything’’ but ‘‘everybody feels something’’ and those feelings strongly influence the power of superstars. When measuring the value of superstars, in addition to characteristics such as experience or awards (Wei,2006), their emotional value must be taken into account. On an aggregate level, this result may offer an explanation for the weak relationship between expert judgments and popular appeal (Holbrook,2005), as the latter is expected to be more strongly affected by the emotional value of stars. 5.2. Fine-grained analysis of the proposed model Next, we analyze the detailed results of our model executed with 𝜀= 0.3and 𝜗= 0.5(the most accurate scenario found previously). In Fig. 4 we report the OD about each superstar, where blue points represent the opinion evolution (each point reflects the opinion of a single spectator agent at a specific time in a [0,1] scale) as a consequence of spectator agents’ interactions along time, represented by the X axis; red points represent the final opinions at each period (including the initial opinions); and green points represent the impact of mass communication processes. We also include subplots with the distribution of final opinions at the end of the last period. An interesting observation from these detailed results is the distinct effect on the opinions of spectator agents’ interactions and mass communication processes. On the one hand, agents’ interaction tend to form a major consensus in the long term. See, for instance, the distribution of final opinions (right subplots of each superstar), which exhibits small differences. On the other hand, mass communication only has a small impact on the dynamics of the opinions. In particular, the opinions resulting from mass communication processes diverge from the average opinion, producing both bad and good opinions in the spectator agents Fig. 3. Sensitivity analysis of OD about superstars in terms of the confidence and repulsion thresholds of the proposed model. (see the green points in Fig. 4). However, this phenomenon only has a short term effect. These results allow us to explain the different behaviors of the audience in both the short and the long term. In the short term, just after a release, the audience may be highly influenced by (aggressive) marketing campaigns. However, these campaigns may be counterbalanced by the spectator’s experience, which is shared with other spectators influencing each other, and this explains the success of the films (and the superstars) in the long term. This phenomenon has been observed in several films and can be interpreted as a consequence of how success film drivers change between shortand long-term box office (HennigThurau et al.,2006). A paradigmatic example of this phenomenon was the first movie ‘‘My bit fat Greek wedding’’, released in 2002. This romantic comedy became ‘‘one of the most profitable films in history’’ that ‘‘went viral when there was not such a thing as going viral’’ (Goldenberg, 2016). Also, in Table 3 we report the OD and the IMDb rankings for the six periods, as well as the RBO value for their comparison. The evolution of these RBO values is represented in Fig. 5. These results show that the RBO value is improved in all the periods, except the second one (although this degradation is small). We conjecture that this is due to the number of films released in this period, most of them having a very large influence. Nevertheless, the RBO of the other four periods shows improvements with respect to the baseline RBO of initial opinions. This suggests that our model is able to adequately represent the OD in this system. We must emphasize that, although these RBO values do not show a total consensus, the ranking of opinions is based on a particular population whereas the IMDb ranking is based on a specialized cinema website. Therefore, certain divergences are expected. Moreover, there exist some differences in the nature of the data. On the one hand, the survey (initial opinions in the model) provides a transversal snapshot of the opinions of this population in a particular moment. On the other hand, the ranking of IMDb is the result of behavioral analysis of users in this website (thus not strictly based on opinions) in a lapse of time. In fact, this reveals one of the main open problems in this field: how to establish a ranking of superstars, which motivates the present work. Even so, the performance of the model is pretty satisfactory according to the RBO values and their evolution.
Expert Systems With Applications 250 (2024) 123750 9 J. Giráldez-Cru et al. Fig. 4. OD about the 16 analyzed superstars in the model executed with 𝜀= 0.3and 𝜗= 0.5. 5.3. Comparison with other methods In this subsection, we finally present a comparison between the proposed method based on the ATBCR model and other state-of-the-art OD models in the literature. In particular, we implement adaptations of our methodology based on the DW model (Deffuant et al.,2000), the DeGroot model (Degroot,1974), and the FJ model (Friedkin & Johnsen, 1990). In order to adapt our methodology to these OD models, we follow the same steps discussed in Section 4, i.e., we implement a multidimensional extension of the specific OD model, we initialize the initial opinions using the real-world data from the existing survey (SuárezVázquez,2015), we model agents’ interactions using a PA network, we model mass communication processes representing film releases, and we measure the accuracy of the output comparing the obtained ranking of superstars versus the real-world ranking of the specialized website IMDb, using the RBO value. The only difference between these adaptations is the underlying OD model that captures how opinions are updated along time. Table 4 reports a comparison on the RBO value obtained for these adaptations, based on the ATBCR, the DW, the DeGroot, and the (two variants of the) FJ models. In most of the six periods analyzed, as well as in the aggregate global average RBO value of the six periods, the proposed methodology based on the ATBCR model is the most accurate one, or it shows a performance very close to the optimal model. In a sensitivity analysis of the DW model adaptation, we found that the best configuration is achieved with 𝜀= 0.1. This surprising value represents a very low confidence system, composed of very stubborn spectator agents, i.e., in general these agents only change their opinions after an interaction with another agent having a very similar opinion, thus their resulting opinions are almost unchanged. We emphasize that, in contrast to this DW-based adaptation, the adaptation based on the ATBCR model returned, as the most accurate configuration, a highly emotional system with a relatively high degree of confidence as well (see Section 5.1). Qualitatively, the DW-based model is unable to capture the OD about superstars, which are expected to be mostly driven by emotional mechanisms (as the ATBCR model is able to capture, as showed above).