Clustering U.S. 2016 presidential candidates through linguistic appraisals
Abstract
Producción Científica
Full text
Clustering U.S. 2016 presidential candidates through linguistic appraisals Raquel Gonz´alez del Pozo1, Jos´e Luis Garc´ıa-Lapresta2, David P´erez-Rom´an3 1PRESAD Research Group, IMUVA, Departamento de Econom´ıa Aplicada, Universidad de Valladolid, Spain [email protected] 2PRESAD Research Group, BORDA Research Unit, IMUVA, Departamento de Econom´ıa Aplicada, Universidad de Valladolid, Spain [email protected] 3PRESAD Research Group, BORDA Research Unit, Departamento de Organizaci´on de Empresas y Comercializaci´on e Investigaci´on de Mercados, Universidad de Valladolid, Spain [email protected] Abstract. During electorate campaigns, linguistic appraisals of presidental candidates given by voters are essential. In order to deal with linguistic appraisals and cluster analysis, this paper presents the results of grouping the United States 2016 presidential candidates using linguistic appraisals collected from a political survey. To do this, we have developed an agglomerative hierarchical clustering procedure based on consensus through the concept of ordinal proximity measure. Keywords: Clustering, Presidential candidates, Ordinal proximity measures, Consensus, Linguistic appraisals 1 Introduction One of the main issues in many research fields is to classify a collection of data into meaningful groups. Cluster analysis is a multivariate method which tries to classify a sample of observations into similar groups regarding given characteristics. This technique is used in several disciplines: Economics, Marketing, Biology, Political Science, Artificial Intelligence, Computer Science, Engineering, etc. In the political context, clustering has hardly considered the views of the electorate. This method has been applied mainly to identify relationships among demographically similar voters, voting tendencies, as well as patterns of political parties support in different elections (see Seabrook [13], Pearson and Cooper [11] and Aleskerov and Nurmi [1], among others). Appraisals of candidates given by voters are essential during presidential races, being easy to find linguistic scales in political surveys for evaluating candidates. In recent years, there have arisen voting systems based on linguistic “This is a post-peer-review, pre-copyedit version of an article published in Kacprzyk, J.; Szmidt, E.; Zadrożny, S.; Atanassov, K.; Krawczak, M. (eds) Advances in Fuzzy Logic and Technology 2017. Advances in Intelligent Systems and Computing, vol. 642. Cham: Springer, 2017, p. 143-153. The final authenticated version is available online at: http://dx.doi.org/10.1007/978-3-319-66824-6_13”.
2 R. Gonz´alez del Pozo, J.L. Garc´ıa Lapresta and D. P´erez-Rom´an appraisals. For instance, Balinski and Laraki [2, 3] introduce the Majority Judgment voting system, where voters assign a linguistic term of the following ordered qualitative scale: ‘reject’, ‘poor’, ‘acceptable’, ‘good’, ‘very good’ and ‘excellent’ to each candidate. However, there is so far little literature addressing cluster analysis applied to linguistic appraisals. The main purpose of this paper is to cluster the United States (U.S.) 2016 presidential candidates taking the linguistic appraisals made by a random representative sample of adults living in the U.S. as our starting point. To do this, we have used the concept of ordinal proximity measure (see Garc´ıa-Lapresta and P´erez-Rom´an [8]), which allows to determine the degree of consensus in a group of agents when a set of alternatives is evaluated through non-necessarily qualitative scales. The rest of the paper is organized as follows. Section 2 contains notation and some applications of ordinal proximity measures to consensus and clustering. Section 3 includes the results of clustering U.S. 2016 presidential candidates through linguistic appraisals. Finally, Section 4 concludes with some remarks. 2 Consensus and clustering In this section we introduce the concept of ordinal proximity measure and we tackle some applications of this concept to consensus and clustering. 2.1 Preliminaries Consider a set of agents A={1, . . . , m}, with m≥2, and a set of alternatives X={x1, . . . , xn}, with n≥2, which have to be appraised. Each agent assigns a linguistic term to every alternative within an ordered qualitative scale L= {l1, . . . , lg}, arranged from the lowest to the highest terms, where the granularity of Lis at least 3 (g≥3). We now recall the notion of ordinal proximity measure, introduced by Garc´ıa- Lapresta and P´erez-Rom´an [8], which is a mapping that assigns an ordinal degree of proximity to each pair of linguistic terms of an ordered qualitative scale. These ordinal degrees of proximity belong to a linear order ∆={δ1, . . . , δh} with δ1 · · · δh, being δ1and δhthe maximum and the minimum degrees of proximity, respectively. The elements of ∆have no meaning, they only represent different degrees of proximity. Definition 1. ([8]) An ordinal proximity measure on Lwith values in ∆is a mapping π:L2−→ ∆, where π(lr, ls) = πrs means the degree of proximity between lrand ls, satisfying the following conditions: 1. Exhaustiveness: For every δ∈∆, there exist lr, ls∈ L such that δ=πrs. 2. Symmetry:πsr =πrs, for all r, s ∈ {1, . . . , g}. 3. Maximum proximity:πrs =δ1⇔r=s, for all r, s ∈ {1, . . . , g}. 4. Monotonicity:πrs πrt and πst πrt, for all r, s, t ∈ {1, . . . , g}such that r < s < t.
Clustering U.S. 2016 presidential candidates through linguistic appraisals 3 Every ordinal proximity measure can be represented by a g×gsymmetric matrix with coefficients in ∆, where the elements in the main diagonal are πrr = δ1,r= 1, . . . , g: π11 · · · π1s· · · π1g ··· ··· ··· ··· ··· πr1· · · πrs · · · πrg ··· ··· ··· ··· ··· πg1· · · πgs · · · πgg . This matrix is called proximity matrix associated with π. 2.2 Consensus To determine the degree of consensus we follow the approach introduced by Garc´ıa-Lapresta and P´erez-Rom´an [8]. We start from the appraisals of alternatives given by agents, collected in a profile, that is a matrix V= v1 1· · · v1 i· · · v1 n · · · · · · · · · · · · · · · va 1· · · va i· · · va n · · · · · · · · · · · · · · · vm 1· · · vm i· · · vm n = (va i) consisting of mrows and ncolumns of linguistic terms, where the element va i∈ L represents the linguistic appraisal given by the agent a∈Ato the alternative xi∈X. According to Garc´ıa-Lapresta and P´erez-Rom´an [10], for measuring the consensus in a group of agents over a set of alternatives, the first step is to consider the degrees of ordinal proximity between the linguistic appraissals over the alternatives. These degrees are arranged in a vector of ordinal degrees δ∈∆p, for some p∈N, from the highest to the lowest degrees (decreasing fashion). In order to avoid loss of information, Garc´ıa-Lapresta and P´erez-Rom´an [10] select the median(s) of the mentioned ordinal degrees in such a way that if the number of ordinal degrees of the vector is odd, the median is unique, δr∈∆, but if the number of ordinal degrees is even, then δhas two medians, δs, δt∈∆ with s≤t. In order to unify this assignment of medians, the authors consider the pair of medians (δr, δr) in the odd case and (δs, δt) in the even case. Given the set of feasibles medians ∆2={(δr, δs)∈∆2|r≤s}, the median operator is the mapping M: ∞ [ p=1 ∆p−→ ∆2 that assigns the corresponding pair of medians to each vector of ordinal degrees. We denote with #Ithe cardinality of Iand with P2(A) = {I⊆A|#I≥2} the family of subsets of at least two agents.
4 R. Gonz´alez del Pozo, J.L. Garc´ıa Lapresta and D. P´erez-Rom´an Definition 2. Given a profile V= (va i), the degree of consensus in a subset of agents I∈ P2(A)over a subset of alternatives ∅ 6=Y⊆Xis defined as C(I, Y ) = M πva i, vb ia,b∈I , a<b xi∈Y!∈∆2. For comparing the degrees of consensus in a group of agents Iwhen they evaluate two subsets of alternatives Y, Z, i.e. C(I, Y )∈∆2versus (I, Z)∈∆2, an appropriate linear order on ∆2is required. In this paper we use the following linear order used by Garc´ıa-Lapresta and P´erez-Rom´an [10]: (δr, δs)(δt, δu)⇔ r+s < t +u or r+s=t+uand s−r≤u−t for all (δr, δs),(δt, δu)∈∆2. It is important to note that the above linear order for ranking the medians is not the only one that can be considered on ∆2. Since the cardinality of ∆2may be low, it is very easy to have ties among the degrees of consensus in different subsets of agents or alternatives. In these cases, we resort to a sequential tie-breaking procedure [2] which consists of removing the median(s) of the respective agents or alternatives that are in a tie, and select again the new median(s) of remaining ordinal degrees until ties are broken. Starting from C(1)(I, Y ) = C(I, Y ), we calculate C(2)(I, Y ) as in C(I, Y ) but after dropping the pair of medians of the list πva i, vb ia,b∈I , a<b xi∈Y , and analogously for C(3)(I, Y ), etc. 2.3 Clustering The goal of clustering methodology is to reduce the number of objects by classifying them into a set of groups with similar features. There are many types of clustering and applications (see Everitt et al. [5]) being the agglomerative hierarchical clustering the method used in this paper. This process starts by assigning each item to its own cluster and then, clusters are sequentially merged according to a given similarity criterion, until all of them end up in the same cluster. The first steps in cluster analysis are to determine the variables and the criterion that will be used for joining clusters or elements. In general, many clustering procedures are based on distances to determine the similarity in a set of elements. However, in this paper we have taken into account the degree of consensus as similarity measure in order to merge clusters. For this purpose, we have considered a similarity function and a sequential similarity vector based on the degree of consensus [7, 9], in such a way that the consensus is measured by means of the degrees of proximity between all pairs of individual appraisals over the evaluated alternatives.
Clustering U.S. 2016 presidential candidates through linguistic appraisals 5 Definition 3. Given a profile V= (va i), the similarity function relative to a subset of agents I∈P(A)\ {∅} SI:P(X)\ {∅}2−→ ∆2 is defined as SI(Y, Z) = (C(I, Y ∪Z),if #(Y∪Z)≥2, (δ1, δ1),if #(Y∪Z)=1. Definition 4. Given a profile V= (va i), the sequential similarity vector relative to a subset of agents I∈P(A)\ {∅}for ∅ 6=Y, Z ⊆Xis defined as SI(Y, Z) = S(1) I(Y, Z), S(2) I(Y, Z), . . . , where S(k) I(Y, Z) = (C(k)(I, Y ∪Z),if #(Y∪Z)≥2, (δ1, δ1),if #(Y∪Z) = 1. The agglomerative hierarchical clustering procedure is related to the one provided by Garc´ıa-Lapresta and P´erez-Rom´an [8]. It consists of a sequential process addressed by the following stages: 1. Given a subset of agents I∈P(A)\ {∅}, the initial clustering is AI 0= {{x1},...,{xn}}. 2. Calculate the similarities between all the pairs of alternatives SI({xi},{xj}) for all xi, xj∈X. 3. Select the two alternatives xi, xj∈Xthat maximize SI(taking into account the corresponding sequential similarity vectors) and construct the first cluster AI 1={xi, xj}. 4. The new clustering is AI 1=AI 0\ {{xi},{xj}}∪ {AI 1}. 5. Calculate the similarities SI(AI 1,{xk}) and take into account the previously computed similarities SI({xk},{xl}), for all {xk},{xl}∈AI 1 6. Select the two alternatives of AIthat maximize SIand to construct the second cluster AI 2 7. Proceed as in previous items until obtaining the next clustering AI 2. The process continues in the same way until obtaining the last cluster, AI n−1={X}.
6 R. Gonz´alez del Pozo, J.L. Garc´ıa Lapresta and D. P´erez-Rom´an 3 Clustering presidential candidates Recently, the issue of clustering presidential candidates of U.S. has been tackled in some publications. For instance, Dowdle et al. [4] gather Republican and Democratic candidates of the parties according to the multiple donor networks, regardless of voters’ appraisals. In this section, taking linguistic appraisals as starting point, we have developed an agglomerative hierarchical clustering in the context of the U.S. 2016 presidential elections. For this purpose, we have considered some data from the January 2016 Political Survey conducted by the Pew Research Center [12], which provides information about political values, domestic policy issues and the interests of U.S. electorate. In particular, we have focused on a question in which a total of 1184 individuals appraised the nine U.S. presidential candidates included in Table 1, in case of being elected in November 2016. To do this, the individuals used the linguistic terms of the qualitative scale contained in Table 2. As has been previously explained in Subsection 2.3, the clustering procedure is based on the similarity between alternatives with respect to a group of agents. Therefore, the candidates (alternatives) are classified into different clusters when they maximize the degree of consensus among the agents. Agents Name 1 Ben Carson 2 Bernie Sanders 3 Chris Christie 4 Donald Trump 5 Hillary Clinton 6 Jeb Bush 7 John Kasich 8 Marco Rubio 9 Ted Cruz Table 1. Candidates. l1Terrible president l2Poor president l3Average president l4Good president l5Great president Table 2. Linguistic terms in L. In order to show the clustering process when the proximities between the linguistic terms of the scale are different, we considered the three most reasonable
Clustering U.S. 2016 presidential candidates through linguistic appraisals 7 cases in accordance with the qualitative scale used in the survey question. First, the uniform case, where the proximities between consecutive linguistic terms are always the same, and two non-uniform cases in which l2and l4are considered symmetric with respect to l3. The ordinal proximity measures and the clusters of candidates related to each cases are shown below. 1. The uniform case is represented in the following upper half proximity matrix, where the subindices of the matrix correspond to the subindices of the δ’s appearing just over the main diagonal and denoting the proximity between consecutive terms of the scale. A2222 = δ1δ2δ3δ4δ5 δ1δ2δ3δ4 δ1δ2δ3 δ1δ2 δ1 This case can be visualized in Figure 1. l1l2l3l4l5 Fig. 1. Ordinal proximity measure with associated matrix A2222. After comparing the different sequential similarity vectors1, we have obtained the following clusters of candidates: AA 0={{1},{2},{3},{4},{5},{6},{7},{8},{9}} AA 1={{1},{2},{3,7},{4},{5},{6},{8},{9}} AA 2={{1},{2},{3,6,7},{4},{5},{8},{9}} AA 3={{1},{2},{3,6,7,8},{4},{5},{9}} AA 4={{1,3,6,7,8},{2},{4},{5},{9}} AA 5={{1,3,6,7,8,9},{2},{4},{5}} AA 6={{1,2,3,6,7,8,9},{4},{5}} AA 7={{1,2,3,5,6,7,8,9},{4}} AA 8={{1,2,3,4,5,6,7,8,9}}. 1The computations for obtaining the corresponding sequential consensus and similarity vectors have been performed with MATLAB.
8 R. Gonz´alez del Pozo, J.L. Garc´ıa Lapresta and D. P´erez-Rom´an 2. The non-uniform case in which l2and l4are considered further away from l3is associated with the matrix A2332 = δ1δ2δ4δ6δ7 δ1δ3δ5δ6 δ1δ3δ4 δ1δ2 δ1 that can be visualized in Figure 2. l1l2l3l4l5 Fig. 2. Ordinal proximity measure with associated matrix A2332. AA 0={{1},{2},{3},{4},{5},{6},{7},{8},{9}} AA 1={{1},{2},{3},{4,7},{5},{6},{8},{9}} AA 2={{1},{2},{3,6},{4,7},{5},{8},{9}} AA 3={{1,4,7},{2},{3,6},{5},{8},{9}} AA 4={{1,4,7},{2},{3,6,8},{5},{9}} AA 5={{1,3,4,6,7,8},{2},{5},{9}} AA 6={{1,3,4,6,7,8,9},{2},{5}} AA 7={{1,2,3,4,6,7,8,9},{5}} AA 8={{1,2,3,4,5,6,7,8,9}}. 3. The non-uniform case in which l2and l4are considered closer to l3is associated with the matrix A3223 = δ1δ3δ5δ6δ7 δ1δ2δ4δ6 δ1δ2δ5 δ1δ3 δ1 that can be visualized in Figure 3.
Clustering U.S. 2016 presidential candidates through linguistic appraisals 9 l1l2l3l4l5 Fig. 3. Ordinal proximity measure with associated matrix A3223 AA 0={{1},{2},{3},{4},{5},{6},{7},{8},{9}} AA 1={{1},{2},{3,7},{4},{5},{6},{8},{9}} AA 2={{1},{2},{3,7,8},{4},{5},{6},{9}} AA 3={{1},{2},{3,6,7,8},{4},{5},{9}} AA 4={{1,3,6,7,8},{2},{4},{5},{9}} AA 5={{1,3,6,7,8,9},{2},{4},{5}} AA 6={{1,2,3,6,7,8,9},{4},{5}} AA 7={{1,2,3,5,6,7,8,9},{4}} AA 8={{1,2,3,4,5,6,7,8,9}}. The clusters of candidates associated with the above ordinal proximity measures are illustrated in Figure 4, Figure 5 and Figure 6, which display hierarchical relationships among presidential candidates. Carson Christie Kasich Bush Rubio Cruz Sanders Clinton Trump Fig. 4. Clustering tree obtained from the matrix A2222