scieee AI-readable full text Open interactive document viewer

Local Adaption and Meaning Creation in Performance Appraisal

Rossing, Morten

Abstract

EconStor is a publication server for scholarly economic literature, provided as a non-commercial public service by the ZBW.

Full text

Rossing, Morten Doctoral Thesis Local Adaption and Meaning Creation in Performance Appraisal PhD Series, No. 29.2013 Provided in Cooperation with: Copenhagen Business School (CBS) Suggested Citation: Rossing, Morten (2013) : Local Adaption and Meaning Creation in Performance Appraisal, PhD Series, No. 29.2013, ISBN 9788792977717, Copenhagen Business School (CBS), Frederiksberg, https://hdl.handle.net/10398/8761 This Version is available at: https://hdl.handle.net/10419/208864 Standard-Nutzungsbedingungen: Die Dokumente auf EconStor dürfen zu eigenen wissenschaftlichen Zwecken und zum Privatgebrauch gespeichert und kopiert werden. Sie dürfen die Dokumente nicht für öffentliche oder kommerzielle Zwecke vervielfältigen, öffentlich ausstellen, öffentlich zugänglich machen, vertreiben oder anderweitig nutzen. Sofern die Verfasser die Dokumente unter Open-Content-Lizenzen (insbesondere CC-Lizenzen) zur Verfügung gestellt haben sollten, gelten abweichend von diesen Nutzungsbedingungen die in der dort genannten Lizenz gewährten Nutzungsrechte. Terms of use: Documents in EconStor may be saved and copied for your personal and scholarly purposes. You are not to copy documents for public or commercial purposes, to exhibit the documents publicly, to make them publicly available on the internet, or to distribute or otherwise use the documents in public. If the documents have been made available under an Open Content Licence (especially Creative Commons Licences), you may exercise further usage rights as specified in the indicated licence. https://creativecommons.org/licenses/by-nc-nd/3.0/ Morten Rossing PhD Series 29.2013 PhD Series 29.2013 Local Adaption and Meaning Creation in Performance Appraisal copenhagen business school handelshøjskolen solbjerg plads 3 dk-2000 frederiksberg danmark www.cbs.dk ISSN 0906-6934 Print ISBN: 978-87-92977-70-0 Online ISBN: 978-87-92977-71-7 Doctoral School of Organisation and Management Studies Local Adaption and Meaning Creation in Performance Appraisal Local Adaption and Meaning Creation in Performance Appraisal "Objectives are not fate; they are direction. They are not commands; they are commitments. They do not determine the future; they are means to mobilize the resources and energies of the business for the making of the future." - Peter Drucker Morten Rossing Local Adaption and Meaning Creation in Performance Appraisal 1st edition 2013 PhD Series 29.2013 © The Author ISSN 0906-6934 Print ISBN: 978-87-92977-70-0 Online ISBN: 978-87-92977-71-7 The Doctoral School of Organisation and Management Studies (OMS) is an interdisciplinary research environment at Copenhagen Business School for PhD students working on theoretical and empirical themes related to the organisation and management of private, public and voluntary organizations. All rights reserved. No parts of this book may be reproduced or transmitted in any form or by any means, electronic or mechanical, including photocopying, recording, or by any information storage or retrieval system, without permission in writing from the publisher. III Acknowledgements There are many people I would like to thank for their help with this piece of work. First, I simply would not have made it without the support of my two supervisors; Associate Professor Christopher John Mathieu, Department of Sociology, Lund University and Associate Professor Magnus Larsson, Department of Organization, Copenhagen Business School. Their persistent support, creative ideas and constructive critique helped me tremendously in all phases of the project. I am grateful for the support provided by Professor Robert Austin, Dean of the Faculty of Business Administration at University of New Brunswick, whose book "Measuring and Managing Performance in Organizations" has been a great inspiration for me, and who supported me with many valuable comments at my closing seminar. Similarly, I am grateful for the constructive and supportive comments and challenging questions I received at my project presentation seminars from Professor Jan Mouritsen, Head of Department of Operations Management at Copenhagen Business School and from Professor Paul Du Gay, Department of Organization, Copenhagen Business School. Also, I would like to thank all faculty and fellow PhD students who have taken time to read and comment on my project presentations during PhD courses and conferences. Many ideas were fostered by these comments. Indeed, I am very grateful to the managers and employees in PharmaComp who stood up for a series of in depth interviews on top of all their daily tasks and who were so enthusiastic about sharing their views on performance appraisal. Likewise, my gratitude goes to colleagues in PharmaComp's Corporate HR who helped me with providing important documentation and sparring, particularly in the early, forming phase of my project. Thank you also to my own management, who showed me the trust to support the idea of embarking on the PhD journey in the first place. Last, but certainly not least, I am indebted to my wife, Camilla, who supported and encouraged me all the way. The project was undertaken while I was working full time in PharmaComp. Without Camilla's support, I would not have dared to embark on the project, and I would for sure not have managed to fulfil it either. IV Preface I became interested in the subject of Performance Appraisal (PA) long before deciding to pursue the opportunity to become an Independent PhD Fellow at Copenhagen Business School. For more than a decade, I have worked as a manager in different private organisations that all subscribed to performance appraisal as a systematised process of setting targets for and evaluating the performance of individuals, for the purpose of improving organisational performance. Both in my role as super-ordinate, as well as in my role as subordinate, I have had mixed experiences with performance appraisal. On the one hand it seems so intuitively logical and right; to agree up front on measurable targets that a later evaluation will be based on, for the benefit of transparency, fairness and wellfounded management decisions about reward, promotion, dismissal and more. On the other hand the tangible experiences are so mixed. Too often I have felt that employees have left appraisal meetings with disappointed expectations because we did not have the same interpretation of how the individual had contributed over the past year. Too often I have felt that employees or managers have been disengaged and unprepared for the process, although the importance of the process in terms of reward allocation and career opportunities should be clear to everyone. And too often I have found that the sum of outcomes on each of an employee's targets did not add up to what my gut feeling told me that the employee had in reality contributed. Sometimes, the employee had done so much more and I felt I had to twist the system to provide the higher rating the employee deserved. And sometimes the opposite was the case. Of course, these personal experiences could have different causes. One reason could be that I was simply lacking the management capabilities required to handle a simple and straight-forward management tool. However, many dialogues with peers, supervisors and employees who all in different ways shared my mixed experiences told me that the cause likely lies somewhere else. More likely, the truth is that it is not easy to conduct performance appraisal in such a way that targets and evaluations remain controlled and transparent, so that the employees are more motivated than they would have been without it and so that visible organisational performance improvement is the natural outcome of the significant resources put into the process. V Since I joined corporate HR in a large pharmaceutical organisation nearly five years ago I have been working with the further development and support of a centrally defined, globally standardised performance appraisal process. This closer intimacy with the practice of performance appraisal seen from the corporate perspective made me curious to learn more about the experiences of other organisations, and the status of current research into the field. I started reading books and articles on performance appraisal and related theoretical themes like measurement theory, motivation theory and justice theory. From this, it became clear to me that my experiences with the process were not unique. In fact, the subject of performance appraisal is a very well-researched phenomenon, in terms of the number of problems related to the process that have been dealt with in research over the past couple of decades. However, I also found that for some reason, large parts of the research into the subject of performance appraisal were strikingly uniform. Uniform in terms of methodology, uniform in terms of empirical subject and uniform in terms of narrowly focusing on one or a few single problems in performance appraisal, and analysis of these based on rational economic models without questioning basic assumptions. This left me with a feeling that the ground would be fertile for digging further into the field of performance appraisal to supplement the large volume of narrow focused literature with a more holistic, comparative contribution. So, I decided to do just that, and the present PhD thesis is the result of the following efforts through a three year period from 2010 to 2013. VI A Brief Reader's Guide This thesis is structured into eleven chapters. In the introduction in chapter 1, I will start by presenting the proposition, the phenomenon, the research question and important definitions. Chapter 2 contains an introduction to the conversation in the literature that I am taking part in. After a high level categorisation of the performance appraisal literature into what I have called a dominating and a radical paradigm, I will present a more detailed analysis of a sample of 125 journal articles from the 20 year period 19902010. I will examine the state of the literature in terms of its empirical, methodological and theoretical focus. Further, I will provide an overview of the more than 50 different problems in PA dealt with in the journal articles. In chapter 3, I will present my methodological approach and my data. This thesis is based on a single case study through a triangulating approach, including participant observation, study of archival material and analysis of a quantitative data set. First and foremost, though, the study is based on 38 qualitative, deep-dive interviews, of which 34 interviews were part of a longitudinal design where dyads of managers and employees in four different workforces were followed during a one year, full performance appraisal cycle. Finally, I will also provide a detailed account of the data coding process through nVivo9, and for documentation and archiving procedures utilised during the study. Chapter 4 provides an introduction to performance appraisal in the case organisation. The historical background and history of performance appraisal is presented in three stages. The first is the introduction of a standard corporate performance appraisal scheme in 2004. The second is the introduction of a supporting IT system in 2006. The third is the more advanced steps planned and fulfilled in recent years. In chapter 5 the macro view of performance appraisal in the case organisation is presented. Based mainly on quantitative analysis of a data set covering all electronically captured performance appraisal ratings from the period 2006-2011, supplemented by analysis of archival material from the case organisation, I will draw up the macro situation in the case organisation on a number of dimensions; the distribution of ratings over time and by workforce, the correlation between hierarchy and rating, and the correlation between rating and retention, wage and bonus. I will VII also introduce the connection between the design of a corporate performance appraisal system and the corporate culture. Finally, I will introduce the approach chosen in different workforces regarding how to balance business and behavioural performance. The purpose of drawing the macro view in this fairly detailed way is to prepare for the subsequent analysis of local adaption and meaning creation. By having a good understanding of the environments within which actors operate and create meaning, we will be better equipped to analyse the meaning creation of managers and employees in performance appraisal. Chapter 6 introduces the theoretical toolbox for the meaning creation analysis. First, I give an introduction to Weick's (1995) sensemaking theory, particularly by relating Weick's seven properties of sensemaking to the process of performance appraisal. This will exemplify how Weick's framework will be helpful when we want to understand meaning creation as a process of enactment of environmental differences within which managers and employees operate. In chapter 6 I will also present a categorisation of the task at hand in four workforces included in the study, against Austin's (1996) two management regimes; management by measurement, which is the regime within which performance appraisal is designed, and the alternative; management by delegation. By this, I will have established the framework against which I will investigate patterns of local adaption and meaning creation in performance appraisal across four workforces. Chapters 7 to 9 contain the detailed analysis of local adaption and meaning creation in performance appraisal in the case organisation across target setting, midyear review and final appraisal. In each of these chapters, a number of challenges in performance appraisal are analysed. In chapter 7 regarding target setting, the challenges of establishing commitment and including behaviour are analysed. In chapter 8 regarding midyear review, the challenges of conveying meaningful feedback, of using rating in feedback, of carelessness and of management disruption are analysed. Chapter 9 regarding final appraisal analyses the challenges of evaluating business performance, evaluating behavioural performance, designing rating distribution policy and of calibration of ratings. Finally, chapter 10 contains the conclusions and implications of the study, and chapter 11 contains important reservations, qualifications and guidance for future research. Appendices are found in Chapter 12, followed by the list of literature. 1. Introduction 2 challenge frames of meaning for managers and employees and provoke important meaning creation processes that affect their perception of the PA process. The paradox is the result of adhering to PA despite the lack of ability to resolve these problems. Further, I claim that one fruitful way to observe and analyse challenges in PA, local adaption and meaning creation processes is to pay attention to patterns of enactment in differing sub-environments in large organisations. One important place to look for such demarcations of environmental differentiation is between workforces. Through the study of local adaption and meaning creation in PA across four different workforces over a full one-year PA-cycle in a case organisation, I found important differences in environmental dynamics and in manager and employee enactment patterns. I propose that understanding these patterns is not only important for our understanding of how PA schemes affect individuals' and organisations' performance. It is also more fundamentally important to our understanding of the dynamics of the problems in PA which have been uncovered, analysed and attempted to be solved over the last 20 years by scholarly research into PA. By using Austin's (1996) separation between management by delegation versus management by measurement, a typology of differences in the task at hand in four workforces in a case organisation sets the stage for an in-depth meaning creation analysis across these workforces throughout an entire one-year PA cycle. Although PA as such is clearly designed to operate within the management by measurement paradigm, comparative findings across workforces show that important differences in dynamics are present. Particularly, Sales and R&D form contrasting cases in terms of characteristics of the task at hand, and characteristics of PA adaption. Thus, the present study shows how a standard corporate PA scheme can be twisted in different directions in different workforces to such an extent that in R&D it can even be argued that the scheme in some ways has more similarities with management by delegation than with management by measurement. This is important, because the clearly dominant approach in the literature on PA is to study large organisations as if they were single entities, while in reality they consist of highly differentiated sub-entities, each adapting standard systems differently in ways that are meaningful in their local contexts. 1. Introduction 3 By reaching a better understanding of this process of local adaption of a standard PA system we also add to the understanding of the paradox in PA. I propose that differences in the task at hand in different workforces form different subenvironments that enable local adaption of the standard PA system. In theoretical terms, local adaption means that particular local practices are established, supported by local meaning creation through stories that make the adapted PA system more meaningful in the local contexts. Thus, while PA schemes are ubiquitous in modern organisations and embedded with many problems, local adaption of PA is an important way for practitioners to establish local meaningful practices and stories about PA that counter the apparent lack of meaning of the standardised PA scheme. This is also an important proposition in a broader theoretical and practice context because the present study of PA is just one example of how one through the study of meaning creation processes can analyse local adaption of standardised management systems. While the present study illustrates how differences in the task at hand can be an important explanatory factor to understand significant differences in local adaption of a standard corporate PA system, there is reason to believe that similar differences would surface if one studied other standard management systems. The corporate PA system is not only enacted differently across workforces. The present study also shows how differences in enactment patterns in themselves articulate different problems in PA in a complex feedback process, so that different workforces experience different problems in PA, deal with them differently and thereby experience different feedback mechanisms. These feedback mechanisms again provoke different new problems in PA, or work to reproduce old problems in PA. To capture these dynamics and interdependencies, I propose the analytical model illustrated in Figure 1-1 below. I call the model the meaning creation feedback model because it shows the interaction between problems in PA, workforce specific enactment, meaning scheme challenge to individuals, meaning creation and a feedback loop to new problem articulation in PA. 1. Introduction 4 Figure 1-1: The meaning creation feedback model I propose that what is needed by both practice and theory is to put aside for a while the ever more subtle analysis of isolated problems in PA and turn attention towards dealing with problems in PA at a more dynamic level where local enactment patterns and problem interdependencies through feedback processes come into focus. One fruitful way to do this is to look at how PA affects the meaning creation of managers and employees working within PA in different ways depending on workforce affiliation, and how these enactment patterns again interplay with articulation of other problems in PA. Thus, although my contribution obviously builds on the existing literature about problems in PA, it attempts to go beyond the more common practice of examining narrow, single problems in PA by introducing a framework for analysing patterns of local adaption and meaning creation through enactment of workforce environmental filters and through feedback processes. 1.2 The Phenomenon PA can be defined as the appraisal rating of individuals' work performance and potentially their behaviours by management, covering a specific time period, applied to all employees or specific groups of employees whose participation is typically mandatory or alternatively motivated by access to extrinsic reward, and where results 1. Introduction 5 in the form of ratings are stored by the organisation to be used for purposes that require differentiation of employees.2 PA processes can vary in the details of their design, but as the definition indicates they also have important elements in common. Like e.g. Armstrong and Baron (2005: 23) do, most schemes operate within a rather generic PA process model based on a sequence of sub-processes ranging from “performance planning” and definition of "objectives" over "measuring", “assessing” and “rating” to “documentation”. Although the term "Performance Appraisal" is of newer origin, formal processes for talent differentiation, which is the core content of the concept, goes much further back. It would be far beyond the purpose of this thesis to give anything like a full account of the history of PA, but a few illustrative historical examples may be helpful for understanding how PA has developed into what it is today. This also serves as a short warm up to the literature review of the last 20 years' research into PA that will be presented in the following chapter. Thus, a very brief genealogy of PA could for example start approximately 1400 years B.C. with the Old Testament, as Wiese and Buckley (1998) do, where Moses evaluates and selects: ...the man who was known to be the most skilled craftsman from the tribes of Israel to build and furnish the tabernacle of the Lord. Later, in the Chinese Han Dynasty from 200 B.C. merit exams were introduced for selection and promotion of public officials, and "Imperial Raters" were appointed to rate high office holders on a nine point scale (Wren, 1994). With industrialisation, PA schemes were beginning to take shape in more modern forms from around 1800. So-called "Silent monitors" were introduced in the cotton mills in the shape of colour coded blocks of wood at each work station that made it visible to everyone what level each individual worker was performing at (Wiese and Buckley, 1998). Wiese and Buckley (1998: 235) specifically identify 1813 as the year when PA formally was introduced in the US. This was done by an army general, who: 2 This definition is an elaboration of Coens & Jenkins' (2002: 13-14) five-point "characteristics" of PA. 1. Introduction 6 Submitted an evaluation of each of his men to the U.S. War Department... The Army General used a global rating, with descriptions of his men such as 'a good-natured man' or 'a knave despised by all'. In his study of administrative history in the US 1869-1901 called "The Republican Era", White (1958: 353) documents how promotional exams were introduced to avoid unjust promotions, by counterweighing: importunate solicitations and coercive influence from the outside and prejudice, favouritism, or corruption in the part of the appointing officers. The documentation of performance was made in the form of "efficiency records" so that: we can ascertain with almost mathematical certainty the proportion of work done by each clerk in the Pension office in point of quantity as well as quality.3 However, a few years later, the evaluation of the efficiency records initiative by a Department of the Interior Committee was less favourable and it was decided to cancel the efficiency records initiative because of inflationary ratings caused by dynamics not unknown to our days' researchers into PA in modern organisations: The elaborate system of efficiency ratings and records, based, as it necessarily must be, upon the individual opinion of the officer giving the marks, has become little more than the mathematically expressed opinion of such officer as to the relative standing of the clerks under his supervision... the head of each division in the great majority of cases giving the highest ratings possible in order to put the clerks in his division ahead of or at least on an equal plane with the clerks of other divisions. (White, 1958: 358) Further, White mentions an illustrative example of how administrative suboptimisation was caused by different departments in the US Navy being responsible for defining their own deliverables, e.g. for the design of the "hull, the engines, the guns and the sail power" (White, 1958: 164). Although each department's targets and plans may have looked fine on paper and although all targets were met for each department, the combined result was disastrous: After the ship Omaha had been commissioned and was ready for sea it was discovered that the several bureaus, working independently, had so completely appropriated the available space as to leave coal room for only four days' steaming... The Omaha was at best merely an old repaired wooden vessel, whose rebuilding cost the full price of an up-to-date steel ship, and which could neither fight nor run away from any ship of a foreign nation. 3 The quote is from an internal government correspondence dated 1879, in White (1958: 356). 1. Introduction 7 In her book "The March of Folly" (Tuchman, 1984) gives a number of similar historical examples from Troy to the Vietnam War of wartime mismanagement, among which some have ties to critiques of PA schemes in that they are dealing with situations where targets are set and aimed for although this has an obvious counterproductive effect on organisational performance. In her view, folly is the "child of power" (Tuchman, 1984: 38) in the sense that powerful people define strategies and targets and cling to them long after it has become evident that this is not in the interest of themselves or the organisations they lead. Turning back to civilian organisations, Chester Barnard in his "The Functions of the Executive" (1938) theorised over how managers can persuade employees to perform to certain standards. In Barnard's analysis, if external incentives are not available in enough quantity to satisfy the need to motivate employees, management must turn to "persuasion". Barnard lists three types of persuasion; creation of coercive conditions, rationalisation of opportunity and inculcation of motives. All three have linkages to PA schemes. First, coercion can come in the shape of exclusion. Grades of exclusion vary and can e.g. be "homicide, outlawing, ostracism, corporal punishment, incarceration, withholding of specific benefits, discharge, etc." (Barnard, 1938: 149). Although Barnard finds it generally accepted that coercion does not to any great extent support performance in organisations of his age, he nevertheless emphasises how setting examples through coercion can be an effective management tool: to create fear among those not directly affected so that they will be disposed to render to an organization certain contributions (Barnard, 1938: 149). This is an argument not so far away from e.g. Kohn's (1999) perception of PA being a punitive system aiming to build motivation from installing fear of low ratings. Second, the "rationalisation of other incentives" equals what Barnard (1938: 150) terms "propaganda", a term with much relevance in the international political context of his book in the late 1930s. This is a process through which leaders: Convince individuals or groups that they 'ought', 'it is to their interest' to perform services or conform to requirements of specific organisations. This second type of persuasion: consists in emphasizing opportunities for satisfaction that are offered, usually in contrast to those available otherwise (Barnard, 1938: 150-152). 1. Introduction 8 It is not hard to see the analogy to the close relation between PA and monetary incentives like bonuses, wage increases etc. where it is also the "Do this to get that" pedagogy that is used. The third and most important type of persuasion is the inculcation of motives. Methods for this includes "Precept, example, suggestion, imitation or emulation, habitual attitudes" etc. These are all elements that can be recognised as possible ingredients in modern PA schemes in the shape of business and behavioural standards, targets and evaluation criteria. However, already in 1938, Barnard was aware of some of the possible pitfalls of incentive schemes: Opportunity for personal prestige as an incentive for one person necessarily involves a relative depression of others (Barnard, 1938: 156). Nevertheless, Barnard (1938: 160) concludes in what can be seen as a strong defence of the need for external incentives and differentiation of employees: No enduring or complex formal organization of any kind seems to have existed without differential material payments.... The same principle applies in principle and practice even more to non-material incentives. The hierarchy of positions, with gradation of honors and privileges, which is the universal accompaniment of all complex organization, is essential to the adjustment of non-material incentives to induce the services of the most able individuals or the most valuable potential contributors to organization, and it is likewise necessary to the maintenance of pride in organization. Lord Wilfred Brown in his classic "Exploration in Management" devotes a chapter to the relationship between "the manager and his subordinates" (Brown, 1960: 72-88). In this chapter, Brown directly relates to issues in employee PA: It is not difficult for a manager to know, or to find out, whether or not one of his subordinates has carried out the prescribed component of his job... It is much more difficult when we come to the discretional component... This seems to me to be the really difficult part of a manager's job. He is called upon to judge the executive work of one of his fellow human beings. I think everybody tends to shrink from such a task. The manager is faced with the fact that his decisions in these circumstances may have a considerable bearing on the future career of his subordinate and on the fortunes of his home and his family. Regarding the basis of assessing subordinates, Brown states that: The assumption is often made that a subordinate can be assessed on quantitative figures alone which indicate his performance... Unfortunately, the assumption is usually not valid; facts and figures are, of course, an aid to assessment, but they are seldom the whole story. 1. Introduction 9 And Brown even touches specifically upon issues in evaluation of sub-ordinates behaviour: The right of a manager to criticize work is clear; but the right and duty to make personal criticism of behaviour because it is such as to interfere with work by setting a bad example, disturbing relationships with others, introducing discord into the company, is not so clear. Modern managers and researchers struggling with theoretical, practical or ethical issues within behavioural target setting and evaluation in PA, can either comfort themselves by, or become worried by, depending on perception, the fact that these issues have been well known for at least 50 years, without much progress being made. Peter Blau (1955) in a case study in his "The Dynamics of Bureaucracy" examined the use of quantitative measurements for performance evaluation in a state employment agency in the US, and identified what he saw as dysfunctional consequences of the same: Statistical records facilitated the job of the superior by providing him with information that he would have had to ascertain laboriously otherwise and by improving his relationships with subordinates. It might be expected, therefore, that superiors favoured the exclusive use of quantitative indices for evaluating subordinates, a method of evaluation that would eliminate personal considerations in accordance with bureaucratic principles. This, however, was not the case. Super ordinates on all levels explained that it would be impossible to judge all aspects of performance on the basis of these indices alone; 'You can't reduce a man to statistic'. The validity of this opinion can hardly be questioned. Blau identified a number of dysfunctional effects of the introduction of statistics and quantitative principles in performance evaluation. First, it undermined the managers' authority because his judgment was no longer enough if he could not support it with quantitative evidence. Second, the relation between statistics and an individual's performance was rarely straight-forward, so the translation of statistics became a source of conflict in itself ("figures can't lie, but liars can figure"). Third, documented performance goals were displacing real value-adding goals as the prime focus of managers and employees. Or, in Blau's terms, there was a "displacement of goals whereby an instrumental value becomes a terminal value". Fourth, the thing that was most annoying to one employee who worked as an interviewer in the employment agency, was that the quantification of employee productivity created: Competition between interviewers to an extent that is - disgusting... they lead to competition and to outright falsification. 1. Introduction 10 Modern PA schemes have close ties to the Management by Objectives (MBO) wave which emerged back in the 1950’s (Drucker, 2007). Although MBO later grew out of fashion, the PA process survived and has become so institutionalised that it is today seen as a mere fact of organisational life - something that organisations have to live with and therefore something they may as well learn to live better with. Today, most large corporations have implemented and maintain standardised PA processes.4 Based on the feedback from people I interviewed, the institutionalisation seem to best be described in anthropological terms as a set of practices and rituals that define what it means to be part of social entities like modern organisations (Douglas, 1986). Or in broader sociological terms as mimetic isomorphism (Powell & DiMaggio, 1983), i.e. as a process of diffusion and isomorphism through a variety of channels like employee movements, standardised advice from consulting firms, business media discourse, higher education curricula etc. to a degree where there is an almost complete lack of fundamental cognitive reflection on why PA is there in the first place. This entails that for managers and employees PA has become a basic element in their perception of what is natural in organisational life. It is not the objective of this study to uncover the dynamics of how this diffusion has happened over time, although that would certainly be an interesting study. However, given the lack of reflection among managers and employees on the core question of why PA is there in the first place, it appears reasonable to conclude that mimetic isomorphism has stronger explanatory value than more rational decision making based institutional diffusion theories would claim, e.g. what Powell and DiMaggio (1983) called coercive pressures for isomorphism, where top management seek increased external legitimacy by fulfilling certain standards set by important stakeholders or authorities. PA is a theme that now plays a role in the work lives of most people in large organisations in the Western world. It has also been the object of broad scholarly interest over time. My database search identified more than a thousand academic journal articles from the twenty year period 1990-2010, where either title, keywords or abstract contained the word combination "Performance Appraisal". Today, a myriad of consulting firm offerings, standard IT solutions, benchmarking institutions, books, 4 Surveys indicate that more than three out of four US business organisations have implemented a formal PA process. (Coens & Jenkins, 2002:35). 1. Introduction 11 training courses and "best practice" schemes are offered to organisations to help them implement, re-implement or improve PA schemes. As such, the concept of PA is not just institutionalised in terms of its dissemination into large organisations throughout the Western world; it has also become big business in itself in which organisations invest heavily, both in terms of costs to advisory services and IT but also in terms of spending significant organisational resources on maintaining and fulfilling the schemes. The spread of PA in modern organisations, the institutionalisation of its design and the significant resources spent on it make it important for both scholars and practitioners to understand what problems there are in PA, if and how they can be solved and more generally how PA affects organisations and individuals within them. The present study focuses on how managers and employees across four workforces create and maintain meaning when they are faced with problems in PA. 1.3 The Research Question and Definitions In the following chapter I will in detail argue for the existence of an important gap in the literature. It is the aim of the present study to contribute to filling that gap in the literature by exploring local adaption and meaning creation processes in PA. Thus, I will examine patterns of how PA is locally adapted and how managers and employees across workforces establish meaning in situations where they are faced with problems in PA that challenge their existing meaning schemes. My research question is therefore: How is managers' and employees' meaning creation affected by problems in PA and what patterns of local adaption and meaning creation can be observed across organisational workforces? By studying local adaption and meaning creation processes I will add to the performance management literature and the on-going conversation about improvement or abandonment of PA. My contribution to the literature will be to extend the understanding of how problems in PA are perceived and dealt with differently across workforces within a corporate one-size-fits-all PA scheme, dependent on the particularities of the task at hand in each workforce. Further, I will add to the literature by showing how these patterns of meaning creation and local 2. The PA Conversation - a Literature Review 18 itself that although it is possible to improve and adjust the process, it is not possible to make it become a positive asset to the organisation. In Coens and Jenkins' analysis, PA schemes generally fail to deliver what they aim for. Coens and Jenkins quote a number of studies to show how organisations working with PA schemes are generally disappointed with the outcome of the efforts. Coens and Jenkins also quote an Industry Week survey showing that only 18% of the respondents considered their own appraisal reviews effective while 48% considered their review "second-guessing sessions" (Coens and Jenkins, 2002: 18). More so, Coens and Jenkins claim that it is simply impossible to find any trustworthy academic study that proves a positive outcome of conducting systematic PA in an organisation. Often, consultants and proponents of PA schemes emphasise the fact that most of the very successful companies in the world are using PA schemes. But it would not be difficult to find highly unsuccessful companies also using PA schemes, so this cannot in itself be taken as any kind of proof. For example, huge American companies like Enron and Arthur Andersen were using PA schemes and relied heavily on extrinsic rewards, but both failed to deliver sustainable organisational performance. Neither of those two companies exists today. For Enron, one commentator even directly blamed the PA system for playing a key role in the disaster: HR ... failed to put “teeth” in its performance management systems that would severely punish (or fire) individuals that kept secrets, took excessive risks or that violated the company's values or ethics. As a result, HR inadvertently sent a message to employees and managers that results, regardless of how they are obtained, are all that matters. (Sullivan, 2002). Maybe surprisingly to some, it is hard to find scientifically based evidence that supports the common perception that organisations using PA schemes are systematically more successful than organisations that do not. Let alone any evidence that if such relation even could be found, it would not be the cause of other parameters like organisation size, ability to attract talent, organisational culture or other possible explanatory factors. In fact, DeNisi and Pritchard (2006) found that academic studies into the effectiveness of PA have significant methodological flaws: ... in 68 of the 70 studies there was no control group for comparison... In fact, they reported in their meta-analysis that, when proper controls were considered, feedback actually had the effect of decreasing subsequent performance in one-third of the 2. The PA Conversation - a Literature Review 19 studies – and that these results were independent of the sign of the feedback received. Thus, the question of direct evidence for the effectiveness of this performance management intervention for improving performance must still await a more definitive answer. Coens and Jenkins develop their argument around 10 proposed false assumptions behind modern PA schemes. Since these arguments form the core of the radical critique of modern PA schemes, I will present these at some length in the following. First, appraisal schemes are assumed to be able to support multiple, often conflicting, objectives at the same time. For example, the objective to provide an honest dialogue about feedback and development on the one hand and the objective of evaluation for bonuses on the other hand. Second, the schemes are built on the assumption that one size fits all - that one uniform setup for dialogue and feedback will serve the needs of all managers and employees, when in fact people have very different needs and strengths that would often be better accommodated in more flexible setups. Third, it is assumed that it is possible to provide employee commitment from a top-down enforced process, when in fact commitment is more often thwarted by such a paternalistic process, with a following decrease in performance. Fourth, performance schemes build on the assumption that it is managers who are responsible for compiling feedback and driving performance, which is detrimental to a more empowering management style where employees themselves take responsibility for compiling feedback and the manager takes a more coaching role. Fifth, it is falsely assumed that it is possible to provide anything near objective, comparable evaluations of employee performance across large complex organisations. In reality, bias is everywhere and ratings will be highly inaccurate and based on incomplete information.8 Sixth, it is often assumed that PA schemes will be helpful when dismissals of employees are required. However, PA schemes are in practice detrimental to the organisation's interests in such cases, since most managers are shy of conflicts and thus tend to avoid giving low-end ratings. Such “too friendly” ratings can be ammunition for the employee as proof of unreasonable dismissal, rather than being useful for the organisation as proof of reasonable dismissal. Seventh, it is assumed that it is motivating to compete for high ratings. However, studies show that most humans see themselves as performing above 8 Throughout this thesis I will use the term 'bias' understood as a "cause to feel or show inclination or prejudice for or against someone or something" (Oxford Dictionary) 2. The PA Conversation - a Literature Review 20 average, so a majority of the employees will in reality be disappointed by the rating they receive, and disappointment is demotivating rather than motivating. Eighth, PA schemes institutionalise in the organisation that feedback on performance and development is a 3-times-a-year standardised and formalised activity, when it would be much better to have a much more frequent, informal feedback culture in the organisation. Ninth, it is assumed that people's motivation is highly dependent on extrinsic factors like performance ratings and bonus, while Coens and Jenkins claim that stimulating intrinsic motivation is a much stronger and lasting way to achieve employee commitment and performance. And lastly, it is assumed that the organisation can be effectively improved by focusing on inspection of the individuals, while it would have much higher impact to look at the systemic level of the organisation and focus on how organisation and processes can be improved. Kohn (1999) draws similar conclusions in his book "Punished by Rewards - the trouble with Gold Stars, Incentive Plans, A's, Praise and other Bribes". Kohn takes his point of departure in research into children’s learning in school and through their upbringing. From this starting point, Kohn also studies the use of rewards in organisations, based on formal appraisal processes. His theoretical perspective is socio-psychological and Kohn sees the appraisal process as based on an oldfashioned, mechanistic view on humans, built on “our reliance on a theory of motivation derived from laboratory animals”. This is with reference to the fact that Kohn sees PA to be a legacy from Skinner's (1974) behaviourism. Behaviourism was focusing on how it is possible to affect human action in an intended direction by rewarding wanted behaviour and punishing unwanted behaviour. Kohn simply rejects the idea that it is possible to motivate people through a "do this and you will get that" doctrine (Kohn, 1999: 3). Whereas many modern managers today have less belief in punishment as a tool, it is much more common to see rewards of many sorts, from praise to bonus, as an important way for management to create motivation among employees. In his interesting study Kohn asks two fundamental questions; Is it right to reward? And is it effective to reward? Kohn finds the answer to both questions to be clear nos. It is quite clear that modern PA schemes are built around a reward and punishment architecture. To put it in Kohn's terms, goal setting is the "do this". The rating and the associated consequences for bonus, promotion, dismissal etc. is the "get that". 2. The PA Conversation - a Literature Review 21 Kohn's critique of appraisal systems and other reward systems is a close kin to a much older discussion among motivation theorists and work psychologists, namely the discussion of the effects of intrinsic and extrinsic motivation, and how the latter affects the former. As such, together with Coens and Jenkins’ critique of the false assumptions behind PA, Kohn's critique of reward-based systems fulfils the picture of how the radical paradigm builds its argument against modern PA schemes. So what is the problem with rewards? First, Kohn argues that rewards are in fact just one side of the coin where the other is punishment. Rewards are punishing in the sense that they are controlling by nature. Rewards direct humans to behave in certain ways that may be detrimental to who they feel they really are or how they feel they can best realise their potential, and this control aspect has in itself punitive character. But even more important, rewards will not be obtained by everyone to the degree hoped for and the effect of this on the individual will be similar to the effect of punishment. The individual will experience that something he deserved has not been granted to him. The more desirable the reward is the heavier the employee will feel punished by not getting it, and the more negative impact will it have on the employee’s work morale. Kohn finds that unexpected rewards are for this reason in fact much less harmful to employee morale, but ...apart from the practical problems of trying to keep people from expecting another reward tomorrow, it is no coincidence that the great majority of rewards are promised in advance. The whole point is to control people's behavior. (Kohn, 1999: 53) Second, Kohn claims that reward schemes are destroying relationships, both horizontally between colleagues and vertically between managers and employees. Often rewards are provided based on relative evaluation of peers. As in my case study, many organisations create a ranking scale that introduces an artificial scarcity of high rankings that then can be distributed among peers, with high-end rankings being most rewarding. The scarcity of high-end ratings and the relativeness will by nature create competition between peers as they will not only see each other as rivals, in fact they will be installed as rivals. According to Kohn, competition among colleagues creates anxiety that typically interferes with performance. But even worse, the majority of the employees will soon realise that they will realistically not be candidates for high rankings and will thus be discouraged from making an extra effort. Can this be dealt with through team rewards? According to Kohn it definitely cannot. Where most people are against collective punishment, collective reward is 2. The PA Conversation - a Literature Review 22 just another variant of the same - exactly as it is at the individual level. Only, at the collective level rewards install even stronger psychological pressures on the individual in the form of group pressure and exposure of weaker individuals as scapegoats for the group's failure to achieve a reward. So, rewards are counterproductive when it comes to the ability to realise a workplace characterised by trust and open communication. Third, Kohn sees rewards as generally ignoring reasons. It is very easy and convenient for the manager to be able to tick off if a target in an employee's performance plan has been met or not. But in today's complex workplace the realisation of a target very often is a result of multiple factors among which the employee's contribution is only one - and often only a minor one: Rewards are... gimmicks, shortcuts, quick fixes that mask problems and ignore reasons. They never look below the surface... A sudden deterioration in performance frequently turns out to be due to problems at home. A chronic record of mediocre performance, meanwhile, may indicate, among other possibilities, that there is something wrong with the job itself... Turning the workplace into a game show [ ] does exactly nothing to solve these underlying problems and bring about meaningful change. (Kohn, 1999: 60-61) Fourth, Kohn states that rewards generally discourage risk-taking. Rewards narrow people's focus so that anything not immediately instrumental in achieving the reward will be ignored. Rewards also have the effect that people will do exactly what is necessary to get the reward, but nothing more. The purpose of working is not to be creative or innovative. The purpose of working is to "get the goodie". In this way, rewards encourage people to focus on obtaining the reward with as little effort as possible, instead of going into a task solution with the purpose of exploring options of getting the best possible outcome: If you have been promised a reward, you come to see the task as something that stands between you and it. The easier the job is, the faster you can be done with it and pick up your prize... Do rewards motivate people? Absolutely. They motivate people to get rewards. (Kohn, 1999: 65-66) Lastly, Kohn finds that extrinsic rewards are not supplements to intrinsic motivational factors; they are directly detrimental to them. Kohn draws on a large number of psychological experimental studies on motivation that show how extrinsic rewards reduce intrinsic motivation. Reward is like a drug that makes employees addictive and destroys their own intrinsically founded commitment to a task. In fact, Kohn believes the effect is so strong that 2. The PA Conversation - a Literature Review 23 Rewarding people might even be regarded as a clever strategy for deliberately undermining interest in something. (Kohn, 1999: 72) Of course, rewards provided as ratings through PA schemes are no better. Kohn finds that they have the same harmful effects on employee motivation as any other reward scheme.9 Through his empirical studies, Kohn claims to demonstrate that not only do rewards not work in the workplace, they also do not work in the classroom and neither do they work to help children become good people. The only thing that really works when it comes to creating engagement, commitment and creativity, according to Kohn, is to create an environment that supports the blossoming of intrinsic motivation. Research from different positions within the paradigms As my literature review of journal articles on PA will show, only a small minority of today's scholars seem to subscribe to the radical paradigm. However, it would be wrong to assert that scholars uncritically position themselves at the opposite end of the attitude dichotomy. Rather, the vast majority of scholars study aspects of and issues with PA schemes, including some of the themes brought forward by Kohn and Coens & Jenkins. They position themselves within the dominant paradigm because they do not fundamentally challenge the legitimacy of appraisal schemes in the first place. They do not agree or even consider the view of the radicals, that it is plainly impossible to achieve a net benefit from conducting PA schemes, no matter how much effort one puts into designing a system that accounts for all the known biases and flaws associated with such schemes. It is fair to say that the radical scholars I have identified are basically concerned with many of the same issues as the more mainstream scholars. The dividing line is whether a scholar believes or does not believe that it is possible to overcome the issues or whether the issues are symptoms of such fundamental aspects of human nature, that PA schemes in any form are doomed to failure. The issues at stake are the same, but the conclusion is more radical. Thus, there is a large body of scholars with varying degrees of scepticism regarding the possibility of harvesting significant net benefits from the conduct of PA, but who 9 For example, Kohn (1993: 42) refers to behavioural research as far back as from the 1960’s which documents that financial reward is not only questionable but in fact directly negative in its impact on human performance. 2. The PA Conversation - a Literature Review 24 are not concerned with fundamentally questioning the legitimacy of the process itself. Rather, they are critically investigating challenges with the theoretical concepts or empirical practice of PA. As we shall see below, their critique takes its departure in many different theoretical perspectives, for example justice theory (Narcisse and Harcourt, 2008. Roch et al., 2007), motivation theory (Buchner, 2007. Kuvaas, 2006), organisational politics (Salimäki and Jämsen, 2010), social cognitive theory (Wood and Marshall, 2008), social learning theory (Tziner et al., 2002), Total Quality Management conceptual thinking (Soltani et al., 2004), discourse analysis (St-Onge et al., 2009) and more. For many authors, optimising the value provided to an organisation by conducting PA schemes is very much about handling rating dysfunctions to secure a fair and transparent rating process. Before turning to my systematic literature review of 125 journal articles on PA, let us take a brief look at some of the most important dysfunctions that has been uncovered through research into the rating process. In his book "Measuring and managing performance in Organizations" Austin (1996) tries to identify which organisational factors allow measurement to work successfully, and which force measurement programs to fail or do damage. Austin dedicates a chapter to what he calls measurement dysfunction, by which he means that measurement in itself becomes the cause of effects that are the opposite of what is valued. According to Austin (1996: 21) a key differentiator is whether measurements are motivational (intended to affect people being measured) or informational (used to build a deeper understanding of what is going on in an organisation). PA is in theory occupied with the former, since target establishment and evaluation is installed to motivate employees to work hard and dedicated towards target realisation. However, a problem is that informational measurement systems are often used to investigate to what degree employees realise the motivational targets (e.g. sales figures, production output, number of customer complaints, etc.). And this gives employees an incentive to try to affect the informational measurement systems so that they operate in favour of proving employees' target fulfilment rather than in favour of showing what "really" goes on. In other words, the instalment of motivational incentives that are to be measured by "objective" informational measurement systems has as a consequence that employees get an incentive to manipulate the measurement systems and thereby undermine the trustworthiness of the output 2. The PA Conversation - a Literature Review 25 generated by the measurements. This can be done in many ways, e.g. when a sales person strategically places an order on the right side of a calendar date, to optimise his own reward scheme. Or when customer complaints are not logged or documented, because this would be detrimental to an employees' target realisation. According to Austin, another key issue is that measurement systems have a built-in tendency to look for measureable indicators. So what is measured is rarely ever the real value of an activity, but an indicator of value, because it is always easier to measure quantity than quality. It is easy to measure if a project delivers on time, if a certain number of cases have been closed or if targets for average turnaround time on case handling have been met. But did the project deliver the value that was anticipated? Or did the agent handle all the easy cases to reach the target, leaving behind the more difficult but maybe more important cases? The solution installed to handle this issue is often even more refined targets that again persistently run into new issues because they again are targets on indicators of value, not of added value in itself. More sophisticated evaluation systems only serve to give more sophisticated dysfunctional reactions: Employees' true output (such as value to the organization) is often intangible and difficult to measure; in its place, organizations choose to measure inputs (such as the amount of effort devoted to the task as measured by counts of interviews performed). (Austin, 1996: 18) Austin sees a strong link between McGregor's (1960) theory X style of management and many performance measurement schemes, because they rely so heavily on extrinsic motivation and share the assumption that humans will attempt to avoid work if they are not coerced and controlled. Austin introduces the concept of measurement cost as something any organisation should consider before installing measurement activities on any working process. Some of the factors affecting measurement costs are the repetitiveness of the task (the more repetitive, the easier to measure), the complexity of the task (the less complex, the easier to measure), newness of the task (the more well-established the easier to measure), the level of specialised knowledge required by the task (the lower level required the easier to measure), interdependence and separability of effort (the less interdependent and the more separability of effort, the easier to measure) and finally also any environmental covariate and other noise (the less of this, the easier to measure). It is not difficult to follow how Austin concludes, that in many modern 2. The PA Conversation - a Literature Review 26 workplaces dominated by highly educated knowledge workers, all these factors will turn out to be in disfavour of measurement. However, measurement schemes are everywhere in such organisations. As the objectivity of employee rating is a cornerstone in almost any PA scheme, other scholars have focused at the subjectivity caused by human biases that are unavoidably present in any human relationship. Some examples of these biases are: managers tend to categorise according to their own prejudice and stereotypes; managers favour people they personally like; managers intentionally or unintentionally discriminate against race, gender or age; managers are shy of conflicts and tend to be too lenient when rating; manager's rating of employees is affected by the rating the manager has received herself; managers tend to cluster their ratings in the middle of the rating scale, not using the extreme ends; managers are affected by the halo effect, causing them to carry perceptions of an employee's performance in one area into perceptions of the employee's performance in other, separate areas; managers (and employees) focus on recent performance, leaving out performance activities not conducted within the last few months up to the PA interview; and finally, managers and employees are prone to attribution errors, so that favourable outcomes are attributed to oneself while unfavourable outcomes are attributed to external factors or someone else (Coens & Jenkins, 2002. Furnham, 2005. Armstrong, 2008. Austin, 1996). Other examples of authors who are critical towards PA without abandoning it altogether are e.g. Amabile (2011) and Frey (1997). From a motivation theory perspective Amabile focuses in her article "Motivation" on particular issues with extrinsic reward schemes in relation to innovation and creativity. Through a typology of personality traits related to "highly creative people" containing elements like e.g. spontaneity, curiosity, openness to change and experience, flexibility, etc. she concludes that creative people are mostly driven by intrinsic motivation while extrinsic motivational pressure often will have a detrimental effect on creativity. Frey (1997) has in his book "Not just for the Money" investigated the effect of extrinsic reward on intrinsic motivation and concludes that under certain conditions there is a clear crowding-out effect in the sense that higher extrinsic reward is detrimental to inner motivation: 2. The PA Conversation - a Literature Review 27 To offer higher pay then makes people less committed to their work, and may reduce their performance (Frey, 1997: ix) Here, the conclusion is the same as Amabile's, although for a broader scope, since Frey is not particularly focusing on highly creative people, but more generally on the effect of extrinsic motivation schemes on human behaviour. Frey's argument is particularly interesting because it takes its point of departure in an economic approach like so many PA text books within the pro-PA dominant paradigm. He does not build his argument on qualitative sociological case studies or non-empirical arguments like some proponents of the radical paradigm tend to do (Bowen, 1995. Fisher, 1994. Townley, 1999). Rather, he investigates the crowding-out effect on the home turf of quantitative economics; through controlled experiments, statistical analysis and mathematical theory building. Through his analysis, he identifies a number of important conditions for the crowding-out effect, including intervention in the shape of "monetary reward" or "regulations". In both cases, the condition under which an employee will experience the crowding-out effect is if the monetary reward or the regulation is perceived as controlling. And as we shall see, PA schemes utilise both elements by installing regulations in the form of targets defining the employees focus, deadlines, approach and priorities, while promising reward, if the employee delivers within the frames of these targets. In short, Frey concludes that a rational agent increases his performance until the marginal benefits equal the marginal cost of so doing. The agent's optimal performance is affected when the principal changes the extent of external intervention. So the art of intervention for management is to identify the optimal level of intervention so that the crowding-out effect is not larger than the positive effect of the intervention. Again, all the above examples of dysfunctions and biases can be investigated with or without accepting or rejecting PA schemes as meaningful tools altogether. As already indicated, most scholars study PA from a critical position to improve or even reform the PA process. But they do this without fundamentally questioning whether it is at all possible to gain a positive net outcome from the process. Table 2-1 below summarises the key differences between the dominant and the radical paradigms. 2. The PA Conversation - a Literature Review 34 Western universities.12 And since my own empirical object is located in a Western context, it will at least culturally fit to the main body of literature in my review. Third, one could criticize the choices I have made for slicing and dicing the body of literature. Why focus on methodology? Why be interested in the theoretical perspective of the author? Why not have gender or age dimensions or some other often used independent variable? My aim was to contribute to the field of PA studies both methodologically and theoretically, to be able to maybe shed some new light on the phenomenon. When I started to read through the material and the must-read monographs about the subject, I found that much of the literature was dealing with the same problem of optimising a process that somehow never seemed to work optimally anyway. So, to be able to come up with something new, I thought that by getting a good overview of what had been done already and then do something different, would optimise my chances to succeed. I hope the reader will find that the following chapters will prove me right. 2.4 Analysing the Empirical Focus of the Literature Before looking at the methodological approaches applied by the authors I will start by turning my attention to what empirical objects the authors are concerned with in the first place. Public vs. Private sector focus First, when the articles are coded for their empirical focus on public vs. private sector, we see that 50% of the empirically based articles are investigating private sector organisations, while 36% are investigating public sector organisations.13 The remaining 14% are investigating populations across both sectors. In a global context this distribution seems to overweight public institutions significantly when compared to the public sector's relative share of the job market in most non-Western countries where welfare state institutions are less developed. 12 E.g. on a list of the world's top 50 business schools compiled yearly by The Economist and Businessweek, all 50 schools are located in Europe and North America (2011). 13 In my coding, I have defined higher education institutions as public sector organisations, even though many universities, particularly in US and UK, are often dependent not only on public funds but also on private tuition fees and donations. Even in privately funded universities, for the employees, there are more similarities in culture and roles to other public universities than there are to more classical for-profit private organisations. 2. The PA Conversation - a Literature Review 35 But even if the distribution fits better to a Western perspective where public sectors are bigger, public sector studies are still overrepresented in the articles. Public employment generally comprises 10-30% of the entire employment base in Western countries, with Scandinavian countries positioned in the top end of the interval (Handler, 2005). An un-weighted average of public employment share of the entire employment base across EU and US was approximately 16% over the period 19802002 (Handler, 2005). Since articles focusing exclusively on public sector make up 36% of all investigated empirical articles, and, as we will see below that European and US populations heavily dominate the empirical focus of the articles, it is reasonable to conclude that public sector studies are over-represented in the PA studies by roughly a factor of two. One reason for this over representation could be that scholars have a preference for conducting empirical studies using populations from their own world: university students and staffs, as we shall see below is in fact the case for some types of studies. Another reason could be that researchers have better access to public sector organisations, although this is not possible to verify through my review. It is beyond the scope of this study to explore this imbalance in detail. However, this could be an interesting thread to take up in other studies. In the context of my own study, the perspective of public vs. private sector organisations becomes most interesting if one sees a difference in PA practices across the two sectors. To get a picture of this, a good place to look is in the articles that have a cross-sector comparative design. Prowse and Prowse (2009) examine research through a literature review of articles that evaluate the aims and methods of appraisal, and the difficulties encountered in the appraisal process. The context is primarily US and UK. It is found that in terms of use of PA as the basis for merit pay, there is little difference between public and private sectors, since 65% of the public vs. 69% of the private organisations having both a merit pay and a PA system had implemented such a link. On the other hand, Prowse and Prowse sees a difference in what the employees in the two sectors see as the top issue with PA schemes: in the public sector it is lack of feedback and awareness, in the private sector it is lack of linkage between effort and performance (Prowse and Prowse, 2009: 74). 2. The PA Conversation - a Literature Review 36 In a Jordanian context, Abu-Doleh and Weir's (2007) study reveals that private organizations’ PA had significantly greater impact than their counterparts in the public sector on key human resource decisions like promotion, retention, lay-offs and for identifying individual training needs, transfers and assignments. In an Indian context, Amba-Rao et al. (2000) examine PA differences between public organisations, private organisations and foreign owned/joint venture organisations. The study concludes that foreign owned/joint venture organisations were conducting the most formalised PA schemes. Further, it was found that public sector organisations were the least likely proponents of a link between PA and pay or disciplinary consequences. Lastly, an interesting finding was that while public sector organisations were focusing heavily on external compliance with legal and regulatory guidelines, private sector organisations were more likely to accord ethical standards (behaviour) greater priority beyond economic efficiency and legal compliance. Unfortunately, only these three of the thirteen articles working with cross-sector populations, makes any conclusions on differences between the sectors. The other ten articles (Millward et al., 2010. Cheng and Cascio, 2009. Sudarsan, 2009. Stark and Poppler, 2009. Wright and Cheung, 2007. Soltani et al., 2006 & 2004. Tziner et al., 2001. Raymark et al., 1999. Snape et al., 1998) have other foci. So, as we have only weak evidence of any differences that would be of relevance also to my own study, I prefer only to conclude two things. First, that there is a significant overweight around a factor of two, of empirical focus in favour of public sector organisations when compared to their relative share of the employment base. My own focus on a private, pharmaceutical sector company will make a little contribution to turn the scale of this imbalance. Second, the scene is still open to researchers who want to make comparative studies of differences in PA schemes and practices across the public and private sectors. Geographical focus If the public/private sector study distribution is somewhat skewed in the articles, the picture is not much different when one turns to the geographical point of view, as illustrated in Figure 2-1 below. 2. The PA Conversation - a Literature Review 37 Figure 2-1: Distribution of empirical focus by continent Figure 2-1 shows how North America with 46%, or almost half of all journal articles selected, dominated the empirical focus of the articles. As runner up Europe reached 21%, Asia 16%, Oceania 7% and various combinations of countries across continents covered the rest. Looking at a country by country level, the top 5 countries in terms of share of the 125 analysed articles were: US 40%, UK 17%, Canada 6%, Australia 6% and China 4%. However, even this strong bias towards North America and Europe at a closer look could be argued to over-estimate the empirical focus of academic articles on nonWestern populations. Since Asia also includes countries we would normally consider western in terms of economic and political system, i.e. here Japan and Israel, and Oceania covers Australia and New Zealand which we would also consider Western, one could argue to see the distribution as skewed with up to 82% Western studies. This needs to be held up against the fact that Western populations only account for approximately 15% of the total world population. So, no matter if one argues that the 3 articles focusing on Japan and the 4 articles focusing on Israel should belong to one or the other category, it is clear that there is a strong and systematic bias in the empirical focus in favour of Western populations. Since my own study places itself in this mainstream tendency by focusing on a Scandinavian context, I can hardly claim to be setting out to do anything really original in terms of geographical focus. The broader question if and how non-Western 2. The PA Conversation - a Literature Review 38 cultural particularities may or may not have an impact on practice of PA in organisations and the way scholars analyse it, I will leave for others to investigate. Workforce focus Turning the focus from the geographical and public/private sector dimensions to the workforce dimension, other interesting insights become available. Each of the 125 articles was coded for their focus on specific or mixed workforces. My taxonomy included five main categories: White collar, blue collar, R&D specific, Management only and HR professionals only. The outcome of the coding is illustrated in Figure 2-2 below. Figure 2-2: Distribution of empirical focus by workforce14 As is evident from Figure 2-2, 65% of the empirical articles were focusing on white collar employees and 19% on managers only. Very few articles focused on blue collar, HR and R&D or on combinations of the categories. Further, one could reasonably argue that the management workforce is most often also white collar work even in e.g. a production or retail company setting with blue collar workers. And HR definitely is a white collar category. Thus, it would be reasonable to conclude that no less than 85% of all empirical articles have a white collar focus. It is obvious that the above documented geographical focus plays a role in this, since the Western economies over the past decades have turned more and 14 In percentage of 91 empirically based articles out of 125 analysed journals 2. The PA Conversation - a Literature Review 39 more into service oriented economies with increasingly smaller proportions of classical blue-collar jobs. However, there is also another bias at stake. Probably for convenience reasons scholars often use higher education students or university staffs as their empirical basis when doing social research, particularly when using experimental or quasi-experimental methods. Thus, three out of the four articles coded as using experimental or quasi-experimental methods in my review, were using university staffs and students as their empirical object (Payne et al., 2009. Cardy et al., 1998. Wagner and Goffin, 1997). Peterson (2001) conducted a large meta-analysis of the implications of the use of students in empirical studies as representatives for other groups. He concluded that not only did students show a slight but statistically significant higher homogeneity than non-students; he also concluded that "effect sizes derived from college student subjects frequently differed from those derived from nonstudent subjects both directionally and in magnitude" (Peterson, 2001: 1). If the heavy focus on white collar workers and managers was not surprising, I still was surprised to see that no attention has been given to cross-workforce comparative studies. In fact, only 8% of the articles were looking at populations covering more than one workforce when studying PA (Brown et al., 2010. Soltani et al., 2006. Brown and Heywood, 2005. Woods, 2003. Shibata, 2002. Tziner et al., 2002. Shibata, 2000. Shaikh, 1995). But of these, none actually performs a comparative study to examine workforce differences. Woods (2003) conducts a cultural comparative study including multiple workforces, but his comparison parameter is national culture rather than workforce, as he sets out to investigate differences in Australian and Singaporean PA practices. Shibata (2002) is working from the same approach, but compares US and Japanese PA practices and underlying cultural causalities. Tziner et al. (2002) use both white collar and blue collar workforces in US and Israel, but his comparison is not between workforces, but between two different approaches to rating scale and performance feedback systems. Shibata (2000) conducts a diachronic comparative analysis of the interrelation between performance rating and wage before and after implementing a new PA system in a Japanese organisation. The other four articles with crossworkforce empirical data are not comparative studies. So, while these comparative 2. The PA Conversation - a Literature Review 40 studies based on cross-workforce populations are interesting, none of them are actually dealing with workforce as the comparative dimension. Summary of empirical review By looking at the empirical focus of my sample of the last twenty years' research into PA a number of findings of interest to my own study appeared. First, there is a significant overweight of public sector studies as opposed to private sector studies, when compared to the share of the full employment base of the two sectors. Since some studies indicate that there are differences in the experiences with PA between the two sectors, this indicates a need for further focus on private sector studies. Second, there is a huge overweight of studies within a Western context. Almost two thirds of all the empirical studies in my sample were made on populations from three countries: US, UK and Canada. Only one of the articles focused on a Scandinavian empirical context, where one could expect to find deviations from the broader AngloAmerican findings, since Scandinavia is often presented as having a more democratic leadership tradition and a less significant power distance than almost any other region in the world. This could leave room for my own study to add to a more balanced picture of PA, although I will not be able to change the overall imbalance between Western and non-Western studies. Third, there is a significant overweight of empirical studies into one single workforce. 85% of all studies were focusing on a white collar empirical population. Only 1% of the articles were focusing on R&D specifically. Even more interestingly, only 8% of the articles were studying more than one workforce, but none of these were in fact comparing experiences between workforces. Thus, my literature study clearly indicates that there is a lack of comparative research into PA practices across workforces, in a non-Anglo-American context and within the private sector. Is this an important gap? In my opinion it is. Multiple authors have investigated particular groups of employees to identify the specific challenges they meet in a PA context. For example Ubeda and Santos (2007) focused on R&D, AbuDoleh and Weir (2007) focused on HR, Yariv (2006) focused on managers and many focused on white collar workers. But comparative studies were surprisingly absent. To improve our understanding of possible differences between workforces it is necessary to take one more step from all these single workforce analyses, and 2. The PA Conversation - a Literature Review 41 conduct a truly comparative study of workforce particularities. Having good access to a large organisation that has implemented a uniform PA system across multiple workforces puts me in a good position to start filling that gap. 2.5 Analysing the Methodological Approach of the Literature The 125 articles were coded for their methodological approach, and the distribution can be seen in Figure 2-3 below. Figure 2-3: Distribution of methodologies utilised in 125 analysed articles Figure 2-3 shows the distribution of methodologies utilised in the 125 articles. The largest category is by far the quantitative studies comprising 43% of all 125 articles. This is not surprising, since it is often referred to as a fact that many top tier journals have a preference for quantitatively based analysis (Marschan-Piekkari & Welch, 2004: Chapter 27). But here we are more concerned with what it means for the type of findings the authors are able to produce. Quantitative methods Quantitative methods rely on numbers and statistical analysis of numbers. In this context the quantification is most often based on survey questionnaires, and sometimes correlated to other quantifiable data, e.g. size of bonus, salary increase or other rewards. The quantitative approach is well suited if one is out to produce statistically reliable tests of hypothesis about already defined problems where a few 2. The PA Conversation - a Literature Review 42 independent variables can be defined and isolated. Often, this approach is chosen based on a desire to test hypothesis about respondents' opinions or attitudes towards certain elements in the PA system under which they operate. For example, Salimäki and Jämsen (2010), Thurston and McNall (2010), Roch et al. (2007), Vest et al. (2000), Waite and Stites-Doe (2000) and Holbrook (1999) all set out to test hypothesis regarding employees' perception of the degree of fairness in the PA scheme under which they work. Salimäki and Jämsen (2010) analyse employee survey responses from three governmental organizations (N=367). The purpose is to see if perceptions of politics and fairness are predictors of whether the pay system was perceived effective in achieving its objectives. The study concludes that there is such a correlation. Thurston and McNall (2010) conduct a nested confirmatory factor analysis of US employee responses (N=188) and investigates if four different categories of perceived justice in the organisation can be correlated to satisfaction with the organisation's PA system, appraisal ratings and employee satisfaction with their managers. The study concludes that these relations can be verified. Roch et al. (2007) conduct two US studies, (N = 208) and (N = 393) to investigate employees' preference for absolute vs. relative rating scales. The studies find that absolute formats are perceived as significantly fairer than relative formats. However, it is also found that significant differences in fairness perceptions also occur within relative and absolute formats. Other examples are Cook and Crossman (2004) and Snape et al. (1998) who both investigate the more general employee attitudes towards PA schemes. Cook and Crossman (2004) conduct a UK survey (N = 382) to test a hypothesis that a person who's role is appraisee will be less satisfied with a PA scheme than a person who is both appraisee and appraiser (typically managers). In the public sector sample examined, there was no significant correlation so the hypothesis could not be verified, although other studies have indicated so (Cook and Crossman, 2004: 533). Snape et al. (1998) examine attitudes on PA in the UK and Hong Kong. Respondents (N = 276) were mature part-time MBA students and graduates; (N = 141) were from Hong Kong and (N = 135) from UK. The study finds that British appraisals tend to be more participative and to place greater emphasis on discussing objectives, 2. The PA Conversation - a Literature Review 43 development and career plans. Hong Kong appraisals on the other hand appear to be more directive and Hong Kong respondents perceive a higher level of ‘negative’ appraiser behaviour. Nevertheless, Hong Kong respondents show stronger support for appraisal’s use for reward and punishment and less support for the objectivessetting and training and development uses than do the British sample. Tziner et al. (2002, 1998), Langan-Fox et al. (1998), Shore et al. (1998) and Greller (1998) all examine the effect of bias in PAs. Tziner et al. (2002) examined a hypothesis that managers' conscientiousness moderates the impact of their attitudes toward PA on their rating behaviour. The hypothesis was verified, as correlations between attitudes and rating behaviour were significantly weaker for managers with high conscientiousness. Langan-Fox et al. (1998) conducted a survey among shop floor employees in an Australian government enterprise (N = 476) where participation in the PA scheme was voluntary for the employees. Here, the bias the researchers were looking for was the prejudiced attitude of the employees when deciding whether to participate in the scheme. It was found that attitudes towards PA in general, perception of consequences of PA, satisfaction with the supervisor, anticipated feedback and awareness of the system were correlated to the decision to adopt or reject the scheme. Other hypothesis handled by utilisation of quantitative methods include employee coping strategies, general process satisfaction, rating effects for employees, relation between PA and stress, rating reliability for employees working in teams etc. As should be evident from these examples, authors utilising quantitative methods often have well defined, narrow and unambiguous questions and/or hypothesis that can be investigated by asking closed questions in a survey format. Qualitative methods Quantitative studies are in many ways contrary to the kind of studies conducted by the authors utilising qualitative methods, which account for 14% of the articles. Qualitative research is research that produce findings without statistical procedures or quantification (Strauss & Corbijn, 1990: 17). These authors are most often embarking on more open questions. They are not interested in setting up hypotheses to be tested through closed question surveys. Rather, their questions and 2. The PA Conversation - a Literature Review 50 First, quantitative studies were significantly more frequent than any other methodology. There were three times as many quantitative studies than qualitative studies in my sample. Second, the group of quantitative studies, experimental studies and quasiexperimental studies were generally characterised by narrow and unambiguous research questions, often formulated as one or more hypothesis testable against closed question survey responses from a representative group of employees. On the other hand, articles concerned with more open ended, exploratory questions were typically utilising a qualitative approach based on deep dive interviews or a case analysis approach. Third, except for only two studies, all empirical articles are concerned with data collected at one given point in time. Only two studies have a longitudinal design. When presenting my study protocol in a later chapter, I will elaborate on why I find this to be very strange since PA in my opinion is an obvious process to study by longitudinal designs. Fourth, only four of the articles are using a mixed-methods approach, although it is generally agreed that methodological triangulation is a way to increase validity in most studies (Yin, 2009. Alvesson and Sköldberg, 2009). Thus, my literature review clearly indicates that there is a lack of longitudinal, qualitatively oriented studies which handle the more complex aspects of PA. Further, if such a design could be combined with a methodological triangulation through the use of supplementary data establishment methods, there seems to be a good chance to create new knowledge about PA. 2.6 Analysing the Theoretical Approach of the Literature Coding for theoretical approach was the most difficult part of the literature review. Not all authors are so nice as to start by stating under which single theoretical approach they intend to conduct their study. In fact, surprisingly many authors do not reflect on what existing theoretical perspectives their contribution relate or contribute to. Others are multi-theoretical in their approach - some by explicit, up-front choice, others de facto by the way of arguing along the way. So the first challenge was to define a distinction between when to code an article as belonging to a theoretical category at 2. The PA Conversation - a Literature Review 51 all. In my opinion, it would be possible to identify elements of theory in any of the 125 articles, even if the author did not explicitly do this. In fact, it would probably be possible even where the author never gave it a thought. However, I find it valuable to distinguish between two types of articles: A) articles that are explicitly relating to theory, either by drawing on it in the analysis or by contributing to theory, or by doing both. And, B) articles that are non-explicit about their relation to existing theory. Although I am aware the term is somewhat misleading, I have, in lack of a better term, used 'non-theoretical' to describe the latter category. Thus, non-theoretical means in the following that no explicit overall theoretical approach at the framework or meta-level is used or referred to in the article. That is, an article may look for general concepts like "best practice", it may investigate "issues", it may examine "motivation effects" or "participant satisfaction" or similar, and it may also propose generalised statements about the possible use of the article findings beyond its own study object, a manoeuvre one could easily justify as being theoretical by nature. However, the author does this without explicitly attaching these reflections to any broader theoretical perspective. Going through the 125 journal articles, two initial key findings became apparent. First, a majority of 64% of the articles were non-theoretical according to the above definition. Second, there was not a handful of dominating theoretical approaches across the articles. Rather, there was a plethora of different theoretical approaches so that I ended up identifying 27 different theoretical approaches across the 44 articles that were coded as belonging to one or more theoretical category. Let us start by looking at the articles coded as theoretical. The distribution of these is illustrated in Figure 2-4 below. 2. The PA Conversation - a Literature Review 52 Figure 2-4: Distribution of article occurrences on theoretical perspectives Figure 2-4 shows how many articles that claim to contribute to each of the 27 identified theoretical perspectives. The figure seems to illustrate a very heterogeneous theoretical composition of the articles. However, at closer look some initial considerations appear to be appropriate. First, there is the question of what defines theory, as opposed to e.g. methodology. Some articles claimed to use grounded theory, which I, despite the reference to theory, coded as a method used in qualitative studies. But what about discourse analysis, then? Although normally seen as a method, I have here taken the liberty to code it as theoretical, because the St-Onge et al. (2009) article not only uses discourse analysis in a qualitative study as a means to work with the material, but also theorises about how the manager discourse can be actively changed to improve the effectiveness of the PA process. Second, the question of what defines a theory is important. For example, is justice theory one theory, or are there in fact two fundamentally differing justice theories; one based within psychology with departure in the justice perception of the individual, and another based in sociology which is relatively and collectively based? I have 2. The PA Conversation - a Literature Review 53 here decided to use the authors' own claim at face value without proposing even more fine grained definitions than the above 27 categories. This is not to ignore the fact that such an analysis would bring improved understanding of the theories in use, but it lies outside the objective of this study. Third, readers acquainted with more than a handful of the theories listed in Figure 2-4 above will probably already have seen that some of the theories are closely related to each other. Before turning towards a more detailed discussion of the most used theories and the findings they produced, it could be beneficial for the purpose of maintaining an overview, to consider how one could group the 27 theories at a metatheory level. This would be beneficial in terms of adding to the understanding of the overall theoretical perspectives into which the 27 theories fit. One way to make such a grouping would be to categorise the theories into three broad perspectives ranging from individual psychology over social psychology and to sociologically oriented theories. In such a construction, the first meta-category would be the broad category of theories based on individual psychology. Roughly split, this category would include expectancy theory, attribution theory, personal construct theory, cognitive processing model theory, cognitive-affective personality system, test theory, self-enhancement theory, conservation of resources theory and attitude dependency theory. The second meta-category would be theories oriented towards social psychological explanatory factors. This category would include implicit theory, accountability theory and social comparison theory. The third meta-category would be theories oriented towards sociological explanatory factors. This category would include TQM, social power theory and possibly culture theory. However, some theories would be harder to put into one of these three boxes. For example, as already mentioned, justice theory can operate at least at the first two levels. Practice theory is also somewhat cutting across the three levels. However, as a rough rule of thumb I believe it can be beneficial to see theoretical approaches to analysing PA to focus at one or sometimes more of the three categories: individual psychology, social psychology and the sociological system/process level. Or in other words; at the level of the individual, the group or the system. 2. The PA Conversation - a Literature Review 54 For the purpose of the present analysis it will not be beneficial to take this categorisation discussion any further. Below, the most used theories and the findings they produced will instead be analysed as they have been named and put forward by the authors. The one theory that sticks out as being the most used is Justice Theory, used in nine articles. Let us start by looking at these articles. Justice Theory Justice theory, or organisational justice theory, is used by Thurston and McNall (2010), Narcisse and Harcourt (2008), Roch et al. (2007), Cook and Crossman (2004), Poon (2004), Holbrook (1999, 2002), Flint (1999) and Korsgaard and Roberson (1995). According to Thurston and McNall (2010) Organizational justice theory is deeply rooted in social exchange theory which takes departure in the assumption that social relationships are viewed as exchange processes in which people make contributions for which they expect certain outcomes; and, individuals evaluate the fairness of these exchanges using information gained through social interactions. (Thurston and McNall, 2010: 204) As a consequence of this assumption and in the context of PA, Justice Theory would in its simplest form claim that there has to be a balance between people's contribution and their share of reward, for people to perceive the exchange as fair. However, later developments of the theory separate between three different forms of justice; distributive, procedural and interactional justice. Distributive justice concerns the distribution of resources among a group of people. In a PA context this could be the performance ratings but it could also be other rewards like bonus, promotion, better assignments or praise. This concept deals with the justice perception of outcomes. Procedural justice concerns the perception of process fairness. The claim here is that people will tend to accept even disappointing ratings if they perceive that the rating process is transparent, consistent and fair, i.e. without bias. Interactional justice concerns the perception of the personal treatment someone receives while being rated. This concerns the quality of the dialogue including the ability of the manager to deliver the rating message in a way that is perceived as respectful to the employee. This type of justice also includes what some authors 2. The PA Conversation - a Literature Review 55 choose to see as a separate form, informational justice, which concerns the procedural explanations for why an outcome happened. As Thurston and McNall (2010: 204) put it: Perceptions of inaccuracy and injustice as well as feelings of dissatisfaction have long plagued PAs and the organizational processes that generate them. After this short introduction to Justice Theory, let us turn to look at some of the key findings that have been generated from this perspective in the articles belonging to this category. Thurston and McNall's (2010) study claims that the different theoretical types of justice perception are distinct but also so highly correlated that it can be discussed how much practical importance the separation has. In the study, justice perceptions were investigated at the general employee-organisation relation level rather than specifically for the PA process. However, the study did show a significant correlation so that employees who rated their perception of justice within the organisation as high, were also the more satisfied with the PA scheme in the organisation. Poon (2004) states that there is evidence that performance ratings are often manipulated for political purposes which in turn are perceived as unfair by the employees. With departure in justice theory she examines the effects of employees’ perceptions of political motives in PA by relating this to their job satisfaction and intention to quit. Her analysis shows that when employees perceived performance ratings to be unjustly manipulated because of raters’ personal bias and intent to punish subordinates, they expressed reduced job satisfaction that, in turn, led to greater intentions to quit their jobs. Holbrook (2002) looks into how the use of justice mechanisms can be conceptualised into the PA interview situation. The study shows how the PA process contains a number of activities that each affects the justice perception of the employee. Specifically, he mentions the evaluation process, the assignment of ratings and the interview between manager and employee. Although not mentioned by Holbrook, one could easily add other activities like goal setting, prioritisation, assignment workload etc. as events that can affect the overall PA process justice perception of the employee. According to Holbrook it is rare to find organisational events that contain all three theoretical forms of justice. But since these are all present in the PA 2. The PA Conversation - a Literature Review 56 process, Holbrook sees the process as an excellent study object for scholars working within Justice Theory. Generally, the relatively high number of articles concerned with justice theory can be seen as an indication of the fact that perceptions of injustice or unfairness appears to be a central issue to many PA scheme implementations. It is not hard to see how a scheme characterised by low justice perception will have a hard time being motivating for the employees. What Justice Theory brings to the table is a systematic approach on how to investigate the elements constituting justice perceptions. Expectancy Theory Expectancy theory is a variant of motivation theory proposed by Vroom (1964) which focuses on how employees' motivation is a function of the perceived reward of conducting a specific task and the expectancy, or confidence, the employee has in his own ability to accomplish the task. In a PA context this means that if conducting the task to a certain quality level is perceived by the employee to be rewarding to himself, then the employee will be motivated to conduct the task. However, the job of the manager is to find out what the employee perceives as "rewarding" enough to find it worthwhile to conduct the task. The employee could be valuing a financial bonus but he could also be more interested in recognition or other less tangible rewards. Second, the employees' expectancy of own capability to conduct the task satisfactorily directly affects motivation. Even if the reward is perceived as attractive, the total motivational effect will be negative if the expectancy factor is negative, i.e. if the employee perceives it as unlikely that he will be able to meet the desired standards. In fact, if expectancy is negative so that the employee finds the task unobtainable, a higher perceived reward will only add to the level of negative motivational effect (Vroom, 1964). With departure in this framework, DeNisi and Pritchard (2006) in a non-empirical contribution define the PA process as a sequence where actions lead to results that are then evaluated and rewarded to a certain level of need satisfaction. It is then proposed that the stronger the perceived linkages between each step in this sequence, the greater will the employee's motivation be. This may sound trivial, but considering the amount of literature on rating and measurement bias, interdependency in outcomes and differences in employee reward preferences, this 2. The PA Conversation - a Literature Review 57 has significant impact on the way one deals with PA in practice. Based on this assumption, DeNisi and Pritchard raise a number of claims. First, if the perceived connection between an individual's contribution and the desired outcome is weak, which would often be the case when more strategic goals are cascaded down through an organisation, the motivation effect will be negative. Second, the perceived ability of the individual to fulfil the required task; be it due to capabilities or resource availability, will directly affect the employee's perception of the link between action and outcome and thus the employee's motivation and performance. Hence, PA systems should include a systematic aid to the employee's perceived ability to fulfil the tasks, to strengthen the perceived link between action and outcome. Third, PA systems should aim to strengthen the perceived link between outcome and evaluation. If the employee evaluation is expected by the employee to have only a weak connection to the outcome obtained, this will have a negative impact on motivation and performance. In other words, the perceived fairness and transparency of the rating system has a direct impact on employee motivation and performance. Fourth, PA systems that are expected to deliver a weak link between evaluation and the reward preference of the individual will have a negative impact on motivation and performance. Smith and Rupp (2004) examine knowledge worker's perception of performance ratings with departure in the expectancy theory claim that for a PA system to have a positive effect on motivation and performance, high performance ratings should entail high merit increases while low performance ratings result in low merit increases. Through a survey to knowledge workers in 5 medium to large US companies, Smith and Rupp found a weak correlation between appraisal rating and merit increases. 58% of the knowledge workers perceived the link between performance rating and merit increase to be weak. In fact, "it appeared that decoupling performance ratings and merit increases is a common practice" (Smith and Rupp, 2004). Smith and Rupp present a range of explanations to this paradox, proposed by the respondents to the survey. The most common explanations were budget constraints, subjective factors, manager bias, external conditions and a general perceived decoupling of employee contribution and outcome. Using expectancy theory, Smith and Rupp claim that this decoupling of performance and reward may in fact be the key to understanding the widespread failure of PA schemes. 2. The PA Conversation - a Literature Review 58 Lee and Son (1998) conducted a longitudinal study in a Korean petrochemical company to investigate employee perception of review quality and utility, and impact on later employee performance. Three independent variables were tested: degree of employee perceived opportunity to participate in a discussion with the manager during the review, the degree of perceived goal clarity, and the degree of perceived inclusion of career themes in the review interview. Somewhat surprisingly, the study concluded that although all three factors were correlated to the employees' later perception of review quality and utility, none of the factors had any significant impact on later performance. Lee and Son suggest that this has to do with a more general fact of the PA practice itself, i.e. that the performance-outcome relationship is weak since performance evaluations for a certain employee have a tendency to be stable over time no matter the fluctuations in actual performance. What expectancy brings to the table is insight into how important it is that performance goals are set in a way so that the employees see them as not only relevant and rewarding, but also realistic to achieve. If individual targets are unrealistic or if the number of targets is so huge that fulfilling them all becomes unrealistic, then the process will have a negative impact on employee motivation. As we shall later see, this is something managers and employees in my own study are in strong agreement on. Motivation Theory Motivation theory is of course a much too broad categorisation of a group of theories that concern human work motivation. Most writings on PA somehow relate to the subject of motivation. However, some relate more specifically to motivational theories. Buchner (2007) investigates in a non-empirical article how applied models of PA do not take full advantage of three key theories of work motivation: Goalsetting Theory, Control Theory and Social Cognitive Theory. He claims that the general predominance of top-down approaches to PA has a negative impact on motivation in all these three perspectives and that it is thus constraining the positive impact of PA schemes. Kuvaas (2006) explores alternative relationships between PA satisfaction and employees' perceptions of own work performance, affective organizational commitment and turnover intention. Through a survey of employees in 64 Norwegian 2. The PA Conversation - a Literature Review 59 savings banks he shows that PA satisfaction is directly related to affective commitment and turnover intention. He also shows how the relationship between PA satisfaction and work performance was moderated by employees’ intrinsic work motivation. The moderation revealed a negative relationship for employees with low intrinsic motivation and a positive relationship for those with high intrinsic motivation. Generally, motivation is a crucial concept in studies of PA because most opponents and proponents of PA schemes claim that it is the effect on motivation that is one of the key arguments for embarking on, or rejecting, the process in the first place. Motivation theory contributes to this discussion by offering a framework for analysing different sources of motivation and how appraisal schemes affect them. Total Quality Management (TQM) A separate small group of three articles specifically deal with the links between TQM and PA. This may at first sight seem to be a rather narrow subject. However, for any sector with keen interests in quality issues - and the pharmaceutical sector certainly is one - it may be of interest to consider how PA schemes combine with a quality focus. Soltani et al. (2004) examine the consistency and congruity of PA systems with quality management precepts in UK-based quality-focused organisations. Soltani et al. claim that a fundamental reason for a purported incompatibility between TQM and PA is the assertion that in the TQM framework most variance in performance is caused by system factors, but that PA schemes focus on person factors rather than on system factors. Soltani et al. found what he calls a "remarkable inconsistency" between the currently dominating PA schemes and TQM demands. Most PA systems fail to meet quality management expectations and contradict TQM assumptions about individual and system performance. In fact, Soltani et al. claim that their study supports Deming's theory that PA schemes are a root cause of quality management problems, because they direct focus so strongly towards the individual, and away from the team and system level. Soltani et al. conclude that, unpalatable though it may be, is that HR performance evaluation in the majority of the TQM-based organisations surveyed is locked into a vicious circle of individual performance, control approach, HR dissatisfaction, and a low degree of success for TQM programmes. (Soltani et al., 2004: 415) 2. The PA Conversation - a Literature Review 66 the managers ”...must recognize that their role has changed from that of judge to that of coach" (Waite and Stites-Doe, 2000: 202). In her study of the introduction of a PA scheme to chairs of university departments, Townley (1999) examines why chairs were so strongly against the scheme. The author found that there was a complete mismatch between what people found was the management rationale behind the scheme as opposed to the management rationale they were operating under in the university world. A somewhat lengthy quote is required at this stage because it so nicely describes the perception of PA schemes seen through the lenses of the radical perspective: Although PA may operate in a range of ways, the underlying model is technocratic. It offers a disengaged, formal rationality involving the identification, specification and measurement of `work', a unit of work, and the components of work - contract, workloads, output, performance, etc. Management is presented with a knowable, standardized, empirically verifiable reality, ordered through universal and abstract regulations that attempt to clarify the relationship between inputs and outputs. Its operation requires no particular type of social setting. It portrays management as a task that can use the same standardized, formalized procedures in an inherently substitutable manner. (Townley, 1999: 298) The university chairs were remarkably aligned in their resistance to the scheme. They could not see how such a scheme could relate to their daily work where it was hard to specify what even the term "work" meant, as responsibility was socially defined rather than given by management. Neither could they define concrete, desired outputs in the form of performance targets that in any meaningful way captured the very independent mode of working characteristic of the employees. And they found it impossible to set up effective and fair measuring mechanisms. Despite the initial resistance in the case organisation, PA has today become a widespread management tool also in universities. But according to Townley, this has happened at a high cost: They [PA schemes] have become normal. But with them comes subtle shifts in the ability to argue about the goods of the tradition. The debate becomes circumscribed, individualized and privatized. Individuals are accountable because they complete annual appraisal forms; an organization is accountable because it has such procedures in place. Through such mechanisms the opportunity for collective engagement and debate about the goods of the tradition and their sustainability is inhibited and diminished. There is a loss of vitality. (Townley, 1999: 302) So, according to Townley, the appraisal schemes bring with them a bureaucratised and individualised, compliance fulfilling management process that has negative 2. The PA Conversation - a Literature Review 67 impact on the more practically grounded and socially oriented coordination and exchange paradigm it has replaced. Bowen (1995) raises the question why society relies on ranking and grading from the first day of school, through the educational system, into the working environment. He argues that PA schemes are creating winners and losers in an organisation to the detriment of overall organisational performance and to general human wellbeing. According to Bowen, ranking schemes rely heavily on a strong belief that extrinsic motivation is what drives human beings, while so much research shows that it is in fact intrinsic motivation that holds a lasting and good effect on performance. But not only are ranking schemes relying on extrinsic motivation - according to Bowen they are in fact destroying intrinsic motivation. Hence, Bowen suggests that ranking should be replaced by a much stronger focus on strengthening co-operation and shared responsibility. Piercy (1994) raises a more instrumental critique against PAs. He accuses management of repeatedly substituting bureaucracy and complex reporting and control systems for the things that really matter to competitive success. According to Pierce, PA schemes are excellent examples of this, since they are so bureaucratised and focused on what he terms "administrative neatness" that they completely lose any connection to the real strategic issues of the organisation. Fisher (1994) comes from a consultant background with years of experience in assisting with implementing appraisal schemes in different organisations. In his article, he is presenting some thoughts on how different schemes can be categorised and what the advantages and pitfalls are for each category. The author is not explicitly departing from PA as a useful management tool as such. But since some of his categories are in practice quite radically departing from the mainstream PA concept of manager-rating-employee, I have included it as an example of a radical approach to PA. Fisher classifies performance schemes on two dimensions. First, he distinguishes between whether they are focused on accountability or development. Second, whether the appraisal is to be carried out by someone hierarchically superior to the appraisee or by a peer. Since the mainstream PA scheme would clearly fall into the accountability/hierarchical combination, and the purely developmental interview would normally be seen as an activity that can be maintained in parallel 2. The PA Conversation - a Literature Review 68 with the PA process, let us take a look at what Fisher has to say about schemes where it is up to peers to conduct the review: The purpose of this type of appraisal is to provide feedback to the appraisee and to help her or him to plan their future self-development. The focus is primarily on the needs of the individual with only a broad and generalized regard being given to the needs of the organization. (Fischer, 1994: 37) Fisher did not find any really good empirical examples of organisations with performance evaluation by peers. However, he mentions academic peer review as something close to this, even in the form of a yearly review by the head of department, if these were able to create an atmosphere where they, ...were acting in the role of academic peer rather than as line manager... and if the interview was not seen as a management exercise but as a personal opportunity to reflect and plan. (Fischer, 1994: 37) Clearly, this is something completely different from the mainstream performance review where there are clear roles between rater and ratee, and where there is a singular, objectified outcome in the form of a ranking on a numeric scale. To round up the presentation of the articles coded as belonging to the radical perspective, let us now turn to Rasch (2004) who also places himself in the radical perspective and who through a literature review gives an account of how the arguments of proponents and opponents of PA schemes contrast. Generally, Rasch finds that proponents of PAs blame implementation flaws as the reason for the widespread dissatisfaction of expectations to PA implementations. Proponents of PA would claim that the process would work if only it was properly used. However, opposing this view are the critics of PA who, ...view the PA process as divisive within the organization, and thus counterproductive to organizational development. They state that the use of ratings in an employee PA establishes arbitrary divisions among employees. Furthermore, they contend that intrinsic rewards are far more effective as a motivator of employees. (Rasch, 2004: 407) Rasch quotes Deming's (1986) proposal that 95% of the employees perform well but that PA schemes are basically designed to address issues that are only relevant for the remaining 5% of the employees. As a logical consequence of this, Rasch suggests that the current practice of PA should be substituted with another scheme consisting of three elements. First, the scheme should build on a focus on individual growth and development through dialogue and without ranking. Second, the scheme 2. The PA Conversation - a Literature Review 69 should include a PA system that deals with the 5% special cases, so that it allows managers to dedicate their evaluation and close target follow up activities on those employees who actually need that. And third, the scheme should focus on developing a continuous improvement mindset in management so that managers can become good coaches to the employees, rather than performance raters. Summary of the radical perspective review Although authors within the radical paradigm all reject the idea that PA schemes will stand a chance to add net value to organisations, they do it by emphasizing different aspects as critical. As we have seen, Coens and Jenkins (2002) emphasize false assumptions while Kohn (1999) emphasizes the criticality of the more narrow perspective of dependence on extrinsic rewards. The authors of the six radical articles in my literature review reached the same conclusion but from other angles. From Waite and Stites-Doe's (2000) TQM position to Townley's (1999) practice argument, through Bowen's (1995) critique of society's dependency on ranking and Piercy's (1994) critique of the technocratic and bureaucratic focus of PA schemes and finally to Fisher's vision of a more development and cooperation oriented manager role, rather than the currently dominating evaluator role. 2.8 Performance Appraisal - What are the Problems? In my literature review I have so far analysed a number of monographs and 125 journal articles to create an overview of the state of research into PA. While the monographs were used to create an overview of the subject and identify a dichotomy of attitudes, the analysis of the 125 articles was based on three different dimensions of the approach of the articles: the empirical, methodological and theoretical approach. I supplemented this by taking a separate look at the six articles belonging to the radical perspective. Through the chapter, I have given numerous examples of the kinds of questions the authors have dealt with and the findings they have presented. It may at this stage be beneficial to take a step back from methodology and attitude, and try to summarize the content of the author's findings as they have been presented above. I will do this under the headline question 'What are the problems with PA?' 2. The PA Conversation - a Literature Review 70 The literature review meta-study of the past 20 years research published in academic journals has shown that only a small minority of today's scholars subscribe to the radical paradigm. The majority of journal articles on PA are of a more pragmatic nature focusing on particular practical and theoretical problems in PA and possible improvements. Thus, my review uncovered more than 50 different practice problems that have been analysed in the literature over the past 20 years. In the literature, some problems in PA are described as specific to the target setting sub-process, e.g. that appraisal schemes remove focus from adding real value to reaching targets (Kohn, 1999), targets are almost always only indirect represents for real organisational value (Austin, 1996), targets become input (effort) rather than output (result) oriented (Soltani et al., 2004), targets are either too ambitious or too easy and in both cases de-motivating (DeNisi and Pritchard, 2006), targets do not incorporate external factors (Furnham, 2005) and targets are quantity rather than quality focused (Deadrick and Gardner, 1999). Other problems are closer connected to the on-going feedback and midyear appraisal sub-process of PA. For example, the informal feedback dialogue between manager and employee is thwarted by appraisal schemes (Waite and Stites-Doe, 2000), feedback is manager centric (Coens and Jenkins, 2002), there is a weak link between effort and feedback (Langan-Fox et al., 1998), there is a general lack of feedback outside the formally required meetings (Prowse and Prowse, 2009), managers are reluctant to give negative feedback (Yariv, 2006) and appraisal schemes Install managers as judges rather than coaches (Waite and Stites-Doe, 2000). Another category of problems in PA concerns the sub-process of finally appraising and rating employees. Judged on the number of contributions, this sub-process of PA seems to have seen particular attention from scholars, who have worked with issues like e.g. that rating processes are subjective (Armstrong, 2008), ratings lack comparability (Coens and Jenkins, 2002), appraisal schemes are in practice counterproductive in dismissals (Coens and Jenkins, 2002), PA schemes build on an extrinsic reward philosophy which is undermining intrinsic motivation (Kohn, 1999), measurement is often more difficult than one would expect (Austin, 1996), there is an inconsequent relation between performance, appraisal and reward (DeNisi and Pritchard, 2006), there are serious issues with perceptions of justice and fairness 2. The PA Conversation - a Literature Review 71 (Thurston and McNall, 2010), relative rating schemes are generally perceived as unfair (Roch et al., 2007), there are differences in employee engagement in the PA process which affects rating but is unrelated to performance (Knippen and Green, 1995), disappointing rating decreases motivation and increases employees' intention to quit (Poon, 2004), focus of PA schemes is on formality rather than utility (Deadrick and Gardner, 1999), and PA schemes (and research) ignores or downplays withinperson variability (Murphy 2008). Further, many scholars have investigated particular biases in the rating process, e.g. that manager's own appraisal affects employee appraisals (Latham et al., 2008), attribution bias (Cardy et al., 1998), gender, race and age bias (Stark and Poppler, 2009), how managers falsely tend to assume performance is stable over time (Lee and Son, 1998), interpersonal affect bias (Varma et al., 2005), manipulation, impression management and politics (Cook, 1995), undeserved reputation bias (Cook, 1995), memory bias (Coens and Jenkins, 2002), central tendency of ratings bias (Bacal, 1999), and bias due to managers' provision of self-serving evaluations (Coens and Jenkins, 2002). Finally, there are scholars who have investigated problems in PA that are of a more cross-process nature, e.g. that appraisal schemes serve conflicting objectives and are built on a one-size-fits-all perception of the organisation (Coens and Jenkins, 2002), PA processes are essentially top-down processes (Coens and Jenkins, 2002), the entire PA process is founded on incomplete information (Raymark et al., 1999), PA undermines intrinsic motivation (Kohn, 1999), focus in PA is on individuals rather than systems and processes (Soltani et al., 2004), PA favours quantity over quality (Deadrick and Gardner, 1999), PA installs competition at risk of deteriorating cooperation (Kohn, 1999), PA discourages innovation and risk taking (Kohn, 1999), PA requires significant resources, particularly from management (Coens and Jenkins, 2002), PA gets heavily politicised (Poon, 2004), cultural differences makes comparison and global use questionable (Woods, 2003), there are typically multiple implementation flaws (Rasch, 2004), PA is based on extrinsic motivation philosophy and has questionable outcome (Bowen, 1995), PA is bureaucratic and control focused (Piercy, 1994), PA is stressful and mental resource requiring for managers and employees (Brown and Benson, 2003), managers are too lenient and focus on avoiding conflicts and therefore undermine the trustworthiness of targets and ratings (Coens and Jenkins, 2002), PA is characterised by a lack of transparency (DeNisi 2. The PA Conversation - a Literature Review 72 and Pritchard, 2006) and last but not least, it is still unproven that PA schemes ultimately or at all increase organisational performance (Kohn, 1999). I have illustrated this process dimension in Figure 2-5 below, where I have listed all the problems referenced through the analysis conducted in this chapter. For the purpose of simplicity, I have ordered the problems by three main process stages that are characteristic for almost any PA scheme: Target setting, feedback & control, and calibration, appraisal & reward. Figure 2-5 Overview of PA problems identified in the literature review The top three boxes in Figure 2-5 summarise the problems primarily connected to one stage of the PA process, as they have been identified and analysed by the authors included in my literature review. The lower box summarises the more 2. The PA Conversation - a Literature Review 73 general, cross-stage problems handled by the authors. To avoid making the illustration too complex, I have only included one reference into the figure for most problems. However, in reality almost all the problems are dealt with by more authors in scope for my literature review. Of course, not all authors would agree about the seriousness or possibly even the relevance of all problems presented here. Some problems have been presented by the authors as contextual - for example Austin (1996) made a point on the particular issues with measuring performance in a knowledge worker context. On the other hand, it is highly unlikely that the list of problems presented here is anything near exhaustive. If one used broader search criteria than I have done in this study, one would likely find more problems described out there. The purpose of this study was not to draw 'the full picture' but rather to provide a comprehensive taste of the challenges and complexities one faces when embarking on studying or implementing PA schemes in organisations. Looking at Figure 2-5, it should be quite clear that installing a successful PA scheme is not a trivial task. It is a resource intensive endeavour with many real and potential problems and risks. Realising this up front is probably a good place to start when embarking on an implementation of a PA scheme in an organisation. One could ask why I have been so focused on the problem side of PA. Would it not be fairer to also dig into all the good things that come out of the process? Partly, this prioritisation is a logical consequence of the nature of my analysis. It is hard to find any academic journal articles that are not problem focused, so a literature review will naturally cover a range of problems. Also, the problem side is an interesting contrast to the obvious fact that my case organisation, together with most other large Western private organisations, at some stage has decided that implementing a PA scheme was a worthwhile investment. I will soon give a more thorough introduction to the rationale behind the implementation and the details of the PA process in the case organisation. At this stage it will suffice to reiterate that there is an interesting tension, a paradox, between the prevalence of performance management schemes and the huge number of problems, risks and costs identified by the last 20 year's research into the subject. This is the tension that leads an author like Pickett (2003), who is by the way himself in favour of conducting PA, to conclude that in too many organisations the process is best characterised as 'the annual fiasco'. It is also this 2. The PA Conversation - a Literature Review 74 tension that made me so interested in the question of local adaption and meaning creation. How do managers and employees create meaning from being part in a process so prevalent in business and so strongly enforced by their organisation, while at the same time being faced with some or all the problems listed in Figure 2-5? 2.9 Summary of Literature Review Based on a literature review of 125 journal articles from 1990-2010 on PA I have documented tendencies and identified possible research gaps within the field from four angles. First, I examined the empirical focus of my sample of articles, I found that there is an overweight of public sector studies as opposed to private sector studies, when compared to the share of the full employment base of the two sectors. Similarly, I found that there is a huge overweight of studies within a Western context with almost two thirds of all the empirical studies in my sample being conducted on populations from three countries: US, UK and Canada. Also, I found a significant overweight of 85% of all empirical studies being conducted on white collar populations. Only 8% of the articles were studying more than one workforce, but none of these compared experiences between workforces. I concluded that this lack of comparative analysis across workforces is a significant gap if we want to improve our understanding of how experiences with PA can differ intra-organisationally. Second, I examined the methodological approaches utilised in my sample of articles. Quantitative studies were significantly more frequent than any other methodologies and there were three times as many quantitative studies than qualitative studies in my article sample. Also, I found that the group of quantitative studies were generally characterised by narrow and unambiguous research questions, often formulated as one or more hypothesis testable against closed question survey responses from a representative group of employees. On the other hand, far fewer articles dealt with more open ended, exploratory questions, and these typically utilised a qualitative approach based on deep dive interviews or a case analysis approach. Only two studies had a longitudinal design and only four of the articles were using a mixedmethods approach. 2. The PA Conversation - a Literature Review 75 Third, I examined the theoretical approach utilised in my sample. Interestingly, the non-theoretical category constituted no less than 64% of the articles, which were generally concerned with investigating specific issues and bias within PA and proposed suggestions for improvement activities. For the articles categorised as theoretical, the picture was very heterogeneous as no less than 27 different theoretical perspectives had been utilised. The most popular was Justice Theory which was used in 9 out of the 44 articles, indicating the significance of the fairness and justice perspective which has been central to the critique of PA in practice. Also, the concept of motivation played a central role in many articles, often with focus on effects on employees' intrinsic motivation of the extrinsic reward paradigms so often linked to PA schemes. Another theoretical category was TQM which was generally seen to distance itself from the tendency of PA schemes to focus on individuals rather than systems and on quantity rather than quality. Finally, it was found that 21 out of the 27 theories utilised in the 44 articles coded as theoretical, were only used in one article. However, we also saw how theories only used once were in some cases closely related to each other. Fourth, I found that out of the 125 articles only six belonged to what I have called the Radical Paradigm. Of course, many authors are critical towards many or a few elements of PA schemes, but only six came to the conclusion that abolishing the scheme rather than improving it would be the recommendable way forward. Fifth, I summarised all the problems identified and handled by the authors included in the literature review and ordered these in accordance with a simple three-stage model of the PA process. This overview illustrated how the past 20 years' research into PA schemes has identified many costs, problems and risks connected to PA. In my mind it is justified to conclude that according to the state of academic studies into the field, there is a clear tension between the prevalence of such schemes and the reality managers and employees meet when fulfilling the process in practice. This seems to be a good place to look for processes of meaning creation. It is to this large and on-going conversation about problems in PA I aim to contribute by shedding light on meaning scheme challenges, local adaption and meaning creation around problems in PA. One important observation when considering the above mentioned paradigmatic dichotomy between the dominant and the radical perspectives is that it is a scholarly rather than a practice dichotomy. This may seem 3. Methodology and Data 82 validity is to utilise a triangulating strategy by analysing multiple sources of evidence and leveraging multiple data establishment methodologies (Yin, 2009). 3.1 Participant Observation By being a manager within corporate HR I am close to the continuing development of the firm-wide conceptual PA solution. I am involved in the strategic development of the underlying IT tools and reporting capabilities that allow correlation analysis to other HR practices, e.g. bonus, salary increases and more. This all gave me a unique opportunity to undertake participant observations of the corporate dialogue and activities. In short, my privileged position as a manager in corporate HR gave me direct access to the PA conceptual machine room. It also gave me the ability to “perceive reality from the viewpoint of someone ‘inside’ the case study rather than external to it” (Yin, 2009: 112). Understanding the corporate politics and lingo improves the ability to conduct relevant analysis of data established. The cost of being a participant observer obviously is the risk of introducing personal bias. This concern was from the start inherent in my work and only became more so along the way, as my understanding of the phenomenon grew and as impressions unavoidably transformed into opinions. As a manager within corporate HR directly involved in the PA process development I of course had my own professional interests and agenda and it was a clear risk that this would affect my ability to perform a balanced analysis based on participant observation alone. Thus, participant observation was from the beginning only thought of to be one among other data establishment methods and it was always planned to be inferior in importance to the much more extended qualitative interviewing approach described below. As it turned out, participant observation became in fact even more limited in its use than originally planned, due to ethical considerations but also because the research question through its refinement process developed in a way where participant observation was a less valuable data establishment source. Below, I will elaborate on this development while introducing the participant observation approach in more detail as it was described in the protocol. 3. Methodology and Data 83 Drawing on Emerson et al. my aim was to produce field notes in an on-going, structured manner and on the basis of these notes, “transform[ing] observation and experience into inspectable texts” (Emerson et al., 1995: xii). I did this by producing a written record of relevant observations made on an on-going basis. By being in a position of what Emerson et al. call immersion (Emerson et al., 1995: 4), I had already walked a significant part of the path that participant observing researchers normally need to go since they often approach a setting not already familiar to them. My position also deviates from the classical participant observer role in another way: Emerson et al. describe how the participant observer never becomes a member in the same sense that those ‘naturally’ in the setting”. (Emerson et al., 1995: 4) Being a member was exactly the position I was in. How did this position affect my field notes? The first thing to realise is that field notes are never just a list of objective facts. They constitute narratives that include not only the event in itself but also the interpretation of the event by the observer. Self-reflection is a requirement to be able to understand how one’s own interpretation affects the accounts and how different or even contrasting interpretations could be possible. This was the approach I decided to follow and it was my aim to reflect on my own role as a participant and observer in all three steps to handle bias rather than aiming for the impossible mission of avoiding bias. Scholars agree that writing field notes is an activity that must be conducted immediately after or as close in time as possible to the event covered (Yin, 2009 and Emerson et al., 1995). Further, I had an advantage over other types of ethnographic observers who often find it awkward to do jotting during the actual observed interaction. This was a perfectly normal activity for me, as I participated in meetings and workshops regarding the PA process. Thus, I developed a simple template to hold my field notes as they unfolded from my jottings. The template was a twocolumn table with my field notes in the first column. Each event constituted a row in the table. Methodologically, I took Emerson et al.’s advice and aimed to keep the field notes in the first column as concrete as possible without “making statements characterizing what people do that rely on generalisations” (Emerson et al., 1995: 32). 3. Methodology and Data 84 In the second column I stated my initial analysis of the notes. This could be concerns, ideas suited for generalisation, methodological reflections, etc. I found such reflections valuable as they are not only a way to bring forward further analytical thoughts early in the process of the study but also because they constitute the first steps towards theorising or relating to theoretical frameworks. As such, they are elements in the analytical coding process which is a necessary step similar to when data is established through qualitative interviews as described below. So, what was I looking for while conducting participant observation? Emerson et al. (1995: 146) among others suggests focusing on “What are people doing?... how exactly do they do this?... what assumptions are they making?... why did I include them?” These questions was a good starting point but as already indicated I added more analytical secondary order notes to capture my own reflections. These notes were not as such conclusions but rather ideas to be utilised in the later analysis. From the start I did not have a very precise idea about what events I would cover with participant observation. Since I work full-time in the organisation and in periods spend significant parts of my time on work related to the PA process, particularly the development and further technological support of the process, it was clear that I could not cover everything in which I was involved. Instead, I initially had the idea that I would cover three main categories of activities by taking field notes. First, I would cover my own experiences from being part of the PA process by being both a department manager for a group of employees, and an employee myself in relation to my manager. By being a middle manager I have of course been a practitioner of the PA process. As a manager I have had years of experience with setting targets, conducting midyear evaluations, performing yearly appraisals, participating in management team calibrations and having tough or happy appraisal conversations with employees. I have also gone through the same as an employee with the managers to whom I have reported over the years in the organisation. Second, I would take notes from work experiences related to the strategic development of the PA process, particularly the part I was myself heavily involved in. Third, I would generally take notes on an ad hoc basis whenever I heard or experienced things that I found important in relation to the study. 3. Methodology and Data 85 Since the participant observation method part of the study methodology was the one that changed the most during the course of the study in terms of purpose, extent and outcome, I will try to explain these changes below. As it turned out, I chose to refrain from most of the systematic participant observation activities due to ethical and practical considerations. Evaluating the PA process conduct of my direct reports or my manager would in the last instance involve stepping out of the confidential dialogue and take the position as an outside observer “judging” the behaviour of not only myself but also my employees or manager, through my analysis of appraisal events of which I was myself part. Although this may be possible to do, I believe it would be very challenging to do so without compromising the trust and confidence required. A founding principle of my design was that I wanted to be open about what I was doing and thoroughly inform involved individuals up front. Doing participant observation without informing my employees or my manager was a no-go, both held against my design principles and because I found it unethical. But taking notes and conducting analysis with their consent would for both them and me be so awkward that the value of the notes would be at risk, and even worse that the relationships would be negatively affected. Taking notes from the general development of the process and the above mentioned ad hoc notes, was less ethically challenging. Of course, getting participant consent would still be required had the note-taking involved quoting other people. But as this would generally not be required as the notes were merely my own, on-going notes on thoughts and reflections along the way, it was sufficient to inform people at a more general level, that I was conducting a study of PA with the organisation as the case object. Since I had access to so much background information, the outcome of my notes turned out to be more supplementary to the two other data establishment methods utilised through the study, rather than a means to break new ground. It became a structured way of reflecting on the things I was myself involved in at the process development macro level, rather than a source to sensemaking at the individual micro level. Besides these ethical and practical considerations I also found that most of the knowledge that this part of the study design would bring could equally as well be derived from qualitative interviews with other pairs of managers and employees, where the outside observer role would not contrast with my own process participant 3. Methodology and Data 86 role. Particularly as the research question was refined and narrowed over the first year of the study, it became more and more clear that the critical data establishment process would be qualitative interviews, and that participant observation would play a less significant role. Nevertheless, understanding the business background of implementing a global PA process in the first place, and the later and on-going rollout and development of the process, has been a major advantage during the study. I already had a good understanding of this before embarking on the study. After all, it was my own experiences with the process at the macro and micro level that made me interested in the subject in the first place. In that sense, my participant observation notes made me a more observant observer, and in this way added to my ability to see and reflect on events with relevance for my study. Thus, although participant observation was from the start of my study a weighty methodological part of the protocol, over time it developed to be less prominent. However, it still was a fruitful way to constantly mirror my findings against my own experiences as they were uncovered through personal experience and observation. 3.2 Archival Records and Documentation Personal experience from a six year long employment history within the case organisation at the start of the study together with participant observations gave me a solid foundation for asking questions. But being an insider in corporate HR also gave me another thing which was invaluable to my study - privileged access to internal documentation. I initially had focus on using this access primarily for two purposes. I wanted to analyse archived material produced in connection with the corporate performance management project that ran from 2004 to 2006. This was the project that first implemented the new, global PA process in the organisation and later rolled out an IT solution to support it. This documentation included presentations, training material, communication material, steering committee meeting minutes, status reports, benefit realisation plans etc. The archival access also covered documentation of later initiatives to further rollout or improve the IT solution. The thorough study of all this material made it possible to draw a picture of what was intended to happen and why. 3. Methodology and Data 87 Second, archival sources were identified that helped me establish the frames within which sensemaking takes place. Within Corporate HR statistical experts over the years had gathered data through their analysis of correlations between appraisal ratings and rewards, promotions and more. It was on the basis of this data combined with my own participant observation that I became able to conclude that the case organisation’s experience with PA was in so many ways congruent with established research on PA: the same macro level purpose, the same bias, the same micro level frustration. Thus, studying archival material and matching the findings to the literature review and the participant observation findings were necessary activities, not only to understand the generalised character of so many features and challenges of the PA scheme in the case organisation, but also to get to the final formulation of a research question with potential to generate new knowledge. I analysed five main categories of archival material. First, documentation from the corporate project which in 2004-06 implemented a global, uniform PA process, named wave I, and subsequently implemented a global PA IT system, named wave II. This material included documentation of the original purpose and business case for the initiatives, steering committee material, project presentations, analysis documentation, communications etc. Second, documentation from the two later initiatives; wave 3 that further rolled out the process and the IT solution globally, and wave IV that added process support in matrix organised parts of the organisation and further rolled out the IT solution to such units globally. Third, material from the 201011 initiative aiming to establish Global Minimum Benchmarks, i.e. a further standardisation of the practical conduct of PA across the global organisation. Fourth, documentation of statistical analyses of rating distributions and correlations to variables such as rewards, position in the organisation etc. These analyses were conducted by corporate HR in the period 2010-2011. Fifth, a data set covering all electronically captured performance appraisal ratings from the period 2006-2011, allowing linear regression analysis of correlations between ratings and elements like hierarchical position and retention ratio. When analysing archival material it is important to observe a consistent source critique. The various PA projects were all managed by corporate HR with certain intentions and success criteria. Corporate agendas are always to some degree political in the sense that there are vested interests and power at play. Hence, I saw 3. Methodology and Data 88 it as important to understand the agenda of senior management and corporate HR when translating the material into statements about what went on and why. 3.3 Interviews In all, I conducted 38 qualitative, longitudinal interviews. By longitudinal I mean a study design of multiple engagements over an extended period of time. I.e. a study design built on following fixed manager-employee dyads over a full year (corresponding to a full PA cycle) and interviewing them multiple times in connection to recurring PA events (target setting, midyear review and final appraisal). Of course, I am aware that even a one year performance cycle is still only a partial image, since employees through their work life in an organisation will go through multiple PA cycles that mutually affect each other, such that performance delivered in one year may affect targets set for the coming year and so forth. However, observing a full one-year PA cycle through multiple interviews with the same people is broadening the analysis significantly compared to most other studies dominated by the tendency to conduct analysis gathered at a single point in time, e.g. at the target setting or at the final appraisal only. Below I will describe in some detail the approach of the interviews and the methodological lessons learned from them. I will start by explaining why I find a longitudinal interview approach such a fruitful way to investigate the PA process. Second, I will discuss how I approached a tricky question that challenges most researchers working in qualitative designs: Who to interview? This will include reflexions on the consequences to the study object and the study results, of my own intervention as researcher. Third, I will discuss the practical planning and conducting of the interviews. Fourth and finally, I will discuss the robustness of the study design. Why a longitudinal design? The literature review of 125 scientific journal articles regarding PA published over the last 20 years (1990-2010) showed an empirical-methodological emphasis on quantitative approaches. Only ten articles used a methodologically triangulating approach through case studies. But what was really striking was that only two out of the 125 studies used a longitudinal setup, although many authors suggested the likely value of conducting longitudinal research on PA in their ‘direction for further 3. Methodology and Data 89 research’ section (Korsgaard and Roberson, 1995. Salimäki and Jämsén. 2010. Lefkowitz, 2000). Thus, the lack of longitudinal research into PA seems more a consequence of lack of resources in terms of time and capacity, than a consequence of lack of longitudinal design benefits. In fact, the benefits of a qualitative, longitudinal research design are in my view significant. Besides the fact that longitudinal design in PA research is rare though recommended, I see at least three important reasons to aim for a longitudinal design when investigating meaning creation in a PA context. First, the appraisal process itself runs over a full year with at least three formal events and many other opportunities for informal feedback events. Thus, to conduct research at any given single point in time will be at a high risk of missing both any overall pattern, but also of missing important fluctuations in attitude developments over time. Second, in a cross-sectional analysis it will be highly likely that one will either entirely miss out on, or get a distorted impression of, important events simply because of time lag between investigation and event. Third, if my proposition is correct that the field of PA is heavily institutionalised so that managers and employees think of the process as a given, this does not encourage reflexion on the process. The degree of institutionalisation must be expected to be a significant filter between what is at face value interpreted from the appraisal process and what would be interpreted through deeper reflexion. To raise the level of reflection would be almost impossible in the course of a single interview or survey research design. It takes time to build confidence between the interviewees and the researcher and to provoke reflexivity around the PA process. Identification of interviewees The interviewees should ideally be identified on the basis of their likely contribution with new insights or reflections around PA. Within such a design, the principles and process for strategic sampling was critical for the success of the study. To identify appropriate candidates, a dialogue with HR Partners in Research & Development, Staffs, Production and Sales was conducted during November and December 2010. I did not look for participants who were particularly high or low performers. Neither did I look for participants that were either strongly for or against the appraisal process design. Rather, I looked for participants who were first of all interested in participating 3. Methodology and Data 90 in my study and spending some time and effort on this and who were reflective and willing to talk openly about their personal experiences around the PA process.17 It was my hope that the participants would also themselves benefit from participation, through personal reflection and discussion with the researcher. It was the initial intention to identify two pairs of managers and employees within each of the four workforces. The reason for looking for pairs of managers and employees is directly derived from my research question about the meaning creation of managers and employees. Differences between any random manager and employee without a direct reference line could be derived from the fact that their experienced performance events differed by nature. The opposite is true for pairs of managers and employees with direct reference lines. Here the manager and employee would at least theoretically observe and evaluate the same performance events and surely participate in the same performance feedback events with each other. Interviewing pairs of managers and employees may uncover varying interpretations and meaning creation processes around identical events. The reason I was looking for two pairs of managers and employees from each workforce was a question of design robustness. If one pair fell out during the study, I would still have another pair left. However, as it turned out, it was only in the Staff area that I interviewed two pairs of managers and employees. In Sales, I interviewed one manager and two of her employees. In R&D and Production, I interviewed one paired manager and employee. The reason for this was partly practical, partly due to more principal issues with the study design, as will be evident from below. In all four workforces, I approached the managers first. I assumed that the managers would be the ones with most time constraints and possible reservations about participating in a study of this duration, so I wanted to let the invitation to employees await the invite acceptance from the managers. In Sales, I was asked from the start if I could accept to interview only one manager since local sales managers were very busy and the sales HR partner wanted to disturb the business as little as possible. I 17 This is in line with Van de Ven's (2007) notion of Engaged Scholarship, i.e. that informants are chosen because they want to engage. 3. Methodology and Data 91 did not have anything against this and accepted the increased risk to my study design that this entailed. In Production I got a yes from one manager but was held waiting by another. She was principally positive about participating, but was very busy in the start of the year and requested to wait to start until March. I accepted this, but as it later turned out she did not have time to participate even from March. At that stage it had become too late to look for a whole new pair of manager and employee in the production organisation, and I thus had to accept to go with only one pair. In R&D, one of two selected managers interestingly rejected participation from a principal standpoint. I will quote the rejection mail at full length below, because it illustrates some important points. R&D Manager: "As our business in our research department is very scientifically and project focused, it may be misleading to follow our PA interviews. Further I think we are far from representative to the degree that I think you could very easily misunderstand our PA (too many scientific issues and too many project/people related matters that it will take hours to get you updated on) and further this could be a too strong influence on our PA interviews. So I am sorry, I cannot prioritize this for now." There are no less than three different arguments against participation hidden in this quote. The first argument is that it would be misleading to follow the PA process in R&D because the business is very scientifically and project focused. This argument is being reinforced by the argument that R&D is not representative and that I would easily misunderstand the process as it is conducted in R&D. It is hard to see why this should be an issue to my design. On the contrary, it was one of the key points in identifying interviewees from four workforces to see how meaning creation within a seemingly standardised process was differing across workforces. The second argument against participation concerns the more common matter of time constraints. "...it will take hours to get you updated on" refers to the fact that the manager would not be able to justify the amount of time she thinks would be required to help me to the right level of knowledge. In all modesty, and based on the interviews I finally did conduct in R&D and elsewhere, I believe this objection does not hold. Of course, there would be targets of such R&D-specific scientific complexity that I would not understand the detailed content of it. But that would not do any harm to my study since I was interested in the meaning creation of the participants from 3. Methodology and Data 98 Figure 3-1 illustrates how the interview audience was stable over the first two interview rounds, while there were four changes at the third interview round. In R&D, the employee was transferred to another organisational unit under another manager. In Staffs, one of the managers swapped jobs with a manager colleague in a sister department. In Production, the employee continued to have the same job content, but the job was transferred to an operational rather than a project oriented unit. In Sales, one of the employees had a new manager due to a re-organisation. For all four affected employees, the new managers took over the responsibility to finalise the 2011 PA cycle, so that four out of six employees were finally appraised by another manager than the one with which they had set the performance targets and held midyear review. Although my design was prepared for organisational changes, it came as a surprise to me that such a large proportion of my interviewees were affected by organisational changes over just one year. The average time a manager or an employee holds a position varies a lot across the organisation, but a change of manager in four out of six cases was a surprise to me. There are both methodological upsides and downsides of these dynamics which I see as inherent in longitudinal studies in dynamic organisations. As an upside, it gives me a good impression of how the PA process is affected by such changes. The downside is of course that the trust and common understanding built over the first two manager interview rounds would not be present to the same extent at the third interview, where the manager had changed. 3.4 Coding: Digital Analysis of Qualitative Interview Data in nVivo9 In all, I conducted 34 interviews with pairs of managers and employees and 4 with representatives from corporate HR and so I ended up with 38 interview transcripts, each at a length of 8-15 pages. To be able to maintain an overview of this amount of qualitative data through a consistent coding and analysis phase, I decided to conduct a digitally supported analysis of the qualitative interview transcripts. This means that I documented the interview transcripts in a format importable to Nvivo9 and that I conducted the coding and analysis in this tool. I defined my code categories, code relations and code hierarchy in the tool, including a coding taxonomy specifying 3. Methodology and Data 99 which codes were descriptive, analytical or merely theoretical markers. The tool also served to document the text searches I conducted during my analysis, thereby allowing me to be more systematic in the analytical steps conducted. While I believe that the right use of software to support my analysis of qualitative data has supported me with structure and methodological rigour, it will never bring any guarantee against missing important points in the material or overseeing important bias in the data or in my own translation of the data. The same kind of methodological reflexivity is required for an analysis conducted with the help of analytical software as without it. To support me in this process I decided to follow Dahler Larsen’s (2002) three rules for using and displaying data. First, I have done my best to quote as precisely as possible. I have done this by being particularly thorough in making interview transcripts and by using the analytical software to cut out text sequences without any editing. One concern here is that most interviews were held in Danish, so that they needed translation before quoted in this thesis. It is of course critically important that this translation did not change the meaning of the quote and I have done my best to secure this. Second, I used the principle of inclusion wherever possible. This means that I worked to make sure that conclusions are based on an exhaustive analysis and display of the underlying data. Third, I aimed for transparency in the sense that data displayed is explained, figures’ elements and dimensions should be unambiguous to the reader, data sorting in presentations is intuitive and easy to understand and so forth. By this, I hope the reader will find the argument compelling and the logic reasonably easy to follow. One sometimes observes that researchers use ‘quantitative language‘, while doing qualitative research. The risk of using quantitative labels in a qualitative context may be even more observable when using digital tools for qualitative research because the software can tend to ‘objectify’ the data. In my opinion, researchers using qualitative methodologies should be very careful with referring to numbers because not much scientific value can be attributed to them.20 The purpose of qualitative 20 As such, the occurrence of something, e.g. a particular attitude towards PA, can in some sense be perceived as a matter of frequency (one occurrence rather than no occurrence). However, my point here is that it is not valid within my design to attribute more explanatory power to a statement confirmed by two interviewees, than to a statement only confirmed by one interviewee. 3. Methodology and Data 100 research is not to give the impression that one can say anything about the distribution of variables across a population. Parts of the coding process I used had much in common with the principles of grounded theory as presented by Strauss and Corbin (1990). I started out by having three high level categories of codes in nVivo9 called "Process", "Reflections" and "Theory". The two former categories were later renamed to "descriptive" and "analytical" code markers, in alignment with the vocabulary of Strauss & Corbin (1990). These two categories were further developed through an open coding strategy where code categories and sub-codes were defined as they presented themselves through the reading and analysis of the texts. In contrast to this, a different approach was used for the theoretical code category, which was defined up front, with departure in Weick's (1995) definition of seven properties of sensemaking. In the descriptive, category I wanted to capture how managers and employees described the process and the targets in their own words. In the analytical, category I wanted to capture how the managers and employees reflected on the process and its inputs and outputs. In the theoretical category I wanted to capture how theoretical sensemaking markers could be attached to interviewee statements, i.e. how statements related to Weick's (1995) vocabulary of sensemaking. From here, I simply started reading the first handful of interviews and added subcategories of codes to the code categories as they appeared in the texts. After the first 3-5 interviews had been coded this way, I had built a much more developed coding structure which kept quite stable through the rest of the study. At this stage I started all over again by coding the first interviews based on this more developed structure so that all interviews were coded almost based on the same coding structure. When I use the term "almost" it refers to the fact that a few coding categories only with relevance for the final appraisal interviews, were not added until at a late stage of the study. This is a consequence of the longitudinal design, since some categories did only become relevant when reading the last group of interviews which concerned the final appraisal event. I wanted to keep my coding structure as "flat" as possible, i.e. without so many layers that it would become too complex to work with. The final coding structure had only three layers. The first layer was the 3-dimensional split between theoretical markers, 3. Methodology and Data 101 descriptive markers and analytical markers. A full coding overview can be seen in Figure 3-2 below. Figure 3-2: Full coding structure for qualitative interviews coding in nVivo9 Figure 3-2 shows the three layers of codes. All text coding is done at level 3, so that levels 1 and 2 are hierarchical categories made to order the detailed coding for analysis purposes. The table's last two columns show the number of sources (interviews) in which each code has been utilised and the total number of occurrences of each code across all sources. For example, the code "Calibration" Level 1: Marker category Level 2: Code category Level 3: Sub code Number Of Sources Coded Number Of Coding References Feedback 20 43 General meeting perception 21 37 Input from others 12 17 Manager handover 12 40 Meeting duration 13 15 Meeting organisation 18 32 Memory bias 7 7 Other bias 9 15 Preparations 27 59 Privacy 13 17 Rating Scale 17 29 Relation to IDP 19 35 Transparancy 17 28 Are targets ambitious 9 16 Behavioral targets 19 29 Manager vs employee influence on targets 16 38 Number of targets 8 9 Target prioritisation 16 23 Target quality 4 7 Target realistic 14 16 Attitude Attitude 22 49 Fairness Fairness 20 35 Improvement suggestions Improvement suggestions 19 26 Researcher's impact Researcher's impact 5 6 Career implications 21 37 Impact on behavior 4 6 Impact on compensation 19 35 Impact on cooperation 3 3 Impact on innovation 7 12 Impact on motivation 28 51 Impact on performance 4 7 Impact on target setting process 1 1 Business performance 27 49 Calibration 22 75 Dependency on others 18 29 Impact of behavior 22 49 Personal relations 10 13 Relative vs absolute evaluation 11 16 Strategic considerations 4 11 Background info 2 22 Purpose of 3P 11 17 Action - focus on why or what in meaning creation 15 22 Event selection - careful selection of examples 15 25 Extraction of cues - control of event selection 13 31 Identity construction - who am I, what is my role 22 75 Plausibility rather than accuracy - a good story is required 8 16 Retrospectivity - definition of individuals part in collective 14 21 Social dimension - credibility standpoint 10 16 Total 38 1277 Analytical Theoretical Process characteristics Impact Performance Assessment Sensemaking Target characteristics Descriptive Other analytical 3. Methodology and Data 102 under the code category "Performance Assessment" under the "Analytical" marker has been used 75 times in all, and in 22 interviews. All in all, 1277 text pieces in the 38 interview transcripts have been assigned to 48 level 3 sub-codes, all in nVivo9. While this surely is a lot of coding, the nVivo tool made it easier to work with as I could with just a click select all code references for a single code, and with just another click jump to each of the reference texts themselves at the exact point where the code was assigned. Below I will explain the code definitions and present some reflections on pros and cons of the approach. I used the descriptive marker category to capture two things. First, I coded for process characteristics as they were described and reflected upon by the managers and employees during the interviews. This included how the recent PA meeting, be it target setting, midyear review or final appraisal, was organised and prepared including time and spatial perceptions. It also included how the meeting was generally perceived by the participants, how performance and general feedback was included in the meeting and whether and how input from others was collected by the manager prior to the meetings. Further, it included perceived bias in the process, general perception of rating scale and its link to rewards, relation between the PA process and the separate process around maintaining an individual development plan. Finally, it included the activities around manager handover of PA in case of organisational changes, the privacy and confidentiality around PA and considerations regarding transparency of the PA process. The second descriptive code category was target characteristics. Within this category I coded text elements that were used to describe the targets against which the later appraisal was supposed to be made. This coding was made by using a number of sub-codes. I soon found that there was a lot to be said by the interviewees about the number of targets. Some interviewees saw it as a critical issue if there were too many targets because this would be a sign of lack of prioritisation and focus. Generally, prioritisation of targets was a key concern to both managers and employees, so I coded for this as well. Another element that kept coming up was whether the targets set were realistic. Sometimes this came up in connection with examples of too ambitious targets. But the opposite also occurred so that targets were sometimes seen as fostering an un-ambitious tick-off process. So I also coded for statements about the level of ambitions in target definition, as well as the broader concept of 3. Methodology and Data 103 target quality, capturing whether the target formulations could hold relevance over the course of the full performance year. Finally, the texts were coded for quotes on the balance between management and employee influence on target formulation. The analytical marker category was used to capture reflections of the interviewees and myself. I captured elements of the interviews where the interviewee revealed his or her attitude towards the PA process. In a separate category, I captured expressions about the perceived fairness of the process, be it either considerations about whether the interviewee found the process fair to herself, but also broader considerations about the general fairness of the process. Improvement suggestions were captured because I found this an interesting way to have the interviewees reflect indirectly about what they saw as most challenging with the process. I simply asked them what they would change if they were given the authority to redesign the process. In a researcher's impact category I captured quotes where the impact of my questioning was directly addresses by the interviewees. My study was not interventionist as such, but I found it interesting to capture thoughts on the impact of my study, typically effects in the form of a more reflectional approach to one's own participation in the process. A number of analytical code categories were further broken down into sub-codes at level 3. Impact was a level 2 category capturing the perceived impacts of ratings by the interviewees. One of these codes was capturing what Career implications ratings have in terms of e.g. promotion, dismissal, opportunities to get interesting tasks or attractive inclusion into talent programs or high status training programs etc. A second code captured expressions about the impact on behaviour of receiving different ratings, i.e. considerations on how ratings is not only an outcome based on behaviour, but is also affecting behaviour as if it was almost a structuration process, to put it in Gidden's (1984) terms. A third code was capturing the impact on compensation of receiving different ratings. This could be in relation to bonus, but also to base salary. A fourth code impact on cooperation captured that some interviewees expressed thoughts on how a relative evaluation of colleagues could affect their motivation to cooperate. A fifth code captured the impact on the motivation to be innovative in relation to the set targets that form the foundation of the PA scheme. A sixth code captured the many thoughts expressed about the impact on motivation that ranking people through a PA scheme has. A seventh code 3. Methodology and Data 104 captured reflections on the impact on future performance of receiving a certain rating. Again, the process could be seen as circular in that the rating is given for past performance but may have effect on future performance as well. Eighth and finally a code captured how one interviewee reflected on the rating's effect on the future target setting process. Performance Assessment was another analytical level 2 code category capturing quotes related to the process of providing ratings of the employees. This category contained the following codes: Business performance captured expressions about how the strict business performance affected the assessment process, while impact of behaviour was moderating the picture by capturing the softer, behavioural inputs that went into the assessment process. The code dependency on others captured reflections on how the individual performance was related to dependency on others, e.g. colleagues, vendors, supervisors, subordinates etc. A special personal relations code was made to capture expressions of how such relations affected the assessment process. Quotes on the calibration process were numerous and in fact the single most used code together with the identity construction code under the Sensemaking code category mentioned below. Finally, a Relative vs. absolute evaluation code captured reflections on the degree of absoluteness vs. relativeness of the assessment process. The last level 2 analytical code category was called other analytical and was broken down into three sub-codes. First, a code called Strategic considerations capturing more strategic thoughts behind the PA process as presented by corporate HR in the early informant interviews. Second, a code was capturing background information about the organisation and its development in relation to PA. This code was used for coding the interviews of HR staff. Third, a code called purpose of PA capturing how some interviewees by their own initiative put forward direct or indirect reflections on the purpose of the PA scheme in the organisation. The differing perceptions of purpose revealed by this made me add an extra question in the last interview round at final appraisal: I asked each interviewee to explain in their own words, what they saw as the purpose of the process seen from the organisation's side. The theoretical marker was used to capture statements that were specifically relevant in a sensemaking analysis context. This coding category and its sub-codes were made up front based on Weick's (1995) seven conceptual properties of sensemaking 3. Methodology and Data 105 and was thus not a result of an open coding process. The first sub-code was action where I captured statements focusing on the action part of the meaning creation process, i.e. the 'what happened and why?' elements. I used the sub-code Event selection to code for statements where the interviewees gave examples of good or bad performance. In close relation to this, the sub-code extraction of cues was capturing statements about control of performance event selection for inclusion in appraisal decision. This of course is a process containing a lot of vested interests and potential power struggles. Not all performance events will be evaluated, so making sense by having the right ones selected becomes crucial for managers and employees. The identity construction sub-code was used to capture statements that indicated self-perception around questions like who am I? And, what is my role? Plausibility rather than accuracy was used as a sub-code to capture the 'good stories' that are sometimes required and often beneficial, when a complex activity like assessing a person's overall individual performance over a full year needs to be communicable. Further, I used the sub-code retrospectivity to capture statements about how the individual's share in any performance event could be limited and defined when the work context of all the interviewees was heavily interdependent on systemic and collective factors. Finally, I used the social dimension sub-code to capture statements about the influence on the PA process of an individual's credibility standpoint within the organisation and particularly in relation to the manager/employee relationship. Generally, the coding structure worked well so that most of the time I found it rather easy to decide which code to use for a specific piece of text. However, I did experience some challenges, which I will elaborate upon below. First, the text pieces did not always present themselves so unambiguously that they fit into one code definition only. Some text pieces had relevance to two or even three different codes at the same time. Particularly, I found the theoretical code category to often overlap with the analytical and descriptive categories. This is hardly surprising since the theoretical code category can be seen as a separate analysis dimension laid down on the text. Or it could be seen as another lens through which the text is read and analysed. The below example, can illustrate this. Manager: "Regarding the ratings I would really like if I could have the feeling that I was the one who decided. That I could give the rating that I wanted to." 3. Methodology and Data 106 The manager is referring to the fact that ratings are calibrated and in other ways affected by other actors than the manager herself. In that sense the piece of text could be marked under the calibration sub-code. However, looking at the quote from a sensemaking perspective, this could be seen as an element of identity construction. The identity as manager is built around influence and the power to make decisions including assessing the performance of the manager's employees. The calibration process limits that power in a way that obviously annoys this manager. To make the picture even more blurred, the quote could also be coded as an improvement suggestion since it could be translated as a proposal for increasing managers' rights to freely appraise as they please. One of the advantages of using digital tools for coding is that it is so easy to drag and drop pieces of texts into code categories, so I simply decided to allow myself to code the same piece of text into multiple sub-codes. This way I avoided losing information at the sub-code level, but I also at the same time introduced a second problem: I ended up with a very large number of codes - more than 1200 codes were attached to the 38 interviews that went through the coding process. Had this been done in the old, manual way, it would have been very difficult to avoid getting lost in this volume of coded text pieces. Again, the digital analysis tool makes it easy to extract all text pieces attached to a code at a mouse click. The number of coded pieces, although time demanding in the analysis process, gave me a richer foundation for the analysis. The third challenge I faced during coding was related to the two already mentioned above. I coded the texts with 48 codes. My intention was to keep the coding tree as flat and simple as possible to avoid the coding process drowning in complex code hierarchies, definitions and interrelations. However, 48 sub-codes is still a lot. Looking at Figure 3-2 one can see that there are a number of sub-codes that were rarely used. Codes that are only used a very few times across a large number of interviews should in my mind at least raise concern if the code is critically important, or if it can be either dismissed or incorporated into another code definition, to keep the coding structure as simple as possible. As an example, the sub-code Impact on target setting process was created while coding one of the early interviews and at that time seemed relevant. As I went through the coding of the other interviews, it 3. Methodology and Data 107 turned out to not be something others were touching upon.21 Still, I decided to keep the code, because I found the reflection interesting. The fourth challenge I experienced in the coding process was in a way in opposition to what I have just expressed. The challenge was not that 48 codes were too many but rather that it was too few. When open coding is used, and it is at the same time a guiding design principle to avoid too many codes, there is a risk that the researcher is twisting quotes into sub codes where it would be fairer to the quote to attach it to a new and separate code definition. An example can illustrate this dilemma. The quote below was coded as preparations because it occurred in the context of discussing how a PA meeting with the manager was prepared. Employee: "I believe all managers in the organisation are well-prepared for these meetings". Interviewer: "Why do you think it is so?" Employee: "Because they have to. It's governance. They cannot allow themselves to go red on this one." Why not create yet another code called "governance" or "manager success criteria"? In this example I chose to stick to the preparations code because I thought that a discussion of governance criteria as a motivator for manager meeting preparation would fall naturally in the context of analysing the meeting preparations. However, I admit that it could easily be argued to do otherwise. I found it important to keep the code structure as simple as possible while still capturing the most significant text perspectives in a structured manner. Keeping this balance comes through compromises and the important thing then is to be consistent in the use of the coding structure through the coding of all texts. 3.5 Procedures for Research and Documentation Besides the above methodological considerations, the research protocol also laid out the more practical procedures for conducting research. These included procedures to secure confidentiality and ethics, establishment of templates, preparation of invitation letters, principles for keeping track of the large volume of established data in a research database and finally procedures for storage and backup of data. Although this part of the research protocol was very stable over the course of the study, some 21 I see this as a side effect of using a grounded theory approach to coding. 3. Methodology and Data 114 qualitative studies through open interviews. Having an open mind towards research questions, methodology, and theory is in my opinion a requirement for being able to conduct interesting qualitative research. If one wants to test a specific issue via closed questions and without the opportunity to adjust along the way, one would be better served by a quantitative approach. But investigating complex individual meaning creation in a highly institutionalised field, does not call for closed questions or quantitative methods. My Interviewees did not see things in one uniform way through the three interview phases. They developed their views and arguments - sometimes even to contradicting positions. And so did I. Conducting a qualitative study in my own organisation requires mastery of balancing the obvious advantages of being an insider with the according risk of prejudice. By spending so much effort and time preparing a detailed research protocol, I have tried my best to make my approach to this dilemma as transparent as possible to the reader. While a chapter on methodology naturally focuses on epistemological questions I find it most fair to the reader to finalise the methodology chapter to be transparent about my ontological standpoint. I consider myself to be a critical realist in the sense that there is a material world outside our minds, but it is complex and difficult to understand, and we therefore need theories to help us understand more, although still not all.24 In a PA context there are real things at stake: real projects and activities that are to be evaluated. There are also real performance ratings and real rating consequences. But the evaluation of an individual's contribution is highly subjective and partly irrational. In fact, the past 20 years of research into evaluation bias and other problems in PA makes this quite evident. 24 See e.g. Archer et al. (1998). Although there are real things at stake, the pre-defined corporate PA process is not uniformly implemented and operated across the organisation. It is mediated by and influenced by sensemaking processes driven by a need for creating meaning. I.e. a need for rationalising what has been done, rather than doing what is rational. 4. Case Introduction 115 4 Case Introduction PharmaComp25 is a large Scandinavian pharmaceutical company. The company is specialised in research and development, production and marketing of medical drugs. Headquartered in Scandinavia, the company has a worldwide presence with production sites and sales affiliates all over the world. The present case study overview is based on preliminary studies of internal documents, my own background knowledge and interviews with key stakeholders in the implementation and operation of PA in the company. These informants have held varying roles over the years since PharmaComp set out to implement PA. For example, an informant was managing the implementation project, one is a central HR statistician utilising the new opportunities for analysing aggregate data and another is responsible for the continued improvement of the global PA process. Over the past 10 years the company has experienced a strong growth rate in terms of sales and also a strong increase in number of employees. The strongest growth markets both in terms of sales and number of employees are outside Europe. In 2004 there was no corporate PA process in the company. Instead, local practices had developed over the years at the discretion of local management. A survey initiated by the corporate HR function in 2004 showed that there existed at least 25 different practices of PA in different parts of PharmaComp. The practices differed in terms of: - Standards (e.g. was there a rating scale or only a free text evaluation? How many grades were there in the rating scale? Was a recommended or ideal rating distribution defined?) - Documentation (e.g. was a template used and how was the template designed and stored? Was data aggregated at higher organisation levels?) - Process (e.g. who were in scope for PA? How many and what steps were included in the process? How many compulsory yearly meetings and appraisals were there?) 25 In the following I will use the term PharmaComp as a pseudonym for the case organisation. 4. Case Introduction 116 The consequence of this state of affairs was that there was no global overview and strategic direction in terms of PA. No one had a clear overview of how to identify talent across the global organisation. There was no IT system to support alignment of the PA process although this was giving way to increasing challenges due to an accelerating internationalisation and a subsequent increased mobility. Thus, performance evaluations were generally not transferable or comparable across organisational borders. In addition, from an HR efficiency point of view, there was an inefficient use of resources since so many different templates and processes were maintained locally and simultaneously. In August 2004 Executive Management decided to launch a corporate PA process. The intention was stated as to: - Establish a common baseline for PA across PharmaComp - Support the strategic focus on internationalisation and mobility - Improve performance data overview and alignment of processes - Secure compliance with the corporate management principles in all units - Create corporate transparency on individual performance - Prepare for an IT system roll out While choosing this approach, PharmaComp relied heavily on a ‘best practice standard’ as defined in the People Capability Maturity Model (SEI/CarnegieMellon). According to this model PA should support the realisation of the unit’s targets and individual performance objectives should be documented, performance should be regularly discussed to identify actions that could improve it, performance problems should be managed, outstanding performance should be recognised or rewarded and PA practices should be institutionalised to ensure they are performed as managed processes (Curtis, et al., 2002). The plan was to first establish a companywide paper based process for PA of all managers and up, in phase 1, and later to implement a supporting IT system in phase 2. The process was named the “People Performance Process”. It was decided not to include the relation to other areas such as rewards (salary increase, long-term incentives, and short-term incentives) in the scope of the first two phases, but leave these more advanced relations to a future phase 3. In Figure 4-1 below, the PA 4. Case Introduction 117 journey is illustrated as it appeared in the project documentation at the outset in 2004. Figure 4-1: PharmaComp’s PA journey as set out in 2004.26 Figure 4-1 illustrates the three-step implementation process for implementing a corporate PA scheme in PharmaComp. Below the three phases will be explained in further detail. 4.1 Phase 1: Establishing a Common PA Process (2004) In phase 1, a set of minimum requirements were defined: a) A common target setting, review and appraisal paper template, b) A common 5-point rating scale, c) A common process and timeline for target setting, review and appraisal (three mandatory yearly events), d) Documentation requirements related to the process and e) Clear roles and responsibilities for managers, local HR and Corporate HR. Below, each of these five core characteristics will be further clarified. Blue collar workers were, not included, partly because of union resistance but also because it was seen as less relevant to manage individual goals in this way for people doing more routine tasks on work shifts. 26 People Performance Management ver.02 (2004) 4. Case Introduction 118 A common template Corporate HR decided to collect all existing local templates as inspiration for the preparation of a new standard template that supported the corporate objectives. Figure 4-2 shows the new template. Figure 4-2: The Corporate PA template The template in Figure 4-2 contains a number of information points to be tracked through the PA process. There was also a reference to the “balanced scorecard perspective” of the business area. The balanced scorecard structure of PharmaComp operates within four perspectives: Financial targets, business process targets, customers/society impact targets, and people and organisation targets. On top of the four balanced scorecard perspectives, two additional perspectives were added; a general “other goals” category and a supporting competency goals category. Then, for each of the three yearly sessions there were distinct columns. In the goal setting column the targets were codified according to a specified format relying on the definition used in the breakdown of the corporate balanced scorecard. 4. Case Introduction 119 First, a Critical Success Factor (CSF) should be referenced. Critical Success Factors are those factors that are most critical for the achievement of business success. CSFs identify the significant things an organisation or an individual must do well to succeed. Thus, the corporate balanced scorecard definition of CSF states that: A CSF does not describe a current situation, but an aimed objective. The objective should be challenging, yet realistic. The balance and priorities set in the CSF are particularly important. Each perspective usually contains 3-6 primary CSFs that state a key aspect of the strategy to be achieved over the next 2+ years.27 Critical Success Factors are the most general level of targets. Examples of critical success factors at the corporate level would be e.g. Realise growth in Operating Profit (Finance perspective) or Enhance winning culture (People & Organisation perspective). The purpose of defining the link to critical success factors and generally following the vocabulary of the balanced scorecard process is to create a strong link to the strategic direction set by senior management. Second, for each target a KPI should be defined. The role of a KPI is to specify a CSF and the KPIs are part of PharmaComp’s management reporting system. The corporate definition of how to create KPIs state: The definition of the KPI should demonstrate how and when the CSF will be achieved, and must be measurable. One or two suitable indicators should be defined for each CSF, demonstrating its achievement. The indicator can be displayed in absolute terms (e.g. $000m), or using other valuation criteria (e.g. employee satisfaction index).28 Examples of KPIs are Operating Profit Growth (Finance perspective) and Customer Satisfaction (Customers & Society perspective). Third, the concrete targets are measuring points for the KPIs, which ideally provide management and employees with a clear idea of what to strive for, and by when. Referring to the KPI examples above, targets could be operating profit growth = 12% or Customer satisfaction = 4.0 on a 5-point rating scale. Each target definition would also require a due date for the target to be measured and evaluated. Fourth and finally, a target weight could optionally be added to codify the relative prioritisation of targets. 27 Corporate definition of Balanced Scorecard elements, Corporate intranet 2011 28 Corporate definition of Balanced Scorecard elements, Corporate intranet 2011 4. Case Introduction 120 The midyear review column should hold the employee and manager comments as discussed through the midyear session. Besides a follow-up on the status of target realisation, the midyear comments could also include rationale for adding new goals, changing existing goals or even deleting goals that had become obsolete during the time passed since target setting. The year-end appraisal column holds the final evaluation by the appraisee (selfappraisal) and the appraiser. The appraiser has the final word in rating each target on the corporate 5-point rating scale (see below), but the appraisee must firstly note her self-appraisal as input to the discussion between the two. As for the midyear review, there is also for the final appraisal a table box that holds the conclusive appraisal comments from both the appraisee and the appraiser. But this time there is an additional field where the manager documents the overall rating of the appraisee’s performance as measured on the corporate 5-point rating scale. For each of the three appraisal sessions there are signature fields for both manager and employee to document that the process has been conducted timely and correctly. The appraisee does not sign to her agreement with the appraiser’s evaluation, but to her own comments, and to verify that she has read and received the appraiser’s feedback. A 5-point rating scale One of the most difficult steps in aligning the many different PA schemes existing prior to the corporate standardisation initiative from 2004 was to establish a compromise for a corporate appraisal rating scale. One of the key people involved in the 2004 standardisation initiative recalls: It was a mixture of pragmatism and best practice. There were all these many rating scales – 3-point, 4-point, 5-point and 10-point scales. There were most that used 5point scales, and also, our later IT solution provider recommended a 5-point scale as most normal. So it became a 5-point scale. But how was the decision to even have a rating scale, i.e. a quantitative rating of employees, taken in the first place? An HR person close to the 2004 initiative explains: It was a given that there should be a rating scale. It should be possible to compare across PharmaComp to be able to deal more systematically with low and high performers. What did a 3 mean here and a 10 there? Nobody could compare across organisational entities. We needed to be able to track our top performers. 4. Case Introduction 121 So, the result was a 5-point scale ranging from does not meet expectations and targets to Outstanding. Below each of the five levels are defined in accordance with the original rating scale definition from the 2004-initiative: Outstanding. An outstanding contribution to the company’s performance. The person delivers substantially more than is required. The performance leaves little or nothing more to be desired. The way to accomplish the target(s) has been significant compared to standard in unit. The reaching of the target(s) will be used as a best practice in the unit. Exceeds expectations and targets. A significant contribution to the company’s performance. The person has exceeded expectations – faster and/or better than set forth and with desired behaviour. The person has during the reaching of the targets anticipated and proactively solved critical problems. Meets expectations and targets. A solid contribution to the company’s performance. The expected results have been achieved within agreed timeframes, available resources and with the desired behaviour. Approaches expectations and targets. The person should improve his contribution to the company’s performance. Performance fails to consistently meet expected results. All targets have not been met and currently, accomplishment of targets is below the standard of the unit. Does not meet expectations and targets. The person’s contribution to the company’s performance is not acceptable. Targets are far from met and behaviour is considered poor or even negative. In addition to these five categories, there was a sixth called not applicable. This was reserved to situations where targets were postponed or not relevant. From the outset, there was no support for enforcing a forced distribution of ratings on the 5-point scale.29 A corporate HR representative involved in the 2004-initiative explains: Where we did not succeed was to build support for a forced distribution. Personally, I was in favour of forced distribution but without absolute definite descriptions in the rating scale. I thought we should be clear: ‘We rank our people, it is a race and someone is first’. But there was no support from top management for a relative scale, probably because they favoured a more cautioned approach with focus on getting basics right first: Is the processes followed and are timelines met? We in HR also had ambitions concerning quality, meaning how things were documented and audits of the process outcome, but we did not get through with that at that stage. So, everyone in scope for PA was rated against the 5-point rating scale, but there was no forced distribution across PharmaComp and no systematic follow-up on differences in rating levels. 29 Forced distribution is a process according to which a specific share of employees must receive a specific rating. E.g. that 5% of all employees must be within the lowest rating category. The purpose of forced distribution is to be "the antidote to the problems of inflated rating and the failure to differentiate that many organisations have installed to help bring the truth into the performance management process" (Grote, 2005) 4. Case Introduction 122 The yearly process cycle of eight steps and three sessions An 8-step process was defined and implemented around a three-session structure, as illustrated in the original process map in Figure 4-3 below. Except for a few minor adjustments, this process has been in operation in PharmaComp since its implementation in 2005. Figure 4-3: Initial PA process definition30 While the appraisal period is defined by the calendar year, the target setting starts on December 1 and ends February 28, after the appraisal period. This is to allow 30 People Performance Management ver.02 (2004) Review (from June 1 to August 31) Schedule session. Obtain relevant, interim performance feedback. Document review in 3P template and sign. Review status on Business Targets. If necessary, agree on revision of targets. Suggest ways to develop and/or improve (under comments in 3P template). Keep copy of updated and signed 3P template. Provide feedback on performance and progress on fulfilment of Business Targets. Enhance motivation. Prepare Execute Finalise Appraiser Joined action Appraisee Self assessment – fill in 3P “appraisee review comments”. Send to appraiser. Target setting (from December 1 to February 28) Schedule session. Prepare own unit targets. Document targets in 3P template and sign. Agree on Business and Competence Targets (use SMART). Draft Business Targets based on unit targets (use SMART). Send to appraiser. Keep copy of signed 3P template. Prepare Execute Finalise Appraiser Joined action Appraisee Appraisal (from December 1 to February 28) Schedule session. Obtain performance feedback. Assess and rate fulfilment of Business Targets. Document rating(s) and comments of appraiser and appraisee and sign. Discuss fulfilment of targets and give appraisee conclusive rating. Self assessment – fill in “appraisee conclusive appraisal comments”. Send to appraiser. Presentation of selfassessment prior to session. Send updated, signed 3P template to local HR function. Keep copy if needed. Provide feedback on performance and fulfilment of Business Targets. Prepare Execute Finalise Appraiser Joined action Appraisee Local HR: Enter conclusive performance appraisal ratings in SAP (Deadline: April 1) Entry criteria: •Knowledge of relevant business priorities Exit criteria: • Filled in 3P Business Targets • Template signed • Optional: Weight Entry criteria: • Signed 3P template with Business Targets Exit criteria: • Filled in 3P review comments • Note revised targets if any • Template signed Entry criteria: • Signed 3P template with review comments Exit criteria: • 3P appraiser conclusive rating filled in • Conclusive comments • Template signed • Optional: appraiser & appraisee rating for each target People Performance Process (3P)People Performance Process (3P) 4. Case Introduction 123 business areas time to define targets before the appraisal year starts, finish the appraisal year before evaluations are done and generally have time to conduct meetings etc. The process requires at least three appraiser-appraisee sessions per year. Below, I will go through each of the three phases and the eight steps in more detail, in accordance with the official documentation held by the company and presented to the managers and employees through training, communication and other material.31 Target setting Step 1: Employee creates PA draft targets To begin target setting the employee prepares a number of targets reflecting the main business objectives for the performance period. Targets must be defined according to SMART32 principles, in this context defined as Significant, Measurable, Ambitious, Realistic and Time bound. Targets can be created and aligned with business objectives in different ways. Targets may be created individually by the employee or be cascaded to the employee from the manager or any other employee. When business targets have been prepared, the employee needs to decide which competence targets may support the business targets, and include these in the PA form. The use of competence targets is optional, unless local guidelines stipulate otherwise. The employee has the opportunity to attach documentation – for instance an Individual Development Plan. Weighting targets is an optional way of aligning appraiser’s and appraisee’s view on the priority of the specific target. Whether a target is weighted high due to the time it takes or due to importance, or both, is a matter of dialogue between appraiser and appraisee. The managers and employees were encouraged to consider the following if using weights: Weights should be defined at target setting in order to align expectations. Weighting is measured in percent and should equal 100%. Weights should be used as guidance in reaching an overall conclusive rating at year end.33 31 People Performance Management ver.02 (2004) and corporate website 32 Doran, G. T. (1981). There's a S.M.A.R.T. way to write management's goals and objectives. Management Review, Volume 70, Issue 11 (AMA FORUM), pp. 35-36. 33 People Performance Management ver.02 (2004) 4. Case Introduction 130 points that units could compare against. These were named "global minimum benchmarks" and were characterised by the following design criteria: They should be auditable so that it was clear what was required. They should be valuable, so that it was clear how the company benefited from the adherence to them. And they should require minimal extra work for line managers. Five global minimum benchmarks were defined and for each, measuring opportunities were established. The five elements were: A) All employees should follow the PA process. The rationale for this element was to ensure that all employees (except where PharmaComp was bound by legal requirements, e.g. union agreements with blue collar workers) had business targets and that they received performance feedback at least twice yearly. B) All employees should receive an introduction to PA. The rationale for this element was that in the past the only formal introduction of employees to PA in PharmaComp consisted of two quiz questions that new employees were made acquainted with through the onboarding process. Corporate HR wanted to secure that a basic employee understanding of the process was established with the present and future employees so that expectations could be set and so that it was signalled to all employees that PA is taken seriously in PharmaComp. C) All managers should be trained in PA. This element was crucial since line management was the owner of the process and in practice had the tough role to identify, document and translate the performance of employees into feedback and ratings. D) All management teams should calibrate their ratings. This element became relevant because it was clear that there were very different practices. In some units, calibration did not take place at all. In others, calibration took place, but in very different ways. Corporate HR saw calibration not only as a method to improve the transparency and fairness of rating distributions across different areas, but also as a learning opportunity for management, who could coach each other, learn from each other and even give input to the rating of each other's employee ratings. E) PA appraisals should focus not only on outcome but also on how those results were accomplished (behaviour). This element is much in line with the strong emphasis on values that top management enforces and which is part of 4. Case Introduction 131 PharmaComp's culture. Delivering strong business results is crucial, but it must not be at the cost of business ethics. Therefore, corporate HR wanted to make sure that all managers and employees had a dialogue on how performance targets could be combined with explicit targets for behaviour. As the reader can imagine it is not a trivial matter to roll these five basic elements out to a global organisation with more than 30.0000 employees in 50+ countries and keeping the elements alive onwards with challenges like employee turnover, intensive growth in employee base and promotion of new managers and coupled with the fact that for line managers there basically are other and higher priorities than HR process compliance. All through 2011 and still into 2012, significant efforts were invested in corporate HR to anchor the global minimum benchmarks. Increased statistically based knowledge and diffusion of use Over the years since PA was implemented in PharmaComp, corporate HR has made different statistically based studies to investigate the PA practice across the organisation. I will go into more details on this in the following chapter. Here, it will suffice to mention that such analyses have produced interesting knowledge about the relation between rating and variables like geographical location, workforce, hierarchical position, bonus payout, retention ratio per performance rating, etc. Such studies have not only uncovered many examples of variation in practice across the organisation, they have also provided an educated platform from which to kick off new corporate initiatives like the above mentioned Global Minimum Benchmarks' initiative. Another factor that has worked to increase the effects of the PA process has been what I would call the diffusion of use. As soon as ratings become available in broader volumes, diffusion to related processes starts mushrooming. This tendency becomes even more pronounced where the data can be automatically incorporated into an IT system, because reuse then becomes much more cost effective. This lowers the entry barrier for reuse compared to the value provided by the reuse. In PharmaComp, PA ratings are now available not only in the PA system itself but also in the corporate ERP system, in the corporate data warehouse, in the corporate manager self-service portal and in the corporate succession planning and talent management system. Most likely, this development of diffusion of use will continue 4. Case Introduction 132 over time, so that PA ratings will be made automatically available to all processes where the data is experienced to provide value to the business. 4.4 Summary This chapter introduced the history and process of PA in PharmaComp. PA started as a corporate project in 2004 with the aim to establish one uniform, global process to substitute the at least 25 local processes that existed at that time. The purpose of PA was to establish a common baseline for performance management across PharmaComp, to support the strategic focus on internationalisation and mobility, to improve performance data overview and alignment of processes, to secure compliance with the corporate management principles in all units and to create corporate transparency on individual performance. PA implementation went through several phases, starting cautiously by defining and implementing a 3 stage, 8 steps paper based process in 2005, moving on to implementing a supporting IT tool in 2007 which in the beginning were aimed primarily at the top layers of the organisation. This was followed by a series of initiatives from 2008 and onwards that further rolled out and improved the IT solution and aimed at increasing the quality and global alignment in the practice of the PA process across the organisation. Lately, corporate HR's focus concerning PA has been to rollout and nurture Global Minimum Benchmarks to secure that there is a solid, common and global foundation for conduction PA across PharmaComp. 5. The Macro View 133 5 The Macro View The purpose of this chapter is to further present and discuss the state of affairs of PA in PharmaComp from a macro view to set the stage for analysing local adaption and meaning creation across workforces. Thus, in this chapter I will present some of the empirical data established during the study of PA in PharmaComp. Particularly, I will draw on a rich quantitative dataset included in the archival material to which I was kindly given access by the Corporate HR Rewards group. The quantitative data set made it possible to draw macro level conclusions about the relation between performance rating and a number of variables like workforce affiliation, hierarchical level, retention and more. Due to limitations in the data set, some parts of the analysis will only be presented at an aggregate level while other parts are broken down into the four workforces in focus of my study. In the analysis of the data I will also draw on the interviews I conducted with representatives from corporate HR early in the study and connect some threads to qualitative interviews with managers and employees where these concern the macro level design of the process in each workforce. Besides helping me in the work to interpret the quantitative data, this material will help to illuminate some of the macro level work that has been done in parts of the organisation to handle various issues and challenges that local management has faced with the PA process. For this purpose, I will also draw on a number of other archival materials from HR departments outside Corporate HR to illustrate local policies within certain parts of the PA process. Finally, along the way I will relate the findings in PharmaComp to the scholarly research agenda as it was previously uncovered in my literature review. The present chapter serves at least two purposes. First, the chapter will further build on the previous case introduction chapter to provide a richer and more detailed picture of the situational circumstances under which local adaption and meaning creation occurs, as it will be analysed in the following chapters. Second, the chapter will serve to illustrate the point that there is no reason to believe that PharmaComp is in any significant way deviant from what existing research has already uncovered through multiple studies in other organisations regarding macro experiences with PA processes. As such, the case is exemplary and holds relevance beyond the boundaries of PharmaComp. This is also important because this 5. The Macro View 134 recognition was one of the reasons why I turned my attention towards local adaption and meaning creation. The case organisation provided me with access to: a) an analysis report from 2012 of various interdependencies between performance rating and organisational impact (Halse, 2012) and; b) a data extract of all centrally and electronically stored PA ratings across the global PharmaComp organisation. The dataset covers a time series of PA ratings from 2006 to 2011.36 The data volume and quality grows over the period as: a) more employees are covered by the PA process and b) PharmaComp's Corporate HR has strengthened its governance concerning collection and electronic aggregation of data. Hence, most of my quantitative analysis below will be made on 2010 ratings, where coverage and data quality is highest.37 Thus, regarding 2010 PA ratings, it is reasonable to say that I am conducting analysis on what is essentially close to a full population dataset.38 One advantage of this is that I can conduct the analysis by using descriptive statistics, without having to cope with the uncertainties of inferential statistics (e.g. sample representativeness and statistical uncertainty), as would have been the case with a randomised sample dataset. Another advantage is that I will have occurrences of all ratings in all workforces which would be difficult to get in a smaller sample since some ratings, particularly the lowest ratings, are rare. The downside of using the full population data set is that I am not only analysing data for the Scandinavian part of the organisation in which I conducted interviews. Though the major parts of R&D and Staffs are located in Scandinavia, the situation is different for Production and Sales. Production has large sites outside Scandinavia and in Sales almost all personnel is located outside Scandinavia. Thus, some caution need to be taken concerning the conclusions made in this quantitative analysis since I have not controlled for possible geographical effects. 36 Each rating corresponds to performance for the year before so that the rating received by an employee in 2011 concerns the employee's 2010 performance. For the remainder of this thesis, when using the term "2010 rating", it means the rating concerning 2010 but communicated to the employee in 2011. 37 Typical data quality issues are: a) some units did not report their PA ratings in due time to Headquarter, b) some units reported incorrect data for some employees, or c) some units reported correct data but changed the rating after reporting to Headquarter without reporting the change. Particularly in the first years of the data set, issue (a) was present. All three mentioned data quality issues are rare regarding 2010 ratings. 38 The 19.843 PA ratings registered for 2010 should be compared to a total number of more than 29.000 employees at the beginning of 2010. From the latter should be drawn employees who left during 2010, employees who were on leave and did not have any PA targets set for 2010 and employees generally not included in the PA process (e.g. blue collar workers). Hence the conclusion that the data set almost covers the full population. 5. The Macro View 135 Below, I will go through the macro analysis in seven steps. First, I will present how the overall number and rating average of employees included in the PA process and whose ratings have been reported to Corporate HR, has developed over the period covered by my dataset. Second, I will break down the dataset to the workforce level and analyse differences. Third, I will present and analyse findings regarding the relation between hierarchical level and performance rating. Fourth, I will present and analyse findings regarding the relation between PA rating and employee retention. Fifth, as the dataset did not include any information about salary increase and size of bonus, I will present and discuss the findings of a previous study in a part of PharmaComp that investigated this relation. Sixth, I will look into the policies for balancing business results against personal behaviour when determining the final appraisal of an employee. Seventh and finally, I will summarise the main points of the quantitative analysis and draw conclusions. 5.1 Performance Ratings at PharmaComp Over Time In 2010, the PA ratings of 19.843 employees across the global organisation were electronically recorded by Corporate HR. As Figure 5-1 below shows, this is the culmination on a trend of growing numbers of recordings. Figure 5-1: Number of employees in dataset and average PA rating, by year Figure 5-1 shows the number of employees and the overall average PA rating per year for the period 2006-2010. The blue bars illustrate the dramatic increase in 5. The Macro View 136 number of employees included in the PA process and whose PA rating is electronically available for analysis due to improved data quality. The electronic PA system is one important data source. As previously explained, the PA process was introduced as a paper based process in 2005 while the electronic PA system was first introduced in 2007. Besides pulling data from the PA system, from 2008 Corporate HR also started to collect data from the paper based PA process from the global organisation. This was done by requesting the units to forward spreadsheets to Corporate HR with PA ratings for employees not using the PA IT system. Over the years this process has been capable of catching a growing share of all ratings and from 2010 the coverage has become really good. The red line in Figure 5-1 illustrates the global average PA rating by performance year. We see a persistent decrease over the years from 2006 to 2009 where after the trend is flat to 2010. Below, I will show how the higher levels in the organisation systematically receive higher ratings than the lower levels. It is therefore no surprise that the average rating decreases as more and more lower level employees are added to the PA process and the systematic collection of ratings. However, over time the top layers in the organisation will constitute a still smaller share of the total population of ratings recorded, and this effect will be relatively smaller, which may explain the flattened curve from 2009. 5.2 Performance Rating Distribution by Workforce Corporate HR is not enforcing or recommending any specific distribution of ratings across the 5-point rating scale: Outstanding (O), Exceeds Expectations (EE), Meets Expectations (ME), Approaches Expectations (AE) and does Not Meet Expectations (NME). However, many business areas define their own guidelines regarding rating distribution. For example, in one part of Staffs the following distribution is used as a guideline: O 0-5%, EE 5-20%, ME 50-75%, AE 5-20% and NME 0-5%. In another illustrative example, Production in 2009 analysed the rating distribution of 2008 and Production Management recommended a different distribution for the following year. 5. The Macro View 137 Figure 5-2: Recommended & realised rating distribution in Production, 2008/09 Figure 5-2 shows the recommended 2009 and realised 2008 PA ratings in Production. The figure shows several things. First, in 2008 Production rated half the share of employees into the two lower rating categories, compared to what Production Management recommended for the following year. So Production Management was not satisfied with the limited use of the lower rating categories. Second, the realised curve was somewhat twisted to the right in the diagram compared to the aspiration for the following year, so that the Exceeds Expectations category had more occurrences than preferred by Production Management. Third, the figure also illustrates that Production Management was very precise in its distribution recommendation for 2009. Although PharmaComp does not operate within a 'forced distribution' paradigm (Grote, 2005), Production Management's recommendation should serve as guidance for calibration meetings in each business unit/functional sub-area.39 Production Management found that the analysis had shown ...great differences in ratings which could not be accounted for in the department/function performance itself. Management saw better alignment of performance ratings across Production as key to minimise rater bias (hard raters vs. soft raters), ensure a uniform understanding of 39 According to a Production HR presentation, the recommended performance distribution for 2009 was "based on PA rating distribution in other areas in PharmaComp compared with overall performance in the area, historical performance in the area (excellent performance = right skewed distribution) as well as books/articles on performance distribution and performance calibration practices." 5. The Macro View 138 the rating scale and to improve the value of the process in terms of creating performance improvement. Management stated that the: ...recommended distribution is not forced but managers should use it for challenging/discussing direct reports and whether the rating distributions for each manager make sense based on overall performance. I.e. a low performing unit having more right skewed performance distribution than a high performing unit should probably be challenged. The terminology used by Production Management is cautiously avoiding the term "forced distribution" (Grote, 2005) and instead use the term "expected distribution". But with a still stronger enforcement to meet the expected distribution, managers and employees sometimes see the difference as subtle. One larger Staff area in 2009-10 conducted an initiative to improve PA. An analysis had shown that there was a significant gap between what senior management in the area termed an "ideal distribution" and the actual distribution of ratings, as illustrated in Figure 5-3 below. Figure 5-3: Recommended and realised rating distribution in Staffs area 2008/09 Figure 5-3 shows the recommended and actual distribution for the Staffs area in 2008 and 2009. The ratings have been grouped into three main clusters; the low-end ratings NME and AE, the medium rating ME and the high-end ratings EE and O. The analysis of the 2008 distribution had clearly shown that managers were very reluctant to use the low-end ratings while they at the same time were using the high-end ratings almost twice as often as senior management wanted. By focusing strongly on PA throughout 2009, management succeeded in bringing the actual distribution much 5. The Macro View 139 closer to the recommended distribution in 2009, although the tendency still was an underweight in low-end ratings and an overweight of high-end ratings. In Sales, the process is somewhat different in that PA targets are primarily containing target indexes for realised sales. A certain percentage index realisation within one target is attached to a specific rating of the target. In theory, everyone could be Outstanding. However, that is very theoretical, since targets are adjusted quarterly based on sales forecasts. When a product has been in market for a while, forecasts are very good and it is hard for salespeople to significantly over perform. When products are new in the market, there is more uncertainty in forecasts, and salespeople can sometimes over perform with several hundred percent. Though, in the following quarters the forecasts will of course be adjusted again to level out the variation seen over the full year. For this reason, and because bonus and PA are not interdependent in Sales as they are in most other parts of PharmaComp, it has not been relevant to Sales management to establish strict recommendations regarding rating distribution. In parts of R&D, management has gone furthest in aligning the relation between distribution and area business performance. Hence, based on the overall performance of an area, a pre-defined average rating with a precision of one decimal, and a predefined number of ratings within each of the five categories on the rating scale, must be met within the area. For example, in one area within R&D, management was told that based on the performance of the area, the average PA rating for all employees should be 3.7. Further, all units should make sure to hit a standard distribution across the rating scale within different job categories. In other words; a forced distribution approach. While this obviously provides management with strong control mechanisms in the rating process throughout the organisation, this would also be a controversial design in the view of scholars presented in the literature review who took a critical stand towards PA schemes, based on the argument that such schemes have a build-in tendency to develop to become bureaucratic and control focused (Piercy, 1994). 8. Meaning Creation in Midyear Review 242 retrospectively create a new meaningful story about the action that has already taken place, namely that the manager has now acquiesced and de facto accepted that the employee will not participate actively. This new story is about how acquiescence is meaningful because the employee is a good performer who just happens to have abandoned PA, and that the result of further conflict escalation would be detrimental to the manager's purpose of securing organisational performance. This story relates to a very fundamental problem in PA as it was depicted by Kohn (1999); that it is still unproven that PA schemes increase organisational performance. In fact, Kohn goes further and claim that PA schemes have an outright deteriorating effect on organisational performance. The manager in the example senses that a performance decrease may be the outcome of insistence on employee engagement in the process. And because organisational performance is the key priority to this manager, it is meaningful for the manager to acquiesce and she accepts what is in reality a partial abandonment of the PA scheme for this single employee. Further, we could argue that it is particularly in R&D we should expect to find such examples of employees' outright rejection of PA. The task at hand in R&D has the worst fit to the management by measurement regime that PA is affiliated with. And R&D has the highest concentration of highly skilled expert employees, often equalling or surpassing the organisational prestige of their managers. Particularly in R&D, we should therefore expect to find employees who not only see PA as meaningless, but who also have the organisational power base to challenge their managers' authority to enforce PA upon them, and to sometimes win that conflict, as was the case in the above example. Acquiescence as creative process redesign The third manager story is about creative process redesign in Sales. One of the Sales managers I interviewed had taken the ultimate consequence of employees' reluctance to prepare for midyear review. She had simply taken systematic and full responsibility for what should, according to the PA guidelines, be the responsibility of the employee: I have prepared the PA before the meeting. None in my team are good at driving their own PA. I have almost stopped asking them to forward their input before the meeting because they don't get it done. I don't know if I have spoiled them... They have become so used to me finding all the numbers and writing it down. 8. Meaning Creation in Midyear Review 243 This third story is related to the first two stories above and to the same problems in PA. However, here we are no longer talking about a single employee who gets special treatment for the sake of maintaining organisational performance. Here, the manager has in effect changed the PA process as it is laid out by corporate HR, by leaving out the process step that concerned employee input. Instead, the manager herself systematically does the meeting preparation work on behalf of the employees. The manager does this simply because she fears that if she did not, the employees would not get it done. We have already seen previously how in Sales PA is not connected to reward for Sales reps and how the instalment of targets were almost entirely top-down controlled. This may help to explain why Sales employees, at least in this department, are particularly reluctant to prepare for PA meetings. But from a governance perspective it would still be unacceptable to the manager to have a whole department missing its PA deadlines, so it is more meaningful to the manager to do the work by herself. This example may at first seem surprising when we consider that Sales is the workforce where we found that the task at hand had the best fit to a management by measurement regime. If the fit is good, why does the manager not prioritise enforcing the corporate process by demanding employee engagement in self-evaluation at Midyear? In a previous example from the target setting process we saw how a Sales employee was completely alienated from her PA because she felt targets were distributed by management without her involvement. A Sales employee elaborates on this: I cannot see the need for this [PA]... I guess it is something management needs to have to feel they are in control. Although I am not sure what they are using it for. I haven't seen the light in this. In Sales, the definition of targets is strictly top-down and the measurement of target realisation is quantified and based on numbers drawn from another system. And further, bonus is allocated purely in accordance with sales realisation, where numbers are beyond discussion. In such a regime, there is little incentive for employees to invest time in adding comments in PA as part of a self-appraisal. What difference should that make? 8. Meaning Creation in Midyear Review 244 However, for the manager, PA process compliance is still required. This explains why the manager in this illustrative, although maybe also extreme, example, is systematically doing the work for the employees in her department. Discussion In the three examples above we have dealt with the concept of preparing for midyear review. However, ideally the preparation for midyear review should not only be about spending an hour on providing input right before the meeting takes place. Preparing effectively for midyear review would require a more continuous effort. Thus, one particularly important element in preparing for the midyear review meeting would be to take notes of important performance events as they unfold. The reason why this is necessary has to do with the recency problem in PA (Coens and Jenkins, 2002). According to Coens and Jenkins this bias could also be called the 'what have you done for me lately' syndrome, because it concerns how feedback, unless very well prepared, will primarily concern what has happened most recently because that is what is most present in the memories of actors. Coens and Jenkins see this as one of the key problems within PA in general, because PA schemes install a calendar controlled paradigm on important feedback from managers to employees. There may be other occasions for feedback, but according to Coens and Jenkins, PA schemes make a differentiation in importance of normal (not committing) feedback and PA (committing) feedback that devalues the former and emphasises the latter. To support on-going note taking - and to be able to compare managers' and employees' selection of performance events in their note taking, I offered a 'diary' template at the first interview with all managers and employees.60 However, only two employees and one manager used the diary. And the manager only did it through the first half of the interview cycle, i.e. until midyear review. To the question whether the notes was a help to the manager during the preparation for the midyear review meeting with the employee, this Staffs manager replied: Yes, I could remember the examples much better because I could run through them right before the meeting... I think she [the employee] maybe felt we got into too much detail with some of the things she considers to be minor. Or maybe because those are some of the things she would prefer we did not take up. 60 See manager and employee diary templates in the appendix 8. Meaning Creation in Midyear Review 245 The ability to draw on tangible examples in some detail was heightened by the manager's diary notes. However, the employee did not necessarily appreciate this higher level of detail. And this manager did not continue with the diary in the second half of the cycle. Like the other managers I interviewed, she gave as reason that she had been too busy with other things and that daily note taking simply is not high enough on the agenda to get it done on a regular basis. Or, to phrase it differently; the manager's diary helped the manager to remember performance events she would otherwise not have remembered. But still, it was less meaningful to continue doing this because the examples were not valued by the employee or they were in other ways not important enough to the process to justify the required manager effort of taking on-going notes. Apparently, and in contrast to what one might have assumed, the PA feedback and midyear review does not require detailed documentation of performance events to be meaningful. In fact, as the above example shows, it is sometimes easier to create meaning from the process without those examples, maybe because maintaining the feedback dialogue on a less tangible level opens opportunities for softer interpretations and less conflict. In a sensemaking perspective this is not surprising, because we would expect plausibility to be of higher importance than accuracy anyways. But in a traditional PA optimisation perspective it is an important observation because it moderates the importance of many of the measurement and bias problems in PA. If accuracy is not necessary, lack of accuracy is less of a problem. And when accuracy is not important, the value of preparing thoroughly for PA meetings by taking detailed notes on an on-going basis becomes less meaningful. Of course, this again raises the question of what the whole purpose of the feedback part of the process is. Is it to improve organisational performance or is it to provide tough, honest manager feedback to individuals based on tangible examples? That these two purposes are not always congruent and maybe sometimes even conflicting, should be clear from the above examples. Through sensemaking, managers create stories that make their translation of the situation and their past action meaningful. And in fact, as examples showed, it is sometimes easier to create meaning from the PA feedback process without detailed examples. 8. Meaning Creation in Midyear Review 246 But if note taking was not a necessary ingredient for managers to create meaningful stories of what is going on at midyear review, why were also employees reluctant to use the diary? Would it not be to the advantage of the employee that all the good performance through the past half year was captured in the midyear review? One Staffs employee explained: I have to admit that I didn't even look at [the diary]. I didn't have a particular need for just another tool. And it is also due to time pressure. I have had more than enough to do. So a new tool which is not urgent doesn't have high priority against daily work tasks. And after 2-3 weeks, you forget all about it. In this quote we see how the employee reasons much along the same lines as the manager above, who also thought that preparing for feedback through note taking is a low priority task that will eventually get squeezed out by other, higher prioritised tasks. Another employee had also forgotten about the diary: R&D employee: "I completely forgot about it. We are so busy in our daily work so we don't really think much about PA" Being busy is one thing, but it is the second part of this employee's statement that is the more interesting in terms of understanding why taking notes aimed to support the PA process is one of those things that is hard to get done. If PA in general is not something this employee has much relation to in daily work, then the PA scheme has become detached from the daily work life of the employee and has as such lost importance for the employee. Things that are not important are easy to forget. A third employee who did not use the diary was even more direct about this interpretation: Staffs employee: "I don't find it important, not at all. I work the same way no matter if there is such a document [PA] or not. But it is important to my manager. That I was told... there have been occasions where I could have used the diary; where my manager has praised me. But I didn't do it. You are always in the middle of something. You run from meeting to meeting. I have been busy." Again, we see reference to being too busy to prioritise conducting on-going documentation of performance events as preparation for PA meetings. But the employee also offers an explanation to why preparing for PA meetings get lower priority than other tasks in a busy schedule; because the employee simply does not find the process itself particularly important. Here, we are back to the compliance fulfilment story we also saw above; if you find that PA is not doing any good to you, it becomes meaningful to reduce the resource efforts you put into the process to the 8. Meaning Creation in Midyear Review 247 absolute minimum, and that is to show up at meetings and sign off meeting conclusions. Above we have seen how managers I interviewed had experiences with employee carelessness in preparation of PA meetings. Although we cannot draw any statistical conclusions based on a qualitative study like the present, it would not be an unreasonable hypothesis that the managers I interviewed are not unique in this respect, although future studies are required to verify the generality of this. In the cases analysed above we saw how some employees were disengaged in PA because they did not find PA important or because they were too busy to spend time on preparing PA meetings. We also saw that when faced with employee carelessness, the managers I interviewed typically acquiesced and accepted a lower quality process and sometimes they also accepted to take over some of the PA process responsibilities that the employees should ideally carry. That employees can be careless and reluctant to spend even minimal efforts on preparing for PA meetings is an important observation because employee alienation and disengagement is threatening the quality of the process. As PA schemes get more connected to other HR processes like reward, promotion, talent identification, dismissal etc., it becomes a paradox that employees feel disengaged and alienated from the process to the degree that they do not find it worthwhile to be well prepared for PA meetings. This study is only a first step in understanding the nature of this paradox and it would be greatly beneficial to see future studies deepen our understanding of this phenomenon. 8.5 The Challenge of Disruption From the above analysis, it should be clear that although sensemaking processes are essentially social and happen in a complex interplay between actors (here, particularly managers and employees), individuals do play a crucial part. If the dyadic relation between a manager and an employee is of key importance to the meaning creation dynamics in PA, we should be interested in looking at what happens if one of these individuals is suddenly substituted by another in the middle of a PA cycle. This is what I call the challenge of disruption. 8. Meaning Creation in Midyear Review 248 As I explained in the methodology chapter, four employees experienced a manager change during the course of the study. For all, the change happened shortly before or after the midyear review. This relational disruption in the middle of the process is a challenge to continuity in mutual agreements and expectations set at the target setting stage. With a new manager, targets and performance may be translated differently. This generates changing conditions for meaning creation for the employee and the manager. Across the workforces, two distinctly different patterns of meaning creation were observed; one was about active change management where the manager operates as a mediator of PA information to her successor. The other was about downplaying the change through taking a manager role as a PA accountant where no mediation is required beyond the PA targets themselves. Disruption and the manager as mediator I asked a Staffs manager if she knew that she was on her way to a new position when she had the midyear meeting with the employee: Yes, and it was a bit funny. At that time I knew I was on my way. So I put a bit more effort into it than I normally would have done at midyear. Simply to document it all so that my successor will have something regarding these people when they reach year-end. The manager is on her way to a new position. Therefore, she will not be the one responsible for making a final rating of the employee when year-end is reached. Hence, the manager cannot count on her own memory about performance events and agreements made at target setting, because she will be somewhere else at that time. For hand-over purposes it becomes meaningful to make a more thorough documentation of status at midyear than the manager would normally do. This way, the manager proactively relates to the problem in PA raised by Raymark et al. (1999) and others, that PA evaluations are characterised by incomplete information. This is a problem, which would be exacerbated by a disruption in management relations. On the other hand, the Staffs employee did not perceive the midyear meeting preparation to be any different than usual: I am trying to be fairly detailed in my own comments anyways. So I don't think there's any need to write more. 8. Meaning Creation in Midyear Review 249 I asked a Production manager also on her way to a new position and who had a similar experience, why the upcoming organisational change meant that this midyear was more thorough than normal. Could it not be that thorough always? Yes it could. But traditionally, it is more about a status on the targets... So this time it was a bit more thorough and I thought that was really good. There are at least two interesting pieces of information in this quote. First, the Production manager tells that a normal midyear is "more about the status on the targets". We could call this the manager role of accountant in the feedback process, because the manager in that scenario is controlling the status on targets similar to the exercise accountants conduct when they audit the books of a company. But in this example, because it was the last session between the manager and the employee, and because the documented outcome should serve as input for the new manager, it was different. The discussion was better prepared and the themes discussed were broader than the usual strict target follow-up at Midyear. We could call this the manager role of mediator, because the manager here is concerned not only with following up on status on targets, but also on mediating detailed information to her successor through the PA system so that the transition to the new manager can be fair and smooth. The second interesting piece of information the above quote gives us is that the changed context and conduct of the midyear meeting is actually perceived to be a good experience by the manager. The more thorough preparation and documentation is perceived as being meaningful because it facilitates a better founded discussion of the performance of the employee and of the transition to the new manager. Disruption and the manager as accountant In Sales, the manager perception of the effect of an organisational change on the PA midyear meeting was different than the above. On the question whether the manager had a discussion with the new manager where they went through all the PAs of the employees affected by the change, the Sales manager replied: "I have not had midyear with X. Her new manager had that... You know, I did make some handover. But it simply became overwhelming for us if we should go double-up on midyear participation." Interviewer: "Is that possible because your targets are very 'hard' in Sales?" Sales Manager: "Yes, exactly." 8. Meaning Creation in Midyear Review 250 This story is different from the previous examples, where the old manager was conducting the midyear review. In this case, the manager change happened in July, but it was the new manager who conducted the midyear review with the employee in August. Because of workload, the old manager could not participate in the meeting. The workload argument is relating to the problem that PA requires significant resources, particularly from management (Coens and Jenkins, 2002). In this case, the problem is handled by letting the new manager take the midyear meetings by herself although she will not be able to base the feedback on any personal knowledge of the employees' targets or performance. The manager perceives it as a reasonable prioritisation because targets are so clearcut in Sales. The logic in this argument is that if you have a Sales target to meet index 100, and any manager can pull your numbers out of an IT system, any manager could also provide the feedback on whether or not you are on target. The Sales employee appeared to share this view: The targets are fixed, you know. They come from above - that's how it is. So we just relate to that as it is... I had put in my comments but basically it is just about looking up the numbers. We may say that the situation in Sales is an extreme case of the accountant role of a manager mentioned above. The advantage of favouring quantifiable targets is that a disruption in manager continuity becomes less of a problem. Just like a company can change its accountant without changing the valuation of its assets, the new manager can look up the numbers and provide feedback to the employee at midyear. However, does the manager role identity as an accountant not leave all the surrounding discussions about what went good or bad, the tangible feedback on behaviour in performance events, the discussion of influence of external factors etc. out of the equation so that midyear risk becoming a strict target status follow up without much opportunity to reflect on reasons and thus opportunities to improve? It could be in some cases, but in this case that conclusion may be too stretched. At least, that risk is moderated by the fact that the new manager in this case was also an experienced manager within Sales who shared an office with the old manager and who knew her new employees quite well already, although not from a manageremployee relationship. 8. Meaning Creation in Midyear Review 251 Discussion In Figure 8-2 below, the two different approaches are depicted in the meaning creation feedback model. The model illustrates how the manager as an accountant is a consequence of downplaying the disruptive element by enacting the perception that PA targets are self-explanatory so that further elaboration is not required, to be able to pass on the PA responsibility to a new manager. In this case, the disruption does not constitute a challenge to meaning schemes. However, with this approach there will be a loss of richness in the feedback the new manager will be able provide because she can only base it on the scarce information held in the PA targets. In the manager as a mediator example we see in the model the enacted assumption is that there is more to PA than what is stated in the targets. Therefore the disruption challenges meaning schemes because there is a fear that the manager change will bring a loss of this undocumented richness. The manager mitigates this risk through mediating additional information, but by doing this she at the same time undermines the perception of PA as being objective and transparent in the sense that there is a direct link between target realisation and performance feedback. Figure 8-2: Mediator or accountant - different handling of dyadic disruption The example from Sales with the manager as an accountant fits well with the categorisation of the task at hand in Sales as being the workforce with the best fit to the management by measurement scheme. When repetitiveness is high and complexity, specialisation and task newness is low, there is a good foundation for establishing a measuring regime to account for individual performance. When PA is only about the targets, and performance on the targets can be directly pulled from an 9. Meaning Creation in Final Appraisal 258 A story of non-causality Let us look at another example of a manager reflecting on the link between targets, performance and evaluation. I asked a Staffs manager if she knew the expression "you get what you measure", and if she thought that holds true for how the employee's incentive to innovate and perform outside the boundaries of her PA targets would be affected by the focus on documented targets installed by the PA process. The Staffs manager answered: I think that is true. But I also think it is about making people feel that they are not just evaluated on their targets but also on what they do outside their targets. Notice that this manager is the same who in one of the quotes above subscribed to a story about a strong link between targets, evaluation and rating. Now, challenged by the possible negative effect of such a rigid target/rating causality regime on innovation and proactivity beyond the targets, the manager moderates the story. Now, it is important that employees understand that they are not just evaluated on their targets but also on what they do outside their targets. I also asked employees about how they saw the connection between innovation and PA: R&D Employee: "I am thinking about it as two opposites. PA - that is where you are static and where you put everything into neat boxes. And then there really is no room for all the ideas you get the day after... And then again, you are not so motivated by the PA targets, because they have to be realistic. It is just a fact that it is easier to be innovative when you have very ambitious targets and have to think along unrealistic paths." So, there is awareness that a too strict causality based relationship between target realisation and final appraisal rating will have negative side effects and maybe even be a barrier to innovation. Although the story of causality is intuitively rational and nicely congruent with the management by measurement doctrine behind PA, reflections on downsides of PA challenges meaning schemes of what is required to secure high performance and employee motivation, and led to moderated stories that were required to create new meaning from this dilemma. This thesis/anti-thesis composition through reflection nicely demonstrates Weick's point about sensemaking being an on-going process where actors continue to establish and maintain stories based on what is most meaningful to themselves. If I was looking for the objective truth I could have asked which is true; is there a strict 9. Meaning Creation in Final Appraisal 259 causal relationship between business target realisation and rating, or is there not. But in a sensemaking perspective that question may not be relevant for anything but to provoke new reflection and possibly new stories. The point is, that there is no way we in a sensemaking perspective can say which of the two answers is most true. In fact, both answers may be true in the sense that sometimes there is a strong causality, but in other cases under other circumstances the manager moderates the causality. It depends. Changing the rules of the game Considering what we have seen so far, we would have expected to see Sales, being the workforce with the best fit to a management by measurement regime, as the area where the story of causality would be most dominating. Sales is also the workforce where we have so far found the most mathematical approach to target setting and performance measurement, and the workforce where we have seen most reluctance to incorporate non-business related targets in PA. Nevertheless, prior to interviewing for final appraisal in Sales, I had been looking at the PA documentation of a Sales employee and noted that most of her targets were evaluated at ME. Still, the final appraisal rating was EE. This made me curious how such a discrepancy could appear in a mathematically based management by measurement system. I asked the Sales manager how this could be: We have performed extremely well on sales of [product y] in the affiliate compared to anchor budget. But the sales targets to sales representatives have been very ambitious. She reached index 100,5. Normally, that would correspond to an ME. But as an affiliate, we have performed to an O on [product y], so we cannot have that all the sales representatives are NME or AE. Therefore, in cooperation with my manager, we have had to make a graduation so that she got an EE. In 2010, this Sales manager's department was in a peculiar situation. Product Y was a new product that was introduced with considerable success and hit the budget targets much faster than anticipated when the budget was finalised in Q3 2009. However, sales representative's targets are set quarterly based on the sales in past quarters and on forecasts. Since the sale of product Y in Q4 2009 had shown impressive growth rates, sales representatives got very ambitious targets for Q1 2010, which they could not meet. Even though this was somewhat compensated in the following quarter, the yearly average ended at the ME level. 9. Meaning Creation in Final Appraisal 260 The reader will recall from the analysis of sensemaking in the target setting process that in Sales there is a strict definition of rating from the realised index. A 100,5 index corresponds to an ME. It is a simple, causal, objective and transparent interrelation where neither the manager nor the employee need discuss the rating because it has been clearly defined in advance what it takes to get a given rating. However, in this case the model did not work. The manager and the organisation had to moderate the target in hindsight so that a 100,5 index realisation was not equal to an ME but to an EE. Following the principles established and communicated at target setting would be a serious challenge to the meaning scheme of managers and employees because the organisation, which is measured on its realisation of the budget, was evaluated to an O for its sale of product Y. How can then the people who actually realised the sale then only be ME or even lower? How can you create meaning from a story where the whole (the organisation) is successful while its parts (the sales representatives) are not, if you believe in a PA system that presupposes that organisational performance is the sum of individual's performance? Management recognised the issue and enacted the change in the environment (the surprisingly high adoption rate of product Y) and retrospectively redefined the link between target and evaluation so that a 100,5 now corresponded to an EE. The link between performance and evaluation has been improved although at the cost of a poorer link between target and evaluation. Thereby a new and more meaningful story is established; that the success of the organisation is still linked to the success of the individual. But one could still reflect on why it is that the organisation performs to an O, while one of the best performing sales representatives only performs to an EE. A more meaningful story may have been established, but it still contains potential challenges to sensemaking. While sales targets constituted 75% of target weight in PA in Sales, the remainder 25% was constituted by other elements. The most important of these are field days (the number of days a sales representative 'produces') and callplan compliance (the degree to which a sales representative realise targeted meetings with pre-defined customers). A Sales manager explained how these targets were evaluated at yearend: Field days and callplan compliance targets are fixed. They are set at year start and they are managed through the sales system and KPI reports. It is all approved by 9. Meaning Creation in Final Appraisal 261 European HQ at year start. They are fixed no matter what. No matter if you get ill or anything. Then you have an explanation for why you didn't achieve it, but the targets remain the same. In our previous analysis of managers' sensemaking in the target setting process we saw how the other three workforces seemed to have varying practices for how to adjust targets through the performance year as things unfolded and assumptions were challenged by the development of things. In contrast to this, we can see from the last two quotes how Sales practices two quite different logics to target evaluation; on the one hand sales targets that are hard numbers with pre-defined numeric success criteria which routinely get adjusted quarterly and where evaluation principally is automated based on index realisation. And on the other hand process targets that get set at year start and cannot be changed no matter how the environment changes or whatever happens through the year, but where again, target realisation is automatically calculated based on documented customer visits. In Sales, the retrospective change of the rules of the game (changing the ratio between sales index and performance evaluation) fills at least parts of the meaning challenging gap between the O rating at the organisational level and the much more modest ME rating of a high performing sales representative that would have been the case without the change. By changing the rules of the game in hindsight, a more meaningful story can be developed, but essentially at the cost of a key foundation in the PA management by measurement regime; the strict causal relation between predefined target criteria and post-performance evaluation. In that sense the Sales case is an interesting example of how the story of causality in Sales is turned into its opposite - a story of non-causality - by a deliberate management action to change the rules of the game in order to be able to maintain meaning. This is illustrated in Figure 9-1 below. 9. Meaning Creation in Final Appraisal 262 Figure 9-1: 'Changing the rules...' as meaning creation feedback process Figure 9-1 illustrates how the provision of transparent and objective ratings is the ideal in the management by measurement regime of PA. This is supported by a story of causality between business targets, individual performance and final appraisal ratings. However, in this case the result is challenging meaning schemes of managers and employees because dynamic changes in the environment have made the causal outcome unacceptable. New meaning is established by changing the rules of the game. However, thereby the provision of transparent and objective ratings is challenged, because there is no longer a clear link between targets, performance and rating. Although the task at hand in Sales has the best fit to the management by measurement doctrine in PA, also in Sales the outcome of this logic can be so challenging to meaning schemes that the logic must be changed to be able to maintain a meaningful outcome. 9.3 The Challenge of Evaluating Behaviour Above we have started our analysis of local adaption and meaning creation in final appraisal rating by focusing on how managers and employees established different stories to maintain meaning of the influence of business performance on rating provision. We also saw that one of the challenges to meaning schemes was the existence of moderators to the precedence of business performance. One particular moderator that all managers persistently mentioned as important was the impact of 9. Meaning Creation in Final Appraisal 263 behaviour. Below, I will look closer at meaning creation processes particularly regarding the inclusion of behavioural evaluation as a component in rating employees for final appraisal. Inclusion or exclusion - two opposing translations in Sales Let us start by looking at the situation in Sales at some length. As we have already seen, Sales is the workforce best fitted to a quantifiable, business target realisation focus. In a previous chapter, we saw how meaning creation through concealment characterised the handling of behaviour in the target setting process in Sales. So, for the performance year I investigated, behavioural targets were not used in Sales. Nevertheless interesting reflections and patterns around inclusion of behaviour in the final appraisal rating surfaced during the last interview round. Above I analysed meaning creation in business performance evaluation in an example in Sales where targets on a particular key product had been set unrealistically high and where none of the interviewed sales representatives had realised sales above the predefined ME level. For one sales representative the average index realisation on her key target product had been 99, corresponding to an AE rating. Although sales constitute 75% of the weight in PA, the employee had nevertheless received an overall rating of EE. With departure in this apparent mismatch I asked the manager how the employee could get an EE when she did not even realise her most important sales targets. The manager explained that this was because the employee had been rated very high on behaviour. Then I asked if Sales management had more generally used the behaviour component to compensate for the fact that they had set business targets too high, and the Sales manager replied: Yes, that's what we did in my opinion. But also, sometimes it is exactly where the targets are too ambitious that you really have to be professional and structured, rather than where the targets are reached more easily. The manager retrospectively establishes a story to create meaning out of why, despite the fact that the business performance according to the predefined criteria only was to an AE, the employee ends up with an EE. The two key points of the story are: a) that Sales management had set unrealistically high targets so that it was impossible for sales representatives to over-perform and therefore they must be compensated through prioritising the weight of business performance lower, and prioritising the weight of behaviour higher, than normally; and b) that it is exactly 9. Meaning Creation in Final Appraisal 264 when the targets are very tough that employees have the opportunity to show their full behavioural potential. Although a more accurate relation between target evaluation and final appraisal would have led to a lower rating, in this case accuracy is trumped by plausibility so that a meaningful translation of what happened can be maintained. The second argument about getting the better behaviour out of employees when targets are stretched could be seen as an argument for always setting unrealistically high targets. However, we know from the literature review that multiple studies would seriously question such a hypothesis. For example, within an expectancy theory framework, DeNisi and Pritchard (2006) argued that perceived inability of an individual to fulfil a required task, whether it is due to lack of capabilities or resources, will directly affect the employee's perception of the link between action and outcome and thus the employee's motivation and performance. So setting unrealistically high targets would according to these theories have a directly negative effect on the employee's motivation and performance. But, as we have now seen many times, meaning creation is not a rational game; it is based on what makes most meaning when actors look back at past actions of themselves and important others, and retrospectively establish plausible stories about what happened, and why it happened. The above example made me curious to find an example where the opposite had happened; i.e. that an employee had over-performed on business targets but was rated lower because targets had been too lenient, based on a behaviour oriented argumentation. A Sales manager explained: That can easily happen. I have one who got an AE. She actually over-performed [on business targets], but that doesn't correspond at all to the behaviour she has shown. Again, this sales manager uses the moderating effect of behaviour in a situation where an employee, despite an unacceptable effort, has been able to over-achieve on her business targets. Giving the employee the rating she should have had according to the predefined business success criteria would not be meaningful to the manager, because it would not correspond to the behaviour the employee had shown. It would have been accurate, in the sense that there would be a clear causal link between target evaluation criteria and actual criteria, but it would not have been plausible. To maintain meaning, the manager set performance criteria aside and 9. Meaning Creation in Final Appraisal 265 gave the AE rating because it more meaningfully represents the manager's translation of what really went on, despite business outcome. We have in a previous chapter seen how Production since 2009 has worked with centrally defined, cascaded behavioural targets. But we also saw that it was not until the 2012 process that this became common in the other workforces. Thus, in the 2011 cycle I followed, it was in all other workforces than Production mostly up to the individual management teams, or even the individual manager, to define what they understood to be good behaviour. A Sales manager defined what constitutes good behaviour this way: For example, this weekend Y was at a conference, and Monday morning there was a mail from her with a full summary of it... That is really good work. You have to remember that. It is objective and tangible. Through this example, the manager illustrates one possible way to show good behavioural performance outside what was agreed in the PA targets. The employee did not have a target in her PA that she should make sure to participate in the conference and deliver a summarising memo immediately after conference. Probably, there is no difference in added business value of making the memo Sunday night rather than Monday morning. It is not about realising PA targets or adding tangible business value. It is much more about attitude and demonstrating extraordinary engagement. By delivering the memo so that it is ready Monday morning the employee demonstrates that she is engaged and ready to pay sacrifices (sales representatives do not get overtime payment) for the benefit of sharing her experiences with her colleagues and the organisation. Employees also have a clear perception of how behaviour is more about the 'how' than the 'what' of performing, even without having tangible behavioural targets in PA as reference: Sales Employee: "I think it is about the way you realise your targets. This year, I had realised one of my targets to an O, but I didn't get it because I had done it in another way than they had preferred... If you looked at the result, it was an O, but my manager went in and evaluated differently." When senior management, as in Production in 2009 and in Sales from 2012, formulate a rating matrix that clearly establishes procedures for how to balance behaviour and business performance, and if this matrix is broadly communicated among managers, we would expect that it would become more plausible, more 9. Meaning Creation in Final Appraisal 266 transparent, more legitimate, easier to communicate and thus more meaningful to the managers, to incorporate a heavier weight to behavioural performance when establishing stories about how they provide final ratings. We would expect managers to enact this change to create new and more meaningful stories where behaviour takes a more prominent place in the process of establishing a final rating. However, contrary to this expectation, one Sales manager I interviewed kept underweighing the importance of behaviour even after the introduction of a rating matrix had clearly stated the principles for weighting the two dimensions: It has to be something very, very significant to change the rating of the targets. It is the business realisation that counts. So for the rating, it is not behaviour. In a way that would also be a bit unfair because people evaluate [behaviour] so differently. Then it is suddenly my manager who has to evaluate if my behaviour is good. And if I had another manager she would have evaluated it differently. I am in favour of making targets measurable. Behaviour is important to talk about, but the targets are much more important. It is the measurability that counts in our rating matrix. It takes a lot to change a business target performance [by behaviour]. And that is also how it ought to be. We see in this quote how the manager, despite the rating matrix which indicate a significant moderating effect of behaviour on final rating, still stick to what we could call the old story of behaviour being of minor importance compared to business performance. Sensemaking is not a fully top-down controllable process. Senior management can introduce a rating matrix but if cascaded PA targets are still almost exclusively about business outcome and a manager values transparency and objectivity highly, and at the same time feels that behavioural evaluation is subjective and highly dependent on the individual manager and even essentially unfair, then it takes more than communicating a rating matrix to make the manager see the alternative story as more meaningful. There may be a change in the environment that allows a different enactment of rating actions, but if the manager still extracts the same cues from the process as she did before the change, her story will remain unchanged and meaning creation will not be challenged to a degree requiring new stories. Discreet inclusion of behavioural evaluation in R&D and Staffs In R&D there was no rating matrix and no centrally determined and cascaded behavioural targets in the 2011 cycle. Nevertheless, as part of the target setting process some managers and employees may have agreed on ‘softer’ targets 9. Meaning Creation in Final Appraisal 267 regarding team contribution and other elements which link to attitude or broader behavioural performance expectations. However, with the lack of a central guidance on the balance between business performance and behavioural performance, also in R&D it is to a large extent up to the local level, or even to the individual manager, to establish a meaningful balance dependent on the manager's perception, what cues the manager extracts and how the manager enacts the environment. An R&D manager explained her view on the importance of behavioural evaluations in the provision of ratings: You could say that if only you realise your targets, in theory it doesn't matter how you do it. But you can do it by being dedicated and by showing lots of energy or you can do the minimum effort required. You know, the final rating is an evaluation of your performance and everything else. Your attitude and things like that. This manager is establishing a two-sided story where the target realisation on the one hand has prevalence so that “in theory it doesn’t matter how you do it”. But this is just theory – another thing is practice; how you do it does actually play a role. The final rating is not just an evaluation of your business performance, but in fact also of "everything else". Although behaviour played either no role at all or an inferior role in PA targets in the area in R&D I investigated, behaviour did play a role when this R&D manager provided final ratings of her employees. A Staffs manager explained how her own experience in her role as employee collided with her experience in her role as manager: I got a 3 this year and was specifically told that I complain too much. That is not something that is described in my targets. So that is a very explicit example of the subjectivity of my own PA. I think it is unfair - or, it depends. It depends on whether you can recognise it yourself... Personally, I found it unpleasant. But on the other hand, I do the same myself. Through the literature review we learned from Coens and Jenkins (2002) that in PA schemes feedback is manager-centric and from Poon (2004) that disappointing ratings have negative effects on employee motivation. Further, we know from Langan-Fox et al. (1998) that there are frequent issues with the link between actual effort and evaluation and we know from Justice Theory studies like Thurston and McNall (2010) that there similarly are frequent issues with perceived fairness in evaluation contexts. The manager feels she is unfairly punished for being critical, but in the final sentence in the quote, the manager at the same time recognises that she does the same thing herself to her employees. 9. Meaning Creation in Final Appraisal 274 The manager is of the perception that senior management believes that managers would, without some sort of coercive governance, rate their employees too high due to managers' tendency towards conflict avoidance. We already heard from Yariv (2006) that managers are reluctant to give negative feedback and from Coens and Jenkins (2002) that managers are generally lenient in evaluations because their preference is to avoid conflict, so it is not an unreasonable assumption that senior management has similar concerns. In fact, this anticipated manager lenience is one of the key reasons listed by a "forced ranking" advocate like Grote (2005) for why senior management should introduce such schemes in the first place; to force management to take tough but necessary decisions that they would otherwise be reluctant to take. This is what Coens and Jenkins (2002) have termed the introduction of "resource scarcity" of top ratings. Since the whole rationale behind the forced distribution is to ensure differentiation of employees, senior management will not be able to accept a general lack of differentiation due to managers' leniency. Therefore they, according to the story of the above quoted manager, install and enforce rating distribution policies. Roch et al. (2007) found that absolute rating schemes are generally perceived as fairer than relative rating schemes. So, one challenge to sensemaking of managers and employees working within a relative distribution paradigm will be to handle possible perceived unfairness. In Staffs, one of the managers I interviewed referred primarily to procedural fairness when reflecting on whether she thought employees had been treated fairly: No, not completely. It is not always transparent to them either; why they get one rating and not another. Some of it gets very subjective... I think we try to do it as objective as we can. But fair and totally objective and transparent - no, that's not how it is. Not at all. Management do their best to be objective, but it would not be meaningful for this manager to claim that the process is completely fair to the employees. This is both due to the inherent subjectivity of ratings which we have dealt with many times above. But it is also because the process is not transparent to the employees; lack of transparency seems in itself to be an element in lack of procedural fairness, because lack of transparency makes it impossible for employees to find out whether the outcome of the process (the rating) is provided in a fair way. 9. Meaning Creation in Final Appraisal 275 One element of fairness concerns whether people feel they are rated in accordance with the same standards across the organisation. One Staffs manager had a particular problem with the way this was practiced in her organisation: I don't think the distribution is relative when you look at VP level and above. So that is really a dilution of the whole concept. As we saw in a previous chapter, the distribution of ratings is direct proportionally connected to the hierarchical level one is working at, so that the higher one is in the hierarchy, the higher is the average rating. The practice of rating distribution that this manager refers to is one where there are different standards; Vice Presidents (VPs) and higher ranking levels are rated according to a standard where relativeness is non-existent - otherwise they could not all get 4s and 5s. In contrast, employees are rated relatively and a certain share of 3s and lower ratings are required and enforced by senior management. Another critique of relative distributions is that they install a more competitive and less collaborative mindset in employees because the failure of others has an indirect positive impact on the evaluation of oneself (Coens and Jenkins, 2002). I asked an HR staff who had been heavily involved in improving the PA process in an area if she had experienced any of that: We talked a lot about absolute versus relative. What is it really? We are somewhere in the middle. We don’t really clarify whether it is absolute or relative. And I think that does make a difference. The SVP showed our ideal distribution to our employees – what we had had and where we were going. And people asked: “does that mean in my department where we are five, one of them has to get a 2?” And [the SVP] said no; this is a distribution over a population of 200+ employees. So there could be three of you who all got a 4 in your department. In this area, the concept of an ideal distribution was introduced. Faced with this, managers were curious to understand how this impacted the requirements for the distribution in each of their own departments. An ideal distribution that has to be realised within a given area is by definition a relative rating scale because you have to compare people to make sure it is the right ones who get into each category. But as we see in the quote, senior management made it clear that the ideal distribution was meant to cover the entire organisation. It did not necessarily apply at the individual department level where there would be room for deviation from the ideal distribution. However, for the ideal distribution to be realised, any department deviating from the ideal distribution would have to be balanced by a counterweighing 9. Meaning Creation in Final Appraisal 276 deviation in other units within the area. This way, the relativeness was moved from the individual level of employees to also include the group level of units, which then again would affect the ratings at the individual level in the affected units. This iterative complexity of rating distribution in any unit deviating from the ideal distribution, or affected by other unit's deviation from the same, is probably what makes the HR staff think that the organisation is somewhere in between relative and absolute distribution frameworks. To enforce the realisation of the ideal distribution and to make sure managers make broad use of the full rating scale, this Staffs area in one year put pressure on managers for them to identify low performing employees who should be rated AE. An HR staff involved in the process explained how the message had been conveyed to managers: The message was; look at your people and see who the worst one is. He gets the 2. So, you know, [the managers] didn’t like that. But on the other side, they found those 2s. I coached some of the managers before giving 2s – for some of them it was hard because it was the first time. This enforcement of identification of low performers is related to what Grote (2005) calls "forced ranking", i.e. the idea that independent of, though often in parallel with, an appraisal rating distribution scheme, manager's should be able to list their employees on a ranking list in order of performance, so that they will be able to identify who is best and who is worst. The enforcement of identifying 2s was initiated by senior management who thought it would be unlikely that managers would have volunteered to move down their own people from the ratings they had originally proposed if senior management had not interfered. With the intensified focus on enforcing an ideal distribution and identifying low performers, there was some concern among middle managers: HR Staff: "One manager said 'this means to me that I can say to my employees: You know I wanted to give you an Exceeds but I can’t because this is the distribution we are working at.' [The Senior management representative] stood up and said: 'If you do that you are failing me as a manager.' So he got an opportunity to really express himself very clearly as a leader." It was very important to the senior management representative in this area to make it clear to middle management that they had a clear responsibility to take ownership of the ratings they provided within the ideal distribution scheme in the area. It would not be tolerated if managers blamed the distribution system for any ratings that may be 9. Meaning Creation in Final Appraisal 277 disappointing to the employees. The distribution was the playing field and the managers had to take personal responsibility for how they rated their own employees within the boundaries of that playing field even when the system forced them to provide lower ratings than they would have done without the system. Blaming the distribution scheme would be an easy way for managers to move any conflict with employees over ratings from the personal manager-employee relationship to the less personal employee-system level, and maybe even installing manager and employee to be on the same side against that system. However, it would also undermine the employees' trust in the system and would thus undermine the system itself. The concept of an ideal distribution rather than a forced distribution was intended to install a less rigid system leaving some room for managers' discretion. Nevertheless, a Staffs manager explained how she practically had understood her superiors' communication about the ideal rating distribution: Last year it was really bad. We were hit by this message that there could only be two in the area that could get a 5 and there should be two 2s in each department. They wanted at least 10% 2s. It is simply not true that 10% are 2s... Then you line up your employees and the ones at the bottom are defined as 2s. Personally, I don't like that. It is not particularly pleasant if those employees have actually reached their targets, but they lose when compared to the next in the line. Again, we see that this manager feels that the degrees of freedom in terms of possible deviation from the ideal distribution are limited by tangible requirements to avoid giving too many top ratings and to enforce identification of low performers. The manager implicitly refers to the method of forced ranking in her description of how she practically went about identifying the required 2s. So, although forced ranking is not formally introduced in PharmaComp, in this managers' department it for all practical purposes is. The manager does not like to have to communicate 2s to employees who "have actually reached their targets". But the manager loyally does it and she also stays loyal to senior management's requirement of taking personal responsibility and not blaming the system. To the employees in Production, Staffs and R&D, the communication about an absolute or relative use of the rating scale is not always perceived as very clear: R&D employee: I have understood that we get rated on an absolute scale. But in reality that is not the case, because in the end we are compared with each other. They want to cover the entire rating scale." 9. Meaning Creation in Final Appraisal 278 When it comes to what consequences employees see from the perceived relativeness of the ratings in terms of the incentive to cooperate, the same R&D employee explained: That the rating is relative works fine for the team players. But for those who are not, I am sure they think about it. [It means] you don't help each other as much, if that could influence your own rating. The relative element in it, I find that peculiar. Where I was working before [joining PharmaComp] they looked at your performance in relation to your own targets. Not in relation to how your colleagues had performed. A Staffs employee also had only a vague idea about the precise evaluation process, but imagined there was a relative distribution based on a mechanistic logic suited to the purpose of distributing rewards: I imagine it is a dictate from above. That 'this is how the [bonus] pool looks like' and then management has got to make the ratings fit into that. Somehow that is a hopeless way of doing it. It would fit on an average unit. But when you have a unit filled with smart people, it is hopeless... People who would get an Exceeds elsewhere, get a Meets here, since it has to fit a normal distribution. That is a bit demotivating. Senior management's initiative to improve the practice of PA in the Staffs area was running over a two-year period and involved many other elements than the introduction of an ideal distribution and enforcement of identification of low performers. It also included training sessions for managers and a closer look at the relation between performance rating and other rewards: HR Staff: "It got better in the years that I was there. We were doing it with bonuses too. It definitely was getting closer... We worked on this for two years while I was there and I think we made a lot of progress. It was huge. From the work group, the management team, the managers, training sessions, and sending articles out to the managers." The area was getting closer to an ideal distribution, as we have also seen was the case in other workforces where stronger enforcement of rating distribution has been implemented over recent years. But to get there, the middle managers had to be brought to the same page, and it is clear that in this area it took a lot of management resources to get there. How did managers and employees experience this intensified focus on fulfilling the ideal distribution? A Staffs manager explained: It is communicated that we do not have a forced distribution. But if you ask me, that's exactly what we have. If we have too many EEs we get challenged. And some [employees] will be moved down when it get reviewed higher in the system. This manager sees a misfit between the communicated flexibility of the system and the reality she experiences during the rating process. It is explicitly communicated 9. Meaning Creation in Final Appraisal 279 that there is not a forced distribution but in reality the manager experiences that positive deviations from the ideal distributions are challenged to a degree so that it in reality is experienced as a forced distribution. Thus, this manager perceives that there is a disconnect between what senior management communicates and what the manager perceives is practiced. Such a disconnect is a challenge to the meaning scheme for the manager who must maintain a meaningful story of what she is actually practising when providing ratings. As we would expect from a sensemaking perspective which is oriented towards interpreting actions, the extraction of cues is in this case clearly much more oriented towards actions (what do I see senior management practice?) than towards symbolic articulation (what do I hear senior management say?). Therefore, the manager's story is that no matter what is communicated, the organisation practices a forced distribution approach to PA rating and therefore the manager's role and responsibility is to keep as close as possible to the ideal distribution when evaluating her employees. Ratings as absolute distribution In contrast to the three other workforces, Sales is less focused on realising a specific rating distribution. I asked a Sales manager if she had to reach a certain distribution when rating her employees: No, we don't have that. You could say that there are quite a lot who get an ME. And because we update the sales targets [quarterly] and sales targets constitute 75% of their PA, it means that it is pretty difficult to get much over 100%. In the part of Sales I investigated, I found that rating calibration was less formal and with less emphasis on reaching a certain rating distribution. Senior management in Sales was apparently less concerned with rating distribution and did not enforce policies to install stricter governance of the same. We have already seen how business targets in Sales were very quantitative and how rating criteria were predefined61 and based on 1:1 meetings between a manager and her superior rather than on larger, cross-unit calibration meetings like in the other workforces. A likely reason for this difference in approach to rating scale would be that Sales is not basing its bonus allocation on PA ratings, but on a parallel bonus system. Thus, 61 However, we have also seen how behaviour and changing circumstances in some cases could moderate or even alter the predefined criteria. 9. Meaning Creation in Final Appraisal 280 in Sales there will not be any need for stronger differentiation of ratings to be able to produce the required differentiation for bonus allocation purposes. I asked a Sales employee if she knew whether her rating was made according to a relative system so that only a certain percentage could be rated EE or above: No, I don't think so... I believe my manager is good at evaluating us individually. She knows us well. In this area of Sales, it appears that not only is rating absolute rather than relative, the manager is also fairly good at making the rating match the perception of the quoted employee, i.e. there is a perceived correspondence between expected outcome (relation between targets and performance) and actual outcome (overall rating). It is exactly this relation that is at risk in a relative rating scenario. One of the Sales managers I interviewed was also quite explicit about the different effects of distributive and procedural fairness in rating provision in her department: I think they perceive [their evaluation] as fair. I think they are happy about the comments I write. Particularly in a year like this where they didn't think the targets they got were fair. According to the analysis of this manager, her employees generally perceive the distribution of ratings as fair. However, the manager is aware that 2011 was a year where the sales representatives in her department had very challenging targets, as we have previously seen, to a degree that was unfair to the employees because it made it close to impossible to realise the targets. The unfairness of the targets became even more of a procedural issue to fairness because, as we have seen, the management team was evaluated against the much lower budgeted sales in the anchor budget and therefore received high ratings and bonus: Sales Manager: "This year we had the challenge that the management and backoffice team went home on index 300 and got a large bonus because we were measured against anchor budget. While the sales team didn't get bonus - that's not fair." In this manager's perception procedural fairness is clearly challenged by the fact that employees are evaluated against other targets than the management team is. In a sensemaking perspective, we can formulate it like this: It is difficult to establish meaningful stories from a situation where procedural fairness perception is undermined by significant and systematic variations in targets so that some (management) can significantly overperform while others (sales representatives) at 9. Meaning Creation in Final Appraisal 281 the same time underperform. However, here the problem was the targets, not the ratings. The rating itself is absolute and perceived both by managers and employees as fair when held up against the documented targets. However, in some cases Sales employees were less impressed with the fairness of the system. These were what we could call "extreme" cases where external factors heavily influenced target realisation but evaluation did not take any account of that. In those cases, distributive fairness perceptions appear to sometimes be negatively affected by the strong causality of an absolute rating regime. For example, one high performing Sales employee had really struggled to meet a target that had been severely out of touch with what was in fact achievable: I think it is difficult to make a fair system... Whether it is fair, I cannot find out. It is so aggravating to achieve 99% on something you have worked on like crazy. That feels unfair. This employee understands why she gets a low rating although she has performed best in class among her peers on a very stretched target. In a target-rating causality perspective the rating may be rationally fair, but considering that the targets lack realism and relation to the relative performance of peers, it "feels unfair". However, in an absolute rating regime, relative performance is not considered. Another employee in Sales had spent the first two months in a year helping a new colleague by handing over customers in an area previous managed by the employee. Although this was done in full agreement with the employee's manager, the effect was that her callplan targets for number of customer visits in her own area could not be met. As previously explained, Sales callplan targets are defined at European level, and there is no process for accommodating for changed priorities in the shape of updated callplan targets. Sales Employee: "I just realised that this means that when we reach 1 March I will have missed 2 months in my own callplan... You know, I could have pulled out and let [the new colleague] take care of herself to be able to concentrate on my own callplan. But you don't do that. Not to a new colleague and not to the business. At some stage it may start to sound like bad excuses, but it is how reality is. These systems try to capture reality but that is just so very difficult sometimes." In this example the assistance to a new colleague is delivered at the cost of the ability to meet own targets. Again, in a relative rating scenario there is the opportunity for management to say that the Sales representative's performance, relative to her colleagues, is better, because despite the lower delivery on callplan she also 9. Meaning Creation in Final Appraisal 282 managed to help her new colleague get established. However, in a strictly absolute rating regime, the relative effort compared to colleagues who did not do the extra handover work, is not taken into account. Discussion Below, the two different approaches to rating are depicted in Figure 9-2. Figure 9-2: Relative or absolute - two different approaches to rating The figure illustrates how in Sales, the practical use of an absolute rating scale disconnects the individual rating from the performance of the ratee's colleagues. This is possible partly because the bonus allocation is disconnected from the PA rating process so that there would not be any consequences to bonus allocation of an uneven distribution of PA ratings. In itself, the missing link to reward is a challenge to meaning schemes because it tends to make sales representatives indifferent to PA. On the other hand, it is possible to create meaning around the fact that a fairly direct link can be maintained between target realisation and PA rating. So, in that sense and in a fairness perspective, the absolute rating scale model has an advantage. In this case unfortunately at the cost of the employee's perception of the relevance of the PA scheme. The three other workforces take the other route through the meaning creation feedback model. Here, the rating process is clearly relative so that ratings are not just dependent on one's performance, but also on the relative performance of others. This 9. Meaning Creation in Final Appraisal 283 challenges meaning schemes of managers and employees in terms of fairness perceptions and perceived effects on incentives to cooperation. In this case, meaning is created around process compliance (management says we must do it like this) and practical requirements for relative talent differentiation to be able to distribute bonus pools in a rationed manner. In the feedback phase, this path articulates a number of PA issues like lack of perceived fairness, lack of transparency in the shape of clear connect between target realisation and final appraisal, and issues with incentives to cooperate and help one's colleagues. The difference in the use of rating regime across workforces is a very significant example of how a standard PA system is locally adapted. Two very different regimes, an absolute and a relative rating process, are used in the same organisation and within a single PA scheme. The use of the absolute rating scheme in Sales matches the fit between the task at hand in Sales and the management by measurement paradigm in this workforce. Being the workforce most inclined towards quantifiable, unquestionable hard facts as the basis of targets and target evaluations, Sales will be able to maintain the clearest link between target, performance and rating within an absolute scheme, because here ratings will not be 'polluted' by the moderating effect of the performance of peers. On the other hand, in the three other workforces where the fit to management by measurement was less evident, there is a weaker case for an absolute rating regime. Further, and very importantly, the missing link to bonus allocation in Sales make it practically possible to maintain an absolute rating regime, while the opposite is the case in the three other workforces, where differentiation of employees is required to be able to distribute reward. 9.5 The Challenge of Calibration Calibration is the process of aligning rating practices across units within a given area of the organisation. In a previous chapter we saw how calibration in 2011 was one out of five prioritised "global minimum benchmarks" that all units in PharmaComp should follow. The purpose of calibration in PharmaComp is at least twofold. First, calibration should support management teams in aligning their rating criteria so that all employees within the area to a reasonable degree are rated against comparable 9. Meaning Creation in Final Appraisal 290 from the predefined targets and the manager will have a hard time getting acceptance of that. Figure 9-3 below illustrates the 2011 calibration process in the area of R&D I investigated. Figure 9-3: The calibration process in R&D No matter if calibration is before or after the employees' self-appraisal and the final appraisal meeting between manager and employee, at some stage the manager has to convey the message to the employee. Since we have already seen how one purpose of the calibration process is to change ratings so that a certain rating distribution is achieved, there will be cases where managers have to convey a rating which is different than what they had suggested before calibration. As we would expect, particularly the communication of disappointing ratings is a challenge to managers: R&D Manager: "You could argue that you could write comments that described really well how good it was to get a 3. You could do that but the issue is that when people get that 3 they stop listening to what else is said. Then it doesn't matter any longer. That message is not taken in by the employee and that is really a shame." Thanks to normal distribution governance and the calibration process, most employees will get a 3. 62 We have previously seen how senior management and HR struggle to persuade employees that 3 is a good rating. Nevertheless, in the experience of the above R&D manager, from the moment the manager communicates a 3-rating, the employee stops listening to what the manager says because of sheer disappointment. Communicating the rating becomes a barrier to conveying the attached story that was supposed to provide meaning to the employee. From a sensemaking perspective this is as we would expect employees to act because people generally make sense of actions (the manager's provision of a 3) and start immediately to extract cues from that. If they perceive 3 to be a 62 A 3 corresponds to an ME on the five point rating scale. [Document text truncated for crawler view.]