scieee AI-readable full text Open interactive document viewer

Impact of audit assurance on the quality of sustainability reporting

Grommes, Alexander

Abstract

EconStor is a publication server for scholarly economic literature, provided as a non-commercial public service by the ZBW.

Full text

Grommes, Alexander Article Impact of audit assurance on the quality of sustainability reporting Junior Management Science (JUMS) Provided in Cooperation with: Junior Management Science e. V. Suggested Citation: Grommes, Alexander (2025) : Impact of audit assurance on the quality of sustainability reporting, Junior Management Science (JUMS), ISSN 2942-1861, Junior Management Science e. V., Planegg, Vol. 10, Iss. 1, pp. 201-235, https://doi.org/10.5282/jums/v10i1pp201-235 This Version is available at: https://hdl.handle.net/10419/313865 Standard-Nutzungsbedingungen: Die Dokumente auf EconStor dürfen zu eigenen wissenschaftlichen Zwecken und zum Privatgebrauch gespeichert und kopiert werden. Sie dürfen die Dokumente nicht für öffentliche oder kommerzielle Zwecke vervielfältigen, öffentlich ausstellen, öffentlich zugänglich machen, vertreiben oder anderweitig nutzen. Sofern die Verfasser die Dokumente unter Open-Content-Lizenzen (insbesondere CC-Lizenzen) zur Verfügung gestellt haben sollten, gelten abweichend von diesen Nutzungsbedingungen die in der dort genannten Lizenz gewährten Nutzungsrechte. Terms of use: Documents in EconStor may be saved and copied for your personal and scholarly purposes. You are not to copy documents for public or commercial purposes, to exhibit the documents publicly, to make them publicly available on the internet, or to distribute or otherwise use the documents in public. If the documents have been made available under an Open Content Licence (especially Creative Commons Licences), you may exercise further usage rights as specified in the indicated licence. https://creativecommons.org/licenses/by/4.0/ Junior Management Science 10(1) (2025) 201-235 Junior Management Science www.jums.academy ISSN: 2942-1861 Editor: DOMINIK VAN AAKEN Advisory Editorial Board: FREDERIK AHLEMANN JAN-PHILIPP AHRENS THOMAS BAHLINGER MARKUS BECKMANN CHRISTOPH BODE SULEIKA BORT ROLF BRÜHL KATRIN BURMEISTER-LAMP CATHERINE CLEOPHAS NILS CRASSELT BENEDIKT DOWNAR RALF ELSAS KERSTIN FEHRE MATTHIAS FINK DAVID FLORYSIAK GUNTHER FRIEDL MARTIN FRIESL FRANZ FUERST WOLFGANG GÜTTEL NINA KATRIN HANSEN ANNE KATARINA HEIDER CHRISTIAN HOFMANN SVEN HÖRNER KATJA HUTTER LUTZ JOHANNING STEPHAN KAISER NADINE KAMMERLANDER ALFRED KIESER NATALIA KLIEWER DODO ZU KNYPHAUSEN-AUFSESS SABINE T. KÖSZEGI ARJAN KOZICA CHRISTIAN KOZIOL MARTIN KREEB TOBIAS KRETSCHMER WERNER KUNZ HANS-ULRICH KÜPPER MICHAEL MEYER JÜRGEN MÜHLBACHER GORDON MÜLLER-SEITZ J. PETER MURMANN ANDREAS OSTERMAIER BURKHARD PEDELL ARTHUR POSCH MARCEL PROKOPCZUK TANJA RABL SASCHA RAITHEL NICOLE RATZINGER-SAKEL ASTRID REICHEL KATJA ROST THOMAS RUSSACK FLORIAN SAHLING MARKO SARSTEDT ANDREAS G. SCHERER STEFAN SCHMID UTE SCHMIEL CHRISTIAN SCHMITZ MARTIN SCHNEIDER MARKUS SCHOLZ LARS SCHWEIZER DAVID SEIDL THORSTEN SELLHORN STEFAN SEURING VIOLETTA SPLITTER ANDREAS SUCHANEK TILL TALAULICAR ANN TANK ORESTIS TERZIDIS ANJA TUSCHKE MATTHIAS UHL CHRISTINE VALLASTER PATRICK VELTE CHRISTIAN VÖGTLIN STEPHAN WAGNER BARBARA E. WEISSENBERGER ISABELL M. WELPE HANNES WINNER THOMAS WRONA THOMAS ZWICK Volume 10, Issue 1, March 2025 JUNIOR MANAGEMENT SCIENCE Julian Anton Meyer, Success Factors and Development Areas for the Implementation of Generative AI in Companies Vincent Alberth-Jan Cremer, Diversity Within Top Management Teams: The Effects of Diversity Within Boards Towards Managerial Attention on Digital Transformation Veronika Timoschenko, Good as Gold or Merely Glitter? Elite Board Members' Impact on Firm Performance Cornelia Kees, Looking Behind the Fading Feminist Façade of #Girlboss Eliza Alena Marie Weitzel, Copreneurial Couples in Startups: A Comprehensive Analysis of Copreneurial Couples in Startups Compared to Classical Businesses Antonia Cichocki, Employment with Autism: What Are Educational and Adaptive Needs of Employers in Austria from the Perspective of Women with LowSymptom Autism? Paula Bao Quiero, "Well, Now They Know": How Mental Illness Identity Management Strategies Influence Leaders' Responses Alexander Grommes, Impact of Audit Assurance on the Quality of Sustainability Reporting Philipp Rittgen,Understanding the Effect of Hedge Fund Activism on the Target Firm -A Qualitative Study on Shareholder Value Gabriel Benedikt Thomas Adams, Energy-Aware Production Planning with Renewable Energy Generation Considering Combined Batteryand Hydrogen-Based Energy Storage Systems 1 24 44 70 95 135 176 201 236 267 Published by Junior Management Science e.V. This is an Open Access article distributed under the terms of the CC-BY-4.0 (Attribution 4.0 International). Open Access funding provided by ZBW. ISSN: 2942-1861 Impact of Audit Assurance on the Quality of Sustainability Reporting Alexander Grommes Catholic University of Eichstätt-Ingolstadt Abstract The subject of sustainability reporting is becoming increasingly important. In consequence of the implementation of the Corporate Sustainability Reporting Directive, a substantial number of companies will be required to have their sustainability reports audited beginning from financial year 2024. This paper examines the influence of external assurance on the quality of those sustainability reports. Therefore, the reports of all DAX and MDAX companies for financial year 2022 are examined using a novel textual analysis approach, to determine the individual report quality. The results demonstrate that there is no statistically significant relationship between assurance level and the quality of sustainability reports. Conversely, it was found that companies that are acting sustainable disclose a higher quantity of information and are more likely to demand voluntary assurance of their reports. These findings offer insights into the implications of assurance on sustainability reporting. Furthermore, the detailed overview of traditional and state-of-the-art textual analysis methods offers researchers a valuable resource for identifying the most appropriate methods to address their individual research questions. Keywords: audit assurance; CSRD; natural language processing; sustainability reporting; textual analysis 1. Introduction 1.1. Motivation Sustainability has been a topic of interest in business and academic research for some time, but now more than ever. The number of companies reporting on sustainability-related issues is growing rapidly, as is the number of scientific publications (e.g. Amel-Zadeh and Serafeim, 2018, p. 87, Lucarelli et al., 2020, p. 5, Guidry and Patten, 2012, p. 81). This is due to an intrinsic interest in sustainability on the part of companies and their stakeholders (Tworzydło et al., 2022, p. 144), but also to regulatory requirements that have been newly imposed and increasingly refined in recent years (H. Christensen et al., 2021, pp. 1178–1179). The Global Reporting Initiative (GRI) reporting framework has emerged as the leading standard for sustainability reporting. As an autonomous entity, the GRI has created guidelines with input from stakeholders across the board, fostering a reliable framework for reporting. Companies are not required to adhere to these guidelines by national lawmakers. Rather, they serve as a common ground for reporting. If adopted, the GRI standards enable standardized reporting and facilitate comparison between companies, regardless of their size, sector, or country of operation (Christofi et al., 2012, pp. 163–164). However, there is more than just voluntary guidelines. In fact, the Directive 2014/95/EU created by the European Union (EU) requires companies to report non-financial information. As a result, public interest entities with over 500 employees must comply with the Non-Financial Reporting Directive (NFRD). Starting in the 2017 fiscal year, these companies are required to disclose information regarding environmental, social, and employee-related matters in their management reports. The purpose of this requirement is to provide stakeholders with a clear understanding of the current state of development and position of companies in these areas (European Union (EU), 2014, pp. 4–5, 8). Not long after the NFRD took effect, the EU revised its sustainability reporting guidelines through the implementation of Directive 2022/2464/EU, also known as the Corporate Sustainability Reporting Directive (CSRD). This directive was introduced to address significant deficiencies in the preDOI: https://doi.org/10.5282/jums/v10i1pp201-235 © The Author(s) 2025. Published by Junior Management Science. This is an Open Access article distributed under the terms of the CC-BY-4.0 (Attribution 4.0 International). Open Access funding provided by ZBW. A. Grommes /Junior Management Science 10(1) (2025) 201-235202 vious requirements, which lacked sufficient depth and scope, and to consider issues such as data comparability and reliability. However, one of the main drivers for change is the limited number of reporting companies. The CSRD will make sustainability reporting mandatory not only for public interest entities but also small and medium-sized companies in the future (European Union (EU), 2022, pp. 19–20). In addition to the new reporting requirements and increased coverage, the CSRD mandates an audit process. Specifically, companies are required to undergo a limited assurance review of their sustainability reports by an external auditor. Under the NFRD, auditors only had to confirm that the related information was published at all. Furthermore, Member States had the option to impose a substantive audit requirement at the national level. The EU aims to establish a consistent link between financial reporting and sustainability reporting by requiring a substantive audit of sustainability reporting conducted by an external auditor, as financial reporting is already subject to a statutory audit. The Commission also reserves the option to take a decision by 2028 to adjust the assurance level from limited assurance to reasonable assurance (European Union (EU), 2022, pp. 34– 35; Velte, 2023, p. 4). The mandatory implementation of sustainability report auditing has the potential to aid the EU Commission’s objectives and enhance the general compliance of sustainability reports with regulatory requirements. There may be more benefits to consider, but the mandatory audit could also create an additional burden. Similarly, elevating the level of assurance from limited to reasonable could either positively impact reporting or cause unnecessary expenses. Sustainability reporting is not limited to the European area. Of the world’s 250 largest companies by revenue (G250), 96 % disclose sustainability information in the form of reports (KPMG, 2022, p. 13). Only 38 of the G250 are companies from EU members. China represents 30 % of the G250 companies and is showing a positive trend towards reporting on sustainability (KPMG, 2022, p. 18). The majority of the remaining non-EU G250 are located in the United States (69), Japan (26) and the United Kingdom (9) (KPMG, 2022, p. 75). All of these states have extensive, but varying, reporting requirements. On a global level, the International Financial Reporting Standards (IFRS) Foundation addresses the issue of sustainability through the implementation of new standards. In this manner, the standards are designed to meet the needs of the stakeholders of reporting companies, such as customers, employees and investors or the natural environment, which can be considered a stakeholder itself (Technical Readiness Working Group (IFRS Foundation), 2021). Two first two exposure drafts have already been issued. IFRS S1General Requirements for Disclosure of Sustainability-related Financial Information is intended to provide general requirements for the disclosure of sustainability-related financial information. IFRS S2 Climate-related Disclosures covers the disclosure of climaterelated risks and opportunities. Both drafts relate to information that is meaningful to the cash flows of companies and thus to the valuation of those companies (International Sustainability Standards Board, 2022a,2022b). With different accounting standards, some mandatory, some voluntary, some national, some international, and with different scope and materiality levels, sustainability reporting is highly diverse. Attempts for homogenization are confronted with ongoing substantive and regulatory dynamics. The implementation of the CSRD regulations could be a crucial step towards improving and harmonizing the reporting landscape. 1.2. Problem definition and objective While the NFRD is currently in effect, the CSRD will find application for the first companies as early as 2024 which means it will impact the fiscal year of 2023 (European Union (EU), 2022, p. 77). It is likely that the application of the Directive will not fully achieve the desired results at first. Similarly, even after the implementation of the NFRD, there remained potential for further improvement (Busco et al., 2022, p. 95), which is one of the reasons why the CSRD was created. In the context of identified weaknesses of the NFRD, the EU Commission directly mentions the role of the audit of reporting and that this should ensure the reliability of the reports (European Union (EU), 2022, p. 19). Additionally, external audits may also increase compliance with the CSRD and other regulations. On the other hand, there are expenses associated with the engagement of audit firms. The increased scope of the audit beyond the financial reporting has a direct impact on the total audit costs (Zaman et al., 2011, p. 190). At the same time, a mandatory auditing requirement does not guarantee audit quality. Previous studies have shown a variety of weaknesses that can occur in the area of auditing (B. Christensen et al., 2016, p. 1671). The purpose of this paper is to investigate the extent to which the audit of sustainability reports can improve the quality of these reports, where quality is primarily expressed in terms of the reports’ compliance with regulatory requirements. Non-financial reporting is heterogeneous and thus provides numerous opportunities for academic research. Due to its actuality, the domain still has some research gaps that can be closed. Since non-financial reporting essentially consists of qualitative reporting in text form, textual analysis methods are particularly helpful in filling these research gaps. This thesis makes several contributions. First, it contributes to the literature in the field of auditing, specifically the auditing of non-financial reporting. This area of auditing, while not entirely new, is considerably less investigated than the area of financial reporting. Second, the thesis also contributes to the literature on European financial reporting requirements. In particular, it connects those two streams of literature. Third, it offers a methodological contribution by providing an upto-date review of textual analysis methods. Finally, a contribution is made by providing evidence on whether auditing improves the quality of non-financial reporting. Stakeholders and other recipients of non-financial reports can assess A. Grommes /Junior Management Science 10(1) (2025) 201-235 203 the value that auditing provides when making investment decisions. The findings can also support the EU Commission’s decision on future assurance level increases. 1.3. Procedure of the work This thesis is organized as follows. Chapter 2 presents the relevant theoretical background. First, it consists of the regulatory framework with a focus on European accounting, in particular the EU taxonomy. Second, the established theories of sustainability reporting and auditing in general are presented. The theoretical application of textual analysis, which is the central instrument of this thesis, concludes the chapter on the theoretical background. The following chapters form the first of the two main parts of this thesis. Chapter 3 discusses the relevant literature in the domains of finance and accounting, while Chapter 4 presents the various methods of textual analysis in terms of their functionality and applicability. These methods are not only distinguished according to their field of application, but also between traditional methods and state-of-theart methods, which utilize the recent technical developments in machine learning and artificial intelligence. On the one hand, the comprehensive presentation of all currently available methods forms the necessary groundwork for the second part of this thesis. On the other hand, it offers a contribution in itself, since it can assist the audience of this thesis in identifying appropriate methods for own research projects in the field of textual analysis. The second part of the thesis involves utilizing textual analysis techniques to assess corporate non-financial reporting. For this purpose, Chapter 5 formulates three related hypotheses. Chapter 6 describes the methodology by discussing the data set, the research design, and the processing of relevant variables trough textual analysis. The results of the analyses are presented in Chapter 7. Finally, Chapter 8 concludes the thesis and discusses the limitations of the work. 2. Theoretical background 2.1. Development of the regulatory framework The climate crisis is one of the greatest challenges of our time. Its negative effects are already being experienced today and will get worse as they become more difficult to mitigate in the future (United Nations, 2022). The majority of United Nations member states are committed to addressing the climate crisis through the Paris Agreement, which aims to limit the increase in global temperature to a maximum of two degrees Celsius above pre-industrial levels, raise overall adaptation capabilities to the impacts of climate change, and shift capital flows to a climate-friendly development (European Union (EU), 2016, p. 5). Europe is contributing through the European Green Deal. This framework includes a program of measures for the necessary transformation. The European Commission has set the goal of making Europe the first continent to become climate neutral by 2050 by reducing greenhouse gas emissions to net zero. The interim target is a reduction of emissions by 55 % until 2030 compared to 1990 levels. The European Green Deal also covers issues such as the sustainable use of consumer goods, with specifications for producers to enable consumers to repair products more easily and make them last longer so that the goods do not have to be replaced. Other parts of the program cover the fields of technology, mobility, food, energy and biodiversity, and various other subjects (European Commission, 2019). Corporate governance is also specifically addressed in the European Green Deal. Companies are still too focused on short-term financial performance rather than sustainable development. Therefore, companies must increasingly disclose their information on sustainability-related issues alongside their annual reporting in order to inform investors about their development in these areas (European Commission, 2019, p. 17). The EU taxonomy is part of the European Green Deal and is designed to accompany and support the transition of the environment to the target state. The taxonomy introduces various instruments to achieve this goal and is also supposed to support the financing of the transition by directing capital flows in a way that is conducive to the transition. Another integral part of the EU taxonomy is corporate disclosure (European Commission, 2020, p. 8). Regulations require two groups of companies in particular to address this issue: Financial market participants1offering financial products within the EU and companies meeting the size criteria of the NFRD. The required content, especially for the second group of companies, is discussed in more detail in Chapter 6.2 of this thesis. The information must be published either in the non-financial section of the consolidated or annual financial statements or as a stand-alone non-financial reporting or sustainability reporting (European Commission, 2020, p. 27). The EU taxonomy encourages, but does not require, companies to obtain assurance from external auditors (European Commission, 2020, p. 37). Even before the introduction of the EU taxonomy, many researchers addressed these substantive issues (Lucarelli et al., 2020, p. 6). More recent research identifies the benefits of the taxonomy mainly in the area of harmonization and investment decision support (Dumrose et al., 2022, p. 2), which is consistent with the reporting objectives of the IFRS. It is also important to consider the full scope of the taxonomy. Beyond the entities directly impacted by the EU taxonomy, other entities are also indirectly affected (Dusík & Bond, 2022, p. 92). Suppliers and customers which do not meet the thresholds for mandatory NFRD reporting do not have to collect environmental data for themselves, but may need to be able to provide it to companies covered by the NFRD for their reporting. In principle, the requirement for more comprehensive reporting also leads to a reduction in information asymmetries. Although this relationship exists in theory, it should not be blindly assumed without evidence and needs to be further 1e.g. Equity funds, exchange-traded funds, real estate funds, pension schemes, venture capital and private equity funds. A. Grommes /Junior Management Science 10(1) (2025) 201-235204 investigated (Breijer and Orij, 2022, p. 350; H. Christensen et al., 2021, p. 1231). The EU taxonomy has one particular strength. By prescribing the narrative that sustainable activities can only be considered as sustainable if they do not harm other sustainable activities, trade-offs between different areas of development cannot be used as a loophole. The benefits of the EU taxonomy presented in academic literature may also extend beyond the European area. The EU serves as a prominent model for implementing regulations, making it probable that legislators outside of Europe will adopt these or raise similar requirements (Bloomberg, 2021; Dusík and Bond, 2022, pp. 93, 96). 2.2. Fundamental theories on sustainability reporting Before dealing with the methodology, the theoretical principles need to be defined. Essentially, the publication of company data is crucial for capital market participants and their investment decisions, with information content and timeliness being particularly important (Ball & Brown, 1968, p. 176). For this thesis, theories that consider voluntary disclosure are most relevant. For a long time, sustainability reporting within non-financial reporting has been on a largely voluntary basis, and companies only disclosed data when the benefits exceeded the costs associated with the disclosure. The NFRD made sustainability reporting mandatory for some companies, but the regulation still allows a great amount of flexibility in the nature and extent of disclosure, which is why voluntary disclosure theories are especially relevant. Voluntary disclosure refers to a company’s decision to publish supplementary information beyond what is required by law. There are various determinants that influence whether and how much voluntary disclosure is made, including firm characteristics, ownership structure, or countryspecific factors (Zamil et al., 2023, pp. 232–235). For the purposes of this thesis, however, the general theories, on which voluntary disclosure is based, are crucial. Agency theory, which is most often applied in the context of voluntary disclosure (Zamil et al. 2002: 239), is closely linked to the well-known principal-agent problem from economics, which is primarily founded on information asymmetries between parties (Arrow, 1963, p. 967). According to agency theory, firms voluntarily disclose information in order to reduce information asymmetries between themselves and their stakeholders and thus facilitate business relationships or capital flows. In addition to the agency theory, the next two theories most commonly used in this context are legitimacy theory and stakeholder theory. Legitimacy theory is concerned with the interaction between companies and their social environment and postulates that companies strive to shape their actions, decisions, and practices so that they are viewed as legitimate and acceptable by the society. Through voluntary disclosure, companies seek to achieve the necessary legitimacy and gain the trust of stakeholders. The theory is founded on the premise that there is a social contract between companies and society. Recently, increased awareness of Corporate Social Responsibility (CSR) concerns has influenced corporate practices in sustainability reporting, and companies have used CSR disclosure to gain legitimacy (Lepore & Pisano, 2023, pp. 56–57). Stakeholder theory, on the other hand, emphasizes that companies are not only beholden to the interests of their owners, but should also take into account the interests of a broader group of stakeholders who are affected by the company’s activities. This theory emphasizes that companies should recognize the expectations, values, and needs of their various stakeholders. Therefore, voluntary disclosure can be seen as an effort to increase transparency and address stakeholder interests and concerns. In addition, there are several other theories on the basis of which voluntary disclosures can be useful for companies (Zamil et al., 2023, p. 239). These theories explain different incentives for companies to voluntarily disclose information. Furthermore, the voluntary disclosure theory considers the costs of disclosure and suggests that information will be voluntarily provided only if the benefits for the company outweigh the costs of disclosure. According to this principle, information that is insignificant or disadvantageous to a company will not be disclosed (Verrecchia, 1983, pp. 179, 192). 2.3. Fundamental theories on audit In addition to voluntary disclosure theories, the principal theories of auditing are particularly relevant for this thesis. Auditing is one of the central areas of accounting. The external verification of financial or non-financial information by an auditor can ensure the reliability of reporting for external stakeholders. In very simplified terms, this is achieved by the audit firm determining the actual financial position and performance of a company through various audit procedures and comparing these with the figures reported in the financial statements (Wagenhofer & Ewert, 2015, pp. 410–411). The exact procedure and structure of the audit is not the focus of this thesis. Instead, the basic theories and related concerns are addressed in order to understand how they relate to the audit of sustainability reporting. Again, the principal-agent theory is a fundamental theory with significant importance. This behavioral theory can be applied to audit firms and their client companies. The principal, representing the company to be audited, hires an agent, the audit firm, to perform an audit of the company’s disclosures. With a predetermined audit fee, the auditing company lacks in motivation to undertake high costs in the form of a detailed audit. Instead, the auditing company seeks to maximize its own benefit by minimizing the audit effort, since the compensation remains the same. The client and other parties seeking audit assurance suffer as a result (Antle, 1982, pp. 503, 508, 512). However, this opportunistic behavior exists only in theory. In practice, other factors also influence audit intensity. For example, the audit result itself is reviewed by other entities, and insufficient audit actions can be sanctioned. Never- A. Grommes /Junior Management Science 10(1) (2025) 201-235 205 theless, it is useful to keep in mind the fundamental problems that arise in auditing and the application of flat fees. Audit firms use various forms of auditing procedures to detect accounting manipulation or unintentional misstatements. The model structure distinguishes between substantive and systematic audit procedures. While substantive audit procedures provide assurance on specific balance sheet items, for instance by sampling, systematic audit procedures provide broader assurance, for example by testing the functionality of internal control systems. In most cases, the desired level of assurance is achieved through a combination of both types of procedures (Wagenhofer & Ewert, 2015, pp. 432–435). Auditing sustainability reporting is unique in that it concerns non-financial reporting. Essentially, nonfinancial reporting provides more qualitative information rather than actual numbers, as it would be in the case for financial reporting. Almost all audit firms refer to the International Standard on Assurance Engagements (ISAE) 3000 (Revised) when performing sustainability reporting audits. This standard specifically covers the audit of information that can be classified as non-financial information (International Auditing and Assurance Standards Board, 2013, p. 5) and is therefore considered an umbrella standard. As the ISAE 3000 (Revised) is applied to a wide range of disclosures, it does not contain explicit audit procedures. Rather, it describes general requirements for audit firms, such as integrity, independence, and professionalism, which are also required for financial audits. It further provides detailed information on the content and scope of the audit firm’s reporting on its engagement. For the actual audit, the standard primarily requires auditors to review the content of the qualitative disclosures for material inconsistencies (International Auditing and Assurance Standards Board, 2013, pp. 20–21). However, the standard does not specify how materiality has to be determined or how audit procedures should be performed. There is no dispute that the audit in itself is a valuable tool. In general, auditing increases the credibility of the information disclosed, as shown, for example, by the fact that firms with audited financial statements pay lower interest rates than comparable firms with unaudited financial statements (Blackwell et al., 1998, pp. 58, 68). It should be noted, however, that the magnitude of such an effect varies depending on whether the information disclosed is favorable or unfavorable for a company. Another variable of particular importance for this thesis is the voluntariness of the audit. The following 2x2 matrix illustrates four possible conditions that financial or non-financial reporting can adopt. According to the attribution theory, financial statement users challenge positive information because it is consistent with the company’s interests. Negative information, on the other hand, is less challenged since it would not be reasonable for companies to misrepresent information that is not in their best interest. This theory is confirmed by practice. In experiments, Coram et al. show that voluntary audit assurance of positive sustainability disclosures has a significant positive effect on the share price. In contrast, no significant results were found for negative disclosures. It can therefore be concluded that financial statement readers have reliability concerns mainly when the disclosed information is positive. These concerns are consistent with attribution theory, which suggests that it is beneficial for firms to voluntarily undergo an external audit when published information is positive, in order to increase reliability of those results, while negative information already carries a higher level of reliability (Coram et al., 2009, pp. 145–148). These results are relevant for the research of this thesis, as the EU member states’ option right and the implementation within Germany allow companies to voluntarily subject their non-financial reporting to an audit. Accordingly, such a voluntary submission can have different motivations: The counteraction of the principal-agent relationship, the creation of a higher reliability of the information, especially if it is positive and therefore, according to the attribution theory, more likely to be doubted by readers of the report, or simply the satisfaction of stakeholders in order not to be at a comparative disadvantage to other companies (Bradbury, 1990, p. 33). The presented theories provide the basis for multiple research streams. They also form the basis for the development of the hypotheses for this thesis presented in Chapter 5. 2.4. Textual analysis in research Research in business economics heavily relies on quantitative methods for gaining new findings. The rationale is clear: countless amounts of data exist in numerical form. Financial statements containing balance sheets and profit and loss statements, stock prices and a vast range of related financial indicators, as well as statistical information on companies, industries, regions, and countries. The amount of numerical data is enormous. When this data is effectively contextualized, new insights can be uncovered. However, how do researchers handle data that is not numbers, but letters? In addition to the balance sheet, every financial statement provides notes. The income statement enables insight into earnings, but the management report covers even more. Stock prices and financial ratios are paired with analysts’ recommendations and company announcements, both written and verbal. Each statistical survey is accompanied by a corresponding text. All of this information is easily overlooked. However, it would be incorrect to state that textual data is not a topic of interest in research at all. In fact, this field of research has been growing in importance for some time. As a result, both earlier and more recent papers cover not only results in this context, but also the methodology on its own (e.g. Bae et al. (2023), Bochkay et al. (2023), Gentzkow et al. (2019), and Loughran and McDonald (2016,2020)). The EU taxonomy and especially the NFRD requirements have greatly increased the volume of non-financial reporting. This new information in form of textual data provides potential for research using textual analysis methods. However, it is important to ensure that the methodology does not take precedence over the actual research question (Bochkay et al., 2023, p. 792; Bae et al., 2023, p. 3). Therefore, before addressing the hypotheses, the methodology of textual analysis will be examined in detail based on the existing literature. A. Grommes /Junior Management Science 10(1) (2025) 201-235206 Table 1: Effects of audit on reliability based on the experiment of Coram et al. (2009, pp. 142–145) Positive information content Negative information content Audit assurance High reliability High reliability No audit assurance Low reliability High reliability This review provides a summary as well as an explanation of current methods and highlights their advantages, disadvantages, and areas of application. This will not only identify the appropriate methods to apply to the research purpose of this thesis. It also provides a valuable contribution as it summarizes the leading research in various fields, particularly in the domain of finance and accounting. Textual analysis has already been used to find evidence in several areas. Chen et al. found that both stock returns and earnings surprises can be predicted by peer-based knowledge on social media. They used one of the simplest methods imaginable: counting negative connoted words in articles written by individual investors on the social media platform Seeking Alpha. The ratio of negative connoted words2to the overall number of words was utilized to determine first a negative sentiment and then a decline in stock returns and even in earnings surprises. This effect increased as the number of negative words increased (H. Chen et al., 2014, pp. 1368– 1369, 1382, 1400). In a more recent study, Sautner et al. measured the extent of corporate exposure to climate change by using an algorithm to count key words3in earning call transcripts that are directly related to this topic. This method captured climate change exposure from the perspective of all key stakeholders, as the earning call transcripts included both shareholder and stakeholder questions as well as management responses (Sautner et al., 2023, pp. 1450–1451, 1492–1493). Using a comparable methodology, Chen and Srinivasan analyzed the 10-K reports of non-tech firms to investigate the relationship between digital activities, firm value, and performance. Specifically, they measured the frequency of digital terms4in the description section of these reports. The authors discovered that non-tech firms have generally increased their digital activities over time, and that greater involvement in digital activities has a positive impact on firm value and stock performance. These findings were made possible by quantifying the degree of digitalization within firms through textual analysis (W. Chen & Srinivasan, 2023, pp. 2, 10, 29, 35). As a final example, in a 2014 study, Purda and Skilliorn analyzed quarterly and annual financial statements for fraudulent activities. A multilevel textual analysis process first 2Examples for negative connoted words are loss, termination, against or impairment. 3The key words with the highest frequency were renewable energy, electric vehicle, clean energy, new energy, climate change and wind power (Sautner et al., 2023, p. 1466). 4Examples for digital terms are analytics, virtual reality, automation, artificial intelligence, big data, data science or digitalization (W. Chen & Srinivasan, 2023, p. 36). sorted words within the sample by frequency, then tested their predictive power using a decision tree-based approach, and finally concluded from the reports the probability that the statements were completely true and did not contain fraud. The algorithm based on textual analysis was able to confirm the presence or absence of fraudulent activity in over 82 % of the reports (Purda & Skillicorn, 2015, pp. 1194, 1197–1200, 1218). The listed research is illustrative for the wide range of possible applications for textual analysis. Ranging from meaningful, market-relevant results in the area of finance, to risk exposure in the example of climate change, to opportunities for companies in the example of technology adaptation, to relevant accounting issues like fraud detection, the application possibilities are unlimited. The textual data examined ranges from individual social media posts to transcribed communications between companies and stakeholders to official corporate disclosures in annual and quarterly financial statements. This demonstrates that textual data can contain relevant information in any conceivable form, regardless of its type and nature. In quantitative research, the approach is usually relatively straightforward. With the evaluation of a sample using statistical methods in direct relation to a hypothesis, researchers intend to obtain significant findings. The procedure in textual analysis is not as simple, as the database is initially qualitative. The information required for research can only be obtained in an exploitable form by means of an appropriate transformation (Loughran & McDonald, 2016, p. 1191). The major difficulty is not that textual data is less structured or presented in a different way, but rather that it has a high dimensionality. The base of the dimensions is defined by the number of different words in a language, and the exponent by the number of words in the text, as shown in Equation 1. When taking a text that consists of only ten words, and it is written in a fictional language that also only possesses ten different words, then this text can have ten billion different dimensions, each of which is different from the others. In reality, texts are much longer than ten words, and languages consist of more than ten different words, so both the exponent and the base, and thus the total number of dimensions, take on an unimaginably high degree (Gentzkow et al., 2019, pp. 535–536). t=nl(1) t=textual data dimensions l=language word options n=length of text in words A. Grommes /Junior Management Science 10(1) (2025) 201-235 207 Humans can handle the high number of dimensions because they are not important when reading text. Words are perceived and interpreted in the context of other words. Accordingly, sentences do not represent a sequence of independent variables. In textual analysis, however, these dimensions are important and must be addressed. At the beginning of any textual analysis, the number of dimensions considered should be drastically reduced in order to deal with the enormous amount of data. This reduction is usually performed within three steps. The first step is to divide the total volume of text into sections suitable for research. In the later course of this work, non-financial reports from different companies are analyzed. It is not necessary to examine the reports of all companies together. Rather it is sufficient to extract information from each report separately. This information can then be used to draw conclusions by applying other research techniques. By analyzing individual texts separately instead of performing one single overall analysis, the number of dimensions can be drastically reduced. In a second step, certain parts of the text can be excluded from the analysis. These are first of all frequently occurring words that maintain the grammatical structure of a text. Such words are important for human readers of a text, but contain little or no information that will emerge in the textual analysis. In addition, for some research it can also be useful to exclude words that occur very rarely in a text. Although these words may contain relevant information, the benefit of gaining this information could be outweighed by the additional effort involved in analyzing these words. In a final step, the stemming method can be used to adjust all words that have the same meaning but are spelled differently. This method unifies differently conjugated words by removing their suffixes. For example, the words connected, connecting, connection, and connections have the same informational meaning. By replacing them with their stem word connect, the dimensional base of the overall text is reduced once again (Porter, 1980, p. 130). The important aspect is to decrease the number of different words with identical information content. With these three steps, the dimensions of a text can be drastically reduced by lowering the base nof the dimension equation (Gentzkow et al., 2019, pp. 537–538). The efforts for simplification are addressing the base n of Equation 1. From a purely mathematical point of view, a reduction of the exponent would be more effective. However, it is not as easy to reduce the exponent. Textual data from a sample can be decreased or simplified at the expense of information loss. But the vocabulary of the language in which a text is written is exogenous. In textual analysis, instead of trying to reduce the exponent, often the entire formula gets modified. The so-called bag of words method ignores the position of words in a text. Alternatively, it only counts whether and how often individual words occur. This implies that the number of possible dimensions only result from the multiplication of the number of words n in the text by the number of possible words of alanguage l. Whereas before, a text of ten words length in a language containing only ten different words already had ten billion dimensions. Using the bag of word method a text in English language5with the same amount of dimensions can contain more than 100,000 words. Thus, ignoring the order of words in texts leads to a massive reduction of the dimensionality of a text (Gentzkow et al., 2019, pp. 539–540). After presenting a general overview of how textual analysis works, the next step is to specify its application areas. Text contains some information, but what kind of valuable information is included and how can it be extracted? Prior research has established various applications of textual analysis for information acquisition, which will be presented in the following literature review for the finance and accounting domain, which does not present all the literature, but the most important in terms of the objective of this thesis. 3. Literature review 3.1. Readability Readability is one of the main areas of use in textual analysis. Depending on the context, the definition of readability varies. Either way, it should somehow determine if text is designed in a way that readers can recognize and comprehend the underlying message (Loughran & McDonald, 2016, p. 1188). More specifically, readability represents the relationship between a text and the cognitive load required to understand it (Martinc et al., 2021, p. 143). Even if a text is generally comprehensible to a reader, a high cognitive load, or in simple terms, a text that is challenging to read, may indicate poor readability. The readability of a text always should be considered in the context of the target audience (Loughran & McDonald, 2020, p. 28). The United States Securities and Exchange Commission (SEC) supports this position (United States Securities and Exchange Commission (SEC), 1998, p. 9). The Plain English Handbook published by the SEC describes the linguistic form in which publications should be made. This includes the annual financial statements and non-financial reporting components. The handbook recommends, among other things, the use of everyday language words and short sentences. It further recommends to perform automatic readability checks by using formulas developed for this purpose, but also manual testing by simple human proofreading of own publications (United States Securities and Exchange Commission (SEC), 1998, pp. 18, 57). This thesis examines non-financial reporting with focus on sustainability reporting within the German market, although European or other international regulators have objectives for publication requirements similar to the SEC. The readability of text has been subject of many studies. Li’s widely cited paper examines the relationship between the readability of annual reports and company earnings. Here, readability was measured by two variables, the so-called Fog Index and the length of the reports. The Fog Index is closely 5There are several underlying bases for determining the number of English words. The following example is calculated using the number of 88,500 English words (Nagy & Anderson, 1984, p. 320). A. Grommes /Junior Management Science 10(1) (2025) 201-235208 related to the topic of readability and will be discussed in Chapter 4.2.1 in more detail. Li found that companies with a low readability score on their reports had worse earnings persistence. On the other hand, companies with easily readable annual reports are more persistent. This correlation could suggest that management is hiding negative information about its company by making the information in the reports more difficult to access (Li, 2008, pp. 222, 225–226, 244–245). Reporting quality can also influence investment allocation. Biddle et al. found that companies with higher reporting quality are less likely to be overor underinvested. Overor underinvestment occurs when the marginal benefit of a capital investment is lower than the marginal cost of that investment. Information asymmetries between management and investors in capital markets are one main reason for such inefficient allocations. Higher reporting quality can reduce the occurrence of adverse selection and its effects. Biddle et al. use the Fog Index to measure reporting quality and find a positive relationship between increasing reporting quality and decreasing overor underinvestment (Biddle et al., 2009, pp. 113–114). Not only the general allocation efficiency can be related to readability, but also individual trading behavior. Miller finds that shares of companies with less readable reports are traded less frequently. This effect is most evident among small investors and can be explained by the higher cost of information acquisition, which is particularly important for such investor groups (Miller, 2010, pp. 2108, 2114, 2138). Lawrence found similar results in that investors are more likely to hold on to shares of companies with more readable financial disclosures. For the increase of one standard deviation in readability, stock returns increase by 91 basis points on average (Lawrence, 2013, pp. 131, 135, 141–142, 144). These two studies use the Fog Index and text length as readability measures. Readability research has also provided insights in the domain of accounting. Chychyla et al. found a relationship between reporting complexity and accounting expertise within companies. They approximate accounting complexity by various parameters, including firm characteristics such as firm size and number of segments covered, as well as financial reporting variables such as the number of words in 10-Ks and their readability. The level of accounting expertise is approximated through boards of directors and audit committees. More specifically, the number of accounting experts6in these functions represent the level of accounting expertise within a company. Chychyla et al. argue that companies with a high level of reporting complexity also have a higher level of accounting expertise. This expertise should counteract the negative effects of accounting complexity and is expected to actively manage accounting complexity (Chychyla et al., 2019, pp. 227–229, 233–236, 247–248). 6An individual is considered an expert if he or she is a certified public accountant (or similar) or has professional experience in relevant areas such as treasury or auditing. Another example of how companies can influence their own reporting is the study by Chakrabarty et al. on the relationship between executive compensation and disclosure transparency. The study concludes that firms with managers receiving higher risk incentives, measured by the stock option compensation, produce less comprehensible 10-K reports. This is because these incentives encourage managers to undergo risky projects with higher rewards, which may not be in line with the company’s strategy. Management attempts to camouflage the undertaking of such projects by making the reporting less readable. Here, readability is assessed through the size of 10-K reports, as evidenced by Loughran and McDonald (2014). The result shows that companies in the top quartile of the stock option vegas7publish reports that are 15.4 % larger. The results were tested for robustness via variables such as firm complexity, while other testing, such as measuring readability via the Fog Index also supported the results (Chakrabarty et al., 2018, pp. 3, 5–7, 10–11, 13, 25). A recent study by Dorfleitner et al. examines readability, among other issues, in a setting similar to the one in this thesis. Dorfleitner et al. investigate the impact of the General Data Protection Regulation (GDPR) on privacy statements. Just like the NFRD, the GDPR was published as a directive in the EU and became binding law in the form of a regulation in 2018. Using methods similar to the Fog method, privacy statements were tested for readability. The result was a worsening of readability due to the introduction of the GDPR. Even when considering the number of words as a readability measure, there was a decrease in readability due to an increase in the length of privacy statements (Dorfleitner et al., 2023, pp. 1–2, 4, 10–12). The papers presented round off the literature review in the field of finance and accounting for the application area of readability, with the research area being constantly expanded and, above all, newer state-of-the-art methods being increasingly used. 3.2. Sentiment analysis In the context of textual analysis in accounting, sentiment analysis finds even more interest than the topic of readability (Bochkay et al., 2023, p. 797). In sentiment analysis, texts are examined to determine whether they have a positive or a negative tone. This has already produced findings in a wide variety of research areas, even in far unrelated fields. For example, Chevalier and Mayzlin examine the sentiment of customer reviews for books in the two largest online bookstores using a differences-in-differences approach and find that reviews with positive (negative) sentiment lead to significantly higher (lower) sales on the respective site. They also find that reviews with negative sentiment have a 7Vega measures how the value of a stock option changes as the volatility of the underlying asset changes (Black & Schloes, 1973, pp. 638–639). The vega parameter is used because it is expected that as volatility increases, management will take on riskier projects in order to increase the value of their own options. A. Grommes /Junior Management Science 10(1) (2025) 201-235 215 but discard them altogether if they occur in a larger fraction of the sample (Hoberg & Phillips, 2010, p. 3782). Finally, it is crucial to consider the text length. The greater the length of the texts to be compared, the higher the likelihood of common words occurring in both texts, thus increasing their similarity (S. Brown & Tucker, 2011, p. 317). Therefore, longer texts will naturally have a higher similarity than shorter ones. This can be counteracted by integrating a correction variable adapted to the respective study (S. Brown & Tucker, 2011, pp. 343–344) or by normalizing the vectors so that all vectors of a study have the same length (Hoberg & Phillips, 2010, p. 3809). Text length itself can also be used as an indicator of similarity. Brown et al. examine spillover effects in qualitative corporate disclosures. Specifically, they find that companies change their disclosures more if the industry leader, direct competitors, or an industry peer with the same auditor receives comments from regulators, even if their own disclosures were not criticized (S. Brown et al., 2018, pp. 623–625). To get to this finding, Brown et al. compared the absolute change in the number of words in risk disclosures from one year to the previous year, assuming that a change in the number of words also indicates a change in the information contained (S. Brown et al., 2018, pp. 631, 636). Chapter 4.2 outlined the fundamental practices of traditional textual analysis. The following chapter establishes the methodological principles of NLP and machine learning for a deeper understanding of the subject before a similar presentation of the state-of-the-art methods is provided in Chapter 4.4. 4.3. Development of natural language processing and machine learning Textual analysis has been utilized in academic research for a long time, but the advantages offered by NLP have brought the field to an advanced stage. Not all state-of-theart methods make use of NLP, but it forms the basis for many newer methods. Therefore, before listing the state-of-the-art methods in the areas of readability, sentiment analysis, and disclosure quantity and similarity, the framework around NLP and related concepts will be further examined. NLP offers the possibility to work with text in various ways. In general, NLP models function by taking given textual data as an input, transforming it, and presenting a new output as a result. In this transformation, NLP models differ in the way they operate. A first category includes rulebased models. The transformation of a text in these models is performed by manually developed rules. For instance, these rules may count the number of certain words or text elements. Keywords can be counted to identify specific contents of a text. Next to them, words defined as complex or the number of sentences can be counted. It can also be helpful to count words that have been previously assigned to categories, such as words with positive or negative sentiment (Bochkay et al., 2023, p. 769). These simple transformations are thus structured similar to traditional models like the Fog Index or traditional sentiment analysis methods. More complex are models like the VSM. Here, the input is modified by representing textual data as a vector, which is then transformed in a rule-based manner in a second step, for example to perform a similarity analysis (Bochkay et al., 2023, p. 771). However, NLP is not restricted to these specific applications, but also facilitates innovative applications utilizing machine learning models that exceed the current ones presented. Machine learning is a process where an algorithm generates a model from pre-existing data. The input data and desired output data are defined, and the model attempts to match the input data to the corresponding output data using defined rules. Once trained on the given training data, the model is expected to be capable of applying this process to new data sets. The training data is typically formed of a small but representative portion of the overall data set to ensure that the machine learning algorithm works well on all data (Zhou, 2021, pp. 3–4). Machine learning is thus classified as artificial intelligence, although, unlike many other artificial intelligence applications, it relies on historical data or training data. From this data, it is possible to identify patterns that can be used for event prediction or classification tasks (Alloghani et al., 2020, pp. 3–4). There are various machine learning models that are applied in the domain of textual analysis. One of them is the Naive Bayes model. The Naive Bayes model is a probabilistic generative model that calculates outputs based on conditional probabilities. It is mainly used for text classification or topic detection, which is an area of analysis that was not mentioned in the traditional methods, listed in Chapter 4.2, but is still used in text analysis. For example, Brown et al. use topic detection to determine whether it is possible to make assumptions about the probability of fraud cases based on company disclosures. They do not refer to measures such as readability or sentiment analysis, but specifically to the content of the disclosures and whether this can be exploited with the help of machine learning techniques to generate robust probabilities for the existence of fraud (N. Brown et al., 2020, pp. 238–239). In the Naive Bayes model, similar to the bag of words method, a text is considered as a vector. However, in this model, the text is represented as a binary vector, so the frequency of one and the same word is not relevant. The model now requires texts as training data. From this data, the model is able to determine the probability of certain parameters, in this case words, occurring in a category. The probabilities determined in the training phase can be transferred to the subsequent prediction phase in order to assign uncategorized texts to the category that the model considers most likely. This is the category of texts representing most similar vectors. The Naive Bayes model gets its name because it is founded on the naive assumption that each probability of the occurrence of a word is independent from the occurrence of any other word. Thus, the joint probability of several specific words occurring is equal to the product of the probabilities of the individual words, which would not be the case in a real setting (Aggarwal, 2018, pp. 123–125). A. Grommes /Junior Management Science 10(1) (2025) 201-235216 Another example of machine learning is the nearest neighbor classifier. This model classifies variables such as text or separate text components into categories, similar to the Naive Bayes model. The major difference is that the nearest neighbor classifier does not require a learning phase, but performs its classification decisions based on the training data only. A variable that is not included in the training data set gets assigned to the category in which an already categorized variable is located, which in the case of textual analysis would be the text from the training data set that was previously categorized and has the smallest deviation, thus representing the nearest neighbor. The deviation is calculated using cosine similarity (Aggarwal, 2018, p. 133). Text regression is another popular machine learning application in textual analysis. Due to the high dimensionality of textual data, as discussed in Chapter 2.4, common regression methods, such as ordinary least squares regression, are not suitable. The large number of different English words, which represents the number of parameters in a regression, usually even exceeds the number of observations in a sample. Instead, nonlinear regressions can be performed. The so called classification and regression trees model is constructed by iterating a text through all available branches of the decision tree. The text is stripped down to the most informative features, which are the words identified as most relevant. This can be done with the help of specific dictionaries. The features are then grouped and iterated through the decision tree, where the algorithm identifies more features that can help with the categorization. At each level of the tree, the algorithm selects features that best contribute to the separation of categories, resulting in increasingly specific criteria for classification. When a new data point, in this case a text outside of the training data, is added for classification, it travels along the branches based on the presence or absence of certain criteria, in this case specific words, and ends up at the end of a branch that determines the classification. In some models, the data point can also land at multiple ends, which are then weighted (Bochkay et al., 2023, pp. 771–772). Such classification and regression trees can be used in sentiment analysis, for example. The classification and regression tree model enables the detection of correlations in the form of a nonlinear regression by dividing variables into domains in which they are homogeneous. Relationships are not measured along one variable, but rather by discrete features. In the case of textual analysis, these features are words and are evaluated within the context of an association, such as sentiment. The models presented so far belong to the supervised machine learning group. This means that the models operate in such a way that the training data is already correctly labeled and the models therefore have a blueprint to which they can refer. In contrast, there is unsupervised machine learning, in which case the model has to recognize patterns without classified training data (Alloghani et al., 2020, p. 4). Topic modeling is one of the unsupervised machine learning models used in textual analysis. The most popular subset is LDA, which was briefly introduced in Chapter 3.3. In this unsupervised machine learning model, the algorithm detects topics within a text by using probabilistic methods to identify words that are related to a topic. The algorithm is therefore used in topic detection within texts, but also for similarity analysis, since texts with similar identified topics are assumed to have a higher similarity (Bochkay et al. 2022: 772). All of these supervised and unsupervised models are traditional models of machine learning. They offer advantages over some other textual analysis methods, but also show their own shortcomings. For example, traditional machine learning models have trouble recognizing complex contexts in the learning process. This can result in an incorrect or insufficient algorithm that does not generate valuable outputs. Another weakness is that users manually define the investigated features. For example, for readability, the number of characters or syllables are defined as examination variables. Finally, it is necessary to train the models, which requires both time and the availability of a suitable training data set. Deep learning methods can overcome these weaknesses (Bochkay et al. 2022: 772-773). At the beginning of this chapter, the functionality of NLP models was described, in which input data, such as textual data, is transformed and subsequently represented in a modified form as output data. The models therefore have three distinct layers: an input layer for the input of the data, a second layer in which a transformation process is applied according to the methods of the models described, and a third output layer for the display of the data. The process of deep learning is comparable, but varies in the middle of the model structure, between the input layer and the output layer. Instead of a straightforward transformation function, a so-called hidden layer is used. This layer, in turn, can consist of several layers that are interconnected (Aggarwal, 2018, p. 326). The hidden layer represents a mathematical function that computes output values from input values. This function itself is composed of several simpler functions, the individual layers. It is the depth of the layers within the model that gives deep learning its name. The more layers a model has, the more complex tasks it can solve. Each layer performs a unique function in the overall transformation (Goodfellow et al., 2016, pp. 5, 8). Figure 2 shows a simplified deep learning model where a specific output parameter is determined from various integer variables. An example use case would be the categorization of textual data. The hidden layer here has a depth of two and can be extended many fold in more complex models. Such models are also called Artificial Neural Networks (ANN) because they resemble the neural network of the human brain (Bochkay et al., 2023, p. 773). The intersections between layers are connected to form a network through which information passes. Deep learning methods consisting of ANNs find application in the processing of visual material such as images and videos, audio material such as audio tracks, and also In the processing of text and speech. Especially in textual analysis, Deep Learning offers enormous possibilities for the future, by considering words and sentences A. Grommes /Junior Management Science 10(1) (2025) 201-235 217 Figure 2: Multilayer neural network example (Aggarwal, 2018, p. 327) in the context of the whole text (LeCun et al., 2015, pp. 436, 442). The ANNs can be designed to process textual data as a simple vector. A more reliable use, taking into account the context of individual words, can be achieved by integrating loops into the structure of the ANNs. Inputs and outputs are not considered as single variables (words), but as dependent sequential variables (sentences or text segments). Such models are useful, for example, in translation applications or the in design of an artificial intelligence which is able to answer questions (Aggarwal, 2018, pp. 342–343, 350–351). In addition, there are models that calculate output via so-called attention. Here, the similarity of the input to different vector series which contain information is calculated first. The higher the similarity, the higher the assigned weight. The sum of the weighted vectors then provides information about the measure of attention, allowing the model to focus on certain parts of the text depending on which information is important for the given input. Vaswani et al. propose that attention weighting provides better output predictions than using loops in recurrent neural networks (Vaswani et al., 2017, pp. 2, 3, 5, 9). The attention mechanism is used in state-of-the-art large language models (LLM). LLMs are NLP models that have large neural network architectures and are trained on large sets of textual data. Open AI’s generative pre-training (GPT) language model ChatGPT is receiving tremendous attention. Immediately after its launch in November 2022, it was the fastest growing consumer application at the time and is now among the top 20 visited websites worldwide with 1.5 billion monthly users (Reuters, 2023, p. 1). Such GPT models are able to perform similarity assessments or to classify texts, but most of all they are known for their ability to answer questions through text generation. The models are built using a combination of unsupervised pre-training and supervised fine-tuning. The unsupervised pre-training is performed by taking a large amount of unlabeled texts as training data and transforming them into outputs using a multi-layer transformer decoder. The multi-layer transformer decoder works according to the attention mechanisms described before, where the transformation takes place over several layers in order to perform a step-by-step refinement of the interpretation of the text, from simple to complex contextual relationships. Finally, the fine-tuning is carried out with labeled textual data for the different tasks of such a model. The combination of unsupervised pre-training using the attention mechanism and supervised learning at the task level helps to bring the performance of GPT models to a new level (Radford et al., 2018, pp. 2–4, 8). Next to ChatGPT is the Bidirectional Encoder Representations from Transformers (BERT). BERT is an LLM with versatile application possibilities. Like ChatGPT, BERT is built in two steps. The first step of pre-training is done using a multilayer bidirectional transformer. The main difference between the models is the direction in which they generate text. In ChatGPT, the text is generated token11 by token or word by word from left to right, as a human would read it. BERT uses the Masked Language Model (MLM) instead. In the MLM, random tokens or words are masked in pre-training, making them unrecognizable to the model, so that they can be predicted based on the surrounding words. The advantage of MLM is that texts are not only generated from left to right so that only the preceding words in the pre-training influence the predictions, but also the following words. The second step of fine-tuning is performed as in the GPT models, using labeled textual data that is required for the particular application. The pre-trained version of BERT is designed so that fine-tuning can be performed by adding a single additional layer to the ANN, allowing it to be built on top of the base model with minimal effort (Devlin et al., 2019, pp. 4171– 4174). Both models are pre-trained with a very large amount of data and thus can be referred to as LLMs. There is no 11 Tokens are sequences of letters that are somewhat similar to words. On average, 100 tokens correspond to approximately 75 English-language words (https://platform.openai.com/tokenizer). A. Grommes /Junior Management Science 10(1) (2025) 201-235218 clear boundary at which NLP models are considered as a LLM. Chelba et al. have designed a benchmark for language modeling that contains one billion words (Chelba et al., 2013, pp. 1–5). When using such a huge benchmark, it is reasonable to describe a model as an LLM. ChatGPT uses the BooksCorpus12 database for pre-training, which contains about one billion words, comparable to the Chelba et al. benchmark. BERT is pre-trained with the BooksCorpus database as well as with English Wikipedia and thus has access to well over three billion words. Both models do not use the Chelba et al. benchmark because it provides a shuffled sentence-level corpus. The sentences and words within the database have been randomly distributed, so that the models that access the benchmark only consider the actual relationship between the words, and not the natural order of those words (Chelba et al., 2013, p. 2). However, this order is fundamentally relevant in the case of ChatGPT and BERT. Throughout the use of BookCorpus and Wikipedia, it is possible for the models to investigate long contiguous sequences (Radford et al., 2018, pp. 4, 5; Devlin et al., 2019, p. 4175). 4.4. State-of-the-art methods 4.4.1. Readability The previous chapter introduced the principles of NLP models. First, rule-based models were briefly presented, followed by machine learning models, including supervised and unsupervised models. Finally, ANNs were discussed and it was shown how they are applied in popular LLMs. Based on this foundation, the state-of-the-art models in the areas of readability, sentiment analysis, and disclosure quantity and similarity are presented now to complete the detailed overview of textual analysis methods. A comparison with the traditional models as well as an in-depth methodological insight then allows to identify the methods suitable for the research questions of this thesis. In the area of readability, the Fog Index has been very popular. Weaknesses such as insufficient applicability to business texts or the classification of words as complex, when they actually tend to be easily readable, have been identified and countered by alternative measures such as the Bog index. Loughran and McDonald find another measure that outperforms traditional readability indices: the 10-K document file size. The authors look at the file size of 10-Ks filed with the SEC. These files are highly standardized and presented in HTML format. Loughran and McDonald find that larger file sizes are associated with lower readability. The results are both strongly correlated with traditional readability measures and consistent with Loughran and McDonald’s definition of readability that higher readability leads to less ambiguity in valuation, which the authors demonstrate in their research. 12 BookCorpus is a database of novel books written by unpublished authors and contains over 11,000 books in genres like romance, history, and adventure (BookCorpus, 2023). File size as a readability measure is as simple as one can imagine and requires very little adjustment before use, making it less prone to errors than other measures. However, it is questionable whether this measure can be applied to other texts such as sustainability reports. Given that readability is a measure of the ease with which value-relevant information can be extracted, file size works mainly on the simplistic premise that a higher quantity of textual data makes it more difficult to extract relevant information (Loughran & McDonald, 2014, pp. 1646, 1650, 1658–1658, 1667–1668). Sustainability reports are very heterogeneous. The large variation in report length could yield significant results in the readability analysis, but it is debatable whether a longer sustainability report increases the difficulty of extracting valuerelevant information or whether it just contains more information. It is also possible for the same information to be presented differently in two reports, one with a brief version and one with a more detailed version. While the longer version may be easier for the reader to understand, it would reduce readability in this model due to the increased file size (Bochkay et al., 2023, p. 779). The approach of Loughran and McDonald is a method that does not take advantage of the developments in the field of NLP and machine learning that are described in Chapter 4.3. However, other state-of-the-art readability methods increasingly rely on them. Machine learning can be utilized to build an accurate model from available components that best fit a particular research question. Table 4lists different features that can be used in readability analysis, divided into categories. Shallow features, such as the length of words or sentences, the ratio of simple words, or the ratio of different terms, are used in traditional models and form the basis of readability analysis. Morphological features capture how words are used in relation to their stem words, and thus can capture complexity and grammatical features. These features get lost in many traditional models because stemming is used to reduce the dimensionality of such words back to their root. Syntactic features capture the frequency of certain word combinations or sentence structure, which also often get lost because of the use of the bag of words approach. Semantic features evaluate readability at the content level, going beyond other methods. For example, the use of many synonyms decreases readability, as a low reading level audience is more likely to require a simplified vocabulary. Another semantic feature is cohesion, which uses similarities to measure how well sentences merge into each other. An easy transition through a higher similarity of the last words of one sentence to the first words of another sentence increases the readability of a text (Madrazo & Pera, 2020, pp. 4–6). Once the analysis methods have been grouped, machine learning can be applied to determine the most effective methods for a particular use case. This can be done by training the model with only one category of analysis methods at a time and then comparing the results. Such a comparison can show, for example, that the shallow features are more accurate than the morphological features, or that the semantic A. Grommes /Junior Management Science 10(1) (2025) 201-235 219 Table 4: Textual features (for definitions of terms see Madrazo and Pera (2020, pp. 4–6) Shallow Features Morphological Features Syntactic Features Semantic Features Word length Inflection ratio Part of speech ratio Synonym usage Sentence length Morphological phenomena frequencies Dependency tree complexity Semantic closeness Ratio of simple terms Cohesion Ratio of different terms features are less accurate than any other group of features. Furthermore, it can also be determined which features within the categories have the greatest influence, so that, for example, the outcome of the shallow features predominantly depends on the result of the individual features word length and sentence length (Madrazo & Pera, 2020, pp. 4–6). Accordingly, the machine learning process does not consist of a single process in which the textual data passes through a very large number of layers. Instead, it is divided into several sub processes, each with a fewer amount of layers. On the one hand, ineffective criteria can be filtered out to reduce the computational effort compared to an all-encompassing process. At the same time, the accuracy of the analysis can be improved by eliminating features that take the wrong path in the learning process and lead to uninformative results. Another way to apply machine learning in readability analysis is to compare not only criteria, but also entire models. Such comparisons would not be feasible without the automation possibilities offered by machine learning due to the high engineering costs. Comparing models works similarly to comparing analysis features. All models are fed the same textual data to ensure comparability, the results are compared with each other in terms of the output or classification accuracy, and only the best model is used in actual research beyond the training data set. To go one step further, the models can then be combined with each other. For example, it may be found that one model, e.g. BERT, has a word prediction accuracy of 90 %, and another model, e.g. GPT, has an accuracy of 85 %, but using both models in combination achieves an even higher accuracy than either model separately, because the strengths of one model compensate for the weaknesses of the other. By having multiple models work together and combining their predictions, an overall more robust and powerful result can be achieved. Combining models can be performed through different approaches, such as categorizing the input text into the category predicted by the majority of models (if more than two models are combined), or weighting the results of the models according to their individual accuracy (Filighera et al., 2019, pp. 335–336, 338–340, 344–345). Readability models are being criticized for not being transferable between different types of text (Bochkay et al., 2023, p. 780). It is not worthwhile evaluating metrics that are of little importance to a text’s target audience (Loughran & McDonald, 2020, p. 28). Schoolbook texts should be accessible to students and adapted to their experience and reading ability. However, financial statements or earning call transcripts are not likely to be read by lower-level students, but by more experienced readers who can be expected to comprehend a certain level of complexity. Martinc et al. show that supervised and unsupervised NLP models are able to assess readability across different audiences. Using methods similar to those of Madrazo et al. and Filighera et al., they prove this by using diverse training data sets. Martinc et al. use text sources, such as educational materials, which are classified by reading ability, age group, or grade, but also large databases, such as Wikipedia. These databases are classified into different readability levels, such as simple, balanced, or normal. Readability can then be measured by an adjustable score that takes into account the reading skills of the respective audience (Martinc et al., 2021, pp. 241, 152, 166–169, 172–175). Technological advances in NLP and machine learning allow for multifaceted readability research, partly through a more efficient evaluation of traditional metrics and partly through new developments. Measuring readability can provide valuable information to companies. It allows them to assess their qualitative disclosures to determine if those disclosures are comprehensible for their stakeholders. At the same time, legislators, internal and external regulators and other readers of disclosures can benefit from the results of such measures. However, the level of readability must be relevant to the research question. Otherwise, it represents nothing more than an indicator without any particular meaning, which at worst is a reflection of the complexity of the company (Loughran & McDonald, 2020, p. 28). 4.4.2. Sentiment analysis Various applications of sentiment analysis were introduced in Chapter 4.2.2. Traditional methods measure sentiment by counting the number of words in a text, which are classified into sentiment categories using dictionaries. The greatest potential for improvement in these methods lies in the improvement of these dictionaries. The creation of a dictionary specifically for the finance and accounting domain, instead of the commonly used H4N, has had a major impact on sentiment analysis research (Loughran & McDonald, 2011, pp. 61–62). The application of machine learning to the field is likely to be even more significant. Machine learning methods eliminate the need for dictionaries to assess sentiment. Instead, the classification of texts or text segments is performed by an algorithm that is trained on data samples to detect the tone of a text. This approach promises a more accurate classification than methods that rely on dictionaries A. Grommes /Junior Management Science 10(1) (2025) 201-235220 (Hartmann et al., 2023, pp. 76, 78). Traditional machine learning works in the same way as supervised learning described in Chapter 4.3. The algorithm receives texts as training data that are pre-labeled with the corresponding sentiment. This approach is used, for example, by Azimi and Agrawal to extract information from the sentiment of 10-Ks. They find that both positive and negative sentiment can predict abnormal returns. Chen et al. already had similar findings in 2014 when analyzing SeekingAlpha articles (see Chapter 2.4), but Azimi and Agrawal’s results differ in that they look at a different dataset with 10-Ks and that they analyze a much larger sample with over 200 million sentences. At the same time, their results are also significant for positive sentiment, in contrast to those of Chen et al. who could not find significant results for other sentiments besides negative tone (Azimi and Agrawal, 2021, pp. 2, 10, 20–21, 32; H. Chen et al., 2014, p. 1337). This might be because the machine learning approach is capturing relationships that are not apparent through the wordbook approach, but there also may be other reasons for this. Machine learning has a significant advantage over dictionary approaches, as it is capable of consistently capturing sentiment over multiple periods. This is possible due to the volume and timeliness of the training data. Dictionaries can only capture a status quo and may have a lack of actuality.13 Therefore, state-of-the-art methods are often preferable to dictionary methods. Nevertheless, traditional dictionary methods can be used if the temporal context does not matter and it is only the occurrence of individual words that is crucial for the research question, or if the cost of the more complex implementation of machine learning exceeds the benefit of more accurate classification (Frankel et al., 2022, pp. 5515, 5522–5524, 5529). Even more accurate than traditional machine learning methods is the application of contextual deep learning to sentiment analysis. Although classical machine learning outperforms the dictionary approach, these methods, such as VSM or Naive Bayes, represent texts as bag of words and are thus subject to the problems described earlier. Algorithms that capture contextual information from word embedding, as in ANNs, can also capture the surrounding context and associated sentiment (Heitermann et al. 2023: 79). This is facilitated through the attention mechanism discussed in Chapter 4.3. Sentiment analysis is also being used in areas other than financial and accounting, such as economics, political science, and medical research. A multidisciplinary study by Colón-Ruiz and Segura-Bedmar finds that LLMs have the highest accuracy for sentiment analysis. While traditional machine learning models, such as VSM, perform well especially with a large amount of training data, LLMs dominate the field, in particular the BERT algorithm. It delivers slightly better results than competing models, but at the expense of 13 An example of this are the results of Long et al. (2023) presented in chapter 3.2, which could only be achieved by the authors creating a new dictionary adapted to time and context. higher computational costs (Colón-Ruiz & Segura-Bedmar, 2020, pp. 1, 5–6, 9–10). Since BERT is a pioneer in the field of LLMs, the model will now be the subject of a more detailed discussion. BERT is a pre-trained model, simplifying its use for users by eliminating the need to navigate through the underlying complexity. In order for BERT to be able to precisely adapt the analysis to research questions, only the fine-tuning of the model has to be performed. Here, BERT can achieve better results than traditional approaches even with only a few hundred training samples (Siano & Wysocki, 2021, pp. 6, 27–28). BERT is pre-trained according to MLM, enabling the model to predict missing masked words or to predict subsequent sentences. BERT is publicly available at no cost. While the algorithm requires high computational power, Google allows free use of online graphics processing units to operate BERT, so the model has few barriers for usage (Siano & Wysocki, 2021, pp. 7, 17, 22). In contrast to dictionary models, the operation of LLMs is more complex and difficult to comprehend. However, it is feasible to verify that such models actually capture sentiment from the context of information in a text, as they are intended to do, by deleting or changing words in manual tests. Siano and Wysocki have performed such tests and found that BERT still performs better than traditional models even when key words that would have influenced sentiment in the wordbook approach are deleted. Although the accuracy decreases, this evaluation indicates that BERT generates its predictions based more on the context of a text than on the word count, as in the case of wordbook approaches. Furthermore, BERT loses much of its predictive power when words in a text get randomized, which again suggests that the model delivers on its promise and, unlike bag of words models, extracts information from the structural organization of a text (Siano & Wysocki, 2021, pp. 20, 25–26). In addition to its many advantages, BERT also has some limitations. The biggest one is probably the limitation of tokens. Currently, texts with a maximum of 512 tokens can be analyzed. A token usually corresponds to a word or, depending on the tokenization, to only a fraction of a word. Therefore, lengthy texts, which would certainly include sustainability reports, cannot be analyzed as easily with BERT. Researchers can apply workarounds by selectively or randomly analyzing individual text components, or by analyzing each text component one at a time in a scrolling pattern. However, advances in machine learning and general technological progress offer hope that computational power will increase and these limitations will fade (Siano & Wysocki, 2021, pp. 9–10, 30). BERT has been used and developed by researchers in various fields. One of the most important developments is FinBERT, a fine-tuned version of BERT specialized for the financial domain. Loughran and McDonald have already identified that the financial domain language differs significantly from general language and have revolutionized textual analysis in this field with their own dictionary (Loughran & McDonald, 2011, pp. 49–50). Huang et al. follow this example A. Grommes /Junior Management Science 10(1) (2025) 201-235 221 by adapting the new state-of-the-art to the financial domain. They do this by pre-training BERT with a large number of texts with financial context like corporate disclosures, financial analyst reports and earnings conference call transcripts. These texts help FinBERT to better process tasks related to financial information. In total, 4.9 billion tokens are used for fine-tuning, which even exceeds the population of the pre-training for the plain BERT version. FinBERT has been compared to other LLMs as well as traditional text analysis methods and outperforms them, as well as untrained BERT, when applied to finance-specific texts, but also when applied to texts related to the Environmental, Social and Governance (ESG) domain (A. Huang et al., 2022, pp. 8–9, 19). Sentiment analysis has benefited from machine learning and NLP developments, which have led to new techniques and models that can significantly improve the accuracy in this field. The third main area of textual analysis considered in this thesis, disclosure quantity and similarity, also benefits from these developments. 4.4.3. Disclosure quantity and similarity While traditional methods determine similarity, using the bag of words approach, researchers are now increasingly utilizing machine learning techniques to determine similarities between texts. In traditional methods, similarity is mainly assessed by overlap in word usage. Matching words or tokens in two texts increase the similarity score. Instead of single words, sequences of words can also be considered. Thus, the similarity score increases when word sequences, usually consisting of two to four words, appear in the texts to be compared. These traditional techniques can be further adapted, for example by applying frequency weighting. Here, less frequent words are given a higher weight under the assumption that they contain more information. The similarity between documents increases, especially when rare words or word sequences overlap (Gaulin & Peng, 2022, pp. 2–3, 12–13). With the help of deep learning, word embedding algorithms are able to recognize similarity in texts without being limited to the occurrence of individual words or word sequences. For this purpose, the algorithm uses a sophisticated method in which it scans the text for predefined words and builds vectors from the surrounding words that are near the searched word. This nearness can be defined by a certain distance, e.g. up to ten words before or after the target word. These words are context words and are used to capture the relationship of the target word to its environment. The vectors of these context words are placed in the same vector space as the searched word, so that both grammatical and semantic relations between words can be captured. This allows for a deeper and more nuanced representation of textual content (Gaulin & Peng, 2022, p. 34). When used in combination with cosine similarity (described in Chapter 4.1.3), word embedding algorithms are a powerful tool for accurately measuring disclosure similarity. Traditional methods based on the bag of words approach can still provide decent results if the research question is primarily based on word choice and less on the context of the texts (Bochkay et al., 2023, p. 781). In summary, however, this area of textual analysis also benefits from the new technical possibilities offered by NLP. The literature review of the finance and accounting domain in Chapter 3 and the comprehensive presentation of traditional and state-of-the-art methods in Chapter 4 form the first main part of this thesis. The insights obtained from this study are significant for the following second main part of the thesis. This section will present a textual analysis application to a case in the accounting domain. This case is covered in the following chapters. 5. Hypotheses development 5.1. The relationship between auditing and compliance with regulatory requirements At the outset of this thesis, it was noted that sustainability reporting varies greatly from company to company. Furthermore, according to the NFRD, EU member states still have the opportunity to opt out of mandatory external audits for sustainability reporting. In addition to heterogeneity in terms of content and structure, the audit of sustainability reports is another criterion for differences in some EU countries, as many companies voluntarily have their reports externally audited with limited assurance, some even with the higher assurance level of reasonable assurance. To address the issue in more detail, the fundamental theories in the areas of sustainability reporting and auditing were examined first, followed by an extensive discussion of the textual analysis methodology. This included a summary of the principles of the methodology and a literature review in the financial and accounting domain. The fundamental theories considered in the area of auditing have been narrowed to the essential principles and have further emphasized the effects of the audit on the reliability of disclosures. Literature indicates that auditors can also serve as intermediaries to support company compliance with regulations. This is because they possess comprehensive knowledge through their activities and their diverse structure (Walker, 2014, pp. 214–215). Therefore, the role of the auditor can not only increase regulatory compliance, but also increase its effectiveness in general (King, 2007, p. 213). In qualitative research, Walker found that companies in the Australian trucking sector that participated in a voluntary compliance program achieved better performance and generated higher social value if they involved auditors in this process (Walker, 2014, pp. 215–216, 221). This example is not directly related to the thesis content wise, but the underlying structural relation is the same. Similar as in the Australian trucking sector case, European companies have the option to involve external auditors in a process voluntarily. Submitting reporting components from the non-financial reporting to an audit imposes additional auditing fees for companies, but can also provide benefits such as potentially increasing the reliability of the reporting in accordance to attribution theory or achieving a higher level of compliance with legal reporting A. Grommes /Junior Management Science 10(1) (2025) 201-235222 requirements, which is desirable for both the company and its stakeholders. From the related example and the theories presented, the following hypothesis can be stated regarding the impact of an external audit on the quality of sustainability reporting: H1: Audit assurance for sustainability reports increases their compliance with regulatory requirements. In Chapter 6.2, this is addressed by a more detailed examination of the requirements of the EU taxonomy as well as currently applicable and forthcoming auditing standards and the determination of the relevant dependent variables. 5.2. The relationship between the extend of corporate sustainability and reporting quantity and audit demand In addition to the influence of the audit on the quality of sustainability reporting, there may be other factors that influence both reporting and the circumstances whether an external audit takes place. Chapter 2.2 discussed the major theories that determine whether and to what extent companies are willing to engage in voluntary reporting. The voluntary disclosure theory suggests that companies will only voluntarily disclose information if their benefits outweigh their costs. There is already much evidence within this theory. For instance, research has shown that firm characteristics significantly drive corporate disclosure. Firm size, for example, has a generally positive impact on disclosure as reporting expertise usually increases with growing firm size, but also because larger firms are subject to greater public exposure and have to legitimate themselves to a greater extent. Other firm characteristics, such as the degree of internationalization, board size, or media exposure, also affect voluntary corporate disclosure (Zamil et al., 2023, pp. 247, 249, 252). Corporate sustainability is a comparatively less studied driver in this context. However, it is possible that sustainable companies report on their sustainable activities overproportionally in order to benefit from it. To address this research gap, the following hypothesis is posed: H2: Companies that are acting sustainable disclose a higher quantity of information in their sustainability reports. The degree of sustainability of companies’ actions here is defined in a simple manner using existing sustainability ratings. The quantity is measured using the file size of sustainability reports in an adjusted form, based on the methodology of Loughran and McDonald (2014). The detailed research design and and use of the variables is presented in section 6.3.2. Reporting requirements have an undeniable influence on this as well. When regulatory bodies impose mandatory disclosures, these disclosures are more likely to be made (Duran & Rodrigo, 2018, p. 14). The implementation of voluntary requirements, such as the GRI, can be a significantly driver in the reporting landscape as well (Dissanayake et al., 2019, pp. 102–103). A less studied influence is the impact of auditing. As voluntary disclosure is performed only if a company’s benefits outweigh the related costs, this theory could also be applied to voluntary audits of disclosures. According to attribution theory, positive information is more likely to be doubted than negative information. Therefore, it would be more reasonable for companies to undergo a voluntary audit if the information contained in their disclosure is predominantly positive, which leads to the following final hypothesis: H3: Companies that are acting sustainable are more likely to demand voluntary assurance of their reports. This hypothesis expands on the voluntary disclosure theory by shifting the focus from disclosure itself to the voluntary submission of voluntary disclosures as well as non-voluntary disclosures to an external audit. H1 represents the central hypothesis of this thesis. The secondary hypotheses H2 and H3 are indirectly related to it and can provide further insights into the research area of sustainability reporting. However, they will only be addressed to a more limited extent. The next chapter first describes the data gathered to address the hypotheses. This is followed by an exposition of the underlying research design. Next, the focus shifts to the determination of the dependent variables in regard to the hypotheses under consideration. Finally, the results are discussed. 6. Data and research design 6.1. Sample data The sustainability reports of a subset of companies required to report under the NFRD, which are large companies within the EU with an average number of at least 500 employees (European Union (EU), 2014, p. 4), are now examined in order to investigate the three hypotheses of this thesis. In total, the requirements of the NFRD affect approximately 10,000 companies. The CSRD will extend the scope of application by including medium-sized companies to approximately 50,000 companies (KPMG, 2022, p. 37), beginning from financial year 2024. Also crucial for the verification of the hypotheses is the distinction that an audit with limited assurance is mandatory under the CSRD, whereas under the NFRD there is still an option at EU member state level to exempt companies from this requirement. Due to time constraints, this thesis does not examine all companies affected by the NFRD, but only a subsample, which consists of all companies listed in the German Stock Index (DAX) and the Midcap DAX (MDAX). This subsample contains 90 companies, which corresponds to about one percent of the overall affected companies, so that the results may not be unconditional replicable at the EU level. Germany was chosen as the country of analysis, as it is the country with the most A. Grommes /Junior Management Science 10(1) (2025) 201-235 223 G250 companies within the EU14 (KPMG, 2022, p. 75) and could therefore take on a pioneering role in reporting issues. Furthermore, Germany has exercised its right to opt out of the mandatory audit of sustainability reporting. Only in this way is it possible to verify the hypotheses presented. The assurance rate in the sample is 77 %, which is slightly higher than the G250 overall (63 %) (KMPG 2022: 24). Appendix B presents additional descriptive statistics on the sample data. The dependent variables are gathered with the help of textual analysis methods outlined earlier in this thesis. Chapter 6.2 discusses the research design and the associated alignment of the dependent variables, while the actual gathering of the variables is described in Chapter 6.3. 6.2. Research design To analyze H1, it is first necessary to define how compliance with regulatory requirements can be assessed. It is then important to determine which of the textual analysis methods, presented in Chapter 4, are appropriate for the assessment. There is no clearly defined benchmark for reporting requirements compliance. This is because the reporting landscape itself is complex, ambiguous and sometimes even contradictory. Interregional standards such as the GRI Reporting Framework are opposed to the first drafts of sustainability standards from the International Sustainability Standards Board, supplemented by national regulations within the individual countries. The EU’s attempt at harmonization further adds to the complexity. The requirements of the NFRD could not provide the desired effects. The inadequate specification of the directive has resulted in a lack of information within the reported data. The options of the EU member states, such as the requirement for an audit, but also the disclosure options that allow to disclosure sustainability reports within or outside the management report, make it difficult to compare information between companies. The recently enacted EU taxonomy regulation imposes further requirements on companies (Velte, 2023, pp. 1–2). Reporting quality cannot only be derived from the regulatory requirements themselves. The relevant auditing standards may also be informative. Auditing standards extensively discuss regulatory requirements and provide guidance to auditors on how they can perform audit procedures. While there are numerous auditing standards covering a variety of areas in financial reporting, the ISAE 3000 (Revised) in particular provides comprehensive coverage of the subject of non-financial reporting. The majority of companies in the sample refer to the standard in various places within their sustainability reports. The ISAE 3000 (Revised) explicitly covers all assurance engagements that do not include historical financial information, which also goes for 14 In Germany there are 13 G250 companies. There are also 13 G250 companies in France, although France has not exercised the option for EU member states to be exempt from the audit, and therefore sustainability reports of companies that meet the size criteria are required to undergo an external audit (Reuters, 2021, p. 2). the non-financial reporting, and has the objective of providing reasonable or limited assurance on that information (International Auditing and Assurance Standards Board, 2013, pp. 5–6). The auditing standard is extensive. It describes the requirements for complying with the standard, including areas such as audit planning, the determination of materiality and the required content of the auditor’s report. Because the ISAE 3000 (Revised) covers such a wide range of topics, it does not provide many specific audit guidelines or requirements in terms of the content of the auditor’s report. However, it does give some guidance for determining when reporting can be considered compliant. First, in its objectives, the standard states that limited or reasonable assurance can be obtained when the subject matter information is free of material misstatement (International Auditing and Assurance Standards Board, 2013, p. 6), as is the case with other auditing standards. It also defines the mandatory characteristics of relevance, completeness, reliability, neutrality and understandability for published information (International Auditing and Assurance Standards Board, 2013, p. 12). Using these characteristics as evaluation criteria, compliance can be more specifically defined. In addition, the ISAE 3000 (Revised) states that inconsistencies indicate material misstatements (International Auditing and Assurance Standards Board, 2013, p. 20), so inconsistent information within sustainability reporting or inconsistencies between financial and non-financial reporting may also indicate a lower level of compliance. Despite its naming, the ISAE 3000 (Revised) is in need of improvement. When it came into force a decade ago, the area of non-financial reporting covered by it was much smaller and less complex. Furthermore, the importance of this information has increased dramatically over the years. As a result, new auditing standards are being developed, that will eventually replace the ISAE 3000 (Revised). For the German market, the Institute of Public Auditors (Institut der Wirtschaftsprüfer: IDW) has published two drafts for new auditing standards. These drafts address the substantive audit of non-financial reporting with reasonable assurance and limited assurance, respectively. The drafts are based on the ISAE 3000 (Revised), but are subject to considerable uncertainties of interpretation and therefore may not be used by auditing firms for current audits, also due to their status as drafts and not as finalized auditing standards (IDW Verlag, 2022b, p. 1; IDW Verlag, 2022a, p. 1). The drafts do, however, reveal a certain direction in which the audit procedures for ensuring the quality of sustainability reporting are being intensified and on what they are based. For example, the drafts IDW EPS 990 and IDW EPS 991 refer to the requirements of the EU taxonomy in many places, starting with the scope of application of the future standards to companies included within the EU taxonomy (IDW Verlag, 2022a, p. 4) to the performance of audit procedures according to the information categories of the EU taxonomy (IDW Verlag, 2022a, pp. 26, 29). Furthermore, the drafts explicitly state that the absence of information required by the EU taxonomy is generally to be considered as a A. Grommes /Junior Management Science 10(1) (2025) 201-235224 material misstatement (IDW Verlag, 2022b, p. 31; IDW Verlag, 2022a, p. 30). Other sections focus on the assessment of the process for the identification of taxonomy-eligible economic activities. Measuring regulatory compliance is challenging, as there are many requirements from different regulatory bodies. The requirements of the NFRD are currently in force but are almost obsolete. The CSRD, which is supposed to replace the NFRD, has not yet come into force. GRI standards exist in parallel and the IFRS Foundation is working on separate new standards. As far as auditing standards are concerned, companies mostly refer to ISAE 3000 (Revised), which is effective but also somewhat outdated. New auditing standards are still being implemented. The EU taxonomy, on the other hand, differs from other regulatory requirements. Its formally adoption in 2021 is relatively recent while it is also already effective for reporting of the recent financial year, 2022 (European Union (EU), 2020, p. 18). The frequent reference of the IDW in new auditing standards underlines the relevance. Due to these factors, this thesis employs the EU taxonomy requirements as a benchmark for regulatory compliance in general. The EU taxonomy has been applied on a mandatory basis for the second time in the last fiscal year of 2022. In this year, non-financial companies were required to report on eligibility and alignment of their activities for the first two of the six taxonomy objectives. At the same time, financial companies were only required to report on the eligibility of their activities, but not on their alignment. The reporting requirements will gradually increase until the financial year 2025, at which point companies will be expected to report fully on all six environmental targets. This reporting includes the identification of eligible activities, an assessment of whether these activities contribute to at least one of the six objectives while not harming any other objective, and the compliance with the minimum safeguards set of the taxonomy. This ensures a consistent identification of activities to be considered sustainable for the purpose of determining the relevant indicators (PricewaterhouseCoopers, 2023, p. 10). The EU taxonomy demands, on the one hand, information on the proportion of a company’s turnover as well as its investment and operating expenditure, which can be classified as sustainable according to the taxonomy (PricewaterhouseCoopers, 2023, p. 23; European Union (EU), 2020, p. 17). Disclosing these metrics provides insight into the current contribution to environmental goals as well as projecting future contributions. On the other hand, qualitative information must also be provided explicitly. Both the computation logic and the key elements of the indicators need to be disclosed. This qualitative information highlights the transition process from taxonomy-eligible activities to taxonomyaligned activities (European Commission, 2022, pp. 7–8, European Commission, 2021, p. 4). The definition of regulatory compliance proves to be difficult due to the many different regulatory bodies involved, although the requirements of the EU taxonomy were identified as a suitable quality characteristic as they are in use today and not going to be superseded by new regulations in the near future. The mandatory requirements for the first two objectives of the taxonomy, climate change mitigation and climate change adaptation, are appropriate for the analysis of this thesis, since in addition to the key figures, qualitative information is explicitly required, which can be evaluated with the help of textual analysis methods. However, it should be noted that these do not represent an exhaustive quality feature of sustainability reporting. Other approaches to measuring the quality of sustainability reporting are also conceivable. In order to analyze H1, it is crucial to identify not only the contents to be considered, but also the method most suitable. The detailed presentation of known methods in Chapter 4 serves this purpose. These methods can be divided into three main categories: readability, sentiment analysis and disclosure quantity and similarity. Within these main categories, the compatible individual methods can then be determined on the basis of the correspondence between the objectives of the method and the research question, as well as on the basis of restrictions, e.g. due to lack of time, computing power or other limited resources. The measure used to assess the quality of non-financial reporting is the extent to which the relevant sections of the reporting comply with the requirements of the EU taxonomy. To assess those extracts, not only their content but also their characteristics are evaluated, more precisely the characteristics that are also listed in the currently relevant auditing standard ISAE 3000 (Revised) and which are also relevant for various other contents in financial and non-financial reporting: relevance, completeness, reliability, neutrality and understandability (International Auditing and Assurance Standards Board, 2013, p. 12). The fulfillment of these characteristics indicates reporting quality. Trying to link the characteristics with textual analysis methods (overview in Table 2), understandability can intuitively be covered by readability measures. A text that is easily readable may not always be understandable. Still, high readability facilitates the reader’s comprehension, while poor readability makes reporting more difficult to understand. Next, neutrality can be assessed by various methods of sentiment analysis by examining whether the extracts show certain sentiments, such as positively formulated language, which indicates a lack of neutrality. Thus, understandability can be analyzed quite well with readability measures and neutrality can be analyzed with sentiment measures. Relevance, completeness and reliability tend to be less intuitive. It is reasonable to argue for analysis methods from the group of disclosure quantity and similarity to compare the extracts with the requirements of the EU taxonomy, but such an approach is likely to be less precise than the assessment of the characteristics of understandability and neutrality, where the methods correspond to the research problem more well. To overcome this problem, the analysis in this thesis is carried out using an exhaustive method through the application of an LLM. In Table 2, it can be seen that LLMs are among the state-of-the-art methods covering all methodological areas. The developments in NLP and machine learning make A. Grommes /Junior Management Science 10(1) (2025) 201-235 231 Figure 6: Relationship between ESG_Score and Assurance_LVL In this analysis, Market_CAP shows a significant correlation (t-statistic: 2.38) with Assurance_LVL, while ESG_Score is just below the significance threshold (t-statistic: 1.93). Based on this result, H3 cannot be confirmed. The correlation between a company’s sustainability and level of audit assurance cannot be statistically proven. However, it appears that larger companies with a higher capitalization are more likely to undergo an external audit. This relationship is also shown in Figure 7. 19 of the sample companies do not provide an external audit of their non-financial reporting.21 Only three of those companies exceed $10 billion in market capitalization, while the mean for the entire sample is $21.2 billion and the median is $8.7 billion (Appendix B). While many small-cap companies also undergo an audit of their non-financial reporting, at the same time a voluntary audit appears to be inevitable once a company reaches a certain market capitalization. Since the audit of non-financial reporting was not mandatory for the sample of German companies in the most recent fiscal year, implementation at the firm level is a tradeoff between costs and benefits (Widmann et al. 2021: 457), similar to the publication of supplementary information according to the voluntary disclosure theory. Audit fees are primarily driven by the size and complexity of firms, but other factors such as a so-called BIG-4 premium also play a role, as firms hope to achieve higher audit quality by hiring the large, prestigious audit firms (Widmann et al., 2021, pp. 473–475, 479). Based on the previous findings, larger firms in particular consider the benefits of an audit to outweigh the associated costs. The reasoning behind this opens up possibilities for further research. At the same time, smaller companies may not be able to afford the voluntary audit, as they do not have the same financial flexibility as larger companies. Neverthe21 One company (Sixt SE) excluded due to incomplete data. less, even the smallest companies included in the sample are listed on the MDAX, meaning that they are still relatively large compared to other companies in terms of total assets or market capitalization, so that this effect is unlikely to be observed in this case. The results of Equation 8and the results of the examination of the control variable Market_CAP should be considered in relation to the conditions of the data set. First, the sample size of 85 is relatively small. Second, the dependent variable Assurance_LVL is not a continuous variable but a categorical variable that distinguishes only between no assurance, limited assurance and reasonable assurance. It is particularly difficult to capture a categorical variable through a linear regression because the outcome of the regression equation provides values that need to be assigned to one of the three assurance levels since there are no intermediate levels. 8. Conclusion Sustainability reporting has emerged as a major element of corporate reporting in recent years. The challenges arising from climate change as well as increasing stakeholder awareness are some of the key driving factors. While a number of companies have been reporting on these issues voluntarily for some time, today almost all of the major companies are getting on board, partly driven by regulatory requirements. The EU taxonomy represents one of the key regulatory foundations in this field. Not only does it tighten reporting requirements, it also encourages real effects by redirecting capital flows and company activities towards more sustainability and reducing practices such as greenwashing (European Union (EU), 2020, p. 14). This thesis merges the content topic of sustainability reporting with the methodological topic of textual analysis. As sustainability reporting is mainly presented in qualitative text A. Grommes /Junior Management Science 10(1) (2025) 201-235232 Figure 7: Relationship between Market_CAP and Assurance_LVL form, this combination fits together quite well. The methodology of textual analysis has become very popular in academic research and among the general public through the introduction of text-generating models such as ChatGPT. This thesis makes several contributions. It contributes to the literature in the area of auditing, especially the auditing of non-financial reporting, and to the literature in the area of European reporting requirements. The two main areas of focus of this thesis are divided, with the first being a literature review and analysis of various textual analysis methods. The literature review focuses primarily on the finance and accounting domain, detailing the textual analysis techniques utilized in prior research. The methodology overview provides a comprehensive examination of the advantages, disadvantages, and limitations of each method. This review contributes to the reader’s understanding of textual analysis capabilities and potential applications, which will help readers identify appropriate methods for their own textual analysis research. A further contribution lies in the results of the investigation of the impact of audit assurance on the quality of sustainability reporting. The results were obtained through a combination of textual analysis methods and multiple linear regressions. The unique aspect of the methodology in this thesis is that the essential data for the analysis is obtained with the help of GPT 3.5, a freely accessible LLM with text comprehension and generation capabilities. The model is used to analyze a defined problem. Ratings or scores have been used to identify genuine economic connections frequently in prior studies. In this case, the LLM is utilized to generate a rating based solely on the text of the non-financial reports of the sampled companies. By constructing a specific prompt, it is possible to determine exactly which factors should be included to create such a rating. For the purpose of this thesis, the compliance of the reporting with the requirements of the EU taxonomy in relation to two specific taxonomy objectives has been defined as a quality feature that defines the rating. Typically, there are no established ratings for such specific objectives. LLMs have been utilized in prior studies in a similar manner. For example, Kim et al. used another GPT 3.5 model to summarize components of corporate disclosure, and found that these summaries generated each had a stronger positive or negative sentiment than the original reporting. Accordingly, GPT appears to be able to filter noise from the texts, improve the information content and present more relevant insights than the original reporting (Kim et al., 2023, pp. 1, 2, 5, 15–16, 19–20, 30). The information content of sustainability reports in this thesis was significantly reduced, with only the GPT_Rating remaining. While similar, the approach here is much more drastic, as Kim et al. eliminated about 70% of the original reporting, while here the entire text was eliminated and replaced by the GPT_Rating as a single number remaining. This significant reduction may account for why the analyses of this thesis did not yield many significant results. The regression results from the analysis of H1 show a negative but insignificant correlation between the level of assurance and the reporting compliance, displayed via the GPT_Rating variable, contrary to the expectations. Additionally, it was found that companies with a higher ESG score also have a higher GPT_Rating and are therefore more likely to meet the requirements of the EU taxonomy. The analysis of H2 indicates a positive correlation between the ESG_Score and the File_Size variable. Thus, sustainable acting companies disclose more information, which supports H2. Finally, the last regression analyses could not confirm the hypothesis of H3, which states that sustainable companies are more likely to have their non-financial reporting externally audited, as the attribution theory would suggest. The research results are subject to a number of limitations. First, the sample size which ranges from 73 to 90 companies, depending on the analysis, is relatively small. This is A. Grommes /Junior Management Science 10(1) (2025) 201-235 233 due to the fact that the sustainability reports had to be manually retrieved from company websites and being further processed for the analysis. In addition, the variable GPT_Rating is prone to error. On one hand, the analysis of this variable is founded solely on sections extracted from sustainability reports, and not on the entire reports themselves. The extraction procedure was based on an intuitive yet untested approach. The second and more influential source of error is the rating of the sections by ChatGPT itself. Due to time constraints, the rating was done using a zero-shot approach. The model was not fine-tuned and has not been fed with training data in advance. Additionally, the results were not subjected to robustness tests. Such robustness testing could be considered in future research, either by comparing the results with existing, externally validated scores, or by comparing them with a portfolio of results from other textual analysis methods. In this thesis, the results present data that are conceivable but not fully comprehensible. Changes in the prompt can affect the output, resulting in different rating results for the same report. This even happens if the prompt is not changed (Appendix E). The cause may be an overload of the GPT 3.5 model utilized, which is optimized for text output and has limits of approximately 8,000 tokens. OpenAI has announced new models capable of capturing an input context of up to 128,000 tokens. These models are designed to produce reproducible outputs, resulting in lower variances. Such models can enhance the quality of analyses and ensure comprehensive reports. However, at the time of performing the analysis (November 2023), the new models were not yet available.22 This thesis has introduced a number of textual analysis methods, ranging from the simplest models with almost no mathematical or technical depth, to models using the latest technical achievements in the field of machine learning. For the analysis of the hypotheses of this thesis, the application of an LLM has proven to be an exhaustive instrument. Despite the emergence and popularity of modern techniques, traditional methods are still commonly utilized in research due to their ease of understanding and intuitive results. However, the future of textual analysis lies in the possibilities presented by the advancing state-of-the-art. Generative LLMs open up numerous possibilities in all fields of research. Textual data is becoming increasingly relevant in the field of accounting. Models like ChatGPT can be particularly helpful in dealing with the growing importance of text in reporting and the associated information overload (Kim et al., 2023, p. 30). Researchers should not overlook these opportunities and utilize the possibilities offered by these and similar models in their research. References Aggarwal, C. (2018). Machine Learning for Text. Springer. 22 For an recent overview see: platform.openai.com/docs/models/overview Agrawal, S., Azar, P., Lo, A., & Singh, T. (2018). Momentum, mean-reversion, and social media: Evidence from StockTwits and Twitter. Journal of Portfolio Management,44, 85–95. Alloghani, M., Al-Jumeily, D., Mustafina, J., Hussain, A., & Aljaaf, A. (2020). A Systematic Review on Supervised and Unsupervised Machine Learning Algorithms for Data Science. In M. Berry, A. Mohamed, & B. Yap (Eds.), Supervised and Unsupervised Learning for Data Science. Springer Nature. Amel-Zadeh, A., & Serafeim, G. (2018). Why and how investors use ESG information: Evidence from a global survey. Financial Analysts Journal,74(3), 87–103. Antle, R. (1982). The Auditor as an Economic Agent. Journal of Accounting Research,20(2), 503–527. Arrow, K. (1963). Uncertainty and the welfare economics of medical care. The American Economic Review,53(5), 941–973. Azimi, M., & Agrawal, A. (2021). Is positive sentiment in corporate annual reports informative? Evidence from deep learning. Review of Asset Pricing Studies,11(4), 762–805. Bae, J., Hung, C., & van Lent, L. (2023). Mobilizing Text As Data [forthcoming].European Accounting Review. Bai, J., Boyson, N., Cao, Y., Liu, M., & Wan, C. (2023). Executives vs. Chatbots: Unmasking Insights through Human-AI Differences in Earnings Conference Q&A. https://ssrn.com/abstract=4480056 Ball, R., & Brown, P. (1968). An Empirical Evaluation of Accounting Income Numbers. Journal of Accounting Research,6(2), 159–178. Biddle, G., Hilary, G., & Verdi, R. (2009). How Does Financial Reporting Quality Relate to Investment Efficiency? Journal of Accounting and Economics,48, 112–131. Black, F., & Schloes, M. (1973). The Pricing of Options and Corporate Liabilities. Journal of Political Economy,81(3), 637–654. Blackwell, D., Noland, T., & Winters, D. (1998). The Value of Auditor Assurance: Evidence from Loan Pricing. Journal of Accounting Research, 36(1), 57–70. Bloomberg. (2021). Applying the EU taxonomy to your investments, how to start? https://www.bloomberg.com/professional/blog/applying -the-eu-taxonomy-to-your-investments-how-to-start/ Bloomberg. (2023). ESG Data. https://www.bloomberg.com/professional /product/esg-data/ Bochkay, K., Brown, S., Leone, A., & Tucker, J. (2023). Textual analysis in accounting: What’s next? Contemporary Accounting Research,40, 765–805. Bonsall, S., Leone, A., Miller, B., & Rennekamp, K. (2017). A Plain English Measure of Financial Reporting Readability. Journal of Accounting and Economics,62(2), 329–357. BookCorpus. (2023). BookCorpus dataset. https://paperswithcode.com/da taset/bookcorpus Bradbury, M. (1990). The Incentives for Voluntary Audit Committee Formation. Journal of Accounting and Public Policy,9, 19–36. Breijer, R., & Orij, R. (2022). The Comparability of Non-Financial Information: An Exploration of the Impact of the Non-Financial Reporting Directive (NFRD, 2014/95/EU). Accounting in Europe,19(2), 332–361. Brown, N., Crowley, R., & Elliott, W. (2020). What are you saying? Using topic to detect financial misreporting. Journal of Accounting Research,58(1), 237–291. Brown, S., & Knechel, W. (2016). Auditor-client compatibility and audit firm selection. Journal of Accounting Research,54(5), 1331–1364. Brown, S., Tian, X., & Tucker, J. (2018). The spillover effect of SEC comment letters on qualitative corporate disclosure: Evidence from the risk factor disclosure. Contemporary Accounting Research,35(2), 622– 656. Brown, S., & Tucker, J. (2011). Large-sample evidence on firms’ year-overyear MD&A modifications. Journal of Accounting Research,49(2), 309–346. Busco, C., D’Eri, A., & Novembre, V. (2022). Corporate Disclosure (N. Linciano, P. Soccorso, & C. Guagliano, Eds.). Springer Nature. Chakrabarty, B., Seetharaman, A., Swanson, Z., & Wang, X. (2018). Management risk incentives and the readability of corporate disclosures. Financial Management,47, 583–616. A. Grommes /Junior Management Science 10(1) (2025) 201-235234 Chelba, C., Mikolov, T., Schuster, M., Ge, Q., Brants, T., Koehn, P., & Robinson, T. (2013). One billion word benchmark for measuring progress in statistical language modeling. Chen, H., De, P., Hu, Y., & Hwang, B. (2014). Wisdom of crowds: The value of stock opinions transmitted through social media. The Review of Financial Studies,27(5), 1983–2020. Chen, W., & Srinivasan, S. (2023). Going digital: implications for firm value and performance [forthcoming].Review of Accounting Studies.ht tps://ssrn.com/abstract=4177947 Chevalier, J. A., & Mayzlin, D. (2006). The effect of Word of Mouth on sales: Online book reviews. Journal of Marketing Research,43(3), 345– 354. Christensen, B., Glover, S., Omer, T., & Shelley, M. (2016). Understanding Audit Quality: Insights from Audit Professionals and Investors. Contemporary Accounting Research,33(4), 1648–1684. Christensen, H., Hail, L., & Leuz, C. (2021). Mandatory CSR and sustainability reporting: economic analysis and literature review. Review of Accounting Studies,26, 1176–1248. Christofi, A., Christofi, C., & Sisaye, S. (2012). Corporate sustainability: historical development and reporting practices. Management Research Review,35(2), 157–172. Chychyla, R., Leone, A., & Minutti-Meza, M. (2019). Complexity of financial reporting standards and accounting expertise. Journal of Accounting and Economics,67, 226–253. Colón-Ruiz, C., & Segura-Bedmar, I. (2020). Comparing deep learning architectures for sentiment analysis on drug reviews. Journal of Biomedical Informatics,110, 103539. Coram, P., Monroe, G., & Woodliff, D. (2009). The Value of Assurance on Voluntary Nonfinancial Disclosure: An Experimental Evaluation. Auditing: A Journal of Practice & Theory,28(1), 137–151. de Kok, T. (2023). Generative LLMs and Textual Analysis in Accounting: (Chat)GPT as Research Assistant? https://ssrn.com/abstract=4 429658 Devlin, J., Chang, M., Lee, K., & Toutanova, K. (2019). BERT: Pre-training of deep bidirectional transformers for language understanding. Proceedings of NAACL-HLT 2019, 4171–4186. Dissanayake, D., Tilt, C., & Qian, W. (2019). Factors influencing sustainability reporting by Sri Lankan companies. Pacific Accounting Review, 31(1), 84–109. Dorfleitner, G., Hornuf, L., & Kreppmeier, J. (2023). Promise not fulfilled: FinTech, data privacy, and the GDPR. Electronic Markets,33(33), 1–29. Drempetic, S., Klein, C., & Zwergel, B. (2020). The Influence of Firm Size on the ESG Score: Corporate Sustainability Ratings Under Review. Journal of Business Ethics,167, 333–360. Dumrose, M., Rink, S., & Eckert, J. (2022). Why Do Firms in Emerging Markets Report? A Stakeholder Theory Approach to Study the Determinants of Non-Financial Disclosure in Latin America. Sustainability,10(9), 1–20. Duran, I., & Rodrigo, P. (2018). Disaggregating confusion? The EU Taxonomy and its relation to ESG rating. Finance Research Letters,48, 1–8. Dusík, J., & Bond, A. (2022). Environmental assessments and sustainable finance frameworks: will the EU Taxonomy change the mindset over the contribution of EIA to sustainable development. Impact Assessment and Project Appraisal,40(2), 90–98. Efretuei, E., & Hussainey, K. (2023). The fog index in accounting research: contributions and challenges. Journal of Applied Accounting Research,24(2), 318–343. European Commission. (2019). Communication from the Commission to the European Parliament, the European Council, the Council, the European Economic and Social Committee and the Committee of the Regions. The European Green Deal. https://eur-lex.europa.e u/resource.html?uri=cellar:b828d165-1c22-11ea-8c1f-01aa75e d71a1.0002.02/DOC_1&format=PDF European Commission. (2020). Taxonomy: Final report of the Technical Expert Group on Sustainable Finance. https://finance.ec.europa.e u/system/files/2020-03/200309-sustainable-finance-teg-final-r eport-taxonomy_en.pdf European Commission. (2021). FAQ: What is the EU Taxonomy Article 8 delegated act and how will it work in practice? https://finance.e c.europa.eu/system/files/2021-07/sustainable-finance-taxono my-article-8-faq_en.pdf European Commission. (2022). FAQs: How should financial and nonfinancial undertakings report Taxonomy-eligible economic activities and assets in accordance with the Taxonomy Regulation Article 8 Disclosures Delegated Act? https://finance.ec.europa.e u/system/files/2022-01/sustainable-finance-taxonomy-article8-report-eligible-activities-assets-faq_en.pdf European Union (EU). (2014). Directive 2014/95/EU. https://eur-lex.euro pa.eu/legal-content/EN/TXT/?uri=CELEX%3A32014L0095 European Union (EU). (2016). Paris Agreement. https://eur-lex.europa.eu /legal-content/EN/TXT/PDF/?uri=CELEX:22016A1019(01) European Union (EU). (2020). Regulation 2020/852. https://eur-lex.euro pa.eu/legal-content/EN/TXT/PDF/?uri=CELEX:32020R0852 European Union (EU). (2022). Directive 2022/2464/EU. https://eur-lex.e uropa.eu/legal-content/EN/TXT/PDF/?uri=CELEX:32022L246 4+ Filighera, A., Steuer, T., & Rensing, C. (2019). Automatic text difficulty estimation using embeddings and neural networks. In M. Scheffel, J. Broisin, V. Pammer-Schindler, A. Ioannou, & J. Schneider (Eds.), Transforming Learning with Meaningful Technologies. Springer International Publishing. https://link.springer.com/chapter/10.10 07/978-3-030-34971-2_17 Frankel, R., Jennings, J., & Lee, J. (2022). Disclosure sentiment: Machine learning vs. dictionary methods. Management Science,68(7), 5514–5532. Frankel, R., Mayew, W., & Sun, Y. (2010). Do pennies matter? Investor relations consequences of small negative earnings surprises. Review of Accounting Studies,15, 220–242. Gaulin, M., & Peng, X. (2022). Semantic vs. literal disclosure similarity. htt ps://ssrn.com/abstract=3971286 Gentzkow, M., Kelly, B., & Taddy, M. (2019). Text as data. Journal of Economic Literature,57(3), 535–574. Goodfellow, I., Bengio, Y., & Courville, A. (2016). Deep Learning. MIT Press. Guay, W., Samuels, D., & Taylor, D. (2016). Guiding through the Fog: Financial statement complexity and voluntary disclosure. Journal of Accounting and Economics,62(2-3), 234–269. Guidry, R., & Patten, D. (2012). Voluntary disclosure theory and financial control variables: An assessment of recent environmental disclosure research. Accounting Forum,36, 81–90. Hartmann, J., Heitmann, M., Siebert, C., & Schamp, C. (2023). More than a Feeling: Accuracy and Application of Sentiment Analysis. International Journal of Research in Marketing,40(1), 75–87. Hoberg, G., & Phillips, G. (2010). Product market synergies and competition in mergers and acquisitions: A text-based analysis. Review of Financial Studies,23(10), 3773–3811. Hu, N., Liang, P., & Yang, X. (2023). Whetting All Your Appetites for Financial Tasks with One Meal from GPT? A Comparison of GPT, FinBERT, and Dictionaries in Evaluating Sentiment Analysis. https://ssrn.c om/abstract=4426455 Huang, A., Lehavy, R., Zang, A., & Zheng, R. (2018). Analyst information discovery and interpretation roles: A topic modeling approach. Management Science,64(6), 2833–2855. Huang, A., Wang, H., & Yang, Y. (2022). FinBERT—A large language model for extracting textual information from financial text [forthcoming].Contemporary Accounting Research. Huang, X., S., T., & Zhang, Y. (2014). Tone Management. The Accounting Review,89(3), 1083–1113. IDW Verlag. (2022a). Entwurf eines IDW Prüfungsstandards: Inhaltliche Prüfung mit begrenzter Sicherheit der nichtfinanziellen (Konzern- )Berichterstattung außerhalb der Abschlussprüfung (IDW EPS 991 (11.2022)). https://www.idw.de/IDW/IDW-Verlautbarung en/IDW-PS/EPS-991-11-2022.pdf IDW Verlag. (2022b). Entwurf eines IDW Prüfungsstandards: Inhaltliche Prüfung mit hinreichender Sicherheit der nichtfinanziellen (Konzern-)Berichterstattung außerhalb der Abschlussprüfung (IDW EPS 990 (11.2022)). https://www.idw.de/IDW/IDW-Verl autbarungen/IDW-PS/EPS-990-11-2022.pdf International Auditing and Assurance Standards Board. (2013). ISAE 3000 (Revised), Assurance Engagements Other than Audits or Reviews of Historical Financial Information. International Framework for A. Grommes /Junior Management Science 10(1) (2025) 201-235 235 Assurance Engagements and Related Conforming Amendments. https://www.ifac.org/_flysystem/azure-private/publications/fil es/ISAE%203000%20Revised%20-%20for%20IAASB.pdf International Sustainability Standards Board. (2022a). [Draft]IFRS S1 General Requirements for Disclosure of Sustainability-related Financial Information. https://www.ifrs.org/content/dam/ifrs/projec t/general-sustainability-related-disclosures/exposure-draft-ifrss1-general-requirements-for-disclosure-of-sustainability-related -financial-information.pdf International Sustainability Standards Board. (2022b). [Draft]IFRS S2 Climate-related Disclosures. https://www.ifrs.org/content/dam /ifrs/project/climate-related-disclosures/issb-exposure-draft-20 22-2-climate-related-disclosures.pdf Kim, A., Muhn, M., & Nikolaev, V. (2023). Bloated Disclosures: Can ChatGPT Help Investors Process Information? https://ssrn.com/abstract =4425527 King, R. (2007). The Regulatory State in an Age of Governance. Palgrave Macmillan. Kothari, S., Li, X., & Short, J. (2009). The effect of disclosures by management, analysts, and business press on cost of capital, return volatility, and analyst forecasts: A study using content analysis. The Accounting Review,84(5), 1639–1670. KPMG. (2022). Big shifts, small steps. Survey of Sustainability Reporting 2022. https://assets.kpmg.com/content/dam/kpmg/sg/pdf/20 22/10/ssr-small-steps-big-shifts.pdf Kutner, M., Nachtsheim, C., Neter, J., & Li, W. (2004). Applied Linear Statistic Models (Fifth). McGraw-Hill Irwin. Lang, M., & Stice-Lawrence, L. (2015). Textual analysis and international financial reporting: Large sample evidence. Journal of Accounting and Economics,60, 110–135. Lawrence, A. (2013). Individual Investors and Financial Disclosure. Journal of Accounting and Economics,56, 130–147. LeCun, Y., Bengio, Y., & Hinton, G. (2015). Deep learning. Nature,521(7553), 436–444. Lepore, L., & Pisano, S. (2023). Environmental disclosure. Critical issues and new trends. Routledge. Li, F. (2008). Annual report readability, current earnings, and earnings persistence. Journal of Accounting and Economics,45, 221–247. Long, C. S., Lucey, B., Xie, Y., & Yarovaya, L. (2023). “I just like the stock”: The role of Reddit sentiment in the GameStop share rally. Financial Review,58, 19–37. Loughran, T., & McDonald, B. (2011). When is a liability not a liability? Textual analysis, dictionaries, and 10-Ks. Journal of Finance,66, 35–65. Loughran, T., & McDonald, B. (2014). Measuring Readability in Financial Disclosures. The Journal of Finance,69, 1643–1671. Loughran, T., & McDonald, B. (2016). Textual Analysis in Accounting and Finance: A Survey. Journal of Accounting Research,54, 1187–1230. Loughran, T., & McDonald, B. (2020). Textual Analysis in Finance. http://d x.doi.org/10.2139/ssrn.3470272 Lucarelli, C., Mazzoli, C., Rancan, M., & Severini, S. (2020). Classification of sustainable activities: EU taxonomy and scientific literature. Sustainability,12(16), 6460. Madrazo, A., & Pera, M. (2020). Is cross-lingual readability assessment possible? Journal of the Association for Information Science and Technology,71(6), 644–656. Martinc, M., Pollak, S., & Robnik-Sikonja, M. (2021). Supervised and unsupervised neural approaches to text readability. Computational Linguistics,47(1), 141–179. McGurk, Z., Nowak, A., & Hall, J. (2020). Stock returns and investor sentiment: textual analysis and social media. Journal of Economics and Finance,44, 458–485. Miller, B. (2010). The Effects of Reporting Complexity on Small and Large Investor Trading. The Accounting Review,85, 2107–2143. Nagy, W., & Anderson, R. (1984). How Many Words Are There in Printed School English? Reading Research Quarterly,19(3), 304–330. Porter, M. (1980). An Algorithm for Suffix Stripping. Programm,14(3), 130– 137. PricewaterhouseCoopers. (2023). EU Taxonomy Reporting 2023. Analysis of the financial and non-financial sector. https://www.pwc.de/de /content/20e6bff9-ea5a-4d03-b375-a6a58f7b8b46/pwc-eu-tax onomy-reporting-2023.pdf Purda, L., & Skillicorn, D. (2015). Accounting Variables, Deception, and a Bag of Words: Assessing the Tools of Fraud Detection. Contemporary Accounting Research,32, 1193–1223. Radford, A., Narasimhan, K., Salimans, T., & Sutskever, I. (2018). Improving language understanding with unsupervised learning (tech. rep.). OpenAI. Reuters. (2021). Regulatory Intelligence. Country Update-France: ESG Reporting. https://www.gide.com/sites/default/files/2021_06_17 _france_esg_guide.pdf Reuters. (2023). ChatGPT’s explosive growth shows first decline in traffic since launch. https://www.reuters.com/technology/booming-tr affic-openais-chatgpt-posts-first-ever-monthly-dip-june-similar web-2023-07-05/ Sautner, Z., van Lent, L., Vilkov, G., & Zhang, R. (2023). Firm-Level Climate Change Exposure. Journal of Finance,78, 1449–1498. Siano, F., & Wysocki, P. (2021). Transfer learning and textual analysis of accounting disclosures: Applying big data methods to small(er) data sets. Accounting Horizons,35(3), 217–244. Sorrosal-Forradellas, M., Barbera-Marine, M., Fabregat-Aibar, L., & Li, X. (2023). A new rating of sustainability based on the Morningstar Sustainability Rating. European Research on Management and Business Economics,29, 1–6. Technical Readiness Working Group (IFRS Foundation). (2021). General Requirements for Disclosure of Sustainability-related Financial Information Prototype. https://www.ifrs.org/content/dam/ifrs/gr oups/trwg/trwg-general-requirements-prototype.pdf Tetlock, P., Saar-Tsechansky, M., & Macskassy, S. (2008). More than words: Quantifying language to measure firms’ fundamentals. Journal of Finance,63, 1437–1467. Tworzydło, D., Gawro´ nski, S., Opolska-Biela´ nska, A., & Lach, M. (2022). Changes in the demand for CSR activities and stakeholder engagement based on research conducted among public relations specialists in Poland, with consideration of the SARS-COV-2 pandemic. Corporate Social Responsibility and Environmental Management,29(1), 135–145. United Nations. (2022). Global Issues. Climate Change. https://www.un.or g/en/global-issues/climate-change United States Securities and Exchange Commission (SEC). (1998). A Plain English Handbook. How to create clear SEC disclosure documents. https://www.sec.gov/pdf/handbook.pdf Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A., Kaiser, Ł., & Polosukhin, I. (2017). Attention is all you need. Proceedings of the 31st International Conference on Neural Information Processing Systems (NIPS 2017). Velte, P. (2023). Prüfung von Nachhaltigkeitsberichten nach der Corporate Sustainability Reporting Directive (CSRD) durch den Wirtschaftsprüfer – Fluch oder Segen? Schmalenbach Impulse,3(1), 1–13. Verrecchia, R. (1983). Discretionary Disclosure. Journal of Accounting and Economics,5, 179–194. Wagenhofer, A., & Ewert, R. (2015). Externe Unternehmensrechnung (Third). Springer Gabler. Walker, C. (2014). Organizational Learning: The Role of Third Party Auditors in Building Compliance and Enforcement Capability. International Journal of Auditing,18, 213–222. Widmann, M., Follert, F., & Wolz, M. (2021). What is it going to cost? Empirical evidence from a systematic literature review of audit fee determinants. Management Review Quarterly,71, 455–489. Zaman, M., Hudaib, M., & Haniffa, R. (2011). Corporate Governance Quality, Audit Fees and Non-Audit Services Fees. Journal of Business Finance & Accounting,38, 165–197. Zamil, I., Ramakrishnan, S., Jamal, N., Hatif, M., & Khatib, S. (2023). Drivers of corporate voluntary disclosure: a systematic review. Journal of Financial Reporting and Accounting,21(2), 232–267. Zhou, Z. (2021). Introduction. In Machine Learning. Springer.