Versification and Authorship Attribution
Abstract
Pro rozpoznávání autorství básnických textů nabízí současná stylometrie řadu metod založených na analýze pestré škály textových rysů (např. frekvence slov, frekvence znakových n-gramů). Jeden podstatný aspekt těchto textů ovšem zůstává stranou, a to jejich stránka versologická. Tato kniha proto na třech korpusech básnických textů (českých, německých a španělských) analyzuje, do jaké míry lze versologické charakteristiky využít jako indikátor autorství. Versologické charakteristiky jsou dále využity při řešení dvou případů sporného autorství (veršované drama The Two Noble Kinsmen a básně připisované Gavriilu Batěnkovovi).
Full text
INSTITUTE OF CZECH LITERATURE KAROLINUM PRESS CZECH LITERATURE STUDIES Contemporary stylometry uses diff erent methods to fi gure out a poem’s author based on features like the frequencies of words and character n-grams. Th ere is, however, one textual fi ngerprint it tends to ignore: versifi cation. Using poetic corpora in three diff erent languages (Czech, German and Spanish), this book asks whether versifi cation features like rhythm patterns and types of rhyme can help determine authorship. It then tests its fi ndings on two unsolved literary mysteries. In the fi rst, we distinguish the parts of the verse play Th e Two Noble Kinsmen written by William Shakespeare from those by his co-author, John Fletcher. In the second, we seek to solve a case of suspected forgery. How authentic is a group of poems attributed to the 19th-century Russian author Gavriil Stepanovich Batenkov? “A valuable and timely contribution to the history, methodology and practical development of stylometry and verse theory.” — IGOR PILSHCHIKOV, professor of Slavic, East European & Eurasian languages & cultures, UCLA PETR PLECHÁČ Versifi cation and Authorship Attribution PETR PLECHÁČ Versifi cation and Authorship Attribution ISBN 978-80-246-4871-2 9 788024 648712
Versification and Authorship Attribution Petr Plecháč Original manuscript reviewed by Mike Kestemont (University of Antwerp) and Igor Pilshchikov (University of California, Los Angeles). INSTITUTE OF CZECH LITERATURE is a part of the Czech Academy of Sciences Na Florenci 1420/3, 110 00 Prague 1, Czech Republic www.ucl.cas.cz KAROLINUM PRESS is a publishing department of Charles University Ovocný trh 560/5, 116 36 Prague 1, Czech Republic www.karolinum.cz Authors © Petr Plecháč, Artjoms Šeļa (chapter 4.2), 2021 © Institute of Czech Literature of the CAS, 2021 © Karolinum Press, 2021 Language review by Debra Shulkes; technical language review by Benjamin Nagy Cover and graphic design DesignIQ Set in the Czech Republic by Karolinum Press First edition Cataloguing-in-Publication Data is available from the National Library of the Czech Republic. This study is the result of research funded by the Czech Science Foundation as part of project GA ČR 17-01723S. This publication was created with the support of Research Development Program RVO 68378068 and published with support from the Czech Academy of Sciences. This work is licensed under a Creative Commons Attribution 4.0 International License (CC BY 4.0), which permits unrestricted use, distribution, and reproduction in any medium, provided the original author and source are credited. https://doi.org/10.14712/9788024648903 ISBN 978-80-7658-027-5 (Institute of Czech Literature of the Czech Academy of Sciences) ISBN 978-80-246-4871-2 (Karolinum Press) ISBN 978-80-7658-028-2 (pdf, Institute of Czech Literature of the Czech Academy of Sciences) ISBN 978-80-246-4890-3 (pdf, Karolinum Press)
Contents Introduction 7 Previous Publications 8 Data and Code 8 HTML version 8 1 Quantitative Approaches to Authorship Attribution 9 1.1 Origins of Stylometry 9 1.2 Searching for the “Golden Feature” 12 1.3 Multivariate Analyses 13 1.4 Support-Vector Machines 19 1.5 Versification-Based Attribution 32 1.6 Summary 34 2 Versification Features 36 2.1 Rhythm 36 2.2 Rhyme 41 2.3 Euphony 42 3 Experiments 43 3.1 Data 43 3.2 Versification-Based Attribution 47 3.3 Comparison with Lexicon-Based Models 56 3.4 Summary 63 4 Applications 69 4.1 The Two Noble Kinsmen 69 4.2 The Case of (Pseudo-)Batenkov: Towards aFormal Proof of Literary Forgery (co-authored by Artjoms Šeļa) 80 References 92
https://doi.org/10.14712/9788024648903.1 Introduction Contemporary stylometry is one of the fastest-growing fields in the computational study of literature. In recent years, anumber of textual characteristics and machine learning techniques have proven highly accurate in distinguishing the texts of different authors. Many of these features like word and character n-gram frequencies amount, however, to what is known as statistical “rare events”, or more precisely, alarge number of rare events (LNRE). As aresult, their analysis calls for fairly large text samples consisting of thousands or tens of thousands of words. Most theoretical studies in stylometry therefore focus on long novels. Poetry is usually omitted although we might expect to find many more cases of disputed authorship among poetic works. At the same time, poetry has anumber of specific versification features that are essentially Boolean or open to only alimited number of values. Some of these—stanza length and rhyme scheme, for example—are subject to the author’s conscious selection and so unsuitable for authorship recognition. In contrast, others like the preference for certain rhythmic configurations or sound frequencies in rhyme may be outside the author’s rational control. Although these characteristics have traditionally been recognised as author-specific (or at least period-specific), they have rarely featured in authorship attribution studies. The goal of this book is to examine the applicability of these versification features to authorship attribution projects. To this end, Idraw on poetic corpora in three different languages (Czech, German and Spanish) and apply this approach to two real-world cases of disputed authorship. Chapter 1 gives abrief history of quantitative methods of authorship attribution with special attention to the methods used in this book. Chapter 2 highlights different ways to capture versification features. Chapter 3 describes experiments with versification-based attribution and compares the methods commonly used in stylometry. Finally, Chapter 4 applies these findings to two actual cases of ambiguous authorship involving Englishand Russian-language texts respectively. In the first case,
8 Introduction Iattempt to determine which parts of the verse play The Two Noble Kinsmen were written by William Shakespeare and which were the work of his co-author, John Fletcher. In the second, working together with Artjoms Šeļa, Iinvestigate the potential forgery of numerous 19th-century Russian poems that were originally attributed to Gavriil Stepanovich Batenkov. These poems first appeared in the 1978 edition of the poet’s collected works, which was compiled by an established literary scholar—the main suspect in this intrigue. Previous Publications Chapter 1 expands on the opening sections of Versification and authorship attribution. Pilot study on Czech, German, Spanish, and English poetry (Plecháč, Bobenhausen and Hammerich 2018). Czech versions of Chapters 1, 2 and 3 were submitted as part of my PhD thesis at Charles University in Prague, Czech Republic in 2019. Data and Code The data and code required to reproduce the analyses in this book can be found at <https://doi.org/10.5281/zenodo.4555250>. HTML version From early 2022, this book will also be available online at <https://versologie.cz/ versification-authorship>.
https://doi.org/10.14712/9788024648903.2 1 Quantitative Approaches to Authorship Attribution 1.1 Origins of Stylometry Many scholars (e.g. Holmes 1998; Juola 2006) trace the origins of stylometry to several passages in aletter written by the British mathematician Augustus De Morgan to Reverend W.Heald on August 18, 1851 (De Morgan 1851/1882). After considering how to distinguish the Pauline epistles actually written by St.Paul from those written by other author(s), De Morgan mused that the average word length measured by the number of characters might give some clue: “If St.Paul’s epistles which begin with Παυλος gave 5.428 and the Hebrews gave 5.516, for instance, Ishould feel quite sure that the Greek of the Hebrews (passing no verdict on whether Paul wrote in Hebrew and another translated) was not from the pen of Paul” (De Morgan 1851/1882: 216; emphasis in the original). Later he complained: “If scholars knew the law of averages as well as mathematicians, it would be easy to raise afew hundred pounds to try this experiment on agrand scale” (De Morgan 1851/1882: 216). In fact, it was not until the end of the 19th century that the American physicist Thomas Corwin Mendenhall raised the money for this experiment. In an initial article entitled “The Characteristic Curve of Composition” (1887), Mendenhall suggested ignoring averages and dealing with overall word length distribution instead. Eventually, thanks to the support of abenefactor, August Hemenway, he applied this method to areal-world case of disputed authorship. The results of that experiment were published in the article “AMechanical Solution to aLiterary Problem” (1901). There, Mendenhall compared the shape of acurve determined by the relative frequencies of words of different lengths in works ascribed to William Shakespeare with equivalent curves for works by Francis Bacon and Christopher Marlowe (FIG. 1.1). Based on the similarities and differences, he cautiously concluded that while Bacon had not written the works in question, there was strong evidence that Marlowe had (Mendenhall 1901: 104–105). The discrepancies between the curves for Shakespeare and Bacon were, however, later found to be due to the comparison of verse texts by the former with non-verse texts by the latter (see Williams 1975).
16 1 Quantitative Approaches to Authorship Attribution word z-score 1.38 1.67 the and of to 's his a in with he that on but for I from as her their not Δ(SCOTT) Δ(BYRON) 0 0.5 1 1.5 2 2.5 word relative frequency Scott: Marmion Scott: The Lady of the Lake Byron: Childe Harold’s Pilgrimage the and of to 's his a in with he that on but for I from as her their not 0 0.02 0.04 0.06 FIG. 1.2: Burrows’ Delta for Walter Scott’s The Lady of the Lake (target text), Walter Scott’s Marmion and George Gordon Byron’s Childe Harold’s Pilgrimage (candidate set). (a) Relative frequencies of the 20 most common words in each text. (b) Relative frequencies of the 20 most common words transformed into z-scores. (c) Absolute values of the differences between the z-scores for each candidate and the target text; the final two columns show the mean values. word z-score the and of to 's his a in with he that on but for I from as her their not −2 −1 0 1 2
1 Quantitative Approaches to Authorship Attribution 17 1.3.2 The Geometric Interpretation of Burrows’ Delta and Its Modifications Argamon (2008) pointed out that the Delta measure that Burrows had stumbled on by intuition was actually the equivalent of measuring the Manhattan distance between two vectors. As such, the entire method could be seen as an instance of nearest neighbour classification or aspecial case of the popular k-nearest neighbour classifier where k= 1. Argamon proceeded from asimple consideration: Since the process was based on candidate ranking, there was no need to divide the sum of differences by the number of analysed words (n). After all, division by aconstant would not affect the ranking. Once the denominator was dropped from formula 1.5, we obtain asimple summary of the absolute values of the z-score differences, i.e. the Manhattan distance (DM; see FIG. 1.3): ( ) ( ) ( ) ( ) = ∆∝ = − ∑ M ,, 1 , t ,t n a b a b aj bj j t t D zt zt (1.6) In the same article, Argamon also suggested amodification of Burrows’ original method, or what he called the quadratic Delta (ΔQ) based on the Euclidean distance (DE) between the given vectors: ( ) ( ) ( ) ( ) = = − ∑2 E ,, 1 ,n a b ai bi i D zt zt tt (1.7) Just as dividing each distance by aconstant did not affect the final ranking in Burrows’ Delta, the same was true for extracting the root in the formula for the Euclidean distance (square root is amonotonically increasing function). The formula for ΔQ was, thus, defined as the square of the Euclidean distance: ( ) ( ) ( ) ( ) = ∆= − ∑ 2 Q ,, 1 , n a b aj bj j t t zt zt (1.8) The cosine Delta (Δ∠; Smith and Aldridge 2011) is another recent popular modification of Burrows’ Delta. It is based on the cosine similarity of vectors, that is, the cosine of the angle θ between them: ( ) ( ) ( ) θ = = ∑,, 1 cos n aj bj j ab zt zt tt‖ ‖‖ ‖ (1.9)
18 1 Quantitative Approaches to Authorship Attribution Since cos( θ ) ∈ [–1, 1], the formula is modifi ed so that—as with Burrows’ Delta and the quadratic Delta—the greater the similarity between two texts, the lower the cosine Delta value and vice versa: ( ) ( ) θ ∠ ∆=−, 1 cos ab tt (1.10) Metrics from the Delta family have been tested across languages and text types with various settings for the number of the most common units (n) and with other features such as lemmata, character n-grams and word n-grams (see, e.g. Eder 2011; Rybicki and Eder 2011; Jannidis et al. 2015). ta tb θ Manhattan distance Euclidean distance Cosine similarity −2 −1.5 −1 −0.5 0 0.5 1 −1.25 −1 −0.75 −0.5 −0.25 0 0.25 0.5 0.75 1 (a) Two-dimensional vector space. FIG. 1.3: Manhattan distance, Euclidean distance and cosine similarity of vectors taand tb.
1 Quantitative Approaches to Authorship Attribution 19 1.4 Support-Vector Machines Outside of the Delta, more sophisticated machine learning methods have gained increasing attention over the last decade or two. Th ese include random forest (e.g. Tabata 2012), naïve Bayes classifi er (e.g. Zhao and Zobel 2005) and above all support-vector machine (SVM) techniques (e.g. Diederich et al. 2003; Koppel and Schler 2004). Th e SVM technique is still probably the most popular in contemporary stylometry although deep-learning methods seem poised to overtake it (see , e.g. Savoy 2020). Th is section outlines the general principles behind SVM. An SVM is asupervised learning technique, which means that its algorithm uses labelled training data to infer aclassifi cation function for new data. Th is key principle can be illustrated with avery simple example based on artifi cial data. Imagine [0;0;0] Manhattan distance Euclidean distance Cosine similarity −3 −2.5 −2 −1.5 −1 −0.5 0 0.5 1 1.5 2 2.5 3 3.5 −1.5 −1.25 −1 −0.75 −0.5 −0.25 0 0.25 0.5 0.75 1 tb ta θ (b) Three-dimensional vector space.
20 1 Quantitative Approaches to Authorship Attribution atarget text t0 and 20 samples from each of two candidates (author 1, author 2). All of the texts are represented by z-scores for the two most common words (“the” and “and”). During the first (learning) phase, the SVM is fed data from author 1 and author 2 (training data). These data are labelled according to author, and the SVM tries to find afunction that correctly separates them by their labels. This is done using ahyperplane—asubspace with one dimension fewer than the original vector space. In our example with its two-dimensional data, this means aone-dimensional space, i.e. aline. During the second phase (classification), the hyperplane inferred from the training data is used to classify the target text. FIG. 1.4a shows that if the data are linearly separable, then an infinite number of potential hyperplanes can separate them correctly. Some of these may attribute the target text to author 1 while others may attribute it to author 2. From all these possibilities, the SVM chooses the hyperplane that maximises the distance to the nearest vectors on each side (also known as the support vectors), as shown in FIG. 1.4b (this is the maximum-margin hyperplane). In this case, the SVM classifies the target text as the work of author 1. Generally, for n-dimensional data, the task is formulated as follows: We are given the training data (x1, y1), (x2, y2), …, (xm, ym) where the first member of each pair denotes the n-dimensional vector xi = (xi,1, xi,2, …, xi,n) and the second member denotes one of two classes to which the vector belongs: yi ∈ {−1, 1}. The goal is to find anormal vector wand aparameter b to define ahyperplane H ⋅+=:0Hbwx (1.11) that separates the vector space into two half-spaces so that each half-space contains only data of the same class and the distance to the nearest vector is maximised. These requirements may be defined formally using the oriented distance d of the vectors xi to hyperplane H. This will be positive for vectors in one half-space and negative for vectors in the other one: ( ) ⋅+ =,i i b dH xw x w‖‖ (1.12) As we have two classes yi ∈ {−1, 1}, the requirement that each half-space contain vectors belonging to the same class may be formulated as:
1 Quantitative Approaches to Authorship Attribution 21 z(”the”) z(”and”) author 1 author 2 target text −4 −2 0 2 −4 −3 −2 −1 0 1 2 3 z(”the”) z(”and”) −4 −2 0 2 −4 −3 −2 −1 0 1 2 3 author 1 author 2 target text FIG. 1.4: A Support-Vector Machine (artificial data). (a) Various possible hyperplanes separating training data from author 1 and author 2. (b) Maximum-margin hyperplane; dashed lines indicate distances to support vectors.
22 1 Quantitative Approaches to Authorship Attribution ⋅+ ∀= > ⋅+ ∀=− < : 1, 0 : 1, 0 i i i i b iy b iy xw w xw w ‖‖ ‖‖ (1.13) This may be simplified as: ⋅+ ∀>,0 i i b iy xw w‖‖ (1.14) Next, we require the maximum possible margin. We therefore try to maximise the Euclidean (non-oriented) distance of the nearest (support) vectors to hyperplane H. All these requirements may be expressed as: ⋅+ ⋅+ ∀> , maxmin where , 0 i i b i i b b iy w xw w xw w ‖‖ ‖‖ (1.15) The number of solutions to this task remains infinite, however, since the direction of vector wis specified but its magnitude ‖w‖ is not. For practical reasons, the magnitude ‖w‖ should be inversely proportional to the Euclidean distance of the support vectors to hyperplane H: ⋅+ = 1 min i i b w xw w‖‖ ‖‖ (1.16) This allows for the simplification of the support vector requirement as follows: ⋅+=1 ibxw (1.17) For all of the vectors, the requirement is therefore: ( ) ∀ ⋅+≥,1 ii iy bxw (1.18) This brings us to abasic statement of the optimisation problem for an SVM: If we are looking for anormal vector w and aparameter b to define thehyperplane H with
1 Quantitative Approaches to Authorship Attribution 23 the maximum possible margin, and if the width of that margin should be inversely proportional to the magnitude ‖w‖ (formula 1.16), then the solution is the minimal possible normal vector w which satisfies inequation 1.18 (see FIG. 1.5). Again for practical reasons, it is not the magnitude ‖w‖ that we minimise but rather its square divided by two: ( ) ∀ ⋅+≥ 2 w, 1 min2 where , 1 b ii iy b w xw ‖‖ (1.19) This task is then solved by Lagrange multipliers (see, e.g. Abney 2007: 117–119). The example above is the simplest instance of the classification of n-dimensional data. In practice, however, we are often faced with more complex issues. Those challenges include (1) linearly inseparable data and (2) the need for classification into more than two classes. feature 1 feature 2 H 2 : w · x + b = −1 H 1 : w · x + b = 1 H: w · x + b = 0 author 1 (y = 1) author 2 ( y = –1) −3 −2.5 −2 −1.5 −1 −0.5 0 −1 −0.5 0 0.5 1 1.5 2 ||w|| 1 ||w|| 1 FIG. 1.5: ASupport-Vector Machine.
24 1 Quantitative Approaches to Authorship Attribution 1.4.1 Linearly Inseparable Data If there is no hyperplane that would correctly separate the classes, one of two approaches is usually employed: (1) the hyperplane condition is relaxed (the soft-margin SVM) or (2) we perform kernel transformation of the data into higher dimensions. Below Iconsider each of these techniques: (1) Asoft-margin SVM tends to be used with data with afairly low noise level. This method relaxes the condition that each half-space must only contain vectors of the same class. Instead, aslack variable ξ is introduced to penalise vectors on the “wrong” side of the hyperplane. Here the goal is to find the hyperplane with the maximum margin and minimum “overlap” of vectors into the half-space of adifferent class. For avector xi occurring in the half-space of adifferent class, ξ i denotes the Euclidean distance xi measured from the side of the margin defined by support vectors of its own class (Hyi) and normalised by the margin width (see FIG. 1.6). For these vectors, thus: ξ ξ ⋅ +− = = ⋅ +− 1 ii i i ii by by wx w w wx ‖‖ ‖‖ (1.20) For other vectors ξ i = 0. The optimisation problem (formula 1.19) is therefore extended to: ( ) ξ ξ ξ ξ = + ∀ ⋅ + ≥− ≥ ∑ 2 w, , 1 1 min2 where , 1 and 0 i m i bi ii i i C iy b w xw ‖‖ (1.21) where C is the penalty parameter of the model. This determines how much it will penalise misclassifications. (2) In kernel transformation, noisy linearly inseparable n-dimensional data are transformed into an (n+k)-dimensional space. In this way, they eventually become linearly separable (the “kernel trick”). As an example, we may consider the
1 Quantitative Approaches to Authorship Attribution 25 transformation of two-dimensional data (FIG. 1.7a) into athree-dimensional space (FIG. 1.7b) where each original vector ( ) =12 ,xxx is converted into ( ) = + 22 121 2 ,,xxx xx' . Since linguistic data tend, however, to include quite afew instances per class and avery high number of dimensions, kernel transformation is not usually required. 1.4.2 Multiclass Classification As we have seen, an SVM is inherently abinary classifier. The most common way to perform multiclass classification is therefore to split the problem into multiple binary tasks. There are two ways that this can be done: the one-vs.-rest strategy and the one-vs.-one strategy. (1) In the one-vs.-rest strategy, aclassification function is constructed for each class in order to separate its data from the rest of the data (kclasses, thus, produce kclassification functions, i.e. khyperplanes). If only one out of all of the kclassification author 1 (y = 1) author 2 ( y = –1) feature 1 feature 2 x1 x2 −3 −2.5 −2 −1.5 −1 −0.5 0 −1 −0.5 0 0.5 1 1.5 2 ||w|| 1 ||w|| 1 ||w|| ξ1 ||w|| ξ2 H 2 : w · x + b = −1 H 1 : w · x + b = 1 H: w · x + b = 0 FIG. 1.6: A soft-margin SVM.
32 1 Quantitative Approaches to Authorship Attribution — While SVM models give different weight to each feature (see Section 1.4.3), in Delta metrics, all these features contribute equally to the classification. An SVM is, thus, theoretically more resistant to data noise. Agood illustration can be seen in FIG. 1.10c where the SVM recognises that feature 2 is irrelevant to the classification. In contrast, Delta metrics would weigh both features equally. As such, Δ and ΔQ would misclassify the lower support vector of the class on the right of the chart since its nearest neighbour is the other class’s support vector. — On the other hand, the SVM approach requires quite alarge number of samples to carry out training. If only limited samples are available for some (or all) of the candidate authors, then we may still solve the task by using the less robust Delta measures. 1.5 Versification-Based Attribution In the previous sections, we saw that stylometry employs awide variety of both techniques and textual features. With the exception of early studies of Shakespeare (see Section 1.1), however, stylometry has not included features from the domain of versification. Yet despite this lack of interest from mainstream stylometry, versification features were taken up in the 20th century in the studies of verse experts associated with the so-called Russian school of metrics. In the early 1920s, for example, Boris Tomashevsky used versification to prove that the ending which Dmitry Zuev claimed to have found to Pushkin’s unfinished poem “The Mermaid” in 1889 was aforgery (Tomashevsky 1923/2008). Elsewhere verse rhythm and rhyme have been used to dispute the authenticity of alleged fragments of the tenth chapter of Eugene Onegin (Lotman and Lotman 1986), to challenge works newly added to Alexander Iliushin’s edition of Gavriil Batenkov’s poems (Shapir 1997, 1998; see Section 4.2 for details) and, above all, in the extensive work of Marina Tarlinskaja on Shakespeare and his contemporaries (Tarlinskaja 1987, 2014). Because of the isolation of these versification-based approaches, however, agulf has opened up between mainstream stylometry with its increasingly advanced methods and these studies, which have remained bound to the simple methods of descriptive statistics. This can be illustrated with an example from Tarlinskaja’s book Shakespeare and the Versification of English Drama, 1561–1642 (2014), which deals with the authorship of the play Henry VIII. Most scholars agree that Henry VIII was acollaborative text in which certain sections were written by John Fletcher (the “A” part) and the remainder were the work of
1 Quantitative Approaches to Authorship Attribution 33 William Shakespeare (the “B” part).6 Tarlinskaja (2014: 140–149) sets out to support this hypothesis with versification-based evidence. She, thus, points out that the two parts have different distributions of “strong syntactic breaks”.7 She also measures the frequencies of these breaks not only in Parts Aand B but also in two other plays from the same period: Fletcher’s Bonduca and Shakespeare’s The Tempest. She finds that within Part A, these breaks occur most frequently after the seventh syllable in aline (disregarding the line’s final syllable) and that the same holds true for Bonduca. In contrast, in Part B and The Tempest, they are most common after the sixth syllable (see FIG. 1.11). In the same way, Tarlinskaja compares the frequencies of monosyllabic words and enjambments (i.e. the lack of a“strong syntactic break”) at the end of lines. Here too she discovers asignificant similarity between Part Aand Bonduca on the one hand and Part B and The Tempest on the other. While these are strong and valid arguments, this analysis does not, in fact, differ substantially from Mendenhall’s approach (cf. Section 1.1). Since his time, however, methods have emerged that are far more reliable and robust than the simple comparison of two measurements. 6 See also the attributions by Spedding and Ingram (Section 1.1). 7 “Astrong syntactic break occurs, for example, at the juncture of sentences, or asentence and aclause, […] between the author’s [speech] and direct speech, […] or between adirect address and the rest of the utterance” (Tarlinskaja 2014: 24). metrical position frequency of strong syntactic breaks Henry VIII — part A Henry VIII — part B Fletcher: Bonduca Shakespeare: The Tempest 2345678910 0 0.1 0.2 0.3 0.4 FIG. 1.11: Frequency of “strong syntactic breaks” after particular syllables (metrical positions) in Parts Aand B of Henry VIII, Fletcher’s Bonduca and Shakespeare’s The Tempest. Source: Tarlinskaja 2014: table B.3.
34 1 Quantitative Approaches to Authorship Attribution 1.6 Summary Authorship attribution, as we have seen, generally relies on the notion that authorship can be determined based on the similarity between the numerical representation of atarget text and the numerical representations of the texts of candidate authors. While 19th-century stylometry used simple quantifications such as word length (Mendenhall), in the years since, the field has turned to far more complex characteristics. At the same time, the understanding of similarity has evolved from the simple comparison of two isolated measures to multidimensional analyses and machine learning methods. Various style markers have been taken into account for these purposes. They include the frequency of words, character n-grams, collocations and parts of speech, to name only afew. Nevertheless, akey aspect of the style of an important literary form— poetry—has almost completely been disregarded. While versification-based features are generally seen as author-specific, they have not been properly tested or used to attribute the authorship of poetic texts. The case for the stylometric study of versification features also has the following support: — Most features measured in stylometry (e.g. words and n-grams) amount to what are known in statistics as “rare events”, or more specifically, large numbers of rare events (LNRE; cf. Baayen 2001). Therefore, fairly large text samples are required. In practice, however, these are rarely available for authorship attribution studies with poetic texts. Usually only asmall number of poems are concerned and not an entire collection. On the other hand, versification features are generally far more frequent. This means that they may be analysed even with significantly smaller samples. — The vocabulary of apoetic text is not determined only by its author and genre/ topic. It may also be affected by poetic metre. Forstall and Scheirer (2010), for example, found an association between metre and the frequencies of certain character n-grams. — Some stylometrists have proposed combining different feature sets within asingle analysis. One example might be most common words + character n-grams + word n-grams (cf. Mikros and Perifanos 2013; Eder 2011). These features are, however, already strongly correlated. Versification, on the other hand, tends to be almost entirely independent of these correlations. We may, thus, expect acombined analysis of lexicon and versification to be more powerful than one of lexicon alone.
1 Quantitative Approaches to Authorship Attribution 35 In the following chapters, Iseek to test the applicability of versification features to modern methods of authorship attribution. To begin, Iexplore this method with Czech, German and Spanish poetry. To the best of my knowledge, this approach has only ever been tested sporadically. Two studies, conducted with small samples of Latin poetry (Forstall, Jacobson and Scheirer 2011) and old Arabic poetry (Al-Falahi, Ramdani and Bellafkih 2017) respectively, both yielded rather unsatisfactory results. There are also some reports of research with Middle Dutch poetry (Kestemont and Haverals 2018) and Portuguese poetry (Mittmann, Pergher and dos Santos 2019). Most recently, versification features have been used with greater success to attribute the authorship of Latin poetry (Nagy 2021). Some of my own attempts to test versification-based features can also be found elsewhere (Plecháč, Bobenhausen and Hammerich 2018; Plecháč and Birnbaum 2019).
https://doi.org/10.14712/9788024648903.3 2 Versification Features 2.1 Rhythm Since the time of Russian formalism, verse studies have distinguished between apoem’s metre (i.e. the abstract pattern of each line) and its rhythm (i.e. the realisation of that metre through particular phonetic units). The relationship between strong (S) and weak (W) metrical positions and particular phonetic qualities is usually not predetermined but rather stochastic. Precisely the same metre may, thus, be achieved in very different ways in particular lines. As an example, we may consider the opening quatrain of the first canto of Karel Hynek Mácha’s Máj, awell-known Czech narrative poem. All of the lines are written in accentual-syllabic iambic tetrameter with a strong ending (S) but the rhythmic realisation through stressed (“1”) and unstressed (“0”) syllables is different in each line: Byl pozdní večer— první máj— rhythm: 0 1 0 1 0 1 0 1 metre: W1 S2 W3 S4 W5 S6 W7 S8 večerní máj— byl lásky čas. rhythm: 1 0 0 1 0 1 0 1 metre: W1 S2 W3 S4 W5 S6 W7 S8 Hrdliččin zval ku lásce hlas, rhythm: 1 0 0 1 1 0 0 1 metre: W1 S2 W3 S4 W5 S6 W7 S8 kde borový zaváněl háj. rhythm: 0 1 0 0 1 0 0 1 metre: W1 S2 W3 S4 W5 S6 W7 S8
2 Versification Features 37 It is widely accepted that the distribution of rhythmic patterns is not random in the works of agiven author. Rather, it is an important part of their individual style. While the choice of metre is often based on general conventions with some metres reserved, for example, for aparticular genre, the overall way that it is achieved (rhythmic style) may be applied as amark of authorship. There are two main methods of capturing rhythmic style in continental European verse studies, and both of them originate in the Russian tradition. They are (1) determining arhythmic profile and (2) measuring the frequencies of rhythmic types.8 2.1.1 Rhythmic Profile Arhythmic profile is avector that tracks the frequency of stressed syllables in particular metrical positions. As an illustration, FIG. 2.1 presents the rhythmic profiles of all lines of iambic tetrameter with a strong ending in (1) Máj, (2) other works by Karel Hynek Mácha and (3)–(5) three books of poetry by alater author, Josef Václav Sládek. FIG. 2.1 captures some important differences between the rhythmic styles of the two authors:9 (1) The initial W1-position is stressed significantly more often in both Mácha samples than it is in Sládek’s works. (2) The S2-position is stressed significantly less often in both Mácha samples than it is in Sládek’s works. (3) The line-ending S8-position is stressed significantly more often in both Mácha samples than it is in Sládek’s works. (4) The W3-position and W5-position tend to be stressed slightly more often in both Mácha samples than they are in Sládek’s works. One disadvantage of the rhythmic profile method is that it completely disregards the context of particular syllables (cf. Dobritsyn 2016). FIG. 2.1 provides no information, for example, about what share of the approximately 13% of stressed syllables in the W5-position appear in monosyllabic words: 8 These features are also known respectively as astress profile and rhythmic forms. 9 For athorough analysis of these differences, see, e.g. Červenka 1998; Červenka and Sgallová 1978; Jirát 1931–1932; Jakobson 1938/1995.
38 2 Versification Features „Kde Vilém můj?“ „Viz“, plavec kní rhythm: 0 1 0 1 1 1 0 0 metre: W1 S2 W3 S4 W5 S6 W7 S8 (K.H.Mácha) And, of course, we face the same question about the remaining share contained in polysyllabic words: Kde borový zaváněl háj rhythm: 0 1 0 0 1 0 0 1 metre: W1 S2 W3 S4 W5 S6 W7 S8 (K.H.Mácha) The most crucial problem, however, relates to so-called extrametrical syllables, i.e. cases where more than one syllable corresponds to asingle metrical position. In Czech accentual-syllabic verse, these instances are rather rare: metrical position frequency of stressed syllables (1) Mácha: Máj (2) Mácha: other poems (3) Sládek: Světlou stopou (4) Sládek: Ze života (5) Sládek: Zlatý máj W1S2W3S4W5S6W7S8 0 0.25 0.5 0.75 1 FIG. 2.1: Rhythmic profiles of all lines of iambic tetrameter with a strong ending in (1) Máj, (2) other works by Karel Hynek Mácha and (3)–(5) three books of poetry by alater author, Josef Václav Sládek.
2 Versification Features 39 Přistoupí strážce alampy zář, rhythm: 1 0 0 1 0 0 1 0 1 metre: W0 S1 W1 S2 └W2┘S3 W3 S4 (K.H.Mácha) They are, however, very common in other syllabic accentual traditions. In English, for instance, we find: Those trackless deeps where many aweary sail rhythm: 0 1 0 1 0 1 0 0 1 0 1 metre: W0 S1 W1 S2 W2 S3 └W3┘ S4 W4 S5 (P. B. Shelley) The same holds true for metrical positions that are left blank (⌀), or what are sometimes called “headless lines”: Stay, the King hath thrown his warder down rhythm: ⌀ 1 0 1 0 1 0 1 0 1 metre: W0 S1 W1 S2 W2 S3 W3 S4 W4 S5 (W.Shakespeare) Since rhythmic profiling assumes metrical position to be abinary variable (achieved through either astressed or unstressed syllable), it has no way to reflect these situations. 2.1.2 Rhythmic Type The rhythmic type of averse line describes the entire bit string that captures its rhythm. According to this approach, apoetic text can be represented based on the frequencies of its rhythmic types. TAB. 2.1 gives an example of one such representation. This is a47-dimensional vector for the entire text of Mácha’s Máj. The rhythmic type method has no difficulty in resolving cases which the rhythmic profile approach cannot handle. Both extrametrical syllables (ranks 31–47, line “Přistoupí strážce…”) and headless lines (ranks 31–47, line “Znovu vmdlobách…”) are processed easily. Moreover, since this method does not focus on particular metrical positions but rather on entire lines, it can also be applied within systems where the
40 2 Versification Features number of positions varies (accentual verse) or where it makes no sense to distinguish them (free verse). On the other hand, the rhythmic type approach may produce rather sparse data. Some author-specific substrings may also end up being divided among alarge number of less common types. 2.1.3 Rhythmic N-Grams Given the limitations outlined in the previous sections, this book proposes using amethod inspired by Forstall, Jacobson and Scheirer (2011) that charts amiddle course. This method involves measuring the frequencies not of entire bit strings but their substrings. The latter are described here as rhythmic n-grams. From averse line with kmetrical positions, Iextract all possible substrings that are of length n and start at the i-th position (i∈ {1, 2, 3, …, k− n + 1}). Ithen measure the frequencies of their rhythmic realisations. This can be illustrated by looking at the frequencies of rhythmic bigrams in Mácha’s Máj (TAB. 2.2). To capture the range of rhythmic variations as fully as possible, Irepresent the samples from my experiments with syllabic (Spanish) and accentual-syllabic (Czech) data through acombination of the frequencies of rhythmic 2-, 3and 4-grams. In the case of the purely accentual (German) samples, Irely on the rhythmic type method for the reasons given in Section 2.1.2. rank rhythmic type relative frequency absolute frequency example 1 10010101 0.2305 80 Večerní máj— byl lásky čas 2 01010101 0.1671 58 Byl pozdní večer— první máj 3 10010100 0.0922 32 Modré se mlhy houpají 4 01010100 0.0605 21 Já zatím hrob mu vyryji 5–6 01000100 0.0519 18 Vzdy zeleněji prosvítá 5–6 10100101 0.0519 18 Břeh je objímal kol akol 7 01000101 0.0432 15 Tam při jezeru vížka ční 8 01001001 0.0288 10 Kde borový zaváněl háj 9–10 10011001 0.0230 8 Hrdliččin zval ku lásce hlas 9–10 10100100 0.0230 8 Dále zeleně zakvítá … 31–47 100100101 0.0029 1 Přistoupí strážce alampy zář 31–47 1010100 0.0029 1 Znovu vmdlobách umírá TAB. 2.1: Rhythmic types of lines of iambic tetrameter with a strong ending in Karel Hynek Mácha’s Máj.
2 Versification Features 41 2.2 Rhyme The peculiarities of rhyme are also generally recognised as author-specific. For my purposes, rhymes are represented as unordered pairs of the following features of both rhyming words: (1) morphological features (for the Czech data, this refers to the first position of the Positional Tag10 (=part of speech); for the German and Spanish data, this is the entire tag produced by the stochastic tagger TreeTagger11), (2) word length measured by the number of syllables, (3) number of syllables after the stressed syllable, (4) final syllable coda, (5) final syllable nucleus, (6) onset of the final syllable + coda of the penultimate syllable (weak rhymes only) and (7) nucleus of the penultimate syllable (weak rhymes only). TAB. 2.3 breaks down the rhymes found in Johann Wolfgang Goethe’s “Wandrers Nachtlied II” according to this schema:12 Über allen Gipfeln Ist Ruh’, In allen Wipfeln 10 See Hajič 2004. 11 See <http://www.cis.uni-muenchen.de/~schmid/tools/ TreeTagger/>. 12 The International Phonetic Alphabet is used to represent sounds throughout this book. Rhythmic realisations 00 01 10 11 000 001 011 101 100 ⌀1 W0S10.4092 0.5533 0.0346 0.0029 S1W10.4611 0.0893 0.4467 0.0029 W1S20.2017 0.7061 0.0836 0.0058 0.0029 S2W20.2104 0.0749 0.6340 0.0490 0.0029 0.0029 W2S30.0432 0.8012 0.1066 0.017 0.0029 0.0029 S3W30.1239 0.0259 0.8242 0.0086 0.0029 0.0144 W3S40.3083 0.6397 0.0345 0.0058 0.0086 0.0029 TAB. 2.2: Rhythmic bigrams of lines of iambic tetrameter with a strong ending in Karel Hynek Mácha’s Máj.
48 3 Experiments was then represented as avector defined by the following versification features (as described in detail in Chapter 2): (1) frequencies of rhythmic 2-, 3and 4-grams for syllabic and accentual syllabic verse (CS, ES); frequencies of the 100 most common rhythmic types for accentual verse (DE); (2) frequencies of morphological, phonetic and rhythmic rhyme characteristics; and (3) frequencies of sounds. Iopted for an SVM as aclassifier using the one-vs.-one strategy for multiclass classification (cf. Section 1.4.2). Implementation took place through the SVC module of the scikit-learn library16 with the following settings (cf. Section 1.4.1): 16 <https://scikit-learn.org/stable/modules/generated/sklearn.svm.SVC.html> Subcorpus Metre(s) Era of birth # of authors Authors (# of samples) CS1 T4 1760–1820 9 Čelakovský, František Ladislav (12); Havelka, Matěj (13); Hněvkovský, Šebestián (11); Kulda, Beneš Metod (27); Nejedlý, Vojtěch (17); Picek, Václav Jaromír (21); Pohan, Václav Alexander (10); Tablic, Bohuslav (16); Vinařický, Karel Alois (15) CS2 I5w 1840–1855 7 Čech, Svatopluk (13); Kvapil, František (11); Mokrý, Otokar (15); Nečas, Jan Evangelista (10); Sládek, Josef Václav (16); H. Uden (17); Vrchlický, Jaroslav (281) CS3 I5w 1860–1870 5 Klášterský, Antonín (64); Kvapil, Jaroslav (19); Leubner, František (10); Machar, Josef Svatopluk (22); Sova, Antonín (15) DE1 F1650–1699 6 Brockes, Barthold Heinrich (51); Drollinger, Carl Friedrich (11); Gottsched, Johann Christoph (29); Kuhlmann, Quirinus (30); Neukirch, Benjamin (21); Tersteegen, Gerhard (25) DE2 F1730–1754 5 Goethe, Johann Wolfgang (46); Jacobi, Johann Georg (12); Müller, Friedrich (15); Pfeffel, Gottlieb Konrad (28); Wieland, Christoph Martin (23) DE3 F1760–1794 7 Bernhardi, Sophie (12); Eichendorff, Joseph von (32); Grillparzer, Franz (52); Müller, Wilhelm (16); Schenkendorf, Max von (10); Schulze, Ernst (19); Tieck, Ludwig (28) ES1 11σ 1500–1560 5 de Acunya, Hernando (10); de Borja, Francisco (17); deCetina, Gutierre (31); de Góngora, Luis (14); de Herrera, Fernando (39) ES2 11σ 1561–1599 6 Argensola, Bartolome (19); de Quevedo, Francisco (63); deRojas, Pedro Soto (15); de Tassis y Peralta, Juan (25); de Ulloa y Pereira, Luis (13); de Vega, Lope (167) TAB. 3.5: Subcorpora details (T4: trochaic tetrameter with both strong and weak endings; I5w: iambic pentameter with weak endings; F: accentual verse; 11σ: hendecasyllabic verse).
3 Experiments 49 — kernel = “linear” (no kernel transformation); — C = 1 (default value of the penalising parameter; different settings had only anegligible impact on results). For other parameters, default values were used. Accuracy for each subcorpus was estimated using leave-one-out cross validation. As there was afairly low number of samples per class, using standard leave-one-out validation might have biased the results since the actual author was only represented by nine samples in the training data while the other authors were each represented by 10 samples. To eliminate this risk, one randomly selected sample was dropped from the training data for every author besides the test sample author. This equalising approach was applied in all of the experiments described in this book, unless indicated otherwise. To achieve more representative results, Irepeated this entire process 30 times with anew random selection of both authors and samples in each iteration. The entire procedure is captured in the following code in Python: ''' A dict contains authors' samples (represented by vectors): samples = { 'author1': [sample1, sample2, …], 'author2': [sample1, sample2, …], … } ''' import random from sklearn.svm import SVC classifier = SVC(kernel='linear', C=1) n_authors = 5 n_samples = 10 n_iterations = 30 for iteration in range(n_iterations): selected_samples = {} correct_classifications = 0 # Select 5 authors/10 samples at random for author in random.sample(samples.keys(), n_authors): selected_samples[author] = random.sample(samples[author], n_samples) # Cross-validation: iteratively select one sample as the test sample for test_author in selected_samples: for i,test_sample in enumerate(selected_samples[test_author]): # Add remaining samples of the test sample author to the training set X = selected_samples[test_author][:i] + selected_samples[test_author][i+1:] y = [test_author] * (n_samples - 1)
50 3 Experiments # Add samples of other authors to the training set but always # drop one sample at random for a in [x for x in selected_samples if x != test_author]: X.extend(random.sample(selected_samples[a], n_samples - 1)) y.extend([a] * (n_samples - 1)) # Train the classifier and classify the test sample classifier.fit(X, y) predicted = classifier.predict([test_sample]) if predicted[0] == test_author: correct_classifications += 1 print('iteration #{0}: accuracy = {1}'.format( iteration + 1, correct_classifications / (n_samples * n_authors) )) The results of cross-validation are given in FIG. 3.2.17 Since each of the 300 values significantly exceeded the random baseline (for five authors represented by 10 samples, each RB = 0.2; cf. Section 1.4.4), Ijudged versification features to be areliable indicator of atext’s authorship. These results, however, differed greatly across the subcorpora. Generally the models fell into two groups: (1) Highly accurate models (CS1–3, ES1) whose medians ranged from 0.94 to 0.96 and lower quartiles ranged from 0.90 to 0.95 and (2) Accurate enough models (DE1–3, ES2) whose medians ranged from 0.74 to 0.82 and lower quartiles ranged from 0.72 to 0.78. There are many possible reasons for these differences, but they are almost impossible to trace since machine learning generally works like a“black box” (we have access to both the input and the output but what’s going on inside is difficult to interpret). However, one plausible explanation may relate to the amount of data. For authors with alarge number of samples—for example, Jaroslav Vrchlický (281 samples), Lope de Vega (168 samples), Francisco de Quevedo (64 samples), Johann Wolfgang Goethe (46 samples), Barthold Heinrich Brockes (51 samples) and Franz Grillparzer (52 samples)—recognition tended to be less accurate than it was for other authors in the same subcorpus (TAB.3.6). If we assume that the larger an author’s body of work (or more precisely, the longer their career), the greater its stylistic variation, this phenomenon is quite intuitive. 17 Unless stated otherwise, all boxplots in this book have the following format: The box shows the interquartile range (Q0,25; Q0,75); the midway line represents the median (Q0,5); and its value is given in the label rounded to two decimal places. Whiskers represent the minimum and maximum of the distribution.
3 Experiments 51 The presence of two of the mentioned prolific authors in the ES2 subcorpus might help explain why its values were lower than those for ES1. In the case of the German subcorpora, we may also consider the impact of aspecific type of versification (accentual verse) or adifferent method of rhythmic analysis (rhythmic types). Moreover avariety of cultural-historical factors may have been significant.18 These factors are, however, beyond the scope of the present work. 3.2.1 Feature Importance Aside from the performance differences across the subcorpora, it is also worth exploring the contribution of particular features. Failing to do this would leave open the possibility that some of the features were completely irrelevant. The option would remain that purely versification-based features yielded no information at all and the classification depended entirely on sound frequencies. In languages with ahighly phonemic orthography like Czech or Spanish, this would basically mean that the 18 It may generally be assumed, for instance, that Romantic poets put more effort into individualising the rhythm of their poems than Baroque poets did. FIG. 3.2: Cross-validation results for versification-based models (30 iterations with random sampling). subcorpus accuracy 0.940.94 0.94 0.94 0.960.96 0.96 0.96 0.960.96 0.96 0.96 0.820.82 0.82 0.82 0.740.74 0.74 0.74 0.840.84 0.84 0.84 0.940.94 0.94 0.94 0.810.81 0.81 0.81 CS1 CS2 CS3 DE1 DE2 DE3 ES1 ES2 0 1 0.25 0.5 0.75 random baseline
52 3 Experiments Čelakovský Havelka Hněv. Kulda Nejedlý Picek Pohan Tablic Vinařický CS1 Čelakovský 0.91 0.06 0.01 Havelka 0.04 0.95 0.01 0.02 Hněvkovský 0.04 0.94 0.04 0.01 0.06 Kulda 0.99 0.01 0.01 Nejedlý 0.06 0.95 0.01 0.04 Picek 0.01 0.02 10.05 0.01 Pohan 0.01 0.03 0.83 0.02 Tablic 0.01 0.01 0.89 Vinařický 0.02 0.94 Čech Kvapil Mokrý Nečas Sládek Uden Vrchlický CS2 Čech 1 0.08 0.01 0.04 0.03 Kvapil 0.9 0.01 0.05 Mokrý 1 Nečas 1 0.02 Sládek 0.01 0.87 0.04 Uden 0.96 Vrchlický 0.01 0.09 0.86 Klášterský Kvapil Leub. Machar Sova CS3 Klášterský 0.9 0.02 Kvapil 0.07 0.97 Leubner 1 Machar 0.02 0.97 0.03 Sova 0.01 0.03 0.96 de Acunya de Borja de Cetina de Góng. de Herrera ES1 de Acunya 0.94 0.08 de Borja 10.01 de Cetina 0.06 0.92 0.01 de Góngora 0.81 de Herrera 0.17 0.99
3 Experiments 53 Argensola de Quev. de Rojas de Tassis de Ulloa de Vega ES2 Argensola 0.8 0.04 0.02 0.04 0.02 0.07 de Quevedo 0.05 0.72 0.16 0.01 0.06 de Rojas 0.01 0.95 0.17 de Tassis y P. 0.09 0.18 0.78 0.02 de Ulloa y P. 0.04 0.04 0.01 0.98 de Vega 0.02 0.02 0.03 0.68 Brockes Droll. Gott. Kuhl. Neu. Terst. DE1 Brockes 0.76 0.06 0.1 0.04 0.05 0.09 Drollinger 0.07 0.84 0.11 0.05 0.02 Gottsched 0.17 0.1 0.77 0.03 0.1 0.06 Kuhlmann 0.88 Neukirch 0.01 0.83 Tersteegen 0.01 0.85 Goethe Jacobi Müller Pfeffel Wie. DE2 Goethe 0.53 0.07 0.01 0.05 Jacobi 0.22 0.83 0.1 0.04 0.02 Müller 0.19 0.03 0.76 0.04 0.09 Pfeffel 0.03 0.13 0.01 0.85 0.08 Wieland 0.03 0.05 0.06 0.76 Bernhardi Eichen. Grill. Müller Schen. Schu. Tieck DE3 Bernhardi 0.97 Eichendorff 0.88 0.1 0.1 0.03 0.05 0.03 Grillparzer 0.7 Müller 0.01 0.03 0.72 0.02 0.04 0.03 Schenkendorf 0.05 0.08 0.1 0.89 0.02 0.05 Schulze 0.02 0.01 0.04 0.01 0.82 0.1 Tieck 0.01 0.06 0.05 0.07 0.06 0.08 0.79 TAB. 3.6: Confusion matrices for versification-based models (relative counts). Rows show the author predicted by the model while columns show the actual author. Individual cells give the relative count of the relevant prediction.
54 3 Experiments classification was determined by acommon stylometric indicator, that is, by character frequencies. To explore how particular features contributed to the classification, Irepeated the set of experiments described above. In lieu of cross-validation, this time all of the data were used to train the model for each of the 30 iterations with the one-vs.-rest strategy (each iteration, thus, constructed five hyperplanes). In this way, up to 30 hyperplanes were constructed for each author (the final number depended on how many times the author was randomly selected). As discussed in Section 1.4.3 (formula 1.11), the separating hyperplane between two classes is defined by anormal vector wand aparameter b. Each iteration iin which author Aoccurred, thus, produced anormal vector wA,i= (wA,i,1, wA,i,2, …, wA,i,m), whose coordinates conveyed information about the importance of particular features. However, rather than the coordinates themselves, which might be either positive or negative, what mattered here was their absolute value. The importance of the j-th feature (j ∈ [1,m]) for the recognition of Ain iteration iwas, thus, assessed based on the value of wA,i,j squared. The overall importance of j to Aacross all N iterations was then assessed by means of ascore calculated as follows: = =∑ 2 ,, , 1 NAi j Aj i w sN (3.1) Finally, for each A, Icollected the 30 features with the highest scores (i.e. the features that generally contributed most to author recognition). As the total number of these features was in the hundreds, Iregrouped them into the categories given in Chapter 2. TAB. 3.7 shows the distributions of the 30 highestscoring features across these groups. r-2-gram r-3-gram r-4-gram rh-pos rh-snds rh-stress rh-word snds-f CS1 Čelakovský 0.07 0.13 0.13 0.1 0.5 0.07 Havelka 0.1 0.1 0.1 0.03 0.43 0.07 0.07 0.1 Hněvkovský 0.07 0.07 0.03 0.1 0.57 0.03 0.13 Kulda 0.13 0.13 0.13 0.13 0.33 0.13 Nejedlý 0.33 0.07 0.13 0.07 0.33 0.03 0.03 0.3 Picek 0.17 0.2 0.23 0.1 0.2 0.03 0.03 0.03 Pohan 0.07 0.17 0.23 0.07 0.43 0.03 Tablic 0.07 0.1 0.13 0.03 0.4 0.13 0.1 0.03 Vinařický 0.1 0.13 0.17 0.03 0.23 0.07 0.07 0.2 CS2 Čech 0.03 0.13 0.63 0.03 0.03 0.13 Kvapil 0.07 0.13 0.23 0.07 0.27 0.03 0.2 Mokrý 0.13 0.2 0.3 0.2 0.03 0.03 0.1 Nečas 0.07 0.13 0.2 0.03 0.33 0.03 0.03 0.17
3 Experiments 55 r-2-gram r-3-gram r-4-gram rh-pos rh-snds rh-stress rh-word snds-f CS2 Sládek 0.03 0.1 0.2 0.07 0.43 0.03 0.03 0.1 Uden 0.03 0.1 0.2 0.1 0.37 0.03 0.17 Vrchlický 0 0.07 0.13 0.1 0.47 0.03 0.03 0.13 CS3 Klášterský 0.2 0.2 0.27 0.03 0.17 0.13 Kvapil 0.1 0.1 0.07 0.13 0.4 0.03 0.17 Leubner 0.13 0.13 0.23 0.1 0.23 0.07 0.1 Machar 0.13 0.13 0.1 0.1 0.17 0.1 0.1 0.17 Sova 0.07 0.17 0.3 0.07 0.3 0.07 ES1 de Acunya 0.03 0.1 0.13 0.33 0.13 0.27 de Borja 0.23 0.27 0.3 0.1 0.1 de Cetina 0.1 0.2 0.3 0.17 0.07 0.17 de Góngora 0.07 0.03 0.1 0.27 0.27 0.03 0.23 de Herrera 0.03 0.1 0.2 0.17 0.23 0.27 ES2 Argensola 0.07 0.23 0.4 0.07 0.23 de Quevedo 0.13 0.1 0.13 0.27 0.13 0.03 0.2 de Rojas 0.17 0.13 0.27 0.17 0.1 0.17 de Tassis y P. 0.07 0.03 0.07 0.2 0.3 0.03 0.3 de Ulloa y P. 0.1 0.33 0.17 0.1 0.17 de Vega 0.13 0.13 0.07 0.27 0.2 0.2 DE1 Brockes 0.23 0.43 0.23 0.03 Drollinger 0.1 0.37 0.3 0.03 0.07 0.17 Gottsched 0.07 0.5 0.23 0.03 0.13 Kuhlmann 0.3 0.2 0.27 0.07 0.2 Neukirch 0.37 0.5 0.03 0.03 0.07 Tersteegen 0.13 0.5 0.13 0.03 0.03 0.2 DE2 Goethe 0.23 0.5 0.07 0.03 0.13 Jacobi 0.27 0.27 0.27 0.03 0.03 0.17 Müller 0.3 0.37 0.17 0.13 Pfeffel 0.2 0.33 0.17 0.17 0.03 0.13 Wieland 0.53 0.33 0.03 0.03 DE3 Bernhardi 0.3 0.33 0.27 0.03 0.07 0.07 Eichendorff 0.23 0.53 0.13 0.03 0.03 Grillparzer 0.3 0.3 0.2 0.03 0.13 Müller 0.33 0.33 0.2 0.03 0.07 0.07 Schenkendorf 0.47 0.17 0.13 0.03 0.07 Schulze 0.37 0.33 0.13 0.17 0.13 Tieck 0.33 0.33 0.07 0.03 0.03 0.23 TAB. 3.7: Feature importance. (1–3) rhythmic n-grams/rhythmic types, (4) morphological characteristics of rhyme, (5) phonic composition of rhyme, (6) stress placement in rhyme, (7) word length in rhyme, (8) sound frequencies. The table shows the share of elements in these categories reflected in the 30 highest-scoring features for each author. The highest value in each row is highlighted in bold.
56 3 Experiments Among the Czech subcorpora, the phonic composition of rhymes tended to be the most prominent category. In contrast, for German works, morphological characteristics played this role, and for the Spanish subcorpora, the results were somewhere in between. Rhythmic characteristics also played an important part in all three corpora. Of the rhythmic n-grams (CS, ES), rhythmic tetragrams were most prominent. The significance of word length and stress placement in rhyme was fairly weak across all the subcorpora. Concerning the stress placement in rhyme, all values were zero in both Spanish subcorpora. The explanation for this was quite simple: one constant of the Spanish hendecasyllable is that the final stress falls on the penultimate syllable: ¡Peñascos Altos, de la mar batidos, rhythm: 0 1 0 1 0 0 0 1 0 1 0 de nubes coronadas las cabezas, rhythm: 0 1 0 0 0 1 0 0 0 1 0 donde se rompen en diversas piezas rhythm: 0 0 0 1 0 0 0 1 0 1 0 cristales espumosos resistidos rhythm: 0 1 0 0 0 1 0 0 0 1 0 (Lope de Vega) There was no exception to this rule across the ES corpus. The null variability of this stress placement on rhyming words, thus, led to its null applicability for classification. On the whole, however, none of the categories appeared dominant and none could be dismissed as irrelevant. 3.3 Comparison with Lexicon-Based Models The goal of the second battery of experiments was to compare the performance of versification-based models with that of models based on standard stylometric features (again for simplicity, these are referred to—albeit imprecisely—as “lexicon-based” models). Through these same tests, Ialso assessed the performance of models combining versification-based and lexicon-based features.
3 Experiments 57 3.3.1 Fine-Tuning Before proceeding with these comparisons, it was necessary to choose adomain (words, lemmata or character n-grams) and the number of types of each feature to be analysed. To find the optimal solution, Ifirst trained and cross-validated many different models and found the best-performing settings. When fine-tuning, it is good practice to employ different datasets to the ones that will be used to measure accuracy. Since in this case, there was no need to limit the poems to any particular metre, plenty of data were available in CS and DE to build alternative subcorpora for validation (denoted here as CS’ and DE’; see TAB. 3.8 for details). This unfortunately was not the case for ES where there was no option but to use ES1 and ES2 themselves for this purpose. The results for those subcorpora, thus, provide only avery general comparison. Subcorpus Era of birth # of authors # of samples CS1’ 1760–1820 32 986 CS2’ 1840–1855 24 1190 CS3’ 1860–1870 27 1476 DE1’ 1650–1699 8 486 DE2’ 1730–1754 10 598 DE3’ 1760–1794 16 1295 TAB. 3.8: Validation of the subcorpora. In training the models, Ifollowed the design sketched above for five randomly selected authors and 10 randomly selected samples (cf. Section 3.2). Over 30 iterations, Ithen performed leave-one-out cross-validation using an SVM with the set of features below: (1) frequencies of the n most common words, (2) frequencies of the n most common lemmata, (3) frequencies of the n most common character bigrams, (4) frequencies of the n most common character trigrams and (5) frequencies of the n most common character tetragrams, where n ∈ {50, 100, 150, …, 2000}. The results (FIG. 3.3) confirmed apattern observed in previous studies, namely that the relationship between the number of types analysed (n) and the attribution accuracy rose sharply, and then, after reaching acertain value, tended to stabilise (cf.Eder 2011; Rybicki-Eder 2011; Smith-Aldridge 2011). While the value appeared similar for
64 3 Experiments FIG. 3.5: Cross-validation of models based on the 150, 250 and 500 most common lemmata (L), versification features (V) and the concatenation of both feature spaces (L ∪ V); 30 iterations with random sampling. CS1 CS2 features accuracy 0.930.93 0.93 0.93 0.940.94 0.94 0.94 0.980.98 0.98 0.98 L V L ∪ V 0.5 0.6 0.7 0.8 0.9 1 accuracy 0.940.94 0.94 0.94 0.940.94 0.94 0.94 0.980.98 0.98 0.98 0.5 0.6 0.7 0.8 0.9 1 features L V L ∪ V accuracy 0.960.96 0.96 0.96 0.940.94 0.94 0.94 11 1 1 0.5 0.6 0.7 0.8 0.9 1 features L V L ∪ V n = 150 n = 250 n = 500 accuracy 0.90.9 0.9 0.9 0.940.94 0.94 0.94 0.980.98 0.98 0.98 0.5 0.6 0.7 0.8 0.9 1 features L V L ∪ V accuracy 0.90.9 0.9 0.9 0.960.96 0.96 0.96 0.980.98 0.98 0.98 0.5 0.6 0.7 0.8 0.9 1 features L V L ∪ V accuracy 0.940.94 0.94 0.94 0.940.94 0.94 0.94 0.980.98 0.98 0.98 0.5 0.6 0.7 0.8 0.9 1 features L V L ∪ V
3 Experiments 65 CS3 DE1 accuracy 0.860.86 0.86 0.86 0.960.96 0.96 0.96 0.980.98 0.98 0.98 0.5 0.6 0.7 0.8 0.9 1 features L V L ∪ V accuracy 0.880.88 0.88 0.88 0.980.98 0.98 0.98 11 1 1 0.5 0.6 0.7 0.8 0.9 1 features L V L ∪ V accuracy 0.88 0.88 0.88 0.96 0.96 0.96 1 1 1 0.5 0.6 0.7 0.8 0.9 1 features L V L ∪ V accuracy 0.940.94 0.94 0.94 0.820.82 0.82 0.82 0.940.94 0.94 0.94 0.5 0.6 0.7 0.8 0.9 1 features L V L ∪ V accuracy 0.980.98 0.98 0.98 0.820.82 0.82 0.82 0.960.96 0.96 0.96 0.5 0.6 0.7 0.8 0.9 1 features L V L ∪ V accuracy 0.980.98 0.98 0.98 0.830.83 0.83 0.83 0.970.97 0.97 0.97 0.5 0.6 0.7 0.8 0.9 1 features L V L ∪ V
66 3 Experiments FIG. 3.5: Cross-validation of models based on the 150, 250 and 500 most common lemmata (L), versification features (V) and the concatenation of both feature spaces (L ∪ V); 30 iterations with random sampling. DE2 DE3 accuracy 0.820.82 0.82 0.82 0.750.75 0.75 0.75 0.860.86 0.86 0.86 0.5 0.6 0.7 0.8 0.9 1 features L V L ∪ V accuracy 0.860.86 0.86 0.86 0.760.76 0.76 0.76 0.90.9 0.9 0.9 0.5 0.6 0.7 0.8 0.9 1 features L V L ∪ V accuracy 0.890.89 0.89 0.89 0.760.76 0.76 0.76 0.920.92 0.92 0.92 0.5 0.6 0.7 0.8 0.9 1 features L V L ∪ V n = 150 n = 250 n = 500 accuracy 0.880.88 0.88 0.88 0.790.79 0.79 0.79 0.90.9 0.9 0.9 0.5 0.6 0.7 0.8 0.9 1 features L V L ∪ V accuracy 0.920.92 0.92 0.92 0.820.82 0.82 0.82 0.930.93 0.93 0.93 0.5 0.6 0.7 0.8 0.9 1 features L V L ∪ V accuracy 0.960.96 0.96 0.96 0.820.82 0.82 0.82 0.970.97 0.97 0.97 0.5 0.6 0.7 0.8 0.9 1 features L V L ∪ V
3 Experiments 67 ES1 ES2 accuracy 0.980.98 0.98 0.98 0.940.94 0.94 0.94 11 1 1 0.5 0.6 0.7 0.8 0.9 1 features L V L ∪ V accuracy 11 1 1 0.940.94 0.94 0.94 11 1 1 0.5 0.6 0.7 0.8 0.9 1 features L V L ∪ V accuracy 11 1 1 0.920.92 0.92 0.92 11 1 1 0.5 0.6 0.7 0.8 0.9 1 features L V L ∪ V accuracy 0.940.94 0.94 0.94 0.820.82 0.82 0.82 0.940.94 0.94 0.94 0.5 0.6 0.7 0.8 0.9 1 features L V L ∪ V accuracy 0.960.96 0.96 0.96 0.840.84 0.84 0.84 0.960.96 0.96 0.96 0.5 0.6 0.7 0.8 0.9 1 features L V L ∪ V accuracy 0.980.98 0.98 0.98 0.860.86 0.86 0.86 11 1 1 0.5 0.6 0.7 0.8 0.9 1 features L V L ∪ V
68 3 Experiments FIG. 3.6: Frequency of ambiguous predictions (lemma-based model predicts adifferent author than versification-based model) per iteration; frequency of correct predictions (both models predict the same author) within all unequivocal predictions; 30 iterations with random sampling. subcorpora frequency of ambiguous predictions 0.120.12 0.12 0.12 0.110.11 0.11 0.11 0.180.18 0.18 0.18 0.20.2 0.2 0.2 0.320.32 0.32 0.32 0.260.26 0.26 0.26 0.080.08 0.08 0.08 0.20.2 0.2 0.2 CS1 CS2 CS3 DE1 DE2 DE3 ES1 ES2 0 0.1 0.2 0.3 0.4 0.5 0.6 frequency of ambiguous predictions 0.080.08 0.08 0.08 0.150.15 0.15 0.15 0.170.17 0.17 0.17 0.160.16 0.16 0.16 0.280.28 0.28 0.28 0.250.25 0.25 0.25 0.080.08 0.08 0.08 0.190.19 0.19 0.19 0 0.1 0.2 0.3 0.4 0.5 0.6 subcorpora CS1 CS2 CS3 DE1 DE2 DE3 ES1 ES2 frequency of ambiguous predictions 0.10.1 0.1 0.1 0.090.09 0.09 0.09 0.160.16 0.16 0.16 0.180.18 0.18 0.18 0.250.25 0.25 0.25 0.20.2 0.2 0.2 0.060.06 0.06 0.06 0.20.2 0.2 0.2 0 0.1 0.2 0.3 0.4 0.5 0.6 subcorpora CS1 CS2 CS3 DE1 DE2 DE3 ES1 ES2 n = 150 n = 250 n = 500 accuracy of unequivocal predictions 11 1 1 11 1 1 11 1 1 0.970.97 0.97 0.97 0.970.97 0.97 0.97 11 1 1 11 1 1 11 1 1 0.6 0.7 0.8 0.9 1 subcorpora CS1 CS2 CS3 DE1 DE2 DE3 ES1 ES2 accuracy of unequivocal predictions 11 1 1 11 1 1 11 1 10.980.98 0.98 0.98 0.970.97 0.97 0.97 11 1 1 11 1 1 11 1 1 0.6 0.7 0.8 0.9 1 subcorpora CS1 CS2 CS3 DE1 DE2 DE3 ES1 ES2 accuracy of unequivocal predictions 11 1 1 11 1 1 11 1 1 11 1 1 0.970.97 0.97 0.97 11 1 1 11 1 1 11 1 1 0.6 0.7 0.8 0.9 1 subcorpora CS1 CS2 CS3 DE1 DE2 DE3 ES1 ES2
https://doi.org/10.14712/9788024648903.5 4 Applications In this final chapter, Iapply the approaches described in this book to two cases of ambiguous or disputed authorship of poetic works. These cases concern English and Russian texts respectively. In the first, Itrace the relative contributions of William Shakespeare and John Fletcher to the play The Two Noble Kinsmen. Then, in the second, Icollaborate with Artjoms Šeļa to investigate the suspected forgery of poems first published in a1978 edition of Gavriil Batenkov’s works. 4.1 The Two Noble Kinsmen The play The Two Noble Kinsmen (TNK) was recorded in the Stationers’ Register in 1634 and published in aquarto edition later that year. In both cases, John Fletcher and William Shakespeare were indicated as the play’s authors. No manuscript has been preserved. Attempts to discern which parts were likely written by each author must therefore rely solely on intratextual indicators. Since the 19th century, researchers have found evidence at various textual levels to suggest that Shakespeare was mostly responsible for Acts 1 and 5 while Fletcher was mostly responsible for Acts 2, 3 and4.20 While there is not much controversy about this general picture, the authorship of certain scenes is still being debated. In what follows, Iseek to contribute to this debate using acombination of versification-based and word-based models. The case of TNK is closely linked to that of another play which was also supposedly co-authored by Shakespeare and Fletcher—The Famous History of the Life of King Henry the Eight. Ihave discussed the authorship of that work elsewhere (Plecháč 2020). Here Ifollow the design of that study and apply the same models to classify passages from TNK. 20 Adetailed history of TNK’s attributions is given in Vickers 2004.
70 4 Applications 4.1.1 History and Related Works The first attempt to provide ascene-by-scene division of TNK between Shakespeare and Fletcher was made by Henry Weber (1812). Based on his observations of enjambments, weak endings, unusual words and metaphors, Weber assigned all of Act 1 and most of Act 5 to Shakespeare and all of Act 2 and most of Acts 3 and 4 to Fletcher (see TAB. 4.1 for details of this and other attributions). Slightly different attributions were proposed by William Spalding (1833) and Samuel Hickson (1847), both of whom relied on observations similar to those of Weber. An important advance came with the publication in the 1874 Transactions of the New Shakspere Society of three articles about the play which instead of merely observing distinctive features sought to quantify them: Frederick Gard Fleay (1874d) measured the number of weak endings and four-feet lines in particular scenes; Frederick James Furnivall (1874c) considered the number of enjambments (the stopt-line test); and John Kells Ingram (1874) applied his weak-ending test (see Section 1.1). All three articles supported Hickson’s division with only one exception—Act 1, scene 2 was now assigned solely to Shakespeare. Just afew years later, Robert Boyle (1882) presented anew theory which claimed that the “Shakespearian” parts had in fact been written by Philip Massinger or—in two cases—by I II P 1 2 3 4 5 1 2 3 4 5 6 Weber 1812 NS S S S S F F F F F F Spalding 1833 NS S S S S F F F F F F Hickson 1847 NS SF S S S S F F F F F Fleay 1874 NS S S S S S F F F F F Boyle 1882 NM M M M M M F F F F F Oliphant 1891 NFSM SM SM SM ? S F F F F F Farnham 1916 NS S S S S F F N F F F Hart 1934 NS S S S S F F F F F F Oras 1953 NS S S N N N F F F F F Hoy 1962 NS S S S S S F F F F F Horton 1987 NS S S N N S F S N F N Matthews-Merriam 1993 N S F Ledger-Merriam 1994 FS S S S ? S F F F F ? Tarlinskaja 2014 NS S S F F F F F Eisen et al. 2017 NS S S S F S F ? F F F TAB. 4.1: Selected attributions of TNK. Sdenotes an attribution of the scene to Shakespeare, F to Fletcher and M to Massinger; N denotes an unassigned scene.
4 Applications 71 III IV V 1 2 3 4 5 6 1 2 3 1 2 3 4 E S S F F F F F F S S F S S N S F F F F F F F F S F S S N S S F F F F F F S S F S S N S S F F F F F F S S F S S N SM SM F F F F F F M M F M M N S S F F F F F F FS FS F S S F S ? F F F F F F S ? F S S N S F F F F F F F F S F S S N S ? F F F F F F S S F S S N S S F F F F F F F FS F S S N S N F N ? F ? ? S S ? S S N F S S N S ? S ? F F S S S S F S S ? SFFFFFFS SF S S N S S F F F F F F S S N S S N Shakespeare and Massinger together. Massinger’s participation was also backed by Henry Oliphant (1891) although he pointed to different scenes to those named by Boyle. Twentieth-century studies generally supported the Shakespeare–Fletcher division that preceded Boyle or else proposed only slight modifications. These works included studies of contractions (Farnham 1916), vocabulary richness (Hart 1934), line endings (Oras 1953) and spelling differences (Hoy 1962). This Shakespeare–Fletcher split has also largely been maintained by more recent scholars. Based on adiscriminant analysis of three sets of function words, Thomas Horton (1987) attributed most scenes in the play to Shakespeare or else left them undecided. Robert Matthews and Thomas Merriam (1993) classified entire acts of TNK using aneural network that had been familiarised with the frequencies of function words in the respective plays of Shakespeare and Fletcher. Ayear later, Merriam reopened the case in astudy with Gerard Ledger which used ahierarchical cluster analysis based on character frequencies; this time the goal was the attribution of particular scenes (Ledger and Merriam 1994). More recently, Marina Tarlinskaja (2014) has applied acomplex versification analysis using features of the kind enumerated in Section1.5. Mark Eisen, Alejandro Riberio, Santiago Segarra and Gabriel Egan (2017) have also used word adjacency networks (Segarra, Eisen and Riberio 2013) to analyse the frequencies of collocations of selected function words in particular scenes of the play.
72 4 Applications 4.1.2 Attribution of Particular Scenes Since the external evidence clearly pointed to Shakespeare and Fletcher’s joint authorship of TNK and previous analyses had ruled out Massinger’s participation on linguistic grounds, Ilimited the candidate set to Shakespeare and Fletcher. Ithen set out to determine the most likely author of particular scenes. To train the models, Iused four plays by Shakespeare and four plays by Fletcher that all dated roughly from the period when TNK was supposedly written (1613–1614). Each scene in these plays was treated as asingle training sample except for those containing fewer than 100 verse lines. This gave me: — Shakespeare: The Tragedy of Coriolanus (5 scenes), The Tragedy of Cymbeline (10scenes), The Winter’s Tale (7 scenes), The Tempest (6 scenes) and — Fletcher: Valentinian (12 scenes), Monsieur Thomas (10 scenes), The Woman’s Prize (14scenes), Bonduca (14 scenes).21 Altogether there were, thus, 28 training samples for Shakespeare and 50 training samples for Fletcher. Having established alarge enough training set, Inow risked employing amethod that might produce rather sparse data: First Iused the frequencies of particular rhythmic types to capture the rhythmic style of the data (cf. Section 2.1.2).22 No rhyme characteristics were considered since all of the plays were written in blank verse and rhymes were, thus, only exceptional. To capture vocabulary, Irelied on word frequencies since words had proven to be amore reliable indicator than lemmata at the pilot testing stage. For both rhythmic types and words, Ilimited the analysis to the 500 most frequent types. An SVM with alinear kernel was used as aclassifier. To estimate the model’s accuracy, Iperformed the following cross-validation: — To avoid overfitting—apotential risk of testing amodel on scenes from the play it was trained with—Idid not perform standard k-fold cross-validation. Instead, Iclassified scenes from each play using amodel trained with the rest of the plays. As such, scenes from Shakespeare’s Coriolanus were classified by amodel 21 For both the training data and the text of TNK itself, Irelied on XML versions of the first editions of the plays, as provided by the EarlyPrint project (https://drama.earlyprint.org). To eliminate spelling variation, regularised spellings (the “reg” attribute of the w-element) were used. All of Shakespeare’s texts came from the First Folio (1623). All of Fletcher’s texts came from the first Beaumont and Fletcher folio (1647), except for Monsieur Thomas for which the 1639 quarto was used. For TNK, Irelied on the 1634 quarto edition. 22 Rhythmic annotation was provided by the Prosodic Python library (https://github.com/ quadrismegistus/prosodic).
4 Applications 73 trained with scenes from the other three plays by Shakespeare and four plays by Fletcher; 27 scenes from Cymbeline were classified similarly and so on. — Since the training data were imbalanced and there was, thus, arisk of bias, Ialigned the number of training samples per author using random selection. — To obtain more representative results, the entire process was repeated 30 times with anew random selection in each iteration; this generated 30 classifications of each scene. — To compare the attribution power of both feature subsets, cross-validation was performed not only on the combined models (500 rhythmic types ∪ 500 words) but also on the versification-based models (500 rhythmic types) and word-based models (500 words) alone. As TAB. 4.2 shows, both versification-based and word-based models proved highly accurate in distinguishing the respective works of Shakespeare and Fletcher. The only issues with the versification-based models were one misattribution of Act 3, scene5 of Cymbeline to Fletcher and two misattributions of Act 5, scene 8 of Valentinian to Shakespeare. In contrast, the word-based models misclassified Act 5, scene 1 of Bondu ca in all 30 iterations. When the two feature sets were merged, however, there were no misclassifications and all models achieved 100% accuracy. FIG. 4.1 presents the results of the application of classifiers to TNK. As with the training samples, testing was limited to scenes with more than 100 lines (12 out of the play’s 24 scenes). Except in the case of Act 4, scene 1, there was astrong consensus among the versification-based, word-based and combined models. Significantly, their Rhyt. type-based models Word-based models Combination models Shakespeare Coriolanus 1 1 1 Cymbeline 0.997 1 1 The Winter’s Tale 1 1 1 The Tempest 1 1 1 Fletcher Valentinian 0.992 1 1 Monsieur Thomas 1 1 1 The Woman’s Prize 1 1 1 Bonduca 1 0.93 1 TAB. 4.2: Accuracy of authorship recognition by models based on (1) the 500 most common rhythmic types, (2) the 500 most common words and (3) 1000-dimensional vectors combining features (1) and (2). Figures show the share of correctly classified scenes over all 30 iterations.
80 4 Applications classification.23 This is also true of line endings (Fleay 1874d). Common strong ending rhythmic types such as 0101010101 (0 = unstressed syllable, 1 = stressed syllable), 0101000101 and 0100010101 were among the most strongly-weighted positive (Shakespearian) features. Similarly, common W-position-terminated rhythmic types such as 01010101010 and 10010100010 appeared among the most strongly-weighted negative (Fletcherian) features. 4.2 The Case of (Pseudo-)Batenkov: Towards a Formal Proof of Literary Forgery (co-authored by Artjoms Šeļa) In 1978, ascholarly monograph about the poetry of G.S.Batenkov (1793–1863) was published in Moscow under the title Poezia dekabrista Gavriila Stepanovicha Batenkova (Iliushin 1978). Its author was A.A.Iliushin. What appeared to be acomplete collection of Batenkov’s poems was appended to the volume. Batenkov, a Russian officer and poet, had fought in the Napoleonic wars and later worked as an engineer and policymaker. His eclectic ideological interests, which ranged from freemasonry and Christian mysticism to political reform, led him to join secret societies and eventually become associated with the Decembrist revolt of 1825. This effectively ended his life as a free citizen of the Empire. He was sentenced to 25years of solitary confinement in the Peter and Paul Fortress in Saint Petersburg and, after serving 20 years, exiled to Siberia. Iliushin, who was both the author of the monograph and the editor of the appendix, was aRussian versification scholar and poetry specialist. He also wrote poetry himself and was known in academic circles for his literary games and imitations. The majority of Batenkov’s late poems (i.e. those written after his release from prison) appeared for the very first time in this collection. There was, however, one major problem: the source of these texts was inaccessible and their origins unverifiable. Iliushin himself referred to amanuscript that was listed as lost in the archives ( Shapir2000). For 20 years, no one publicly questioned the authenticity of these poems. This all changed when the scholar M.I.Shapir published aseries of studies in the late 1990s that showed that there were indeed grounds for doubt. Shapir (1997, 1998) conducted 23 This appraisal is based on the mean value for feature importance in 30 combined models trained with 100-line samples taken from the training set (four plays by Shakespeare, four plays by Fletcher).
4 Applications 81 an extensive quantitative analysis of the poems in the controversial section of Batenkov’s work (we refer to these texts as the “disputed poems”). To this end, he meticulously examined every linguistic level—prosody, metrics, morphology, syntax and semantics—and pointed out many significant differences between these texts and Batenkov’s known works. Among the issues Shapir observed in the disputed poems were their abundance of inexact rhymes, overly archaic morphology, discrepancies in the use of pronouns and conjunctions and some possible anachronisms. To date, his analysis remains one of the most impressive non-computational authorship attribution studies of Russian poetry. This research convinced many scholars that the disputed poems were in fact forgeries (Gasparov and Tarlinskaja 2008; Tarlinskaja 2014). Indeed, in the years since, this consensus has become so strong that the editors of arecently published collection of Iliushin’s original poems did not hesitate to include all of the disputed poems in the volume (Iliushin 2020). However, this interpretation is at odds with Shapir’s own conclusion: having uncovered significant differences at some textual levels but striking similarities at others, he judged that there was not enough evidence to draw any conclusions about the origins of the disputed poems. This reasoning led Shapir to an important generalisation about the limitations of using formal and linguistic methods to determine authorship. If, as he argued, we cannot trace the identity of an author based on various levels of linguistic features, then the concept of the “author” who makes linguistic choices that are unique and recognizable is nothing more than ascholarly construct. From amodern-day perspective, Shapir’s strong statements lack methodological support. Compared with other scholars who have used versification features for authorship attribution (Tomashevsky 1923/2008; Lotman and Lotman 1986; Tarlinskaja 2014), Shapir dramatically increased the number of textual levels under investigation. Nevertheless, his analysis remained univariate: all of the levels were treated in isolation and the features were compared one by one. It might be said, then, that Shapir’s inquiry was multivariate in scope but he lacked the tools to deal with multivariate and seemingly contradictory signals. As aresult, he could not estimate the compound authorial signal in either Batenkov’s known works or the disputed poems. Key questions went unaddressed: How important were the differences in the frequency of inexact rhymes or function words compared, say, with similarities in the rhythmic structure of iambic tetrameter and use of formulae? In the final part of this book, we return to this question that Shapir left unsolved. Our aim is to reach amore definitive conclusion about the authorship of the disputed poems using amultivariate approach that combines lexical and versification features. We break the problem down into the following experiments:
82 4 Applications — We first test the general performance of our approach using 19th-century Russian poetry data. — We then formulate the task as averification problem. The goal here is not to find the most probable candidate from afinite set but rather to verify the likelihood that Batenkov’s poems and the disputed poems were produced by asingle author. — Finally, we compare the disputed poems not only to Batenkov’s established works but also to Iliushin’s own poems. The task is, thus, reformulated as aclassification problem. 4.2.1 Features Afull-scale replication of Shapir’s study cannot be undertaken with large corpora because of the limitations of automated text analysis and scansion. We therefore confine our analysis to three levels: — Vocabulary modelled by lemmata frequencies (with lemmatisation provided by MyStem 3.1, https://yandex.ru/dev/mystem/); — Morphology modelled indirectly by character 3-grams (excluding punctuation and including blank spaces); — Versification modelled by the rhyme features described in Section 2.2 (rhyme recognition provided by RhymeTagger (Plecháč 2018); IPA transcription provided by Espeak, http://espeak.sourceforge.net/). We do not consider rhythmic features because of the scarcity of lines in any particular metre in the data for either Batenkov or pseudo-Batenkov. 4.2.2 Fine-Tuning Our first goal is to determine the most efficient feature space. To do this, we train multiple models with the following sets: (1) frequencies of the n most common lemmata (L), (2) frequencies of the n most common character 3-grams (G), (3) frequencies of the n most common lemmata and the n most common character 3-grams (LG) and (4) frequencies of the n most common lemmata and the n most common character 3-grams enriched with rhyme characteristics (LGR).
4 Applications 83 This is done for 40 different values of the most common types: n ∈ {50, 100, 150, …, 2000}. Here we use acorpus of Russian poems whose composition dates to the 1820s. This is partitioned into 200-line samples. Multiple poems can be combined in asingle sample, and no poem contributes to more than one sample. This generates: — 19 samples by Yevgeny Baratynsky, — 23 samples by Mikhail Lermontov, — 60 samples by Alexander Pushkin, — 12 samples by Pyotr Vyazemsky, — 36 samples by Nikolay Yazykov and — 11 samples by Vasily Zhukovsky. We apply the two different classifiers that will be used in subsequent experiments: linear SVM and cosine Delta. To train the models, we follow the design laid out in Section 3.2. with five randomly selected authors and 10 randomly selected samples. Over 30 iterations, we perform cross-validation for the SVM model and nearest neighbour classification with the Delta approach. The results can be seen in FIG. 4.4. The performance is similar to those recorded for other languages (Chapter 3): for all of the feature sets, accuracy generally increases to approximately the level of the 1000 most common types. At that point, it stabilises. For both classifiers, the LG combination tends to significantly outperform both L and G on their own. Even greater accuracy is almost always achieved, however, when rhyme features are also taken into account (LGR). In the next set of experiments, we therefore retain LGR-based models and choose the 1000 most common types as the optimal level. 4.2.3 The One-Class Problem (Authorship Verification) So far all the tasks we have considered in this book have involved authorship classification. In this situation, there is aclosed set of candidates {A1, A2, A3, …, An} and the goal is to determine which one is most likely the author of the text(s) X.In contrast, authorship verification deals with adifferent scenario. Here it is not possible to determine aclosed set that we are sure includes the real author. The goal is instead to decide whether acertain Ais or is not the author of X. The Batenkov case needs to be treated first and foremost as averification problem. If there are doubts about the origin of the disputed texts, then we first need to
84 4 Applications determine how likely it is that Batenkov himself wrote them regardless of Iliushin’s status as apotential author. Here we loosely apply the unmasking technique (Koppel and Schler 2004; Koppel et al. 2007). In its classic version, this technique makes aseries of pairwise SVM classifications between same-author and other-author samples. It then iteratively drops the most distinctive features from the learning process. Compared to other verification techniques such as those based on entropy or deep learning (Halvani et al. 2019), unmasking stands out for its clear assumptions and production of interpretable results. Unmasking assumes that text samples from the same author will share deeper similarities than the samples of two different authors. In the former case, there may still be differences but they will emerge from high-level features such as theme, chronology or genre and not from the underlying style. Moreover, such features will inevitably be exploited by machine classification. That is why the original unmasking method relies on several stages of classification: in each iteration, acertain number of the most distinctive features are dropped and the classification is performed again. Given their underlying similarity, same-author samples should, thus, quickly become indistinguishable from one another while other-author samples retain their differences across many iterations. This is because their “distinctiveness” is distributed over many features and not concentrated in afew high-performing ones. FIG. 4.4: Accuracy of (a) the SVM model and (b) the cosine Delta model with the most common lemmata (L), the most common character 3-grams (G), the L and G combination (LG) and the L and G combination enriched by rhyme features (LGR) across different levels of the most common types. number of most common types accuracy L G LG LGR 50 200 350 500 650 800 950 1100 1250 1400 1550 1700 1850 2000 0.4 0.5 0.6 0.7 0.8 0.9 number of most common types accuracy L G LG LGR 50 200 350 500 650 800 950 1100 1250 1400 1550 1700 1850 2000 0.4 0.5 0.6 0.7 0.8 0.9 (b)(a)
4 Applications 85 Since multiple poems can be combined in asingle sample and no poem contributes to more than one sample, there is no reason to suppose that any high-level features distinguish the works of asingle author. We therefore tweak the classic unmasking process by asking asimple question: Can the known Batenkov poems be distinguished from the disputed poems in apairwise SVM classification? To gauge the accuracy of this technique, we also test it on acontrol group of works published by other Russian poets in the 1840s and 1850s (i.e. the period when the majority of the disputed texts had allegedly been written). Like the Batenkov poems and the disputed poems, these works are divided into 100-line samples. (A200 line size would generate only three samples from both Batenkov’s work and the disputed poems). This produces: — 13 samples by Mikhail Lermontov, — 14 samples by Fyodor Tyutchev, — 18 samples by Pyotr Vyazemsky, — 15 samples by Nikolay Yazykov, — six samples by Gavriil Batenkov and — six samples from the disputed poems. We then follow the four steps below: (1) Randomly select 12 samples from each of the four “control” authors. (2) Randomly split each group of 12 samples in half. These two groups are the A-samples and B-samples. (3) Use the A-samples and the LGR feature set to train SVM models for each possible pair of “control” authors (i.e. Lermontov vs. Tyutchev, Lermontov vs. Vyazemsky, through to Vyazemsky vs. Yazykov). Perform leave-one-out cross-validation of each model. (4) Train the SVM models with the LGR feature set for each “control” author using his own A-samples and B-samples as separate classes (i.e. Lermontov (A) vs. Lermontov (B), Tyutchev (A) vs. Tyutchev (B), Vyazemsky (A) vs. Vyazemsky (B), Yazykov (A) vs. Yazykov(B)). Perform leave-one-out cross-validation of each model. We repeat this entire process 30 times for each quantity of the most common types: n ∈ {50, 100, 150, …, 1000}. Anew set of randomly selected samples is used in each iteration. For each n, we therefore obtain 4 × 30 = 120 accuracy estimations for samples written by the same author and ( ) 4 2 × 30 = 180 accuracy estimations for samples written by different authors. Finally, for each n, we also cross-validate the Batenkov poems against the disputed poems model.
86 4 Applications FIG. 4.5 shows the results. Th e “control” authors behave as might be expected. Th e median classifi cation accuracy for same-author pairs (A-samples vs. B-samples) hover s around 50%, meaning that on average they are indistinguishable for aclassifi er. At the same time, accuracy remain s high for the pairwise classifi cation of diff erent authors as well. Th e dashed line in FIG 4.5. represents the classifi cation accuracy for Batenkov poems vs. disputed poems. Without exception, this line follow s the general trend for texts from two diff erent sources. Although these results seem fairly convincing on their own, we wish to go one step further and interpret them in terms of probabilities. As there appear s to be no significant divergence among diff erent quantities of the most common words (except perhaps when using the lowest values to classify diff erent authors), we merge all of the values to obtain accuracy estimations for: (1) same-author classifi cations, (2)diff erent-author classifi cations and (3) Batenkov poems vs. disputed poems classifi cations. AMann-Whitney test24 show s that the probability of these outcomes if Batenkov was not theauthor of the disputed poems is 0.9265 (U= 111, n1 = 3600, n2 = 20). In contrast, if Batenkov was the author, the probability is less than 10–14 (U= 60618, n1 = 2400, n2 = 20). 24 As there are always 12 samples, there are only 12 possible outcomes of cross-validation. The variable in question is, thus, not continuous but discrete. We therefore opt for the non-parametric Mann-Whitney test over the perhaps more expected t-test. FIG. 4.5: Accuracy of pairwise classifications for different quantities of the most common feature types. Boxplots depict the median, the interquartile range (box) and the 5th-to-95th percentile range (whiskers). number of most common types accuracy different authors same author Batenkov vs. Pseudo-Batenkov 50 100 150 200 250 300 350 400 450 500 550 600 650 700 750 800 850 900 950 1000 0 0.25 0.5 0.75 1
4 Applications 87 4.2.4 The Two-Class Problem (Batenkov vs. Iliushin) There is, however, afly in the ointment. As we have observed, Batenkov’s poems spanned the 1810s to the 1860s with asignificant gap from 1825 to 1846 when he was in solitary confinement (see FIG. 4.6 for amore detailed depiction of this output). The disputed poems date almost entirely from the period after his imprisonment. We therefore cannot rule out ascenario also raised by Shapir: during Batenkov’s confinement, there might have been adramatic change in his writing style which would explain the irregularities in the disputed poems. To address this objection, the disputed poems have to be compared with Batenkov’s later poems alone. Unfortunately, there are not enough data to perform apairwise SVM experiment with only the poems that Batenkov published after his release. We therefore need to switch to the less data-hungry Delta method. We depend here especially on the cosine variation, which has proven to be the most reliable technique with our “control” authors. The problem is, thus, reframed as aclassification task. To begin, we increase the sample size to 200 lines. This produces the following numbers of samples per author: FIG. 4.6: Batenkov’s poems and the disputed poems according to their (supposed) composition dates. decade number of lines Batenkov disputed poems 1810 1820 1830 1840 1850 1860 0 100 200 300 400 500 600
88 4 Applications — Mikhail Lermontov (8), — Fyodor Tyutchev (8), — Pyotr Vyazemsky (12), — Nikolay Yazykov (8), — Gavriil Batenkov (2) and — disputed poems (3). Over multiple experiments with different feature space settings, the disputed poems remain clustered with Batenkov’s poems. This does not say much about the Iliushin hypothesis, however, since the suspected author is not included in the candidate set (if, on the other hand, the disputed poems and Batenkov’s poems did not cluster together, this might be interpreted as strong evidence of aforgery). Although Iliushin never published any poems under his own name, preferring to mask his authorship of non-academic works, several texts have been attributed to him by consensus. These include Дедушка и девушка (published as an anonymous poem), Michele Trivolis— Максим Грек and Добрый вампир (both published under the name Y.F.Sidorin) and Тайная дочь декабриста Бесстужева… (the so-called PseudoGrigo riev, which was presented as awork by the poet A.Grigoriev, 1822–1864). All of these are long narrative poems from which it is possible to extract asample comparable to those used in our past experiments. Now we add the (apparent) Iliushin samples to the corpus and perform another battery of experiments. The quantity of most common types is set to 1000 for both lemmata and character 3-grams. To verify the robustness of these results, we perform 10,000 classifications; in each iteration, 0–1000 types of each feature are dropped from the classification (both the quantity of types and the features themselves are randomly selected). The results are summarised in aconfusion matrix (TAB. 4.3). Batenkov Iliushin Lermontov disputed poems Tyutchev Vyazemsky Yazykov Batenkov 1 0.06 Iliushin 0.99 0.01 0.21 0.09 Lermontov 0.89 0.03 disputed poems 0.02 0.73 Tyutchev 0.03 0.95 0.01 Vyazemsky 0.01 0.04 0.01 0.89 Yazykov 0.01 0.01 1 TAB. 4.3: Confusion matrix (relative counts). Rows represent the author predicted by the model while columns represent the actual author. Individual cells show the relative number of predictions in each case.
4 Applications 89 In over 20% of the vector spaces, one sample of the disputed poems appears to be closer to Iliushin’s poems than to the other disputed poem samples. This is completely unlike the pattern with the other authors, which showed only minimal variation across the predictions. Interestingly enough, all of these “misattributions” of the disputed poems to Iliushin concern just two of his samples. These are both poems published under the name Y.F.Sidorin. This, in turn, raises aquestion: Do these works differ somehow from the other two Iliushin samples? There are indeed several differences beginning at the level of metre. The Sidorini - an poems are written in iambic pentameter, one of the most common metres in Russian poetry in the first half of the 19th century; in contrast, Дедушка и девушка is loosely trochaic and “PseudoGrigoriev” is dactylic. Clearly, vocabulary, morphology and rhyme structure can all be profoundly affected by the choice of metre as well. Acloser look at the Sidorinian poems yields even more information. FIG. 4.7 shows the cosine distances across various quantities (50, 100, 150, …, 2000) of the most common types when the disputed poems are compared with (i) the Sidorinian poems and (ii) Batenkov’s own poems published between the 1840s and the 1860s. In all of the number of most common types cosine distance to the disputed poems Batenkov Ilyushin 50 200 350 500 650 800 950 1100 1250 1400 1550 1700 1850 2000 1.075 1.1 1.125 1.15 1.175 1.2 1.225 FIG. 4.7: Cosine distances between the disputed poems and (1) the Sidorinian poems (Iliushin) and (2) Batenkov’s own poems published between the 1840s and the 1860s for different quantities of the most common types.
96 References Navarro-Colorado, B. (2017). Ametrical scansion system for fixed-metre Spanish poetry. Digital Scholarship in the Humanities, 33(1), 112–127. doi: 10.1093/llc/fqx009. Navarro-Colorado, B., Ribes-Lafoz, M. and Sánchez, N. (2016). Metrical annotation of alarge corpus of Spanish sonnets. Representation, scansion and evaluation. Proceedings of the Tenth International Conference on Language Resources and Evaluation (LREC 2016). Portorož. Oliphant, E.H.C. (1891). The Works of Beaumont and Fletcher. Englische Studien 15, 321–360. Oras, A. (1953). “Extra monosyllables” in Henry VIII and the problem of authorship. Journal of English and Germanic Philology, 52, 198–213. Platt, J. (1999). Probabilistic outputs for support vector machines and comparisons to regularized likelihood methods. Advances in Large Margin Classifiers, 10(3), 61–74. Plecháč, P. (2016). Czech verse processing system KVĚTA: Phonetic and metrical components. Glottotheory, 7, 159–174. doi: 10.1515/glot-2016-0013. Plecháč, P. (2018). Acollocation-driven method of discovering rhymes (in Czech, English, and French poetry). In M.Fidler and V.Cvrček (eds.), Taming the Corpus. From Inflection and Lexis to Interpretation. Cham: Springer, 79–95. doi: 10.1007/978-3-319-98017-1_5. Plecháč, P. (2020). Relative contributions of Shakespeare and Fletcher in Henry VIII: An analysis based on most frequent words and most frequent rhythmic patterns. Digital Scholarship in the Humanities [advance articles]. doi: 10.1093/llc/fqaa032. Plecháč, P. and Birnbaum, D. (2019). Assessing the reliability of stress as afeature of authorship attribution in syllabic and accentual syllabic verse. In P.Plecháč, B.P.Scherr, T.Skulacheva, H.Bermúdez-Sabel, R.Kolár (eds.), Quantitative Approaches to Versification. Prague: ICL CAS, 201–210. Plecháč, P., Bobenhausen, K. and Hammerich, B. (2018). Versification and authorship attribution. Pilot study on Czech, German, Spanish, and English Poetry. Studia Metrica et Poetica, 5(2), 29–54. doi: 10.12697/smp.2018.5.2.02. Plecháč, P., Kolár, R. (2015). The Corpus of Czech Verse. Studia Metrica et Poetica, 2(1), 107–118. doi: 10.12697/smp.2015.2.1.05. Porter, J.D. (2018). Popularity/Prestige. Pamphlets of Stanford Literary Lab, 17. Rybicki, J. and Eder, M. (2011). Deeper Delta across genres and languages: Do we really need the most frequent words? Literary and Linguistic Computing, 26(3), 315–321. doi: 10.1093/llc/fqr031. Savoy, J. (2020). Advanced models for stylometric application. In: Machine Learning Methods for Stylometry. Authorship Attribution and Author Profiling. Cham: Springer, 153–187. doi: 10.1007/978-3-030-53360-1_7. Schmid, H. (1994). Probabilistic part-of-speech tagging using decision trees. In Proceedings of International Conference on New Methods in Language Processing. Manchester. Segarra, S., Eisen, M. and Riberio, A. (2013). Authorship attribution using function words adjacency networks. In International Conference on Acoustics, Speech and Signal Processing (ICASSP). IEEE, 5563–5567. doi: 10.1109/ICASSP.2013.6638728. Shapir, M.I. (1997). Fenomen Batenkova iproblema mistifikatsii (lingvistikhovedcheskij aspekt 1–2). Philologica, 4, 85–144. Shapir, M.I. (1998). Fenomen Batenkova iproblema mistifikatsii (lingvistikhovedcheskij aspekt 3–4). Philologica, 5, 49–132. Shapir, M.I. (2000). Fenomen Batenkova iproblema mistifikatsii. In M.I.Shapir. Universum versus: Iazyk— stikh— smysl vrusskoi poezii XVIII-XIX vekov, 1. Moscow: Yazyki russkoi kultury, 335–443.
References 97 Sherman, L.A. (1888). Some observations upon sentence-length in English prose. The University of Nebraska Studies, 1(4), 337–366. Skoumalová, H. (2011). Porovnání úspěšnosti tagování korpusu. In V.Petkevič and A.Rosen (eds.), Korpusová lingvistika Praha 2011/3. Gramatika aznačkování korpusů. Prague: Nakladatelství Lidové Noviny, 199–207. Smith, P.W.H. and Aldridge, W. (2011). Improving authorship attribution: Optimizing Burrows’ Delta method. Journal of Quantitative Linguistics, 18(1), 63–88. doi: 10.1080/09296174.2011.533591. Sobchuk O. (2018). Charting Artistic Evolution: An Essay in Theory. Dissertation. Tartu: University of Tartu Press, 2018. Spalding, W. (1833). ALetter on Shakespeare’s Authorship of The Two Noble Kinsmen. Edinburgh: A. & C.Black. Spedding, J. (1850). Who wrote Shakespeare’s Henry VIII? The Gentlemen’s Magazine, 115–123. Spoustová, D., Hajič, J., Votrubec, J., Krbec, P. and Květoň, P. (2007). The best of two worlds: Cooperation of statistical and rule-based taggers for Czech. Proceedings of the Workshop on Balto-Slavonic Natural Language Processing. ACL, 67–74. Tabata, T. (2012). Approaching Dickens’ style through random forests. Digital Humanities 2012: Conference Abstracts. Hamburg: Universität Hamburg, 388–391. Tarlinskaja, M. (1987). Shakespeare’s Verse: Iambic Pentameter and the Poet’s Idiosyncrasies. New York: Peter Lang. Tarlinskaja, M. (2014). Shakespeare and the Versification of English Drama, 1561–1642. Farnham et al.: Ashgate. Thorndike, A. H. (1901). The Influence of Beaumont and Fletcher on Shakespeare. Worcester: Oliver B. Wood. Tomashevsky, B.V. (1923/2008). Pjatistopnyj jamb Pushkina. In Izbrannye raboty ostikhe. Moscow & Sankt Peterburg: Akademija, 140–242. Vickers, B. (2004). Shakespeare, Co-Author. AHistorical Study of Five Collaborative Plays. Oxford: Oxford University Press. doi: 10.1093/acprof:oso/9780199269167.001.0001. Weber, H. (1812). Observations on the participation of Shakespeare in The Two Noble Kinsmen. In H.Weber (ed.), The Works of Beaumont and Fletcher in Fourteen Volumes 13. Edinburgh: J.Ballantyne & Co., 151–169. Williams, C.B. (1975). Mendenhall’s studies of word-length distribution in the works of Shakespeare and Bacon. Biometrika, 62(1), 207–212. doi: 10.1093/biomet/62.1.207. Yule, G.U. (1939). On sentence-length as astatistical characteristic of style in prose, with application to two cases of disputed authorship. Biometrika, 30, 363–390. Yule, G.U. (1944). The Statistical Study of Literary Vocabulary. Cambridge: Cambridge University Press. Zhao, Y. and Zobel, J. (2005). Effective and scalable authorship attribution using function words. In G.G.Lee et al. (eds.), Information Retrieval Technology. AIRS 2005. Lecture Notes in Computer Science. Berlin and Heidelberg: Springer, 174–189. doi: 10.1007/11562382_14. Zipf, G.K. (1932). Selected Studies on the Principle of Relative Frequency in Language. Cambridge, MA: Harvard Uniersity Press.