scieee Open visual document viewer

A framework for the Comparative analysis of text summarization techniques

Ghosh, Trijit

Abstract

We see that with the boom of information technology and IOT (Internet of things), the size of information which is basically data is increasing at an alarming rate. This information can always be harnessed and if channeled into the right direction, we can always find meaningful information. But the problem is this data is not always numerical and there would be problems where the data would be completely textual, and some meaning has to be derived from it. If one would have to go through these texts manually, it would take hours or even days to get a concise and meaningful information out of the text. This is where a need for an automatic summarizer arises easing manual intervention, reducing time and cost but at the same time retaining the key information held by these texts. In the recent years, new methods and approaches have been developed which would help us to do so. These approaches are implemented in lot of domains, for example, Search engines provide snippets as document previews, while news websites produce shortened descriptions of news subjects, usually as headlines, to make surfing easier. Broadly speaking, there are mainly two ways of text summarization – extractive and abstractive summarization. Extractive summarization is the approach in which important sections of the whole text are filtered out to form the condensed form of the text. While the abstractive summarization is the approach in which the text as a whole is interpreted and examined and after discerning the meaning of the text, sentences are generated by the model itself describing the important points in a concise way.

Full text

i A amewo k o he Compa a i e analysis o ex summa iza ion echniques T iji Ghosh Disse a ion p esen ed as pa ial equi emen o ob aining he mas e ’s deg ee in Da a Science and Ad anced Analy ics 2 Ins i u o Supe io de Es a ís ica e Ges ão de In o mação Uni e sidade No a de Lisboa A FRAMEWORK FOR THE COMPARATIVE ANALYSIS OF TEXT SUMMARIZATION TECHNIQUES by T iji Ghosh (M20170009) Disse a ion p esen ed as pa ial equi emen o ob aining he mas e ’s deg ee in Da a Science and Ad anced Analy ics Ad iso / Co Ad iso : P o esso Rica do Rei; P o esso Robe o Hen iques July 2021 3 ACKNOWLEDGEMENTS I would i s like o hank my hesis ad iso P o esso Doc o Rica do Rei o he NOVA In o ma ion Managemen School a Uni e sidade NOVA de Lisboa as he was he one who challenged me o a heme as inno a i e as he one ha ga e mo o o his mas e ’s hesis on ex summa iza ion. I wan o hank him o encou aging me and mo i a ing me e en when he ime o de o e o his mas e ’s hesis was no wha was wan ed and expec ed. Finally, a big hank you o my pa en s o encou aging me and gi ing me he chance o do his mas e ’s in a eas as in e es ing and exci ing as ad anced analy ics a e. Tha ga e me he possibili y o ha e a ca ee ha I ha e d eamed o . I would no be possible wi hou hei suppo . 4 5 Con en s 1. In oduc ion ...................................................................................................................................... 8 1.1. Backg ound ............................................................................................................................... 8 1.2. Mo i a ion ................................................................................................................................. 8 1.3. Objec i e .................................................................................................................................... 8 2. EXTRACTIVE SUMMARIZATION ............................................................................................... 9 2.1. In e media e Rep esen a ion ............................................................................................ 9 2.2. Sen ence Sco e ......................................................................................................................... 9 2.3. Summa y Sen ences Selec ion ........................................................................................... 9 3. TOPIC REPRESENTATION APPROACHES .......................................................................... 11 3.1. Topic Wo ds .......................................................................................................................... 11 3.2. F equency-d i en App oaches ....................................................................................... 11 3.3. La en Seman ic Analysis ................................................................................................. 14 3.4. Bayesian Topic Models ...................................................................................................... 14 3.5. BERT ......................................................................................................................................... 15 4. THE IMPACT OF CONTEXT IN SUMMARIZATION ........................................................... 20 4.1. Web Summa iza ion ........................................................................................................... 20 4.2. Scien i ic A icles Summa iza ion ................................................................................. 20 4.3. Email Summa iza ion ........................................................................................................ 21 5. METHODOLOGY ........................................................................................................................... 35 5.1. Design Sea ch Resea ch .................................................................................................... 35 5.2. S a egy ................................................................................................................................... 37 6. PROPOSAL o a amewo k on scena ios o ex summa iza ion echniques ...... 38 6.1. PROPOSAL .............................................................................................................................. 38 6.2. VALIDATION .......................................................................................................................... 38 7. CONCLUSIONS ............................................................................................................................... 46 8. Re e ences ...................................................................................................................................... 47 6 Lis o Tables Table 1 ............................................................................................................................................................... 21 Table 2 ............................................................................................................................................................... 27 Table 3 ............................................................................................................................................................... 38 Table 4 ............................................................................................................................................................... 42 Table 5 ............................................................................................................................................................... 43 7 Lis o igu es Figu e 1 : Weigh ed Te ms /s Speci ici y .......................................................................................... 13 Figu e 2 : Be Embeddings ...................................................................................................................... 16 Figu e 3 : A chi ec u e o BERT .............................................................................................................. 17 Figu e 4 : Encode s and Decode s ......................................................................................................... 17 Figu e 5 : O e all p e- aining and ine- uning p ocedu es o BERT ..................................... 18 Figu e 6 : Fine Tuning phase .................................................................................................................... 18 Figu e 7 : Accu acy o BERTbase on Masked LM and Le - o-Righ ............................................. 19 Figu e 8 : P ecision and Recall o di e en ex iles ...................................................................... 40 Figu e 9 : P ecision, Recall and F-Measu e o di e en alues o k applying LSA ............ 41 Figu e 10 : O e all Compa ison o he me hods ............................................................................... 44 8 1. INTRODUCTION 1.1. BACKGROUND We see ha wi h he boom o in o ma ion echnology and IOT (In e ne o hings), he size o in o ma ion which is basically da a is inc easing a an ala ming a e. This in o ma ion can always be ha nessed and i channeled in o he igh di ec ion, we can always ind meaning ul in o ma ion. Bu he p oblem is his da a is no always nume ical and he e would be p oblems whe e he da a would be comple ely ex ual, and some meaning has o be de i ed om i . I one would ha e o go h ough hese ex s manually, i would ake hou s o e en days o ge a concise and meaning ul in o ma ion ou o he ex . This is whe e a need o an au oma ic summa ize a ises easing manual in e en ion, educing ime and cos bu a he same ime e aining he key in o ma ion held by hese ex s. In he ecen yea s, new me hods and app oaches ha e been de eloped which would help us o do so. These app oaches a e implemen ed in lo o domains, o example, Sea ch engines p o ide snippe s as documen p e iews, while news websi es p oduce sho ened desc ip ions o news subjec s, usually as headlines, o make su ing easie . B oadly speaking, he e a e mainly wo ways o ex summa iza ion – ex ac i e and abs ac i e summa iza ion. Ex ac i e summa iza ion is he app oach in which impo an sec ions o he whole ex a e il e ed ou o o m he condensed o m o he ex . While he abs ac i e summa iza ion is he app oach in which he ex as a whole is in e p e ed and examined and a e disce ning he meaning o he ex , sen ences a e gene a ed by he model i sel desc ibing he impo an poin s in a concise way. 1.2. MOTIVATION As he In e ne has g own in popula i y, a as amoun o in o ma ion has become a ailable. Summa izing as amoun s o ex is challenging o humans. In his age o in o ma ion o e load, au oma ic summa izing echnologies a e in high demand. We will y o ocus on a ious ex ac ion me hodologies o single and mul i-documen summa iza ion in his hesis. Some o he mos o en used me hods, such as opic ep esen a ion app oaches, equency- d i en me hods, g aph-based and machine lea ning echniques, will be desc ibed. E en hough i is di icul o ho oughly explain all o he many algo i hms and app oaches in his hesis, we will y o gi e a good o e iew o ecen ends and b eak h oughs in au oma ic summa izing me hods, as well as discuss he s a e-o - he-a and compa e a ious ways. 1.3. OBJECTIVE The objec i e o his pape is o p o ide compa a i e analysis o di e en echniques o ex summa iza ion used in di e en scena ios, me hodology - o de ine a se o analysis pa ame e ha can allow us o classi y di e en echniques e.g., complexi y, accu acy and speed. 9 2. EXTRACTIVE SUMMARIZATION Ex ac i e summa iza ion, as men ioned abo e, chooses pe inen subse s om he ex gi en, based on some me ics and is hus combined o a condensed o m. To unde s and how summa iza ion sys ems wo k, we desc ibe h ee, ai ly, independen asks which all summa ize s pe o m: 1) Build a ansi ional depic ion o he inpu ex which communica es he mos impo an aspec s o he ex . 2) Sco e he sen ences suppo ing he ep esen a ion. 3) Choose a summa y consis ing o a ie y o ex s. 2.1. INTERMEDIATE REPRESENTATION All hese summa iza ion echniques will de elop some in e media e ep esen a ions o he gi en inpu ex suppo ed ce ain me ics and disce n he impo an sen ences suppo ed hose me ics. The e a e wo o ms o app oaches suppo ed he ep esen a ion: opic ep esen a ion and indica o ep esen a ion. Topic ep esen a ion me hods emodel he ex in o a ansi ional cha ac e iza ion and examine he opic(s) gi en wi hin he ex . This me hod di e s in e ms o o mula ion and he eby, complexi y, and a e di ided in o equency-d i en app oaches, opic wo d app oaches, la en seman ic analysis and Bayesian opic models. We ake a deepe look in o opic ep esen a ion app oaches wi hin he ollowing sec ions. Indica o ep esen a ion app oaches desc ibe e e y sen ence as a lis ing o ea u es (indica o s) o impo ance like sen ence leng h, posi ion wi hin he documen , ha ing ce ain ph ases, e c. 2.2. SENTENCE SCORE When he inpu ex is ans o med in o a o m which he model in e p e s, a sco e is assigned o e e y sen ence based on ha me ic. This sco e is jus a ep esen a ion o how impo an a sen ence is. These indica o s (me ic) a e de i ed om ma hema ical o machine lea ning models. Once he sco e o e e y o he sen ences a e ob ained, hey a e, hen, agg ega ed o ed in o a unc ion which anks hese sen ences based on he sco es ob ained. hey' e also e e ed o as indica o weigh s. 2.3. SUMMARY SENTENCES SELECTION A e he sen ences a e sco ed and anked suppo ed he a ious me ics o indica o s, he impo an sen ences a e il e ed ou . These p ocesses can use di e en algo i hms o sepa a e he sen ences suppo ed he anks and a ew o he sco es as an example edundancy sco e. Some me hods use g eedy algo i hms o sea ch ou he simples sen ences ep esen ing he essence o he ex . This me hod will no always be he mos e ec i e app oach which igno es 16 Figu e 2 : Be Embeddings The inpu ex is i s ed in o he ex p ocessing embeddings namely: - Posi ion Embeddings Segmen Embeddings Token Embeddings 17 Figu e 3 : A chi ec u e o BERT The ou pu o hese embedding is hen ed o he BERT laye which consis s o ans o me s. Figu e 4 : Encode s and Decode s A ans o me can consis o 12/24 blocks o encode s wi h 12/16 a en ion heads and 110/340 million pa ame e s, namely BERTbase/BERTla ge espec i ely. I looked closely, each ans o me can be a se o encode s and decode s as shown abo e in he diag am. I is he e ha he ou pu o he ans o me is ed o he summa iza ion laye . P e- aining has been qui e c ucial when i came o language models. The e ha e been 18 applica ions o hese models such as na u al language in e ence and pa aph asing. This pape (De lin, Chang, Lee, & Tou ano a, 2018), mainly alks abou ine uning app oach wi h he p oposal o BERT as desc ibed ea lie . The eason i is called bidi ec ional is because unlike o he models whe e he ex is ead om le o igh o om igh o le , he BERT app oach eads he sen ences in bo h he di ec ions and ies o unde s and he con ex o he wo ds bo h o i s le and igh . Figu e 5 : O e all p e- aining and ine- uning p ocedu es o BERT In he abo e diag am, he a chi ec u e o he BERT is shown whe e apa om he ou pu laye s, bo h use he same a chi ec u e. Figu e 6 : Fine Tuning phase Sou ce: (De lin, Chang, Lee, & Tou ano a, 2018) The abo e igu e shows ha by wo king on he ine- uning phase o he model, he model was able o achie e an accu acy by an o e whelming ma gin o 4.5% and 7% p io o i s s a e o he a . 19 Figu e 7 : Accu acy o BERTbase on Masked LM and Le - o-Righ Sou ce: (De lin, Chang, Lee, & Tou ano a, 2018) Al hough i is o be men ioned ha his bi-di ec ional app oach akes longe han unidi ec ional app oach which is wha has been shown in he abo e diag am whe e he BERT MLM me hod con e ges slowe han le o igh bu accu acy wise, i is well ahead o he o he app oach. This bi-di ec ional app oach is, de ini ely, use ul al hough a bi slow and makes he applica ion o BERT o a wide ange o ields. 20 4. THE IMPACT OF CONTEXT IN SUMMARIZATION I is qui e e iden ha a con ex is e y use ul when i comes o unde s anding he in o ma ion p esen ed in a ex . In he same way, i models can be empowe ed wi h he in o ma ion o con ex , hen a summa ize sys em would be able o p une he co ec in o ma ion o summa ize he ex . Fo example, acco ding o (Allahya i, e al., 2017), when summa izing blogs, he deba es o commen s ha ollow he blog pos a e use ul esou ces o de e mining which po ions o he blog a e c i ical and in iguing. The e is a signi ican quan i y o in o ma ion a ailable in scien i ic pape summa ies, such as published a icles and con e ence in o ma ion, ha can be used o highligh essen ial sen ences in he o iginal wo k. 4.1. WEB SUMMARIZATION I we look a he web pages, we will see ha hey ha e lo o objec s which is no always possible o summa ize o example, pic u es, gi s, and some unwan ed ma e ials like ad e isemen s which is so no ele an o he in o ma ion gi en in he pages. Fo hose kinds o si ua ions, i will be help ul o use he links which di ec us o he page o be summa ized. These links would gi e he model a knowledge o he con ex and would be help ul o p o ide imp o ed summa y o he page. In (Ami ay & Pa is, 2000), whe e hey alk abou websi e summu iza ion o he i s ime, hey came up wi h a concep called “THE INCOMMONSENSE SYSTEM”. This model inco po a es a web c awling sys em which c awls all he links ha link back o he cu en page and his is how he con ex is de i ed. Since hen, a lo o di e en algo i hms ha e been de eloped which was based on he abo e p inciple bu has beeen imp o ed. 4.2. SCIENTIFIC ARTICLES SUMMARIZATION In (Mei & Zhai, 2008), a summu iza ion p oblem ela ed o scien i ic pape s is s udied and p esen ed. Needless, o say wi h each passing yea , new disco e ies a e being made and new esea ch pape s a e being published e e y yea . So wi h his, he daun ing ask o b ie ing a scien i ic pape s becomes eally challenging. Specially when i comes o including a con ex in he summa y which e e ences a mul i ude o di e en esea ch pape s. In o de o sol e his p oblem, hey came wi h a concep o impac based summu iza ion as men ioned in (Mei & Zhai, 2008). This me hod le e ages on he impo ance o sen ence sco e which has been used in he o iginal pape using he KL di e gence me hod (i.e., inding he simila i y be ween a sen ence and he language model). The conclusions made in his pape a e subs an ial po ing ha his me hod was use ul and could be used o u u e b ei ing models o scien i ic pape s.They p oposed a language model ha gi es a p obabili y o each wo d in he ci a ion con ex sen ences. They hen sco e he impo ance o sen ences in he o iginal pape . . 21 4.3. EMAIL SUMMARIZATION When i comes o email summa iza ion, he ex s become a bi di e en . In o de o ind a con ex , a whole chain mail h eads a e o be p ocessed o de e mine he s o y. In (Nenko a & Bagga, 2003), hey discuss a ew me hods o how he con e sa ional na u e o he ex can be used o ga he con ex . Thei me hods does p o ide a conclusi e e idence as o how use ul his me hod could be and his u he , his me hod could be enhanced by using some isualiza ion echniques as well. While in (Rambow, Sh es ha, Chen, & Lau idsen, 2004), hey ake a bi o a di e en app oach whe e hey de e mine impo an ea u es o summa izing he email, whe e each ea u e could be one h ead o se e al h eads o email. (Newman & Bli ze , 2003) discusses a whole new app oach o summa izing an email. They le e age he clus e ing algo i hm o sol e he p oblem o email summa iza ion. In hei pape hey alk abou ew clus e ing algo i hms which help hem see all he h eads o he email as a whole and pos applica ion o he men ioned algo i hms, an o e iew is o med. Below, Table 2. ci e he jou nals o he e e ences om which some echniques we e analyzed and u he mo e, hei bene i s o de ec s a e gi en. Table 2. Illus a es he a ious me hods which we e explained in he a icles men ioned in Table 1 22 21 Table 1 Jou nal/Con e ence Sub Topic Yea s (xxxx- xxyy) # A icles #A icles Techniques Bene i s #A icles Techniques D awbacks 2017 In e na ional Con e ence on Compu ing Me hodologies and Communica ion (ICCMC) Tex Summa iza ion 2017 Au oma ic ex summa iza ion by local sco ing and anking o imp o ing cohe ence I has made use o sen ence ea u e me ics o sco e sen ences like “sen ence o sen ence cohesion”, “wo d equency”, i makes use o me ics which migh igno e aluable in o ma ion and he eby making he summa y less meaning ul. Fo example, me ic like “leng h o sen ence” is used o a oid selec ing oo sho o oo long sen ences o he documen . Doing his, a imes, he model migh o e look some in o ma ion which migh ha e been p esen in hose sen ences bu we e no aken in o accoun while summa izing he ex . 22 2017 In e na ional Con e ence on Big Da a, IoT and Da a Science (BID) Tex Summa iza ion 2017 Au oma ic ex summa iza ion o news a icles The lexical chain gene a ion p oposed by Silbe and McCoy algo i hm has linea un ime complexi y. Fu he , ce ain issues we e esol ed in bo h algo i hms by implemen ing p onoun esolu ion and enhanced sen ence sco ing o le e age he s uc u e o news a icles. One o he lexical chain gene a ion algo i hm adop ed was p oposed by (Ba zilay & Elhadad, 2000) has exponen ial un ime complexi y. A i icial In elligence Re iew a chi e Volume 47 Issue 1, Janua y 2017 Pages 1- 66 Tex Summa iza ion 2017 Recen au oma ic ex summa iza ion echniques: a su ey This pape alks abou a a ie y o echniques which ha e hei own bene i s. 1. T ained Summa ize and La en Sema ic Analysis – uses a modi ied co pus-based app oach and a TRM echnique based on la en seman ic analysis. The summa ize is based on a unc ion ha assesses main sen ences/wo ds o hings like loca ion, keywo d, likeness o i le, and cen ali y in o de o gene a e summa ies. The sco e unc ion is op imized using s ochas ic echniques such as gene ic algo i hms. 2. In o ma ion e ie al pe o mance has g ea ly imp o ed The e a e some d awbacks o he app oaches used which a e as ollows:T ained Summa ize and La en Sema ic analysis – he summa ies gene a ed a e no e y consis en wi h opic and he sen ences don’ co ela e so much a imes. Fea u e weigh s o sco e unc ion p oduced by GA ail o consis en ly gi e he bes esul s o he es co pus. Ob aining he app op ia e dimension educ ion a io and explaining LSA e ec s a e ough in he LSA+TRM echnique. Mo eo e , he ime complexi y o compu e SVD is qui e high. 2. Using a sen ence-based abs ac ion echnique o ex ac da a – In his app oach, only casual cohe ence is conside ed whe eas 23 because o he use o a sen ence-based abs ac ion echnique. Sen ences ha ep esen he cen al no ion a e linked. 3. Unde s anding and summa izing documen s using a documen concep la ice – In compa ison o exis ing sen ence g ouping and sen ence sco ing algo i hms, he sugges ed app oach pe o ms excep ionally well. 4. Sen ence ex ac ion using ex summa iza ion based on con ex and s a is ics – This space and ime which a e in e - ela ed links which makes sense ou o he documen as a whole a e also equi ed o ep esen ing beha io al con ex . 3. Th ough he use o a documen concep la ice, i is possible o comp ehend and summa ize ex – ime complexi y o gene a ing a DCI is high because i conside s all possible combina ions. 4. Sen ence ex ac ion h ough con ex ual and s a is ical based summa iza ion ex . 30 documen summa ies h ough non - nega i e ma ix ac o iza ion his app oach could no na u ally ca ch he meaning o seman ic ea u es ha a e highly spa se and ha e a limi ed iew o meaning. As a esul , summa iza ion sys ems based on LSA a e unable o choose meaning ul ph ases. As a esul , elemen s o seman ic ea u e ec o s in he sugges ed me hod exclusi ely con ain non-nega i e alues and a e also ex emely spa se, allowing seman ic cha ac e is ics o be easily ead. A sen ence can be ep esen ed by a linea combina ion o ce ain signi ican seman ic elemen s. As a esul , sub opics in a documen can be quickly iden i ied, and he e's a be e p obabili y o ex ac ing ele an lines. A me hod o picking ph ases o cons uc gene al documen summa ies is sugges ed using NMF, in which a con en is i s p e-p ocessed and hen summa ized. To gene a e a non-nega i e seman ic ea u e ma ix, NMF is used o a e m-by-sen ence ma ix. Fo each sen ence, gene ic ele ance is calcula ed, which indica es how much a sen ence explains. Que y based summa iza ion o mul iple documen s by applying eg ession models 2011 Ouyang Y Ouyang, Li, Li, & Lu, 2011 (Ouyang, Li, Li, & Lu, 2011) sugges ed a me hod o anking ph ases in que y-based summa iza ion o nume ous manusc ip s using eg ession models. Th ee que y- dependen ea u es, such as named-en i y ma ching, wo d-ma ching, and seman ic ma ching, and ou que y-independen ea u es, such as sen ence posi ion, named en i y, wo d TF-IDF, and s op-wo d penal y, a e used in his me hodology o choose main sen ences in que y-based summa iza ion o mul iple documen s. To begin wi h, human summa ies gene a e " alse" aining da a. Then, using di e en me hods based on he N-g am me hodology ha calcula e "nea ly ue" ele ance a ings o ph ases a e c ea ed and analyzed using his aining da a and hei collec ion o ex s, and a mapping unc ion is lea ned using his aining da a ia a collec ion o p e iously speci ied ea u es o sen ences. Then, using his lea ned unc ion, he signi icance o sen ences in he es da a is es ima ed. An e icien da a collec ion o aining da a o lea ning eg ession models equi es wo hings: (a) an app op ia e g oup o opics wi h co ec ly handw i en summa ies, and (b) a good app oach o compu ing he ele ance o wo ds. The Maximal Ma ginal Rele ance (MMR) echnique is used o emo e edundancy om he summa y. Au oma ic ex summa iza ion using MR, GA, 2009 Mohamed Abdel Fa ah, Fuji Ren (Fa ah and Ren 2009) Wi h he use o a ew s a is ical ea u es, (Fa ah & Ren, 2009) sugges ed an app oach o imp o e con en selec ion in au oma ic ex summa iza ion. As a ainable summa ize , his me hod elies on dis inc s a is ical aspec s in each sen ence o gene a e summa ies. Posi ion o Sen ence (Pos), + e keywo d, - e keywo d, + e 31 FFNN, GMM and PNN based models keywo d, + e keywo d, + e keywo d, + e keywo d, + e keywo d, + e keywo d, + e keywo d, + e keywo d, + e keywo d, R2T, Cen ali y o Sen ence (Cen), P esence o Name En i y in Sen ence (PNE), P esence o Numbe s in Sen ence (PN), Bushy Pa h o Sen ence (BP), Rela i e Leng h o Sen ence (RL), and Agg ega e Simila i y (AS) a e all measu es o sen ence simila i y. Gene ic Algo i hm (GA) and Ma hema ical Reg ession (MR) models ha e been ained o acqui e an op imal mix o ea u e weigh s by mixing all o hese ea u es. Fo sen ence ca ego iza ion, eed o wa d neu al ne wo ks (FFNN) and p obabilis ic neu al ne wo ks (PNN) a e u ilized. Some ex ea u es, such as he + e and - e keywo ds, a e language-dependen , while eigh o he s a e no . All o he abo e-men ioned a iables a e aken in o accoun when calcula ing a sen ence's weigh ed sco e unc ion. All sen ences in a documen a e anked in dec easing o de o hei sco es, and a highly sco ed clus e o sen ences is u ilized o gene a e a summa y o he con en using a ious comp ession a es (10, 20, 30 pe cen used he e). Conclusion: The esul s demons a e ha ea u e BP is he mos essen ial ex ea u e since i p oduces he bes esul s, while ea u e PND p oduces he wo s esul s because s a is ical da a is absen om eligious and poli ical pieces. Because i could model a bi a y densi ies, he GMM app oach p oduced he bes esul s o all he s a egies. Maximum co e age and minimum edundancy in summa iza ion o ex 2011 Algulie , R. M Algulie , Aliguliye , Haji ahimo a, & Mehdiye , 2011 (Algulie , Aliguliye , Haji ahimo a, & Mehdiye , 2011) in oduced an unsupe ised summa izing model o gene ic ex as an In ege Linea P og amming p oblem (ILP) ha immedia ely de ec s essen ial sen ences om he a icle as well as he ull a icle's ele an in o ma ion. Maximum Co e age and Minimum Redundancy is he name o his s a egy (MCMR). This me hod aims o imp o e h ee key aspec s o a summa y: (a) ele ance, (b) edundancy, and (c) leng h. A subse o sen ences om he documen collec ion's ele an ex is picked. Then, using NGD-based simila i y (No malized Google Dis ance) and cosine simila i y, simila i y be ween he summa y and he documen collec ion is compu ed, and his simila i y mus be maximized. An objec i e unc ion is de eloped and mus be maximized o ensu e ha he summa y con ains he impo an con en ound in he documen collec ion and ha he summa y does no con ain a signi ican numbe o ph ases ha communica e he same in o ma ion. A he same ime, he leng h o he summa y mus be limi ed. Las ly, an empi ical unc ion is c ea ed by linea ly combining he cosine simila i y- 32 based and NGD-based simila i y empi ical unc ions, and his combined empi ical unc ion mus be maximized. This echnique o summa izing is inco po a ed as an op imiza ion p oblem ha aims o ind a global solu ion o he p oblem. The B anch & Bound algo i hm (B&B) and he Bina y Swa m Op imiza ion me hod a e he algo i hms used o add ess he ILP p oblem. Conclusion: This me hod, which combines MCMR wi h he B&B algo i hm, su passes all o he s. I demons a es ha summa izing ou comes a e dependen on simila i y measu emen s. I is also p o ed h ough es s ha using cosine simila i y and NGD-based simila i y me ics oge he p oduces be e esul s han using hem sepa a ely. Summa iza ion o documen s h ough a p og essi e echnique o selec ion o sen ences 2013 Ouyang Y Ouyang, Li, Zhang, Li, & Lu, 2013 (Ouyang, Li, Zhang, Li, & Lu, 2013) p oposed a new p og essi e echnique o gene a ing a summa y based on he selec ion o "no el and salien " sen ences. Subsuming ela ionship be ween wo sen ences, i.e., an i egula ela ionship be ween sen ences ha shows he le el o ecommenda ion o one ph ase by ano he . In o de o asce ain he link be ween wo ph ases, he ela ionship be ween hei concep s mus be ound. The associa ion be ween concep s is hen disco e ed by using a co e age-based measu e o disco e he ela ionship be ween wo ds. A Di ec Acyclic G aph (DAG) is used o o ganize all o he wo ds ha appea in he ound wo d ela ions. A p og essi e s a egy o sen ence selec ion is c ea ed on he basis o an asymme ic ela ionship be ween sen ences, in which a sen ence is ei he picked as a no el gene al s a emen o as a suppo ing sen ence. The ollowing wo me hods a e used o choose new and ele an sen ences in his me hod: (a) disco e ed concep s a e only included du ing he assessmen o sen ence ele ance o assu e sen ence o iginali y, and (b) o now, he ela ionship be ween sen ences is used o imp o e he saliency measu e. To implemen his s a egy, a andom walk on he DAG om he cen al node o i s nea by nodes is pe o med, wi h he goal o co e ing he cen al wo ds i s and hen eaching he g ea es amoun o wo ds ia wo d ela ions. Redundancy is elimina ed by punishing epe i i e wo ds, esul ing in esh concep s being in oduced each ime a new ph ase is chosen. Conclusion: In e ms o gene a ing summa ies wi h imp o ed saliency and co e age, he P og essi e sys em su passes he adi ional Sequen ial app oach. E alua ion o 2013 Fe ei a, Ra ael; Cab al, Luciano de Fe ei a, e al., In he ecen decade, (Fe ei a, e al., 2013) inco po a ed i een sco ing echniques ha had been e e enced in he esea ch. ROUGE (Lin 2004) is used o quan i a i e 33 sen ence sco ing me hods o ex ac i e summa iza ion o ex Souza; Lins, Ra ael Duei e; Sil a, Gab iel Pe ei a e; F ei as, F ed; Ca alcan i, Geo ge D.C.; Lima, Rinaldo; Simske, S e en J.; Fa a o, Luciano 20132013 e alua ion, while he numbe o sen ences ha a e simila among he machine- gene a ed and human-made summa ies is coun ed o quali a i e e alua ion. The p ocessing ime o each algo i hm is aken in o accoun . Wo d sco ing, sen ence sco ing, and g aph sco ing app oaches a e used o pick ele an sen ences. The mos essen ial e ms a e gi en sco es in he wo d sco ing app oach. Wo d equency, TF/IDF, uppe case, p ope noun, wo d co-occu ence, and lexical simila i y a e among he app oaches used o sco e wo ds. The p ope ies o sen ences a e examined in he sen ence sco ing app oach. The exis ence o cues, nume ical da a, sen ence leng h, sen ence posi ion, and sen ence cen ali y a e all ac o s in sen ence sco ing. Sco es a e de e mined using he g aph sco ing app oach by looking a he ela ionships be ween sen ences. Tex ank, bushy pa h o he node, and agg ega e simila i y a e all g aph sco ing app oaches. The six common conce ns o s op wo ds, s uc u al ans o ma ion, compa able seman ics, ambigui y, edundancy, and co- e e ence a e hen explo ed, along wi h some sugges ions o ad ancing sen ence sco e ou comes. Explo ing co ela ions among mul iple e ms h ough a g aph-based summa ize , GRAPHSUM 2013 Ba alis, Elena; Caglie o, Luca; Maho o, Naeem; Fio i, Alessand o Ba alis, Caglie o, Maho o, & Fio i, 2013 GRAPHSUM, a new g aph-based, gene al-pu pose summa ize o summa izing nume ous documen s, was p oposed by (Ba alis, Caglie o, Maho o, & Fio i, 2013). This me hod in es iga es and applies associa ion ules, a da a mining me hodology o inding connec ions be ween se e al e ms. I is no elian on sophis ica ed seman ic models (like axonomies o on ologies). The documen collec ion is o ganized as a ansac ional da ase a e p ep ocessing so ha associa ion ule mining may be conduc ed on i . Then, om he ansac ional da ase , equen ly ecu ing i emse s wi h high co ela ions among he e ms a e iden i ied, and a co ela ion g aph is cons uc ed om hese e ms, which will aid in he selec ion o signi ican lines o he summa y. The Ap io i algo i hm is used o mine equen ly ecu ing i emse s, and he suppo measu e is employed o his job. The li measu e indica es he in ensi y o ela ionship be ween wo e ms and is used o e alua e posi i e o nega i e connec ions be ween commonly used wo ds. A a ia ion o he classic PageRank g aph anking algo i hm is used o de e mine he ele ance o he g aph nodes. The g aph nodes ha ha e a signi ican numbe o posi i e co ela ions a e placed i s , while hose ha ha e a nega i e connec ion wi h he adjacen nodes a e penalized. Fo summa y c ea ion, he sen ences ha a e he mos app op ia e o he co ela ion 34 g aph and ha e a high ele ance sco e a e picked. The g eedy algo i hm is employed o selec sen ences in his case. GRAPHSUM pe o ms be e o e a wide ange o s a e-o - he-a echniques, some o which ely hea ily on highly de eloped seman ic- based models o complica ed language p ocesses. Inco po a ing a ious le els o language analysis o ackling edundancy in ex summa iza ion 2013 Elena Llo e , Manuel Paloma Llo e & Paloma , 2013 (Llo e & Paloma , 2013) p o ided a me hod o de ec ing edundan in o ma ion based on h ee laye s o language analysis: lexical, syn ac ic, and seman ic. Cosine simila i y is u ilized in he lexical based echnique o de ec simila i y be ween sen ences in wo sou ces. Those sen ences ha ha e a cosine simila i y g ea e han a ce ain h eshold a e conside ed epe i i e, and hey a e all elimina ed. In a syn ac ic- based me hod, en ailmen ela ions a e compu ed be ween pai s o ph ases o de e mine whe he he meaning o one sen ence can be deduced om he meaning o he o he sen ence. I a posi i e en ailmen is ob ained, he second sen ence is deemed supe luous and elimina ed. Sen ence alignmen is de e mined a he documen le el be ween a se o linked documen s using a open sou ce a ailable Champollian Tool Ki in a seman ic-based manne . Syn ac ic and seman ic echniques a e p e e able han lexical app oaches ha ely on cosine simila i y. Tex summa iza ion can be done in wo ways. Be o e he ma e ial is summa ized, unnecessa y sen ences a e dele ed in he i s echnique. The se o use ul sen ences is hen gi en o he summa iza ion sys em, which uses s a is ical ( e m equency) and linguis ic (code quan i y p inciple) ac o s o selec essen ial sen ences, as well as a summa y. 35 5. METHODOLOGY This s udy will be conduc ed h ough a quali a i e esea ch, based on a well-s uc u ed pa ame e o compa ing he di e en echniques adop ed o ex summa iza ion echniques. This is he basis o compa e di e en echniques and shed ligh on which echniques would be mo e use ul in (i any) pa icula si ua ions. Wha a e hei d awbacks and ad an ages? 5.1. DESIGN SEARCH RESEARCH Design Science Resea ch is a o m o in es iga ion ha en ails building o imp o ing some hing in a no el way in esponse o a speci ic challenge. The ques o a solu ion based on ex ensi e scien i ic in es iga ion ensu es ha he inal p oposed a i ac is cohe en and c edible. A c ucial phase ha should no be o e looked is good communica ion o he inished p oduc (He ne , Ma ch, Pa k, & Ram, 2004). Each o he six key s ages o DSR me hodology, as shown in Figu e 5, will be discussed in g ea e de ail igh away. Figu e 3. DSR Me hod Adap a ion (Pe e s, Tuunanen, Ro henbe ge , & Cha e jee, 2007) Iden i y p oblem and mo i a ion De ine he esea ch challenge in de ail and jus i y he impo ance o a solu ion. Begin by es ablishing a es able heo y ha leads o a esea ch p oblem by demons a ing o s akeholde s he alue o an e ec i e solu ion and wha hey will gain om i s esul (Pe e s, Tuunanen, Ro henbe ge , & Cha e jee, 2007). De ine objec i es and a solu ion 36 Clea ly de ine goals (quan i a i e o quali a i e) o es ablish he ounda ion o a solu ion based on he p oblem cha ac e iza ion and wha can and canno be done (Pe e s, Tuunanen, Ro henbe ge , & Cha e jee, 2007). Design and De elopmen The goal o he design and de elopmen s ages is o c ea e knowledge h ough he design and de elopmen o he a i ac i sel (G ego & He ne , 2013). This could be accomplished by b eaking down he majo scien i ic p oblem in o smalle componen s (He ne , Ma ch, Pa k, & Ram, 2004). To ha e an e ec i e/ clea s uc u e in he nex phase, i is necessa y o ha e a clea g asp o he solu ion alue and o de end i wi h some heo e ical ounda ion (Pe e s, Tuunanen, Ro henbe ge , & Cha e jee, 2007). A solu ion ha mus mee business equi emen s (He ne , Ma ch, Pa k, & Ram, 2004). To gain he app op ia e heo e ical basis, i is essen ial o do esea ch and collec knowledge abou he p esen s a us o he p oblem and exis ing solu ions, as well as o analyze di ec and indi ec solu ions and hei e icacy (Pe e s, Tuunanen, Ro henbe ge , & Cha e jee, 2007). Wi h he knowledge, i is possible o de elop a solu ion o mee esea ch and, as a esul , business objec i es, as well as o deba e he use ulness o he sugges ed a i ac (He ne , Ma ch, Pa k, & Ram, 2004). E alua ion To ce i y an a i ac 's e icacy, i mus be pu o use o p esen ed o s akeholde s (Pe e s, Tuunanen, Ro henbe ge , & Cha e jee, 2007), which mus be suppo ed by a clea speci ica ion o e alua ion me hodologies ha a e sui able o he si ua ion a hand and a e based on indus y equi emen s. Because he majo i y o claims on he inal solu ion a e ela ed o pe o mance issues, alignmen wi h business needs is c i ical (He ne , Ma ch, Pa k, & Ram, 2004). Compa ing wha alls unde he pu iew o he mas e 's hesis wi h wha could be obse ed in i s p ac ical implemen a ion is one echnique o e alua e how he answe ma ches he ini ial challenge (Pe e s, Tuunanen, Ro henbe ge , & Cha e jee, 2007). Al hough i is c i ical o emphasize ha he p ima y goal is o "iden i y how well an a i ac wo ks" a he han " heo ize o p o e any hing abou why he a i ac wo ks" (He ne , Ma ch, Pa k, & Ram, 2004). A he end o his phase, i should be de e mined whe he he a i ac is eady o be sha ed wi h he es o he wo ld, o whe he mo e e o should be spen imp o ing i o make i mo e e ec i e/aligned wi h he o iginal p oblems (Pe e s, Tuunanen, Ro henbe ge , & Cha e jee, 2007). Communica ion While eleasing he inal a i ac o he public is a s ep in he igh di ec ion, i 's also c i ical o le people know how unique and success ul he a i ac is in sol ing he highligh ed p oblems (Pe e s, Tuunanen, Ro henbe ge , & Cha e jee, 2007). I is c i ical o discuss how he a i ac was c ea ed and he e iew p ocess ha led o i s alida ion h oughou his communica ion (He ne , Ma ch, Pa k, & Ram, 2004). 37 I should be con eyed o echnical and managemen audiences in o de o ga he inpu o enhance he solu ion, bo h in e ms o business and echnology, o u u e implemen a ions (He ne , Ma ch, Pa k, & Ram, 2004). 5.2. STRATEGY P oblem The e a e nume ous ex summa izing app oaches, each wi h i s own se o bene i s and d awbacks. Some a e mo e compu a ionally complex han o he s, while o he s ha e only been implemen ed in speci ic languages. Some u ilize mo e s a is ical measu es o quan i a i ely add ess summa izing p oblems, while o he s mo e ex ac i e in na u e. Despi e all o hese possibili ies, he e is no single app oach o me hodology ha can be used on any ype o ex . We need o know which s a egy o echnique o use in a ious si ua ions. Objec i e A e s a ing he opic, ou goal in his pape will be o esea ch and assess se e al s a egies, as well as o desc ibe hei bene i s and d awbacks, as well as he si ua ions and ci cums ances in which hey migh be employed. In he same case, no all me hods would pe o m he same. As a esul , we would do ou bes o p oduce a ai compa ison and highligh he echniques' o me hodology' limi a ions. Design and De elopmen Ini ially, a numbe o esea ch publica ions on ex summa iza ion app oaches we e examined. Some o he s a egies o examining i s algo i hm, ime complexi y, he da a i was implemen ed on, how e icien he algo i hm is, how use ul he gene a ed summa y is, and whe he i was an abs ac i e o ex ac i e based me hodology ha e been de ailed in dep h abo e. 38 6. PROPOSAL OF A FRAMEWORK ON SCENARIOS OF TEXT SUMMARIZATION TECHNIQUES 6.1. PROPOSAL Al hough ex summa iza ion has a as numbe o echniques o o e , i was no possible o co e all o hose he e and a such only ew we e selec ed, which we e s udied he e a o emen ioned in he abo e ables. The ollowing able below compa es hose abo e echniques in e ms o accu acy and ime complexi y, applicabili y. Al hough his able does no gi e a ai compa ison since, all hese echniques we e no applied on he same documen and o he same si ua ions. Table 3 Techniques Pa ame e s Accu acy Speed Applicabili y on di e en language Scena ios applicable The lexical chain gene a ion Accu acy is be e Has linea un ime complexi y Fo example, Bengali Al hough his me hod can be applied o mul iple si ua ions, mos esea ch pape s s a e i s main applicabili y in Wo ld Wide Web. La en Seman ic Analysis Ce ain combina ions show di e en accu acy men ioned below Linea Time complexi y Fo example, Bengali, Hindi LSA now scales o ca. 100 million-wo d co po a by la ge compu e memo y and new algo i hms. Que y based summa iza ion o mul iple documen s by Resul s demons a e ha o compu ing he impo ance o In The speed a ies wi h documen s explained in de ail Al hough, any pape ela ed o his echnique has no summa izing esea ch pape s o a speci ic domain, biomedical documen s o be e accu acy 39 applying eg ession models compa ison o aining o sco e and classi ying models, eg ession models pe o m be e . below ye been applied o o he language. Bu his echnique should no ha e any issues ( echnical) i applied o o he language. In summa izing ex , maximum scope and desi ed minimal epe i ion Accu acy is 97% acco ding o (HoudaOu aida, Oma Nouali, & PhilippeBlache, 2014). Al hough his is jus one sample. Compu a ional ime is p opo ional o O(X*Y) whe e X and Y a e di e en e ms in he dis ance ma ix used o disce n he simila i y. Tes ed in languages like A abic, Czech, English, F ench, G eek, Heb ew and Hindi single- and mul i-documen summa iza ion. In bo h asks, documen s a e spli in o sen ences in p ep ocessing E olu iona y op imiza ion algo i hm o summa izing mul iple documen s Accu acy is usually good i he algo i hm is un making su e ha he whole sea ch space is co e ed and no s uck a local maxima Time complexi y o hese algo i hms is usually p e y high as i has o make su e ha he whole sea ch space is co e ed du ing he un- ime. This p oposed me hod has no ye been applied in o he languages. Digi al a chi es o go e nmen al documen s The lexical chain gene a ion - Wo d Sense Disambigua ion (WSD) accu acy is be e . The algo i hm p oposed by Silbe and McCoy has linea un ime complexi y. Tes ed in di e en languages apa om English. Fo example, Bengali. Al hough his me hod can be applied o mul iple si ua ions, 46 7. CONCLUSIONS As he In e ne has g own in popula i y, a as amoun o in o ma ion has become a ailable. Summa izing as amoun s o ex is challenging o humans. In his age o in o ma ion o e load, au oma ic summa izing echnologies a e in high demand. Va ious ex ac ion me hodologies o single and mul i-documen summa iza ion we e highligh ed in his esea ch. Topic ep esen a ion app oaches, equency-d i en me hods, g aph- based and machine lea ning echniques we e desc ibed as some o he mos o en u ilized me hodologies. Al hough i is impossible o elucida e all o he many me hods and app oaches in my hesis, i does p o ide a good o e iew o ecen ends and ad ancemen s in au oma ic summa izing me hods and desc ibes he cu en s a e-o - he-a in his ield. Limi a ions One o he main limi a ions o his epo is ha i wasn’ alida ed by lo o people gi en he ewe numbe o expe s in his ield. Wi h ha goes he unsaid, ha his pape doesn’ documen all he NLP echniques, which is qui e a b oad ield. 47 8. REFERENCES A., N., & K., M. (2012). A Su ey o Tex Summa iza ion Techniques. Em A. C., & Z. C., Mining Tex Da a (pp. 43-76). Bos on, MA: Sp inge , Bos on, MA. Algulie , R. M., Aliguliye , R. M., Haji ahimo a, M. S., & Mehdiye , C. A. (2011). MCMR: Maximum co e age and minimum edundan ex summa iza ion model. Expe Sys ems wi h Applica ions , 14514-14522. Allahya i, M., Pou iyeh, S., Asse i, M., Sa aei, S., T ippe, E. D., Gu ie ez, J. B., & Kochu , K. (2017). Tex Summa iza ion Techniques: A B ie Su ey. Ami ay, E., & Pa is, C. (2000). Au oma ically Summa ising Web Si es - Is The e A Way A ound I ? CIKM00: P oceedings o he nin h in e na ional con e ence on In o ma ion and knowledge managemen (pp. 173–179). Associa ion o Compu ing Machine yNew Yo kNYUni ed S a es. Ba alis, E., Caglie o, L., Maho o, N., & Fio i, A. (2013). G aphSum: Disco e ing co ela ions among mul iple e ms o g aph-based summa iza ion. Em In o ma ion Sciences (pp. 96-109). Ba zilay, R., & Elhadad, M. (2000). Using Lexical Chains o Tex Summa iza ion. Ca enini, G., Ng, R. T., & Zhou, X. (2008). Summa izing Emails wi h Con e sa ional Cohesion and Subjec i i y. (pp. 353–361). Associa ion o Compu a ional Linguis ics. Dee wes e , S., Dumais, S. T., Fu nas, G. W., Landaue , T. K., & Ha shman, R. (1990). Indexing by La en Seman ic Analysis. Jou nal o he Ame ican Socie y o In o ma ion Science . De lin, J., Chang, M.-W., Lee, K., & Tou ano a, K. (2018). BERT: P e- aining o Deep Bidi ec ional T ans o me s o Language Unde s anding. Fa ah, M. A., & Ren, F. (2009). GA, MR, FFNN, PNN and GMM based models o au oma ic ex summa iza ion. Compu e Speech & Language , 126-144. Fe ei a, R., Cab al, L. d., Lins, R. D., Sil a, G. P., F ei as, F., Ca alcan i, G. D., . . . Fa a o, L. (2013). Assessing sen ence sco ing echniques o ex ac i e ex summa iza ion. Em Expe Sys ems wi h Applica ions (pp. 5755-5764). Gong, Y., & Liu, X. (2001). Gene ic ex summa iza ion using ele ance measu e and la en seman ic analysis. P oceedings o he 24 h annual in e na ional ACM SIGIR con e ence on Resea ch and de elopmen in in o ma ion e ie al (pp. 19-25). SIGIR '01. HoudaOu aida, Oma Nouali, & PhilippeBlache. (2014). Minimum edundancy and maximum ele ance o single and mul i-documen A abic ex summa iza ion. Jou nal o King Saud Uni e si y - Compu e and In o ma ion Sciences , 450-461. III, H. D., & Ma cu, D. (2006). Bayesian Que y-Focused Summa iza ion. P oceedings o he 21s In e na ional Con e ence on Compu a ional Linguis ics and he 44 h annual mee ing o he Associa ion o Compu a ional Linguis ics (pp. 305-312). ACL-44. 48 Jones, K. S. (2004). A s a is ical in e p e a ion o e m speci ici y. Jou nal o Documen a ion Volume 60 Numbe 5 , 493-502. Ko, Y., & Seo, J. (2004). Lea ning wi h Unlabeled Da a o Tex Ca ego iza ion Using a Boo s apping and a Fea u e P ojec ion Technique. P oceedings o he 42nd Annual Mee ing o he Associa ion o Compu a ional Linguis ics (ACL-04) , (pp. 255–262). Ko, Y., & Seo, J. (2008). An e ec i e sen ence-ex ac ion echnique using con ex ual in o ma ion and s a is ical app oaches o ex summa iza ion. Em Pa e n Recogni ion Le e s (pp. 1366-1371). Lee, J.-H., Pa k, S., Ahn, C.-M., & Kim, D. (2009). Au oma ic gene ic documen summa iza ion based on non-nega i e. In o ma ion P ocessing and Managemen , 20-34. Llo e , E., & Paloma , M. (2013). Tackling edundancy in ex summa iza ion h ough di e en le els o language analysis. Em Compu e S anda ds & In e aces (pp. 507-518). Luhn, H. P. (1958). The Au oma ic C ea ion o Li e a u e Abs ac s. IBM JOURNAL APRIL 1958 . Mei, Q., & Zhai, C. (2008). Gene a ing Impac -Based Summa ies o Scien i ic Li e a u e. P oceedings o ACL-08: HLT (pp. 816–824). Columbus: Associa ion o Compu a ional Linguis ics. Nenko a, A., & Bagga, A. (2003). Facili a ing email h ead access by ex ac i e summa y gene a ion. Recen Ad ances in Na u al Language P ocessing III: Selec ed pape s om RANLP 2003 , pp. 287-. Newman, P. S., & Bli ze , J. C. (2003). Summa izing A chi ed Discussions: A Beginning. P oceedings o he 8 h in e na ional con e ence on In elligen use in e aces (pp. 273– 276). IUI '03. Ouyang, Y., Li, W., Li, S., & Lu, Q. (2011). Applying eg ession models o que y- ocused mul i- documen summa iza ion. In o ma ion P ocessing & Managemen , 227-237. Ouyang, Y., Li, W., Zhang, R., Li, S., & Lu, Q. (2013). A p og essi e sen ence selec ion s a egy o documen summa iza ion. In o ma ion P ocessing & Managemen , 213-221. Rambow, O., Sh es ha, L., Chen, J., & Lau idsen, C. (2004). Summa izing Email Th eads. P oceedings o HLT-NAACL 2004: Sho Pape s (pp. 105–108). HLT-NAACL-Sho '04. Sal on, G., & Buckley, C. (1988). Te m-weigh ing app oaches in au oma ic ex e ie al. In o ma ion P ocessing and Managemen . W.K.Chan, S. (2006). Beyond keywo d and cue-ph ase ma ching: A sen ence-based abs ac ion echnique o in o ma ion ex ac ion. Em Decision Suppo Sys ems (pp. 759-777). Ye, S., Chua, T.-S., Kan, M.-Y., & Qiu, L. (2007). Documen concep la ice o ex unde s anding and summa iza ion. Em In o ma ion P ocessing & Managemen (pp. 1643-1662). Yeh, J.-Y., Ke, H.-R., Yang, W.-P., & Meng, I.-H. (2005). Tex summa iza ion using a ainable summa ize and la en seman ic analysis. Em In o ma ion P ocessing & Managemen (pp. 75-95). Page | i