scieee Science in your language
[en] (orig)

A framework for the Comparative analysis of text summarization techniques

Abstract

We see that with the boom of information technology and IOT (Internet of things), the size of information which is basically data is increasing at an alarming rate. This information can always be harnessed and if channeled into the right direction, we can always find meaningful information. But the problem is this data is not always numerical and there would be problems where the data would be completely textual, and some meaning has to be derived from it. If one would have to go through these texts manually, it would take hours or even days to get a concise and meaningful information out of the text. This is where a need for an automatic summarizer arises easing manual intervention, reducing time and cost but at the same time retaining the key information held by these texts. In the recent years, new methods and approaches have been developed which would help us to do so. These approaches are implemented in lot of domains, for example, Search engines provide snippets as document previews, while news websites produce shortened descriptions of news subjects, usually as headlines, to make surfing easier. Broadly speaking, there are mainly two ways of text summarization – extractive and abstractive summarization. Extractive summarization is the approach in which important sections of the whole text are filtered out to form the condensed form of the text. While the abstractive summarization is the approach in which the text as a whole is interpreted and examined and after discerning the meaning of the text, sentences are generated by the model itself describing the important points in a concise way.

Read accessible full text

A framework for the Comparative analysis of text summarization techniques

Author: Ghosh, Trijit
Year: 2022
Source: https://run.unl.pt/bitstream/10362/136208/1/TCDMAA0147.pdf
i
A amewo k o he Compa a i e analysis o ex
summa iza ion echniques
T iji Ghosh
Disse a ion p esen ed as pa ial equi emen o
ob aining he mas e ’s deg ee in Da a Science and
Ad anced Analy ics
2
Ins i u o Supe io de Es a ís ica e Ges ão de In o mação
Uni e sidade No a de Lisboa
A FRAMEWORK FOR THE COMPARATIVE ANALYSIS OF TEXT
SUMMARIZATION TECHNIQUES
by
T iji Ghosh
(M20170009)
Disse a ion p esen ed as pa ial equi emen o ob aining he mas e ’s deg ee in Da a Science
and Ad anced Analy ics
Ad iso / Co Ad iso : P o esso Rica do Rei; P o esso Robe o Hen iques
July 2021
3
ACKNOWLEDGEMENTS
I would i s like o hank my hesis ad iso P o esso Doc o Rica do Rei o he NOVA
In o ma ion Managemen School a Uni e sidade NOVA de Lisboa as he was he one who
challenged me o a heme as inno a i e as he one ha ga e mo o o his mas e ’s hesis on ex
summa iza ion. I wan o hank him o encou aging me and mo i a ing me e en when he ime
o de o e o his mas e ’s hesis was no wha was wan ed and expec ed.
Finally, a big hank you o my pa en s o encou aging me and gi ing me he chance o do his
mas e ’s in a eas as in e es ing and exci ing as ad anced analy ics a e. Tha ga e me he
possibili y o ha e a ca ee ha I ha e d eamed o . I would no be possible wi hou hei
suppo .
4
5
Con en s
1. In oduc ion ...................................................................................................................................... 8
1.1. Backg ound ............................................................................................................................... 8
1.2. Mo i a ion ................................................................................................................................. 8
1.3. Objec i e .................................................................................................................................... 8
2. EXTRACTIVE SUMMARIZATION ............................................................................................... 9
2.1. In e media e Rep esen a ion ............................................................................................ 9
2.2. Sen ence Sco e ......................................................................................................................... 9
2.3. Summa y Sen ences Selec ion ........................................................................................... 9
3. TOPIC REPRESENTATION APPROACHES .......................................................................... 11
3.1. Topic Wo ds .......................................................................................................................... 11
3.2. F equency-d i en App oaches ....................................................................................... 11
3.3. La en Seman ic Analysis ................................................................................................. 14
3.4. Bayesian Topic Models ...................................................................................................... 14
3.5. BERT ......................................................................................................................................... 15
4. THE IMPACT OF CONTEXT IN SUMMARIZATION ........................................................... 20
4.1. Web Summa iza ion ........................................................................................................... 20
4.2. Scien i ic A icles Summa iza ion ................................................................................. 20
4.3. Email Summa iza ion ........................................................................................................ 21
5. METHODOLOGY ........................................................................................................................... 35
5.1. Design Sea ch Resea ch .................................................................................................... 35
5.2. S a egy ................................................................................................................................... 37
6. PROPOSAL o a amewo k on scena ios o ex summa iza ion echniques ...... 38
6.1. PROPOSAL .............................................................................................................................. 38
6.2. VALIDATION .......................................................................................................................... 38
7. CONCLUSIONS ............................................................................................................................... 46
8. Re e ences ...................................................................................................................................... 47

6
Lis o Tables
Table 1 ............................................................................................................................................................... 21
Table 2 ............................................................................................................................................................... 27
Table 3 ............................................................................................................................................................... 38
Table 4 ............................................................................................................................................................... 42
Table 5 ............................................................................................................................................................... 43
7
Lis o igu es
Figu e 1 : Weigh ed Te ms /s Speci ici y .......................................................................................... 13
Figu e 2 : Be Embeddings ...................................................................................................................... 16
Figu e 3 : A chi ec u e o BERT .............................................................................................................. 17
Figu e 4 : Encode s and Decode s ......................................................................................................... 17
Figu e 5 : O e all p e- aining and ine- uning p ocedu es o BERT ..................................... 18
Figu e 6 : Fine Tuning phase .................................................................................................................... 18
Figu e 7 : Accu acy o BERTbase on Masked LM and Le - o-Righ ............................................. 19
Figu e 8 : P ecision and Recall o di e en ex iles ...................................................................... 40
Figu e 9 : P ecision, Recall and F-Measu e o di e en alues o k applying LSA ............ 41
Figu e 10 : O e all Compa ison o he me hods ............................................................................... 44
8
1. INTRODUCTION
1.1. BACKGROUND
We see ha wi h he boom o in o ma ion echnology and IOT (In e ne o hings), he size o
in o ma ion which is basically da a is inc easing a an ala ming a e. This in o ma ion can
always be ha nessed and i channeled in o he igh di ec ion, we can always ind meaning ul
in o ma ion. Bu he p oblem is his da a is no always nume ical and he e would be
p oblems whe e he da a would be comple ely ex ual, and some meaning has o be de i ed
om i . I one would ha e o go h ough hese ex s manually, i would ake hou s o e en
days o ge a concise and meaning ul in o ma ion ou o he ex . This is whe e a need o an
au oma ic summa ize a ises easing manual in e en ion, educing ime and cos bu a he
same ime e aining he key in o ma ion held by hese ex s. In he ecen yea s, new me hods
and app oaches ha e been de eloped which would help us o do so. These app oaches a e
implemen ed in lo o domains, o example, Sea ch engines p o ide snippe s as documen
p e iews, while news websi es p oduce sho ened desc ip ions o news subjec s, usually as
headlines, o make su ing easie .
B oadly speaking, he e a e mainly wo ways o ex summa iza ion – ex ac i e and
abs ac i e summa iza ion. Ex ac i e summa iza ion is he app oach in which impo an
sec ions o he whole ex a e il e ed ou o o m he condensed o m o he ex . While he
abs ac i e summa iza ion is he app oach in which he ex as a whole is in e p e ed and
examined and a e disce ning he meaning o he ex , sen ences a e gene a ed by he model
i sel desc ibing he impo an poin s in a concise way.
1.2. MOTIVATION
As he In e ne has g own in popula i y, a as amoun o in o ma ion has become a ailable.
Summa izing as amoun s o ex is challenging o humans. In his age o in o ma ion
o e load, au oma ic summa izing echnologies a e in high demand. We will y o ocus on
a ious ex ac ion me hodologies o single and mul i-documen summa iza ion in his hesis.
Some o he mos o en used me hods, such as opic ep esen a ion app oaches, equency-
d i en me hods, g aph-based and machine lea ning echniques, will be desc ibed. E en
hough i is di icul o ho oughly explain all o he many algo i hms and app oaches in his
hesis, we will y o gi e a good o e iew o ecen ends and b eak h oughs in au oma ic
summa izing me hods, as well as discuss he s a e-o - he-a and compa e a ious ways.
1.3. OBJECTIVE
The objec i e o his pape is o p o ide compa a i e analysis o di e en echniques o ex
summa iza ion used in di e en scena ios, me hodology - o de ine a se o analysis pa ame e
ha can allow us o classi y di e en echniques e.g., complexi y, accu acy and speed.
9
2. EXTRACTIVE SUMMARIZATION
Ex ac i e summa iza ion, as men ioned abo e, chooses pe inen subse s om he ex gi en,
based on some me ics and is hus combined o a condensed o m. To unde s and how
summa iza ion sys ems wo k, we desc ibe h ee, ai ly, independen asks which all
summa ize s pe o m:
1) Build a ansi ional depic ion o he inpu ex which communica es he mos impo an
aspec s o he ex .
2) Sco e he sen ences suppo ing he ep esen a ion.
3) Choose a summa y consis ing o a ie y o ex s.
2.1. INTERMEDIATE REPRESENTATION
All hese summa iza ion echniques will de elop some in e media e ep esen a ions o he
gi en inpu ex suppo ed ce ain me ics and disce n he impo an sen ences suppo ed hose
me ics. The e a e wo o ms o app oaches suppo ed he ep esen a ion: opic ep esen a ion
and indica o ep esen a ion.
Topic ep esen a ion me hods emodel he ex in o a ansi ional cha ac e iza ion and
examine he opic(s) gi en wi hin he ex . This me hod di e s in e ms o o mula ion and
he eby, complexi y, and a e di ided in o equency-d i en app oaches, opic wo d
app oaches, la en seman ic analysis and Bayesian opic models. We ake a deepe look in o
opic ep esen a ion app oaches wi hin he ollowing sec ions.
Indica o ep esen a ion app oaches desc ibe e e y sen ence as a lis ing o ea u es
(indica o s) o impo ance like sen ence leng h, posi ion wi hin he documen , ha ing ce ain
ph ases, e c.
2.2. SENTENCE SCORE
When he inpu ex is ans o med in o a o m which he model in e p e s, a sco e is assigned
o e e y sen ence based on ha me ic. This sco e is jus a ep esen a ion o how impo an a
sen ence is. These indica o s (me ic) a e de i ed om ma hema ical o machine lea ning
models. Once he sco e o e e y o he sen ences a e ob ained, hey a e, hen, agg ega ed o
ed in o a unc ion which anks hese sen ences based on he sco es ob ained. hey' e also
e e ed o as indica o weigh s.
2.3. SUMMARY SENTENCES SELECTION
A e he sen ences a e sco ed and anked suppo ed he a ious me ics o indica o s, he
impo an sen ences a e il e ed ou . These p ocesses can use di e en algo i hms o sepa a e
he sen ences suppo ed he anks and a ew o he sco es as an example edundancy sco e.
Some me hods use g eedy algo i hms o sea ch ou he simples sen ences ep esen ing he
essence o he ex . This me hod will no always be he mos e ec i e app oach which igno es
16
Figu e 2 : Be Embeddings
The inpu ex is i s ed in o he ex p ocessing embeddings namely: -
Posi ion Embeddings
Segmen Embeddings
Token Embeddings

17
Figu e 3 : A chi ec u e o BERT
The ou pu o hese embedding is hen ed o he BERT laye which consis s o ans o me s.
Figu e 4 : Encode s and Decode s
A ans o me can consis o 12/24 blocks o encode s wi h 12/16 a en ion heads and 110/340
million pa ame e s, namely BERTbase/BERTla ge espec i ely. I looked closely, each
ans o me can be a se o encode s and decode s as shown abo e in he diag am. I is he e
ha he ou pu o he ans o me is ed o he summa iza ion laye .
P e- aining has been qui e c ucial when i came o language models. The e ha e been
18
applica ions o hese models such as na u al language in e ence and pa aph asing.
This pape (De lin, Chang, Lee, & Tou ano a, 2018), mainly alks abou ine uning app oach
wi h he p oposal o BERT as desc ibed ea lie . The eason i is called bidi ec ional is because
unlike o he models whe e he ex is ead om le o igh o om igh o le , he BERT
app oach eads he sen ences in bo h he di ec ions and ies o unde s and he con ex o he
wo ds bo h o i s le and igh .
Figu e 5 : O e all p e- aining and ine- uning p ocedu es o BERT
In he abo e diag am, he a chi ec u e o he BERT is shown whe e apa om he ou pu
laye s, bo h use he same a chi ec u e.
Figu e 6 : Fine Tuning phase
Sou ce: (De lin, Chang, Lee, & Tou ano a, 2018)
The abo e igu e shows ha by wo king on he ine- uning phase o he model, he model was
able o achie e an accu acy by an o e whelming ma gin o 4.5% and 7% p io o i s s a e o
he a .
19
Figu e 7 : Accu acy o BERTbase on Masked LM and Le - o-Righ
Sou ce: (De lin, Chang, Lee, & Tou ano a, 2018)
Al hough i is o be men ioned ha his bi-di ec ional app oach akes longe han
unidi ec ional app oach which is wha has been shown in he abo e diag am whe e he BERT
MLM me hod con e ges slowe han le o igh bu accu acy wise, i is well ahead o he
o he app oach.
This bi-di ec ional app oach is, de ini ely, use ul al hough a bi slow and makes he
applica ion o BERT o a wide ange o ields.
20
4. THE IMPACT OF CONTEXT IN SUMMARIZATION
I is qui e e iden ha a con ex is e y use ul when i comes o unde s anding he in o ma ion
p esen ed in a ex . In he same way, i models can be empowe ed wi h he in o ma ion o
con ex , hen a summa ize sys em would be able o p une he co ec in o ma ion o
summa ize he ex . Fo example, acco ding o (Allahya i, e al., 2017), when summa izing
blogs, he deba es o commen s ha ollow he blog pos a e use ul esou ces o de e mining
which po ions o he blog a e c i ical and in iguing. The e is a signi ican quan i y o
in o ma ion a ailable in scien i ic pape summa ies, such as published a icles and con e ence
in o ma ion, ha can be used o highligh essen ial sen ences in he o iginal wo k.
4.1. WEB SUMMARIZATION
I we look a he web pages, we will see ha hey ha e lo o objec s which is no always
possible o summa ize o example, pic u es, gi s, and some unwan ed ma e ials like
ad e isemen s which is so no ele an o he in o ma ion gi en in he pages. Fo hose kinds
o si ua ions, i will be help ul o use he links which di ec us o he page o be summa ized.
These links would gi e he model a knowledge o he con ex and would be help ul o p o ide
imp o ed summa y o he page. In (Ami ay & Pa is, 2000), whe e hey alk abou websi e
summu iza ion o he i s ime, hey came up wi h a concep called “THE INCOMMONSENSE
SYSTEM”. This model inco po a es a web c awling sys em which c awls all he links ha link
back o he cu en page and his is how he con ex is de i ed. Since hen, a lo o di e en
algo i hms ha e been de eloped which was based on he abo e p inciple bu has beeen
imp o ed.
4.2. SCIENTIFIC ARTICLES SUMMARIZATION
In (Mei & Zhai, 2008), a summu iza ion p oblem ela ed o scien i ic pape s is s udied and
p esen ed. Needless, o say wi h each passing yea , new disco e ies a e being made and new
esea ch pape s a e being published e e y yea . So wi h his, he daun ing ask o b ie ing a
scien i ic pape s becomes eally challenging. Specially when i comes o including a con ex
in he summa y which e e ences a mul i ude o di e en esea ch pape s. In o de o sol e
his p oblem, hey came wi h a concep o impac based summu iza ion as men ioned in (Mei
& Zhai, 2008). This me hod le e ages on he impo ance o sen ence sco e which has been
used in he o iginal pape using he KL di e gence me hod (i.e., inding he simila i y
be ween a sen ence and he language model). The conclusions made in his pape a e
subs an ial po ing ha his me hod was use ul and could be used o u u e b ei ing models
o scien i ic pape s.They p oposed a language model ha gi es a p obabili y o each wo d in
he ci a ion con ex sen ences. They hen sco e he impo ance o sen ences in he o iginal
pape .
.
21
4.3. EMAIL SUMMARIZATION
When i comes o email summa iza ion, he ex s become a bi di e en . In o de o ind a
con ex , a whole chain mail h eads a e o be p ocessed o de e mine he s o y. In (Nenko a &
Bagga, 2003), hey discuss a ew me hods o how he con e sa ional na u e o he ex can be
used o ga he con ex . Thei me hods does p o ide a conclusi e e idence as o how use ul
his me hod could be and his u he , his me hod could be enhanced by using some
isualiza ion echniques as well. While in (Rambow, Sh es ha, Chen, & Lau idsen, 2004),
hey ake a bi o a di e en app oach whe e hey de e mine impo an ea u es o
summa izing he email, whe e each ea u e could be one h ead o se e al h eads o email.
(Newman & Bli ze , 2003) discusses a whole new app oach o summa izing an email. They
le e age he clus e ing algo i hm o sol e he p oblem o email summa iza ion. In hei pape
hey alk abou ew clus e ing algo i hms which help hem see all he h eads o he email as a
whole and pos applica ion o he men ioned algo i hms, an o e iew is o med.
Below, Table 2. ci e he jou nals o he e e ences om which some echniques we e analyzed
and u he mo e, hei bene i s o de ec s a e gi en. Table 2. Illus a es he a ious me hods
which we e explained in he a icles men ioned in Table 1

22
21
Table 1
Jou nal/Con e ence
Sub Topic
Yea s
(xxxx-
xxyy)
# A icles
#A icles
Techniques Bene i s
#A icles
Techniques D awbacks
2017 In e na ional
Con e ence on
Compu ing
Me hodologies and
Communica ion
(ICCMC)
Tex
Summa iza ion
2017
Au oma ic ex
summa iza ion
by local
sco ing and
anking o
imp o ing
cohe ence
I has made use o sen ence ea u e
me ics o sco e sen ences like “sen ence
o sen ence cohesion”, “wo d equency”,
i makes use o me ics which migh igno e
aluable in o ma ion and he eby making he
summa y less meaning ul.
Fo example, me ic like “leng h o sen ence”
is used o a oid selec ing oo sho o oo
long sen ences o he documen . Doing his,
a imes, he model migh o e look some
in o ma ion which migh ha e been p esen
in hose sen ences bu we e no aken in o
accoun while summa izing he ex .
22
2017 In e na ional
Con e ence on Big
Da a, IoT and Da a
Science (BID)
Tex
Summa iza ion
2017
Au oma ic ex
summa iza ion
o news
a icles
The lexical chain gene a ion p oposed
by Silbe and McCoy algo i hm has
linea un ime complexi y. Fu he ,
ce ain issues we e esol ed in bo h
algo i hms by
implemen ing p onoun esolu ion and
enhanced sen ence sco ing o le e age
he s uc u e o news a icles.
One o he lexical chain gene a ion algo i hm
adop ed was p oposed by
(Ba zilay & Elhadad, 2000) has exponen ial
un ime complexi y.
A i icial In elligence
Re iew a chi e
Volume 47 Issue 1,
Janua y 2017 Pages 1-
66
Tex
Summa iza ion
2017
Recen
au oma ic ex
summa iza ion
echniques: a
su ey
This pape alks abou a a ie y o
echniques which ha e hei own
bene i s. 1. T ained Summa ize and
La en Sema ic Analysis – uses a
modi ied co pus-based app oach and a
TRM echnique based on la en seman ic
analysis. The summa ize is based on a
unc ion ha assesses main
sen ences/wo ds o hings like loca ion,
keywo d, likeness o i le, and cen ali y
in o de o gene a e summa ies. The
sco e unc ion is op imized using
s ochas ic echniques such as gene ic
algo i hms. 2. In o ma ion e ie al
pe o mance has g ea ly imp o ed
The e a e some d awbacks o he app oaches
used which a e as ollows:T ained
Summa ize and La en Sema ic analysis –
he summa ies gene a ed a e no e y
consis en wi h opic and he sen ences don’
co ela e so much a imes. Fea u e weigh s
o sco e unc ion p oduced by GA ail o
consis en ly gi e he bes esul s o he es
co pus. Ob aining he app op ia e dimension
educ ion a io and explaining LSA e ec s a e
ough in he LSA+TRM echnique. Mo eo e ,
he ime complexi y o compu e SVD is qui e
high. 2. Using a sen ence-based abs ac ion
echnique o ex ac da a – In his app oach,
only casual cohe ence is conside ed whe eas
23
because o he use o a sen ence-based
abs ac ion echnique. Sen ences ha
ep esen he cen al no ion a e linked.
3. Unde s anding and summa izing
documen s using a documen concep
la ice – In compa ison o exis ing
sen ence g ouping and sen ence sco ing
algo i hms, he sugges ed app oach
pe o ms excep ionally well. 4. Sen ence
ex ac ion using ex summa iza ion
based on con ex and s a is ics – This
space and ime which a e in e - ela ed links
which makes sense ou o he documen as a
whole a e also equi ed o ep esen ing
beha io al con ex . 3. Th ough he use o a
documen concep la ice, i is possible o
comp ehend and summa ize ex – ime
complexi y o gene a ing a DCI is high
because i conside s all possible
combina ions. 4. Sen ence ex ac ion h ough
con ex ual and s a is ical based
summa iza ion ex .
30
documen
summa ies
h ough non -
nega i e
ma ix
ac o iza ion
his app oach could no na u ally ca ch he meaning o seman ic ea u es ha a e
highly spa se and ha e a limi ed iew o meaning. As a esul , summa iza ion sys ems
based on LSA a e unable o choose meaning ul ph ases. As a esul , elemen s o
seman ic ea u e ec o s in he sugges ed me hod exclusi ely con ain non-nega i e
alues and a e also ex emely spa se, allowing seman ic cha ac e is ics o be easily
ead. A sen ence can be ep esen ed by a linea combina ion o ce ain signi ican
seman ic elemen s. As a esul , sub opics in a documen can be quickly iden i ied, and
he e's a be e p obabili y o ex ac ing ele an lines. A me hod o picking ph ases
o cons uc gene al documen summa ies is sugges ed using NMF, in which a con en
is i s p e-p ocessed and hen summa ized. To gene a e a non-nega i e seman ic
ea u e ma ix, NMF is used o a e m-by-sen ence ma ix. Fo each sen ence, gene ic
ele ance is calcula ed, which indica es how much a sen ence explains.
Que y based
summa iza ion
o mul iple
documen s by
applying
eg ession
models
2011
Ouyang Y
Ouyang, Li, Li, &
Lu, 2011
(Ouyang, Li, Li, & Lu, 2011) sugges ed a me hod o anking ph ases in que y-based
summa iza ion o nume ous manusc ip s using eg ession models. Th ee que y-
dependen ea u es, such as named-en i y ma ching, wo d-ma ching, and seman ic
ma ching, and ou que y-independen ea u es, such as sen ence posi ion, named
en i y, wo d TF-IDF, and s op-wo d penal y, a e used in his me hodology o choose
main sen ences in que y-based summa iza ion o mul iple documen s. To begin wi h,
human summa ies gene a e " alse" aining da a. Then, using di e en me hods based
on he N-g am me hodology ha calcula e "nea ly ue" ele ance a ings o ph ases
a e c ea ed and analyzed using his aining da a and hei collec ion o ex s, and a
mapping unc ion is lea ned using his aining da a ia a collec ion o p e iously
speci ied ea u es o sen ences. Then, using his lea ned unc ion, he signi icance o
sen ences in he es da a is es ima ed. An e icien da a collec ion o aining da a o
lea ning eg ession models equi es wo hings: (a) an app op ia e g oup o opics
wi h co ec ly handw i en summa ies, and (b) a good app oach o compu ing he
ele ance o wo ds. The Maximal Ma ginal Rele ance (MMR) echnique is used o
emo e edundancy om he summa y.
Au oma ic ex
summa iza ion
using MR, GA,
2009
Mohamed Abdel
Fa ah, Fuji Ren
(Fa ah and Ren
2009)
Wi h he use o a ew s a is ical ea u es, (Fa ah & Ren, 2009) sugges ed an app oach
o imp o e con en selec ion in au oma ic ex summa iza ion. As a ainable
summa ize , his me hod elies on dis inc s a is ical aspec s in each sen ence o
gene a e summa ies. Posi ion o Sen ence (Pos), + e keywo d, - e keywo d, + e

31
FFNN, GMM
and PNN based
models
keywo d, + e keywo d, + e keywo d, + e keywo d, + e keywo d, + e keywo d, + e
keywo d, + e keywo d, + e keywo d, R2T, Cen ali y o Sen ence (Cen), P esence o
Name En i y in Sen ence (PNE), P esence o Numbe s in Sen ence (PN), Bushy Pa h o
Sen ence (BP), Rela i e Leng h o Sen ence (RL), and Agg ega e Simila i y (AS) a e all
measu es o sen ence simila i y. Gene ic Algo i hm (GA) and Ma hema ical Reg ession
(MR) models ha e been ained o acqui e an op imal mix o ea u e weigh s by
mixing all o hese ea u es. Fo sen ence ca ego iza ion, eed o wa d neu al
ne wo ks (FFNN) and p obabilis ic neu al ne wo ks (PNN) a e u ilized. Some ex
ea u es, such as he + e and - e keywo ds, a e language-dependen , while eigh
o he s a e no . All o he abo e-men ioned a iables a e aken in o accoun when
calcula ing a sen ence's weigh ed sco e unc ion. All sen ences in a documen a e
anked in dec easing o de o hei sco es, and a highly sco ed clus e o sen ences is
u ilized o gene a e a summa y o he con en using a ious comp ession a es (10,
20, 30 pe cen used he e). Conclusion: The esul s demons a e ha ea u e BP is he
mos essen ial ex ea u e since i p oduces he bes esul s, while ea u e PND
p oduces he wo s esul s because s a is ical da a is absen om eligious and
poli ical pieces. Because i could model a bi a y densi ies, he GMM app oach
p oduced he bes esul s o all he s a egies.
Maximum
co e age and
minimum
edundancy in
summa iza ion
o ex
2011
Algulie , R. M
Algulie ,
Aliguliye ,
Haji ahimo a, &
Mehdiye , 2011
(Algulie , Aliguliye , Haji ahimo a, & Mehdiye , 2011) in oduced an unsupe ised
summa izing model o gene ic ex as an In ege Linea P og amming p oblem (ILP)
ha immedia ely de ec s essen ial sen ences om he a icle as well as he ull
a icle's ele an in o ma ion. Maximum Co e age and Minimum Redundancy is he
name o his s a egy (MCMR). This me hod aims o imp o e h ee key aspec s o a
summa y: (a) ele ance, (b) edundancy, and (c) leng h. A subse o sen ences om
he documen collec ion's ele an ex is picked. Then, using NGD-based simila i y
(No malized Google Dis ance) and cosine simila i y, simila i y be ween he summa y
and he documen collec ion is compu ed, and his simila i y mus be maximized. An
objec i e unc ion is de eloped and mus be maximized o ensu e ha he summa y
con ains he impo an con en ound in he documen collec ion and ha he
summa y does no con ain a signi ican numbe o ph ases ha communica e he
same in o ma ion. A he same ime, he leng h o he summa y mus be limi ed.
Las ly, an empi ical unc ion is c ea ed by linea ly combining he cosine simila i y-
32
based and NGD-based simila i y empi ical unc ions, and his combined empi ical
unc ion mus be maximized. This echnique o summa izing is inco po a ed as an
op imiza ion p oblem ha aims o ind a global solu ion o he p oblem. The B anch &
Bound algo i hm (B&B) and he Bina y Swa m Op imiza ion me hod a e he
algo i hms used o add ess he ILP p oblem. Conclusion: This me hod, which
combines MCMR wi h he B&B algo i hm, su passes all o he s. I demons a es ha
summa izing ou comes a e dependen on simila i y measu emen s. I is also p o ed
h ough es s ha using cosine simila i y and NGD-based simila i y me ics oge he
p oduces be e esul s han using hem sepa a ely.
Summa iza ion
o documen s
h ough a
p og essi e
echnique o
selec ion o
sen ences
2013
Ouyang Y
Ouyang, Li, Zhang,
Li, & Lu, 2013
(Ouyang, Li, Zhang, Li, & Lu, 2013) p oposed a new p og essi e echnique o
gene a ing a summa y based on he selec ion o "no el and salien " sen ences.
Subsuming ela ionship be ween wo sen ences, i.e., an i egula ela ionship be ween
sen ences ha shows he le el o ecommenda ion o one ph ase by ano he . In o de
o asce ain he link be ween wo ph ases, he ela ionship be ween hei concep s
mus be ound. The associa ion be ween concep s is hen disco e ed by using a
co e age-based measu e o disco e he ela ionship be ween wo ds. A Di ec Acyclic
G aph (DAG) is used o o ganize all o he wo ds ha appea in he ound wo d
ela ions. A p og essi e s a egy o sen ence selec ion is c ea ed on he basis o an
asymme ic ela ionship be ween sen ences, in which a sen ence is ei he picked as a
no el gene al s a emen o as a suppo ing sen ence. The ollowing wo me hods a e
used o choose new and ele an sen ences in his me hod: (a) disco e ed concep s
a e only included du ing he assessmen o sen ence ele ance o assu e sen ence
o iginali y, and (b) o now, he ela ionship be ween sen ences is used o imp o e he
saliency measu e. To implemen his s a egy, a andom walk on he DAG om he
cen al node o i s nea by nodes is pe o med, wi h he goal o co e ing he cen al
wo ds i s and hen eaching he g ea es amoun o wo ds ia wo d ela ions.
Redundancy is elimina ed by punishing epe i i e wo ds, esul ing in esh concep s
being in oduced each ime a new ph ase is chosen. Conclusion: In e ms o gene a ing
summa ies wi h imp o ed saliency and co e age, he P og essi e sys em su passes
he adi ional Sequen ial app oach.
E alua ion o
2013
Fe ei a, Ra ael;
Cab al, Luciano de
Fe ei a, e al.,
In he ecen decade, (Fe ei a, e al., 2013) inco po a ed i een sco ing echniques
ha had been e e enced in he esea ch. ROUGE (Lin 2004) is used o quan i a i e
33
sen ence
sco ing
me hods o
ex ac i e
summa iza ion
o ex
Souza; Lins, Ra ael
Duei e; Sil a,
Gab iel Pe ei a e;
F ei as, F ed;
Ca alcan i, Geo ge
D.C.; Lima,
Rinaldo; Simske,
S e en J.; Fa a o,
Luciano
20132013
e alua ion, while he numbe o sen ences ha a e simila among he machine-
gene a ed and human-made summa ies is coun ed o quali a i e e alua ion. The
p ocessing ime o each algo i hm is aken in o accoun . Wo d sco ing, sen ence
sco ing, and g aph sco ing app oaches a e used o pick ele an sen ences. The mos
essen ial e ms a e gi en sco es in he wo d sco ing app oach. Wo d equency,
TF/IDF, uppe case, p ope noun, wo d co-occu ence, and lexical simila i y a e
among he app oaches used o sco e wo ds. The p ope ies o sen ences a e examined
in he sen ence sco ing app oach. The exis ence o cues, nume ical da a, sen ence
leng h, sen ence posi ion, and sen ence cen ali y a e all ac o s in sen ence sco ing.
Sco es a e de e mined using he g aph sco ing app oach by looking a he
ela ionships be ween sen ences. Tex ank, bushy pa h o he node, and agg ega e
simila i y a e all g aph sco ing app oaches. The six common conce ns o s op wo ds,
s uc u al ans o ma ion, compa able seman ics, ambigui y, edundancy, and co-
e e ence a e hen explo ed, along wi h some sugges ions o ad ancing sen ence
sco e ou comes.
Explo ing
co ela ions
among
mul iple e ms
h ough a
g aph-based
summa ize ,
GRAPHSUM
2013
Ba alis, Elena;
Caglie o, Luca;
Maho o, Naeem;
Fio i, Alessand o
Ba alis, Caglie o,
Maho o, & Fio i,
2013
GRAPHSUM, a new g aph-based, gene al-pu pose summa ize o summa izing
nume ous documen s, was p oposed by (Ba alis, Caglie o, Maho o, & Fio i, 2013).
This me hod in es iga es and applies associa ion ules, a da a mining me hodology o
inding connec ions be ween se e al e ms. I is no elian on sophis ica ed seman ic
models (like axonomies o on ologies). The documen collec ion is o ganized as a
ansac ional da ase a e p ep ocessing so ha associa ion ule mining may be
conduc ed on i . Then, om he ansac ional da ase , equen ly ecu ing i emse s
wi h high co ela ions among he e ms a e iden i ied, and a co ela ion g aph is
cons uc ed om hese e ms, which will aid in he selec ion o signi ican lines o he
summa y. The Ap io i algo i hm is used o mine equen ly ecu ing i emse s, and
he suppo measu e is employed o his job. The li measu e indica es he in ensi y
o ela ionship be ween wo e ms and is used o e alua e posi i e o nega i e
connec ions be ween commonly used wo ds. A a ia ion o he classic PageRank
g aph anking algo i hm is used o de e mine he ele ance o he g aph nodes. The
g aph nodes ha ha e a signi ican numbe o posi i e co ela ions a e placed i s ,
while hose ha ha e a nega i e connec ion wi h he adjacen nodes a e penalized.
Fo summa y c ea ion, he sen ences ha a e he mos app op ia e o he co ela ion
34
g aph and ha e a high ele ance sco e a e picked. The g eedy algo i hm is employed
o selec sen ences in his case. GRAPHSUM pe o ms be e o e a wide ange o
s a e-o - he-a echniques, some o which ely hea ily on highly de eloped seman ic-
based models o complica ed language p ocesses.
Inco po a ing
a ious le els
o language
analysis o
ackling
edundancy in
ex
summa iza ion
2013
Elena Llo e ,
Manuel Paloma
Llo e & Paloma ,
2013
(Llo e & Paloma , 2013) p o ided a me hod o de ec ing edundan in o ma ion
based on h ee laye s o language analysis: lexical, syn ac ic, and seman ic. Cosine
simila i y is u ilized in he lexical based echnique o de ec simila i y be ween
sen ences in wo sou ces. Those sen ences ha ha e a cosine simila i y g ea e han a
ce ain h eshold a e conside ed epe i i e, and hey a e all elimina ed. In a syn ac ic-
based me hod, en ailmen ela ions a e compu ed be ween pai s o ph ases o
de e mine whe he he meaning o one sen ence can be deduced om he meaning o
he o he sen ence. I a posi i e en ailmen is ob ained, he second sen ence is deemed
supe luous and elimina ed. Sen ence alignmen is de e mined a he documen le el
be ween a se o linked documen s using a open sou ce a ailable Champollian Tool Ki
in a seman ic-based manne . Syn ac ic and seman ic echniques a e p e e able han
lexical app oaches ha ely on cosine simila i y. Tex summa iza ion can be done in
wo ways. Be o e he ma e ial is summa ized, unnecessa y sen ences a e dele ed in
he i s echnique. The se o use ul sen ences is hen gi en o he summa iza ion
sys em, which uses s a is ical ( e m equency) and linguis ic (code quan i y
p inciple) ac o s o selec essen ial sen ences, as well as a summa y.
35
5. METHODOLOGY
This s udy will be conduc ed h ough a quali a i e esea ch, based on a well-s uc u ed
pa ame e o compa ing he di e en echniques adop ed o ex summa iza ion echniques.
This is he basis o compa e di e en echniques and shed ligh on which echniques would be
mo e use ul in (i any) pa icula si ua ions. Wha a e hei d awbacks and ad an ages?
5.1. DESIGN SEARCH RESEARCH
Design Science Resea ch is a o m o in es iga ion ha en ails building o imp o ing
some hing in a no el way in esponse o a speci ic challenge.
The ques o a solu ion based on ex ensi e scien i ic in es iga ion ensu es ha he inal
p oposed a i ac is cohe en and c edible. A c ucial phase ha should no be o e looked is
good communica ion o he inished p oduc (He ne , Ma ch, Pa k, & Ram, 2004).
Each o he six key s ages o DSR me hodology, as shown in Figu e 5, will be discussed in
g ea e de ail igh away.
Figu e 3. DSR Me hod Adap a ion (Pe e s, Tuunanen, Ro henbe ge , & Cha e jee, 2007)
Iden i y p oblem and mo i a ion
De ine he esea ch challenge in de ail and jus i y he impo ance o a solu ion.
Begin by es ablishing a es able heo y ha leads o a esea ch p oblem by demons a ing o
s akeholde s he alue o an e ec i e solu ion and wha hey will gain om i s esul (Pe e s,
Tuunanen, Ro henbe ge , & Cha e jee, 2007).
De ine objec i es and a solu ion

36
Clea ly de ine goals (quan i a i e o quali a i e) o es ablish he ounda ion o a solu ion
based on he p oblem cha ac e iza ion and wha can and canno be done (Pe e s, Tuunanen,
Ro henbe ge , & Cha e jee, 2007).
Design and De elopmen
The goal o he design and de elopmen s ages is o c ea e knowledge h ough he design and
de elopmen o he a i ac i sel (G ego & He ne , 2013). This could be accomplished by
b eaking down he majo scien i ic p oblem in o smalle componen s (He ne , Ma ch, Pa k,
& Ram, 2004). To ha e an e ec i e/ clea s uc u e in he nex phase, i is necessa y o ha e a
clea g asp o he solu ion alue and o de end i wi h some heo e ical ounda ion (Pe e s,
Tuunanen, Ro henbe ge , & Cha e jee, 2007). A solu ion ha mus mee business
equi emen s (He ne , Ma ch, Pa k, & Ram, 2004).
To gain he app op ia e heo e ical basis, i is essen ial o do esea ch and collec knowledge
abou he p esen s a us o he p oblem and exis ing solu ions, as well as o analyze di ec and
indi ec solu ions and hei e icacy (Pe e s, Tuunanen, Ro henbe ge , & Cha e jee, 2007).
Wi h he knowledge, i is possible o de elop a solu ion o mee esea ch and, as a esul ,
business objec i es, as well as o deba e he use ulness o he sugges ed a i ac (He ne ,
Ma ch, Pa k, & Ram, 2004).
E alua ion
To ce i y an a i ac 's e icacy, i mus be pu o use o p esen ed o s akeholde s (Pe e s,
Tuunanen, Ro henbe ge , & Cha e jee, 2007), which mus be suppo ed by a clea
speci ica ion o e alua ion me hodologies ha a e sui able o he si ua ion a hand and a e
based on indus y equi emen s. Because he majo i y o claims on he inal solu ion a e
ela ed o pe o mance issues, alignmen wi h business needs is c i ical (He ne , Ma ch, Pa k,
& Ram, 2004).
Compa ing wha alls unde he pu iew o he mas e 's hesis wi h wha could be obse ed in
i s p ac ical implemen a ion is one echnique o e alua e how he answe ma ches he ini ial
challenge (Pe e s, Tuunanen, Ro henbe ge , & Cha e jee, 2007).
Al hough i is c i ical o emphasize ha he p ima y goal is o "iden i y how well an a i ac
wo ks" a he han " heo ize o p o e any hing abou why he a i ac wo ks" (He ne , Ma ch,
Pa k, & Ram, 2004).
A he end o his phase, i should be de e mined whe he he a i ac is eady o be sha ed
wi h he es o he wo ld, o whe he mo e e o should be spen imp o ing i o make i
mo e e ec i e/aligned wi h he o iginal p oblems (Pe e s, Tuunanen, Ro henbe ge , &
Cha e jee, 2007).
Communica ion
While eleasing he inal a i ac o he public is a s ep in he igh di ec ion, i 's also c i ical o
le people know how unique and success ul he a i ac is in sol ing he highligh ed p oblems
(Pe e s, Tuunanen, Ro henbe ge , & Cha e jee, 2007). I is c i ical o discuss how he
a i ac was c ea ed and he e iew p ocess ha led o i s alida ion h oughou his
communica ion (He ne , Ma ch, Pa k, & Ram, 2004).
37
I should be con eyed o echnical and managemen audiences in o de o ga he inpu o
enhance he solu ion, bo h in e ms o business and echnology, o u u e implemen a ions
(He ne , Ma ch, Pa k, & Ram, 2004).
5.2. STRATEGY
P oblem
The e a e nume ous ex summa izing app oaches, each wi h i s own se o bene i s and
d awbacks. Some a e mo e compu a ionally complex han o he s, while o he s ha e only been
implemen ed in speci ic languages. Some u ilize mo e s a is ical measu es o quan i a i ely
add ess summa izing p oblems, while o he s mo e ex ac i e in na u e. Despi e all o hese
possibili ies, he e is no single app oach o me hodology ha can be used on any ype o ex .
We need o know which s a egy o echnique o use in a ious si ua ions.
Objec i e
A e s a ing he opic, ou goal in his pape will be o esea ch and assess se e al s a egies,
as well as o desc ibe hei bene i s and d awbacks, as well as he si ua ions and ci cums ances
in which hey migh be employed. In he same case, no all me hods would pe o m he same.
As a esul , we would do ou bes o p oduce a ai compa ison and highligh he echniques'
o me hodology' limi a ions.
Design and De elopmen
Ini ially, a numbe o esea ch publica ions on ex summa iza ion app oaches we e examined.
Some o he s a egies o examining i s algo i hm, ime complexi y, he da a i was
implemen ed on, how e icien he algo i hm is, how use ul he gene a ed summa y is, and
whe he i was an abs ac i e o ex ac i e based me hodology ha e been de ailed in dep h
abo e.
38
6. PROPOSAL OF A FRAMEWORK ON SCENARIOS OF TEXT SUMMARIZATION TECHNIQUES
6.1. PROPOSAL
Al hough ex summa iza ion has a as numbe o echniques o o e , i was no possible o co e all o hose he e and a such only ew we e
selec ed, which we e s udied he e a o emen ioned in he abo e ables. The ollowing able below compa es hose abo e echniques in e ms o
accu acy and ime complexi y, applicabili y. Al hough his able does no gi e a ai compa ison since, all hese echniques we e no applied on he
same documen and o he same si ua ions.
Table 3
Techniques
Pa ame e s
Accu acy
Speed
Applicabili y on
di e en language
Scena ios applicable
The lexical chain
gene a ion
Accu acy is be e
Has linea un ime
complexi y
Fo example,
Bengali
Al hough his me hod can be
applied o mul iple si ua ions,
mos esea ch pape s s a e i s
main applicabili y in Wo ld
Wide Web.
La en Seman ic
Analysis
Ce ain combina ions
show di e en
accu acy men ioned
below
Linea Time
complexi y
Fo example,
Bengali, Hindi
LSA now scales o ca. 100
million-wo d co po a by la ge
compu e memo y and new
algo i hms.
Que y based
summa iza ion o
mul iple documen s by
Resul s demons a e
ha o compu ing
he impo ance o In
The speed a ies
wi h documen s
explained in de ail
Al hough, any pape
ela ed o his
echnique has no
summa izing esea ch pape s
o a speci ic domain, biomedical
documen s o be e accu acy
39
applying eg ession
models
compa ison o
aining o sco e and
classi ying models,
eg ession models
pe o m be e .
below
ye been applied o
o he language. Bu
his echnique
should no ha e any
issues ( echnical) i
applied o o he
language.
In summa izing ex ,
maximum scope and
desi ed minimal
epe i ion
Accu acy is 97%
acco ding o
(HoudaOu aida,
Oma Nouali, &
PhilippeBlache,
2014). Al hough his
is jus one sample.
Compu a ional ime
is p opo ional o
O(X*Y) whe e X and
Y a e di e en e ms
in he dis ance
ma ix used o
disce n he
simila i y.
Tes ed in languages
like A abic, Czech,
English, F ench,
G eek, Heb ew and
Hindi
single- and mul i-documen
summa iza ion. In bo h asks,
documen s a e spli in o
sen ences in p ep ocessing
E olu iona y
op imiza ion algo i hm
o summa izing
mul iple documen s
Accu acy is usually
good i he algo i hm
is un making su e
ha he whole sea ch
space is co e ed and
no s uck a local
maxima
Time complexi y o
hese algo i hms is
usually p e y high
as i has o make
su e ha he whole
sea ch space is
co e ed du ing he
un- ime.
This p oposed
me hod has no ye
been applied in
o he languages.
Digi al a chi es o
go e nmen al documen s
The lexical chain gene a ion - Wo d Sense Disambigua ion (WSD) accu acy is be e . The algo i hm p oposed by Silbe and McCoy has linea un
ime complexi y. Tes ed in di e en languages apa om English. Fo example, Bengali. Al hough his me hod can be applied o mul iple si ua ions,
46
7. CONCLUSIONS
As he In e ne has g own in popula i y, a as amoun o in o ma ion has become a ailable.
Summa izing as amoun s o ex is challenging o humans. In his age o in o ma ion o e load,
au oma ic summa izing echnologies a e in high demand.
Va ious ex ac ion me hodologies o single and mul i-documen summa iza ion we e
highligh ed in his esea ch. Topic ep esen a ion app oaches, equency-d i en me hods, g aph-
based and machine lea ning echniques we e desc ibed as some o he mos o en u ilized
me hodologies. Al hough i is impossible o elucida e all o he many me hods and app oaches in
my hesis, i does p o ide a good o e iew o ecen ends and ad ancemen s in au oma ic
summa izing me hods and desc ibes he cu en s a e-o - he-a in his ield.
Limi a ions
One o he main limi a ions o his epo is ha i wasn’ alida ed by lo o people gi en he
ewe numbe o expe s in his ield. Wi h ha goes he unsaid, ha his pape doesn’
documen all he NLP echniques, which is qui e a b oad ield.

47
8. REFERENCES
A., N., & K., M. (2012). A Su ey o Tex Summa iza ion Techniques. Em A. C., & Z. C.,
Mining Tex
Da a
(pp. 43-76). Bos on, MA: Sp inge , Bos on, MA.
Algulie , R. M., Aliguliye , R. M., Haji ahimo a, M. S., & Mehdiye , C. A. (2011). MCMR: Maximum
co e age and minimum edundan ex summa iza ion model.
Expe Sys ems wi h
Applica ions
, 14514-14522.
Allahya i, M., Pou iyeh, S., Asse i, M., Sa aei, S., T ippe, E. D., Gu ie ez, J. B., & Kochu , K. (2017).
Tex Summa iza ion Techniques: A B ie Su ey.
Ami ay, E., & Pa is, C. (2000). Au oma ically Summa ising Web Si es - Is The e A Way A ound I ?
CIKM00: P oceedings o he nin h in e na ional con e ence on In o ma ion and
knowledge managemen
(pp. 173–179). Associa ion o Compu ing Machine yNew
Yo kNYUni ed S a es.
Ba alis, E., Caglie o, L., Maho o, N., & Fio i, A. (2013). G aphSum: Disco e ing co ela ions among
mul iple e ms o g aph-based summa iza ion. Em
In o ma ion Sciences
(pp. 96-109).
Ba zilay, R., & Elhadad, M. (2000). Using Lexical Chains o Tex Summa iza ion.
Ca enini, G., Ng, R. T., & Zhou, X. (2008). Summa izing Emails wi h Con e sa ional Cohesion and
Subjec i i y. (pp. 353–361). Associa ion o Compu a ional Linguis ics.
Dee wes e , S., Dumais, S. T., Fu nas, G. W., Landaue , T. K., & Ha shman, R. (1990). Indexing by
La en Seman ic Analysis.
Jou nal o he Ame ican Socie y o In o ma ion Science
.
De lin, J., Chang, M.-W., Lee, K., & Tou ano a, K. (2018). BERT: P e- aining o Deep Bidi ec ional
T ans o me s o Language Unde s anding.
Fa ah, M. A., & Ren, F. (2009). GA, MR, FFNN, PNN and GMM based models o au oma ic ex
summa iza ion.
Compu e Speech & Language
, 126-144.
Fe ei a, R., Cab al, L. d., Lins, R. D., Sil a, G. P., F ei as, F., Ca alcan i, G. D., . . . Fa a o, L. (2013).
Assessing sen ence sco ing echniques o ex ac i e ex summa iza ion. Em
Expe
Sys ems wi h Applica ions
(pp. 5755-5764).
Gong, Y., & Liu, X. (2001). Gene ic ex summa iza ion using ele ance measu e and la en
seman ic analysis.
P oceedings o he 24 h annual in e na ional ACM SIGIR con e ence
on Resea ch and de elopmen in in o ma ion e ie al
(pp. 19-25). SIGIR '01.
HoudaOu aida, Oma Nouali, & PhilippeBlache. (2014). Minimum edundancy and maximum
ele ance o single and mul i-documen A abic ex summa iza ion.
Jou nal o King Saud
Uni e si y - Compu e and In o ma ion Sciences
, 450-461.
III, H. D., & Ma cu, D. (2006). Bayesian Que y-Focused Summa iza ion.
P oceedings o he 21s
In e na ional Con e ence on Compu a ional Linguis ics and he 44 h annual mee ing o
he Associa ion o Compu a ional Linguis ics
(pp. 305-312). ACL-44.
48
Jones, K. S. (2004). A s a is ical in e p e a ion o e m speci ici y.
Jou nal o Documen a ion
Volume 60 Numbe 5
, 493-502.
Ko, Y., & Seo, J. (2004). Lea ning wi h Unlabeled Da a o Tex Ca ego iza ion Using a
Boo s apping and a Fea u e P ojec ion Technique.
P oceedings o he 42nd Annual
Mee ing o he Associa ion o Compu a ional Linguis ics (ACL-04)
, (pp. 255–262).
Ko, Y., & Seo, J. (2008). An e ec i e sen ence-ex ac ion echnique using con ex ual in o ma ion
and s a is ical app oaches o ex summa iza ion. Em
Pa e n Recogni ion Le e s
(pp.
1366-1371).
Lee, J.-H., Pa k, S., Ahn, C.-M., & Kim, D. (2009). Au oma ic gene ic documen summa iza ion
based on non-nega i e.
In o ma ion P ocessing and Managemen
, 20-34.
Llo e , E., & Paloma , M. (2013). Tackling edundancy in ex summa iza ion h ough di e en
le els o language analysis. Em
Compu e S anda ds & In e aces
(pp. 507-518).
Luhn, H. P. (1958). The Au oma ic C ea ion o Li e a u e Abs ac s.
IBM JOURNAL APRIL 1958
.
Mei, Q., & Zhai, C. (2008). Gene a ing Impac -Based Summa ies o Scien i ic Li e a u e.
P oceedings o ACL-08: HLT
(pp. 816–824). Columbus: Associa ion o Compu a ional
Linguis ics.
Nenko a, A., & Bagga, A. (2003). Facili a ing email h ead access by ex ac i e summa y
gene a ion.
Recen Ad ances in Na u al Language P ocessing III: Selec ed pape s om
RANLP 2003
, pp. 287-.
Newman, P. S., & Bli ze , J. C. (2003). Summa izing A chi ed Discussions: A Beginning.
P oceedings o he 8 h in e na ional con e ence on In elligen use in e aces
(pp. 273–
276). IUI '03.
Ouyang, Y., Li, W., Li, S., & Lu, Q. (2011). Applying eg ession models o que y- ocused mul i-
documen summa iza ion.
In o ma ion P ocessing & Managemen
, 227-237.
Ouyang, Y., Li, W., Zhang, R., Li, S., & Lu, Q. (2013). A p og essi e sen ence selec ion s a egy o
documen summa iza ion.
In o ma ion P ocessing & Managemen
, 213-221.
Rambow, O., Sh es ha, L., Chen, J., & Lau idsen, C. (2004). Summa izing Email Th eads.
P oceedings o HLT-NAACL 2004: Sho Pape s
(pp. 105–108). HLT-NAACL-Sho '04.
Sal on, G., & Buckley, C. (1988). Te m-weigh ing app oaches in au oma ic ex e ie al.
In o ma ion P ocessing and Managemen
.
W.K.Chan, S. (2006). Beyond keywo d and cue-ph ase ma ching: A sen ence-based abs ac ion
echnique o in o ma ion ex ac ion. Em
Decision Suppo Sys ems
(pp. 759-777).
Ye, S., Chua, T.-S., Kan, M.-Y., & Qiu, L. (2007). Documen concep la ice o ex unde s anding
and summa iza ion. Em
In o ma ion P ocessing & Managemen
(pp. 1643-1662).
Yeh, J.-Y., Ke, H.-R., Yang, W.-P., & Meng, I.-H. (2005). Tex summa iza ion using a ainable
summa ize and la en seman ic analysis. Em
In o ma ion P ocessing & Managemen
(pp. 75-95).
Page | i