scieee Science in your language
[en] (orig)

FEMsum: A flexible eclectic multitask summarizer architecture evaluated in multidocument tasks

Abstract

This article describes two types of summarization approaches integrated in a flexible architecture for multitask summarization. The first type is based on the use of lexical features, while the second one is grounded on syntactic and semantic information. All the approaches have been evaluated in experiments where, given a set of documents, they are expected to produce summaries answering a user need (expressed by a query) in a reduced set of relevant textual fragments. Their performance is analyzed in two different tasks: written news and scientific oral presentations.

Read accessible full text

FEMsum: A flexible eclectic multitask summarizer architecture evaluated in multidocument tasks

Author: Fuentes Fort, Maria,Rodríguez Hontoria, Horacio,Turmo Borras, Jorge
Year: 2007
Source: https://upcommons.upc.edu/bitstream/2117/87422/1/R06-41.pdf
FEMsum: A Flexible Eclec ic Mul i ask
Summa ize A chi ec u e e alua ed in
Mul idocumen Tasks ?
.
Ma ia Fuen es, Ho acio Rod ´ıguez, Jo di Tu mo
TALP Resea ch Cen e , Uni e si a Poli `ecnica de Ca alunya Ba celona, Spain
Abs ac
This a icle desc ibes wo ypes o summa iza ion app oaches in eg a ed in a lexible
a chi ec u e o mul i ask summa iza ion. The i s ype is based on he use o lexical
ea u es, while he second one is g ounded on syn ac ic and seman ic in o ma ion.
All he app oaches ha e been e alua ed in expe imen s whe e, gi en a se o docu-
men s, hey a e expec ed o p oduce summa ies answe ing a use need (exp essed
by a que y) in a educed se o ele an ex ual agmen s. Thei pe o mance is
analyzed in wo di e en asks: w i en news and scien i ic o al p esen a ions.
Keywo ds: Tex Summa iza ion, Spon aneous Speech Summa iza ion.
1 In oduc ion
La ge amoun s o digi al in o ma ion a e p oduced on a daily basis in he con-
ex o human in e ac ion e en s, such as news epo s, scien i ic p esen a ions,
mee ings, e c. Documen s gene a ed om hese e en s can be o di e en na-
u e in e ms o hei media (e.g., w i en, audio, ideo), hei domain (e.g.,
jou nalism, poli ics, esea ch, business), o he scena io hey o igina e om
(e.g., newspape , elephone con e sa ion, poli ical speech, business mee ing,
?This esea ch has been pa ially unded by he Eu opean Commission (CHIL
p ojec , IST-2004506969) and he Spanish go e nmen (TIN-2004-0171-E). We
would like o hank all CHIL e alua o s. We a e especially g a e ul o Daniel Fe ´es,
En ique Al onseca and Rose Sau ´ı.
Email add esses: [email p o ec ed] (Ma ia Fuen es),
[email p o ec ed] (Ho acio Rod ´ıguez), [email p o ec ed] (Jo di Tu mo).
P ep in submi ed o Else ie 19 Decembe 2006
o scien i ic con e ence). In his con ex , au oma ic summa iza ion can de i-
ni ely help deal wi h he inc easing amoun o a ailable in o ma ion.
Au oma ic documen summa iza ion s ongly depends no only on he o iginal
documen ea u es, bu also on he use needs (e.g., size o he summa y,
ou pu media, con en ela ed o a que y). Cu en s a e-o - he-a wo k in
he ield e lec s such a iabili y. Mos o he esea ch is based on English
newspape and newswi e da a, like he wo k de eloped in he con ex o he
Documen Unde s anding Con e ence 1(DUC), in bo h o i s wo e sions:
single documen (SDS) and mul idocumen summa iza ion (MDS). Less e o
has been de o ed o summa izing spon aneous speech, mos o he esea ch
ocusing on b oadcas news, ypically ead aloud om a w i en ex . Cu en
wo k on o al p esen a ions ends o be based on one single documen , he
speech ansc ip , [7], [5], [2], al hough some o he wo k ocus on di ec ly
summa izing he speech signal o lec u es [8].
S udies such as Sh ibe g [12] show ha o al communica ion is ha de o p o-
cess han w i en ex . Fo ha eason, we p opose using a mul idocumen
summa ize capable o handling documen s om di e en media ypes. Com-
bining documen s om di e en media can help coun e ac no only he di -
icul ies in he p ocessing o o al communica ion, bu also hose e o s in o-
duced by au oma ic speech ecognize s (ASRs). The cu en wo d e o a e
(WER) in speake -independen ASRs is a ound 25% on close- alking mic o-
phone, and 50% on a -dis ance. In his a icle we p esen FEMsum, a lexible
eclec ic mul i ask summa ize which will be e alua ed in wo di e en sce-
na ios: w i en news and scien i ic o al p esen a ions. FEMsum is a lexible
sys em capable o dealing wi h di e en summa iza ion asks by combining
in o ma ion om documen s o di e en so , i a ailable, and ha akes in o
accoun he use needs as well as he pa icula ea u es o he documen s o
be summa ized.
Nex sec ion gi es an o e iew o FEMsum’s gene al a chi ec u e. Sec ion 3
desc ibes he componen s o ou ool. We will see ha combining hem in
di e en ways leads o dis inc summa ize app oaches, ocusing he e o e on
MDS asks. In pa icula , wo di e en app oaches ha e been explo ed he e. A
i s one, based on lexical in o ma ion (LEX), and a second one, which uses a
iche seman ic ep esen a ion in o de o a oid edundancy and imp o e he
cohesion o he esul ing summa y (SEM). Sec ion 4 discusses he expe imen al
esul s ob ained om e alua ing bo h app oaches in di e en use -o ien ed,
que y- ocused mul idocumen summa iza ion asks. Finally, we conclude in
Sec ion 5 wi h he conclusions om he esea ch.
1h p:/www-nlpi .nis .go /p ojec s/duc/
2
2 Func ional O e iew o FEMsum’s A chi ec u e
The au oma ic summa iza ion sys em p esen ed he e is a highly modula
and pa ame e izable sys em able o deal wi h di e en in o ma ion needs. An
o e iew o he sys em ea u ing i s basic unc ionali ies and global a chi ec-
u e is depic ed in Figu e 1. The unc ional equi emen s a e se by means o a
pa ame e se spli ed in o inpu and ou pu se ings. Inpu se ings conce n
he cha ac e is ics o he documen s o be summa ized, while ou pu se ings
apply o he con en and p esen a ion o he summa y.
Inpu se ings include he ollowing pa ame e s:
•Domain. The summa ize can be domain independen o domain es ic ed.
In he second case, addi ional knowledge sou ces can be included, such as:
lis s o equen colloca ions om scien i ic pape s [5]; o a se o gaze ee s
o help inc ease he accu acy and ob ain ine classes in he named en i ies
classi ica ion ask es ic ed o he geog aphical domain.
•Documen s uc u e. The sys em can ake in o accoun in o ma ion de i ed
om he documen s uc u e (posi ion, i le, sec ions, o a ailable ags).
•Language. Cu en ly English and Spanish a e suppo ed. The linguis ic p o-
cessing pe o mance depends hea ily on his pa ame e .
•Media. Di e en media a e conside ed depending on he scena io o be deal
wi h ( ideo, audio, well w i en ex o any kind o ex ual documen ).
•Uni . Bo h single documen s and collec ions o ela ed documen s a e used
o SDS and MDS espec i ely.
•Gen e. We conside bo h gen e independen and dependen : jou nalis ic o
scien i ic (pape s, spon aneous speech, au ho no es and slides) op ions.
Ou pu se ings include:
•Con en . In o de o ex ac he ele an agmen s om he inpu doc-
umen s, he sys em can ake in o accoun he wo ds om he associa ed
que y (in case he e is a na u al language ques ion o lis o keywo ds), o
all he wo ds in he collec ion (i he e is no such que y).
•Size. Numbe o wo ds o he summa y.
•Documen es ic ion. Used o il e ing ou some ypes o documen s, such
as hose coming om a speci ic media o gen e.
•Ou pu o ma . The summa y can be p esen ed o he use as ex , syn he-
sized oice om ex , o as an audio/ ideo eco ded segmen .
3
3 FEMsum componen s
In o de o achie e he unc ionali ies p esen ed abo e, he sys em is o ga-
nized in h ee main componen s (see Figu e 1): Rele an In o ma ion De ec o
(RID), Con en Ex ac o (CE), and Summa y Compose (SC). In addi ion,
he e is a Que y P ocessing componen (QP). No all he componen s a e
needed o all he app oaches. In ac , in he expe imen s epo ed he e only
RID and SC a e always used.
3.1 Rele an In o ma ion De ec o
The RID module p o ides a anked se o ele an Tex Uni s (TU). The
de ini ion o TU depends basically on he inpu media. Fo ins ance, in he
case o well w i en ex , he TU is he sen ence. Two di e en s a egies a e
used, depending on he exis ence o a use que y (a NL ques ion o a lis o
keywo ds). I such a que y exis s, o each documen se he p onoun e e ence
is sol ed, he ex is lemma ized and indexed, and a Passage Re ie al (PR)
so wa e (JIRS [6] in he epo ed expe imen s) is used o ob ain he mos
ele an TUs. The sys em e ie es he passages wi h he highes simila i y
be ween he la ges n-g am o he que y and he one in he passage. RID
e u ns NTUs om passages ela ed o he que y. The de aul alue o Nis
no ixed, bu i is he numbe o TUs om passages selec ed in some o he
execu ions based on pa icula use need. I he summa y is no que y-d i en,
he sys em can use all TUs in he documen se , o pe o m a selec ion and
anking using a *id me ic o e he whole se .
3.2 Con en Ex ac o
As can be seen in Figu e 2, he CE, consis s o h ee componen s: a Linguis ic
P ocesso (LP), a Candida es Simila i y Ma ix Gene a o (CSMG), and a
Candida es Selec o (CS). Inpu o CE is he se o NTUs p o ided by RID.
All hese TUs a e p ocessed by LP. Then, CSMG compu es he simila i ies
among hem, and he mos app opia e ones a e p oposed by CE o be pa o
he summa y.
3.2.1 Linguis ic P ocesso
The LP is illus a ed in Figu e 3. I consis s o a pipeline o gene al pu pose NL
p ocesso s pe o ming: okeniza ion, POS agging, lemma iza ion, ine g ained
4
named en i ies ecogni ion and classi ica ion (NERC), syn ac ic pa sing, se-
man ic labeling (wi h Wo dNe synse s, Magnini’s domain ma ke s, and Eu-
oWo dNe Top Concep On ology labels), discou se ma ke anno a ion, and
seman ic analysis. Some o hese ools a e language dependen (English and
Spanish), while o he s a e gene al ools uned o a speci ic language. The same
ools a e used o he linguis ic p ocessing o he RID esul and o he que y
(QP) –when needed. The speci ic ools o be used in each case depend on he
ype o TU in ol ed and he inpu language (see [4] o mo e de ails).
Tools used o Spanish include:
•F eeLing, which pe o ms okeniza ion, mo phological analysis, POS ag-
ging, lemma iza ion, and pa ial pa sing.
•ABIONET, a NERC on basic MUC ca ego ies (pe son, loca ion, o ganiza-
ion, o he s).
•Eu oWo dNe , used o ob ain he lis o synse s (wi hou a emp ing Wo d
Sense Disambigua ion), a lis o hype nyms o each synse up o he op o
he axonomy, and he Top Concep On ology class.
Tools used o English include:
•TnT, a s a is ical POS agge .
•Wo dNe lemma ize 2.0.
•ABIONET.
•Wo dNe
•A modi ied e sion o he Collins’ pa se which pe o ms ull pa sing and
obus de ec ion o e bal p edica e a gumen s.
•Alembic, a NERC wi h MUC classes, used in o de o boos ABIONET
pe o mance.
As a esul , sen ences a e en iched wi h lexical (sen ) and syn ac ic (sin ) lan-
guage dependen ep esen a ions. Fo each sen ence, i s syn ac ic cons i uen
s uc u e (including head speci ica ion) and he syn ac ic ela ions be ween
i s cons i uen s (subjec , di ec and indi ec objec , modi ie s) a e ob ained.
F om sen and sin , a seman ic ep esen a ion o he sen ence is p oduced,
he en i onmen (en ). The in o ma ion in each o hese componen s is he
ollowing:
Sen p o ides lexical in o ma ion o each wo d: o m, lemma, POS ag, se-
man ic class o NE, lis o WN o EWN synse s and, whene e possible,
de i a ional in o ma ion.
Sin con ains wo lis s: one eco ding he syn ac ic cons i uen s uc u e (ba-
sically nominal, p eposi ional, and e bal ph ases), and he o he ep esen -
ing he dependencies be ween hese cons i uen s.
En is a seman ic-ne wo k-like ep esen a ion compu ed using a p ocess ha
ex ac s he seman ic uni s (nodes) and he seman ic ela ions (edges) hold-
5

ing be ween he di e en okens in sen . Uni and ela ion ypes belong o an
on ology o abou 100 seman ic classes (as pe son, ci y, ac ion, magni ude,
e c.), and 25 ela ions be ween hem (mos ly bina y, as ime o e en , ac-
o o ac ion, loca ion o e en , e c.). Bo h classes and ela ions a e ela ed
by axonomic links (see [4] o de ails) allowing o inhe i ance. Table 1 p o-
ides an example o a sen ence en i onmen (en ), and Figu e 4 gi es he
comple e ep esen a ion o ano he sen ence.
3.2.2 Candida es Simila i y Ma ix Gene a o
CSMG is in cha ge o compu ing he simila i y ma ix among candida es. Fo
ha pu pose, i uses he en i onmen o each candida e TU (sen ence in he
epo ed expe imen s). En i onmen s a e ans o med in o labeled di ec ed
g aph ep esen a ion, whe e nodes a e assigned o posi ions in he sen ence
and labeled wi h he co esponding oken, and edges a e assigned o p edica es
(a dummy node, 0, is used o ep esen ing una y p edica es). Only una y
and bina y p edica es a e used. Figu e 5 is he g aph ep esen a ion o he
en i onmen in Table 1.
On op o his ep esen a ion, a ich panoply o lexico-seman ic p oximi y mea-
su es be ween sen ences ha e been buil . Each measu e combines wo compo-
nen s:
•A lexical componen which includes he se o common okens, i.e. hose
occu ing in bo h sen ences. The size o his se and he s eng h o he
compa ibili y links be ween i s membe s a e used o de ining he measu e.
A lexible way o measu ing oken-le el compa ibili y has been empi ically
se , anging om wo d- o m iden i y, lemma iden i y, o e lapping o Wo d-
Ne synse s, app oxima e s ing ma ching be ween Named En i ies e c. Fo
ins ance, ”Romano P odi” is lexically compa ible wi h ”R. P odi” wi h a
sco e o 0.5 and wi h ”P odi” wi h a sco e o 0.41. ”I aly” and ”I alian” a e
also compa ible wi h sco e 0.7.
•A seman ic componen , compu ed o e he subg aphs co esponding o he
se o lexically compa ible nodes. Fou di e en measu es ha e been de ined:
·S ic o e lapping o una y p edica es.
·S ic o e lapping o bina y p edica es.
·Loose o e lapping o una y p edica es.
·Loose o e lapping o bina y p edica es.
The loose e sions allow a elaxed ma ching o p edica es by climbing up in
he on ology o p edica es, e.g. p o ided ha A and B a e lexically compa -
ible, i en ci y(A) can ma ch i en p ope place(B),loca ion(B) o en i y(B).
Ob iously, loose o e lapping implies a penal y on he sco e.
Se e al ways o combining he simple sco es ha e been conside ed and es ed.
6
Once an app op ia e measu e has been selec ed, we can compu e he simila i y
be ween e e y sen ence pai .
3.2.3 Candida es Selec o
In o de o selec he candida es, h ee c i e ia ha e been aken in o accoun :
•Rele ance (wi h espec o he que y o any o he elemen )
•Densi y and cohesion
•An i- edundancy
CS p oceeds in he ollowing s eps:
Le Sim be he simila i y ma ix, Candida es a lis o candida e TUs, and
Summa y an o de ed lis o TUs o be included in he summa y.
(1) Se Candida es o he lis p o ided by RID componen .
(2) Se Summa y o he emp y lis .
(3) Se Sim o he ma ix con aining he simila i y alues be ween membe s
om Candida es.
(4) Fo each candida e in Candida es, compu e a sco e ha akes in o ac-
coun he ini ial ele ance sco e and he alues in Sim. The sco e used
is based on PageRank, as used by Mihalcea and Ta au [10], bu wi hou
making he dis inc ion be ween inpu and ou pu links.
(5) So Candida es by his sco e.
(6) Append he mos sco ed candida e ( he head o he lis ) o he Summa y
and emo e i om Candida es.
(7) In o de o p e en o e lapping, he S% TUs mos simila (using Sim) o
he one selec ed in he p e ious s ep a e emo ed as well om Candida es.
The R% leas sco ed TUs a e also emo ed om Candida es.
(8) I Candida es is no emp y go o 4.
3.3 Summa y Compose
Fo he summa y composi ion, wo di e en app oaches ha e been explo ed.
The i s one is based on lexical in o ma ion (LEX). The second one uses a
iche seman ic ep esen a ion in o de o a oid edundancy and o imp o e
he cohesion o he esul ing summa y (SEM).
The inpu se o candida e TUs in he LEX app oach consis s o hose TUs
p e iously de ec ed by he RID componen as ele an acco ding o he opic.
In con as , in he SEM app oach, he inpu se consis s o hose TUs ex-
ac ed by he CE componen . Summa y TUs a e selec ed by ele ance un il
7
he desi ed summa y size is achie ed. Fo each selec ed TU, i is checked
whe he he p e ious sen ence in he o iginal documen is also a candida e.
I posi i e, bo h a e added o he Summa y in he o de hey appea in he
o iginal documen .
4 Que y o ien ed MDS e alua ion
In his sec ion, we desc ibe he expe imen s we ca ied ou in o de o e alua e
ou sys em’s pe o mance when dealing wi h di e en que y-o ien ed MDS
asks. The e alua ion is pe o med on wo asks. On one hand, Sec ion 4.1
epo s he expe imen s and esul s o FEMsum in he in e na ional DUC
2006 e alua ion o w i en news scena io. On he o he hand, Sec ion 4.2
analyzes he esul s ob ained in a simila ask wi hin he amewo k o he
CHIL 2p ojec o scien i ic o al p esen a ion documen s.
The asks in he DUC 2006 and CHIL amewo ks di e in he ollowing
aspec s:
•que y (complex s. lis o keywo ds),
•summa y leng h (250 s. 100),
•numbe o inpu documen (25 s. abou 4),
•e alua ion es (50 opics s 20, 10 opics x 2 que ies).
•numbe o manual summa y models (4 abs ac based s. 3 ex ac based).
•domain (jou nalis ic s scien i ic),
•inpu media (well w i en ex s. aw ex comming om: spon aneous
speech, and documen s ela ed o an o al p esen a ion),
•gen e (w i en jou nalism s scien i ic spon aneous speech, esea ch pape s,
and au ho no es o slides om a p esen a ion).
4.1 Expe imen s in w i en news scena io
This sec ion desc ibes he DUC 2006 e alua ion amewo k, ou app oach
se ings o each kind o summa y p oduced in ou DUC 2006 pa icipa ion,
and he esul s ob ained.
4.1.1 E alua ion F amewo k
Fo he DUC 2006 e alua ion, we we e p o ided wi h 50 opics which had been
selec ed o be used as es da a. Each opic had assigned a clus e o 25 ela ed
2h p://chil.se e .de/se le /is/101/
8
ex ual news documen s, as well as a s a emen desc ibing he in o ma ion ha
could be answe ed using his documen clus e . The opic s a emen could be
in he o m o a ques ion o se o ela ed ques ions and can include backg ound
in o ma ion ha he assesso has conside ed would cla i y his/he in o ma ion
need. Fo each opic 4 manual summa ies we e p oduced a NIST.
The DUC baseline was a simply sys em ha e u ned all he leading sen-
ences (up o 250 wo ds) o he mos ecen documen . All 34 DUC 2006
pa icipa ing sys ems and he baseline we e e alua ed a wo le els: manu-
ally (Linguis ic quali y and Responsi eness) and au oma ically (me ics om
he package ROUGE, he Recall-O ien ed Unde s udy o Gis ing E alua ion
[9]). Manual e alua ion sco ed each aspec o a gi en summa y as 1: e y poo ,
2:poo , 3:accep able, 4:good, o 5: e y good. In addi ion, ou un is one o
he 21 pa icipan sys ems ha we e also manually e alua ed by means o he
py amid me hod [11].
4.1.2 FEMsum se ings in DUC 2006
Ou goal in pa icipa ing a DUC was o e alua e a numbe o aspec s o
ou sys em. We he e o e submi ed h ee di e en kinds o au oma ic sum-
ma ies in a single un: one lexically based (LEX), and wo seman ically based
(SEM150, SEM250). Ou o he 50 summa ies we we e expec ed o submi , 7
we e p oduced using he LEX app oach, 13 by means o he SEM150 s a egy,
and 30 by using he SEM250 one. Ou sys em was assigned he iden i ica ion
numbe 19.
Gi en he que y, a common c ucial s ep in all he app oaches is o de ec
he mos ele an TUs (sen ences in his expe imen s). We decided o ix a
maximum numbe o sen ences de ec ed as ele an by RID. Fo ha eason
we use he co pus o sen ences de ec ed as pa o a manual summa y in DUC
2005 p oposed by Copeck and Spakowicz [3]. Analyzing P ecision and Recall
Nwas empi ically ixed in a maximum o 250.
In he LEX app oach, ele an sen ences a e de ec ed by RID and hen SC is
applied o ob ain he summa ies. On he o he hand, in bo h SEM app oaches
he ini ial c i e ia o sen ence ele ance is ha sen ences om a same documen
a e conside ed o ha e a simila ele ance, independen ly o he RID sco e. In
he SEM250 s a egy, all he sen ences om he RID ou pu a e aken as CE
inpu , whe eas in SEM150 he inpu o CE is he clus e o 150 sen ences
om he i s documen s in he se . SEM150 ends o educe he numbe o
documen s whose con en is candida e o appea in he summa y.
9
Table 1. Sample o en i onmen buil om a sen ence.
Table 2. DUC FEMsum manually e alua ion o linguis ic quali y sco es by
app oach.
Table 3. DUC FEMsum manual con en esponsi eness sco e.
Table 4. DUC Con en esponsi eness sco es by app oach.
Table 5. DUC Con en esponsi eness sco es dis ibu ion by app oach
Table 6. DUC ROUGE measu es when conside ing 4 manual summa ies as
e e ences.
Table 7. CHIL ROUGE measu es when conside ed 3 manual summa ies as
e e ences.
Table 8. CHIL esponsi eness conside ing 3 human models when e alua ing
au oma ic summa ies and 2 when e alua ing human summa ies.
Table 9. CHIL esponsi eness sco es dis ibu ion by au oma ic sys em.
Figu e 1. FEMsum Global A chi ec u e
Figu e 2. Con en Ex ac o modul
Figu e 3. Linguis ic P ocesso cons i uen s
Figu e 4. Sample o a sen ence analysis and he co esponden en i onmen
ep eseen a ion
Figu e 5. Sample o he g aph ep esen a ion o an En i onmen
16

Table 1
Sample o en i onmen buil om a sen ence
”Romano P odi 1is 2 he 3p ime 4minis e 5o 6I aly 7”
i en p ope pe son(1), en i y has quali y(2), en i y(5), i en coun y(7),
quali y(4), which en i y(2,1), which quali y(2,5), mod(5,7), mod(5,4)
Table 2
DUC FEMsum manually e alua ion o linguis ic quali y sco es by app oach.
G amma icali y Non- edundancy Re e en ial cla i y Focus
FEMsum mean FEMsum mean FEMsum mean FEMsum mean
LEX 3,14 3,45 2,43 4,02 2,43 2,83 3,29 3,73
S150 3,00 3,60 4,15 4,19 3,08 3,09 3,77 3,84
S250 3,33 3,59 4,23 4,27 2,77 3,12 3,20 3,42
Table 3
DUC FEMsum manual con en esponsi eness sco e.
Sys em(ID) Sco e Mean Dis ance
Human(A-J) 4,75 2,19
Bes (27) 3,08 1,83
FEMsum(19) 2,60 0,04
+www(30) 2,58 0,02
Baseline(1) 2,04 -0,52
Mean(2-35) 2,56 S de 0,28
Table 4
DUC Con en esponsi eness sco es by app oach.
Mean(1-35) FEMsum Mean Dis ance
LEX 2,36 2,29 -0,07
SEM150 2,55 2,92 0,37
SEM250 2,58 2,53 -0,05
Table 5
DUC Con en esponsi eness sco es dis ibu ion by app oach
1: Ve y Poo 2: Poo 3: Accep able 4: Good 5: Ve y Good
LEX 14% 43% 43% 0% 0%
SEM150 7,5% 31% 31% 23% 7,5%
SEM250 6,7% 50% 26,7% 16,7% 0%
17
Table 6
DUC ROUGE measu es when conside ing 4 manual summa ies as e e ences.
Bes (24) FEMsum(19) +www(30) Baseline(1)
R-2 0,095 0,076 0,067 0,050
R-SU4 0,155 0,131 0,122 0,098
Table 7
CHIL ROUGE measu es when conside ed 3 manual summa ies as e e ences.
SDS LEX LEXnoT +www SEM
ROUGE-1 0,293 0,309 0,312 0,333 0,323
ROUGE-2 0,060 0,092 0,102 0,089 0,073
ROUGE-3 0,029 0,056 0,064 0,052 0,032
ROUGE-4 0,019 0,043 0,050 0,043 0,021
ROUGE-L 0,256 0,272 0,279 0,289 0,280
ROUGE-W1.2 0,089 0,098 0,100 0,104 0,098
ROUGE-S1 0,057 0,088 0,097 0,087 0,067
ROUGE-S4 0,064 0,089 0,095 0,094 0,073
ROUGE-S9 0,069 0,095 0,102 0,103 0,083
ROUGE-SU1 0,136 0,162 0,169 0,168 0,152
ROUGE-SU4 0,102 0,126 0,132 0,134 0,115
ROUGE-SU9 0,090 0,116 0,122 0,124 0,105
Table 8
CHIL esponsi eness conside ing 3 human models when e alua ing au oma ic sum-
ma ies and 2 when e alua ing human summa ies.
M1 M2 M3 SDS LEX LEXnoT +www SEM
3,625 3,400 3,375 1,250 1,775 2,025 1,800 1,800
Table 9
CHIL esponsi eness sco es dis ibu ion by au oma ic sys em.
M1 M2 M3 SDS LEX LEXnoT +www SEM
1: Ve y Poo 0% 0% 0% 70% 40% 15% 30% 35%
2: Poo 10% 5% 10% 25% 25% 50% 50% 40%
3: Accep able 20% 35% 30% 5% 35% 30% 15% 20%
4: Good 40% 40% 45% 0% 0% 5% 5% 5%
5: Ve y Good 30% 20% 15% 0% 0% 0% 0% 0%
18
pape s
slides
no es
aw
ex
Que y
Con en
Ex ac o
Summa y
Summa y
Compose
ASR
ansc ip
Use Need
OUTPUT Se ings
CONTENT
DOC_RESTRICT
SIZE
FORMAT
No ansc ip s
Tex
Num. wo ds
Keywo ds
Gene ic
Speech/ ideo
In o ma ion
Rele an
De ec o
NLs a emen
All
Sin hesizedTex
MDS
Independen
P esen a ion
Tex
English
No used
Used
SDS
Scien i ic
DOC_STRUCTURE
LANGUAGE
MEDIA
UNIT
Spanish
GENRE
Jou nalis ic
INPUT Se ings
DOMAIN
Independen
Scien i ic
News
Fig. 1. FEMsum Global A chi ec u e
Linguis ic
P ocesso
sen
en
sin
Tex Uni s Selec o
Candida es Rele an
Tex Uni s
sim
CONTENT EXTRACTOR
Candida es
Simila i y Ma ix
Gene a o
Fig. 2. Con en Ex ac o modul
Tokenize POS
Tagge Lemma ize NERC Syn ac ic
Chunke
Seman ic
Tagge Anno a o
DM Seman ic
Analize
Discou se
Ma ke s
Wo dNe
Use
Need
Tex
Uni s
sen
sin
en
LINGUISTIC PROCESSOR
Fig. 3. Linguis ic P ocesso cons i uen s
19
Fig. 4. Sample o a sen ence analysis and he co esponden en i onmen ep eseen-
a ion
Romano P odi
1
is
2
p ime
4
minis e
57
I aly
0
which_en i y which_quali y
mod mod
en i y_has_quali y
i_en_p ope _pe son
quali y
en i y
i_en_coun y
Fig. 5. Sample o he g aph ep esen a ion o an En i onmen
20