scieee Open visual document viewer

FEMsum: A flexible eclectic multitask summarizer architecture evaluated in multidocument tasks

Fuentes Fort, Maria,Rodríguez Hontoria, Horacio,Turmo Borras, Jorge

Abstract

This article describes two types of summarization approaches integrated in a flexible architecture for multitask summarization. The first type is based on the use of lexical features, while the second one is grounded on syntactic and semantic information. All the approaches have been evaluated in experiments where, given a set of documents, they are expected to produce summaries answering a user need (expressed by a query) in a reduced set of relevant textual fragments. Their performance is analyzed in two different tasks: written news and scientific oral presentations.

Full text

FEMsum: A Flexible Eclec ic Mul i ask Summa ize A chi ec u e e alua ed in Mul idocumen Tasks ? . Ma ia Fuen es, Ho acio Rod ´ıguez, Jo di Tu mo TALP Resea ch Cen e , Uni e si a Poli `ecnica de Ca alunya Ba celona, Spain Abs ac This a icle desc ibes wo ypes o summa iza ion app oaches in eg a ed in a lexible a chi ec u e o mul i ask summa iza ion. The i s ype is based on he use o lexical ea u es, while he second one is g ounded on syn ac ic and seman ic in o ma ion. All he app oaches ha e been e alua ed in expe imen s whe e, gi en a se o docu- men s, hey a e expec ed o p oduce summa ies answe ing a use need (exp essed by a que y) in a educed se o ele an ex ual agmen s. Thei pe o mance is analyzed in wo di e en asks: w i en news and scien i ic o al p esen a ions. Keywo ds: Tex Summa iza ion, Spon aneous Speech Summa iza ion. 1 In oduc ion La ge amoun s o digi al in o ma ion a e p oduced on a daily basis in he con- ex o human in e ac ion e en s, such as news epo s, scien i ic p esen a ions, mee ings, e c. Documen s gene a ed om hese e en s can be o di e en na- u e in e ms o hei media (e.g., w i en, audio, ideo), hei domain (e.g., jou nalism, poli ics, esea ch, business), o he scena io hey o igina e om (e.g., newspape , elephone con e sa ion, poli ical speech, business mee ing, ?This esea ch has been pa ially unded by he Eu opean Commission (CHIL p ojec , IST-2004506969) and he Spanish go e nmen (TIN-2004-0171-E). We would like o hank all CHIL e alua o s. We a e especially g a e ul o Daniel Fe ´es, En ique Al onseca and Rose Sau ´ı. Email add esses: [email p o ec ed] (Ma ia Fuen es), [email p o ec ed] (Ho acio Rod ´ıguez), [email p o ec ed] (Jo di Tu mo). P ep in submi ed o Else ie 19 Decembe 2006 o scien i ic con e ence). In his con ex , au oma ic summa iza ion can de i- ni ely help deal wi h he inc easing amoun o a ailable in o ma ion. Au oma ic documen summa iza ion s ongly depends no only on he o iginal documen ea u es, bu also on he use needs (e.g., size o he summa y, ou pu media, con en ela ed o a que y). Cu en s a e-o - he-a wo k in he ield e lec s such a iabili y. Mos o he esea ch is based on English newspape and newswi e da a, like he wo k de eloped in he con ex o he Documen Unde s anding Con e ence 1(DUC), in bo h o i s wo e sions: single documen (SDS) and mul idocumen summa iza ion (MDS). Less e o has been de o ed o summa izing spon aneous speech, mos o he esea ch ocusing on b oadcas news, ypically ead aloud om a w i en ex . Cu en wo k on o al p esen a ions ends o be based on one single documen , he speech ansc ip , [7], [5], [2], al hough some o he wo k ocus on di ec ly summa izing he speech signal o lec u es [8]. S udies such as Sh ibe g [12] show ha o al communica ion is ha de o p o- cess han w i en ex . Fo ha eason, we p opose using a mul idocumen summa ize capable o handling documen s om di e en media ypes. Com- bining documen s om di e en media can help coun e ac no only he di - icul ies in he p ocessing o o al communica ion, bu also hose e o s in o- duced by au oma ic speech ecognize s (ASRs). The cu en wo d e o a e (WER) in speake -independen ASRs is a ound 25% on close- alking mic o- phone, and 50% on a -dis ance. In his a icle we p esen FEMsum, a lexible eclec ic mul i ask summa ize which will be e alua ed in wo di e en sce- na ios: w i en news and scien i ic o al p esen a ions. FEMsum is a lexible sys em capable o dealing wi h di e en summa iza ion asks by combining in o ma ion om documen s o di e en so , i a ailable, and ha akes in o accoun he use needs as well as he pa icula ea u es o he documen s o be summa ized. Nex sec ion gi es an o e iew o FEMsum’s gene al a chi ec u e. Sec ion 3 desc ibes he componen s o ou ool. We will see ha combining hem in di e en ways leads o dis inc summa ize app oaches, ocusing he e o e on MDS asks. In pa icula , wo di e en app oaches ha e been explo ed he e. A i s one, based on lexical in o ma ion (LEX), and a second one, which uses a iche seman ic ep esen a ion in o de o a oid edundancy and imp o e he cohesion o he esul ing summa y (SEM). Sec ion 4 discusses he expe imen al esul s ob ained om e alua ing bo h app oaches in di e en use -o ien ed, que y- ocused mul idocumen summa iza ion asks. Finally, we conclude in Sec ion 5 wi h he conclusions om he esea ch. 1h p:/www-nlpi .nis .go /p ojec s/duc/ 2 2 Func ional O e iew o FEMsum’s A chi ec u e The au oma ic summa iza ion sys em p esen ed he e is a highly modula and pa ame e izable sys em able o deal wi h di e en in o ma ion needs. An o e iew o he sys em ea u ing i s basic unc ionali ies and global a chi ec- u e is depic ed in Figu e 1. The unc ional equi emen s a e se by means o a pa ame e se spli ed in o inpu and ou pu se ings. Inpu se ings conce n he cha ac e is ics o he documen s o be summa ized, while ou pu se ings apply o he con en and p esen a ion o he summa y. Inpu se ings include he ollowing pa ame e s: •Domain. The summa ize can be domain independen o domain es ic ed. In he second case, addi ional knowledge sou ces can be included, such as: lis s o equen colloca ions om scien i ic pape s [5]; o a se o gaze ee s o help inc ease he accu acy and ob ain ine classes in he named en i ies classi ica ion ask es ic ed o he geog aphical domain. •Documen s uc u e. The sys em can ake in o accoun in o ma ion de i ed om he documen s uc u e (posi ion, i le, sec ions, o a ailable ags). •Language. Cu en ly English and Spanish a e suppo ed. The linguis ic p o- cessing pe o mance depends hea ily on his pa ame e . •Media. Di e en media a e conside ed depending on he scena io o be deal wi h ( ideo, audio, well w i en ex o any kind o ex ual documen ). •Uni . Bo h single documen s and collec ions o ela ed documen s a e used o SDS and MDS espec i ely. •Gen e. We conside bo h gen e independen and dependen : jou nalis ic o scien i ic (pape s, spon aneous speech, au ho no es and slides) op ions. Ou pu se ings include: •Con en . In o de o ex ac he ele an agmen s om he inpu doc- umen s, he sys em can ake in o accoun he wo ds om he associa ed que y (in case he e is a na u al language ques ion o lis o keywo ds), o all he wo ds in he collec ion (i he e is no such que y). •Size. Numbe o wo ds o he summa y. •Documen es ic ion. Used o il e ing ou some ypes o documen s, such as hose coming om a speci ic media o gen e. •Ou pu o ma . The summa y can be p esen ed o he use as ex , syn he- sized oice om ex , o as an audio/ ideo eco ded segmen . 3 3 FEMsum componen s In o de o achie e he unc ionali ies p esen ed abo e, he sys em is o ga- nized in h ee main componen s (see Figu e 1): Rele an In o ma ion De ec o (RID), Con en Ex ac o (CE), and Summa y Compose (SC). In addi ion, he e is a Que y P ocessing componen (QP). No all he componen s a e needed o all he app oaches. In ac , in he expe imen s epo ed he e only RID and SC a e always used. 3.1 Rele an In o ma ion De ec o The RID module p o ides a anked se o ele an Tex Uni s (TU). The de ini ion o TU depends basically on he inpu media. Fo ins ance, in he case o well w i en ex , he TU is he sen ence. Two di e en s a egies a e used, depending on he exis ence o a use que y (a NL ques ion o a lis o keywo ds). I such a que y exis s, o each documen se he p onoun e e ence is sol ed, he ex is lemma ized and indexed, and a Passage Re ie al (PR) so wa e (JIRS [6] in he epo ed expe imen s) is used o ob ain he mos ele an TUs. The sys em e ie es he passages wi h he highes simila i y be ween he la ges n-g am o he que y and he one in he passage. RID e u ns NTUs om passages ela ed o he que y. The de aul alue o Nis no ixed, bu i is he numbe o TUs om passages selec ed in some o he execu ions based on pa icula use need. I he summa y is no que y-d i en, he sys em can use all TUs in he documen se , o pe o m a selec ion and anking using a *id me ic o e he whole se . 3.2 Con en Ex ac o As can be seen in Figu e 2, he CE, consis s o h ee componen s: a Linguis ic P ocesso (LP), a Candida es Simila i y Ma ix Gene a o (CSMG), and a Candida es Selec o (CS). Inpu o CE is he se o NTUs p o ided by RID. All hese TUs a e p ocessed by LP. Then, CSMG compu es he simila i ies among hem, and he mos app opia e ones a e p oposed by CE o be pa o he summa y. 3.2.1 Linguis ic P ocesso The LP is illus a ed in Figu e 3. I consis s o a pipeline o gene al pu pose NL p ocesso s pe o ming: okeniza ion, POS agging, lemma iza ion, ine g ained 4 named en i ies ecogni ion and classi ica ion (NERC), syn ac ic pa sing, se- man ic labeling (wi h Wo dNe synse s, Magnini’s domain ma ke s, and Eu- oWo dNe Top Concep On ology labels), discou se ma ke anno a ion, and seman ic analysis. Some o hese ools a e language dependen (English and Spanish), while o he s a e gene al ools uned o a speci ic language. The same ools a e used o he linguis ic p ocessing o he RID esul and o he que y (QP) –when needed. The speci ic ools o be used in each case depend on he ype o TU in ol ed and he inpu language (see [4] o mo e de ails). Tools used o Spanish include: •F eeLing, which pe o ms okeniza ion, mo phological analysis, POS ag- ging, lemma iza ion, and pa ial pa sing. •ABIONET, a NERC on basic MUC ca ego ies (pe son, loca ion, o ganiza- ion, o he s). •Eu oWo dNe , used o ob ain he lis o synse s (wi hou a emp ing Wo d Sense Disambigua ion), a lis o hype nyms o each synse up o he op o he axonomy, and he Top Concep On ology class. Tools used o English include: •TnT, a s a is ical POS agge . •Wo dNe lemma ize 2.0. •ABIONET. •Wo dNe •A modi ied e sion o he Collins’ pa se which pe o ms ull pa sing and obus de ec ion o e bal p edica e a gumen s. •Alembic, a NERC wi h MUC classes, used in o de o boos ABIONET pe o mance. As a esul , sen ences a e en iched wi h lexical (sen ) and syn ac ic (sin ) lan- guage dependen ep esen a ions. Fo each sen ence, i s syn ac ic cons i uen s uc u e (including head speci ica ion) and he syn ac ic ela ions be ween i s cons i uen s (subjec , di ec and indi ec objec , modi ie s) a e ob ained. F om sen and sin , a seman ic ep esen a ion o he sen ence is p oduced, he en i onmen (en ). The in o ma ion in each o hese componen s is he ollowing: Sen p o ides lexical in o ma ion o each wo d: o m, lemma, POS ag, se- man ic class o NE, lis o WN o EWN synse s and, whene e possible, de i a ional in o ma ion. Sin con ains wo lis s: one eco ding he syn ac ic cons i uen s uc u e (ba- sically nominal, p eposi ional, and e bal ph ases), and he o he ep esen - ing he dependencies be ween hese cons i uen s. En is a seman ic-ne wo k-like ep esen a ion compu ed using a p ocess ha ex ac s he seman ic uni s (nodes) and he seman ic ela ions (edges) hold- 5 ing be ween he di e en okens in sen . Uni and ela ion ypes belong o an on ology o abou 100 seman ic classes (as pe son, ci y, ac ion, magni ude, e c.), and 25 ela ions be ween hem (mos ly bina y, as ime o e en , ac- o o ac ion, loca ion o e en , e c.). Bo h classes and ela ions a e ela ed by axonomic links (see [4] o de ails) allowing o inhe i ance. Table 1 p o- ides an example o a sen ence en i onmen (en ), and Figu e 4 gi es he comple e ep esen a ion o ano he sen ence. 3.2.2 Candida es Simila i y Ma ix Gene a o CSMG is in cha ge o compu ing he simila i y ma ix among candida es. Fo ha pu pose, i uses he en i onmen o each candida e TU (sen ence in he epo ed expe imen s). En i onmen s a e ans o med in o labeled di ec ed g aph ep esen a ion, whe e nodes a e assigned o posi ions in he sen ence and labeled wi h he co esponding oken, and edges a e assigned o p edica es (a dummy node, 0, is used o ep esen ing una y p edica es). Only una y and bina y p edica es a e used. Figu e 5 is he g aph ep esen a ion o he en i onmen in Table 1. On op o his ep esen a ion, a ich panoply o lexico-seman ic p oximi y mea- su es be ween sen ences ha e been buil . Each measu e combines wo compo- nen s: •A lexical componen which includes he se o common okens, i.e. hose occu ing in bo h sen ences. The size o his se and he s eng h o he compa ibili y links be ween i s membe s a e used o de ining he measu e. A lexible way o measu ing oken-le el compa ibili y has been empi ically se , anging om wo d- o m iden i y, lemma iden i y, o e lapping o Wo d- Ne synse s, app oxima e s ing ma ching be ween Named En i ies e c. Fo ins ance, ”Romano P odi” is lexically compa ible wi h ”R. P odi” wi h a sco e o 0.5 and wi h ”P odi” wi h a sco e o 0.41. ”I aly” and ”I alian” a e also compa ible wi h sco e 0.7. •A seman ic componen , compu ed o e he subg aphs co esponding o he se o lexically compa ible nodes. Fou di e en measu es ha e been de ined: ·S ic o e lapping o una y p edica es. ·S ic o e lapping o bina y p edica es. ·Loose o e lapping o una y p edica es. ·Loose o e lapping o bina y p edica es. The loose e sions allow a elaxed ma ching o p edica es by climbing up in he on ology o p edica es, e.g. p o ided ha A and B a e lexically compa - ible, i en ci y(A) can ma ch i en p ope place(B),loca ion(B) o en i y(B). Ob iously, loose o e lapping implies a penal y on he sco e. Se e al ways o combining he simple sco es ha e been conside ed and es ed. 6 Once an app op ia e measu e has been selec ed, we can compu e he simila i y be ween e e y sen ence pai . 3.2.3 Candida es Selec o In o de o selec he candida es, h ee c i e ia ha e been aken in o accoun : •Rele ance (wi h espec o he que y o any o he elemen ) •Densi y and cohesion •An i- edundancy CS p oceeds in he ollowing s eps: Le Sim be he simila i y ma ix, Candida es a lis o candida e TUs, and Summa y an o de ed lis o TUs o be included in he summa y. (1) Se Candida es o he lis p o ided by RID componen . (2) Se Summa y o he emp y lis . (3) Se Sim o he ma ix con aining he simila i y alues be ween membe s om Candida es. (4) Fo each candida e in Candida es, compu e a sco e ha akes in o ac- coun he ini ial ele ance sco e and he alues in Sim. The sco e used is based on PageRank, as used by Mihalcea and Ta au [10], bu wi hou making he dis inc ion be ween inpu and ou pu links. (5) So Candida es by his sco e. (6) Append he mos sco ed candida e ( he head o he lis ) o he Summa y and emo e i om Candida es. (7) In o de o p e en o e lapping, he S% TUs mos simila (using Sim) o he one selec ed in he p e ious s ep a e emo ed as well om Candida es. The R% leas sco ed TUs a e also emo ed om Candida es. (8) I Candida es is no emp y go o 4. 3.3 Summa y Compose Fo he summa y composi ion, wo di e en app oaches ha e been explo ed. The i s one is based on lexical in o ma ion (LEX). The second one uses a iche seman ic ep esen a ion in o de o a oid edundancy and o imp o e he cohesion o he esul ing summa y (SEM). The inpu se o candida e TUs in he LEX app oach consis s o hose TUs p e iously de ec ed by he RID componen as ele an acco ding o he opic. In con as , in he SEM app oach, he inpu se consis s o hose TUs ex- ac ed by he CE componen . Summa y TUs a e selec ed by ele ance un il 7 he desi ed summa y size is achie ed. Fo each selec ed TU, i is checked whe he he p e ious sen ence in he o iginal documen is also a candida e. I posi i e, bo h a e added o he Summa y in he o de hey appea in he o iginal documen . 4 Que y o ien ed MDS e alua ion In his sec ion, we desc ibe he expe imen s we ca ied ou in o de o e alua e ou sys em’s pe o mance when dealing wi h di e en que y-o ien ed MDS asks. The e alua ion is pe o med on wo asks. On one hand, Sec ion 4.1 epo s he expe imen s and esul s o FEMsum in he in e na ional DUC 2006 e alua ion o w i en news scena io. On he o he hand, Sec ion 4.2 analyzes he esul s ob ained in a simila ask wi hin he amewo k o he CHIL 2p ojec o scien i ic o al p esen a ion documen s. The asks in he DUC 2006 and CHIL amewo ks di e in he ollowing aspec s: •que y (complex s. lis o keywo ds), •summa y leng h (250 s. 100), •numbe o inpu documen (25 s. abou 4), •e alua ion es (50 opics s 20, 10 opics x 2 que ies). •numbe o manual summa y models (4 abs ac based s. 3 ex ac based). •domain (jou nalis ic s scien i ic), •inpu media (well w i en ex s. aw ex comming om: spon aneous speech, and documen s ela ed o an o al p esen a ion), •gen e (w i en jou nalism s scien i ic spon aneous speech, esea ch pape s, and au ho no es o slides om a p esen a ion). 4.1 Expe imen s in w i en news scena io This sec ion desc ibes he DUC 2006 e alua ion amewo k, ou app oach se ings o each kind o summa y p oduced in ou DUC 2006 pa icipa ion, and he esul s ob ained. 4.1.1 E alua ion F amewo k Fo he DUC 2006 e alua ion, we we e p o ided wi h 50 opics which had been selec ed o be used as es da a. Each opic had assigned a clus e o 25 ela ed 2h p://chil.se e .de/se le /is/101/ 8 ex ual news documen s, as well as a s a emen desc ibing he in o ma ion ha could be answe ed using his documen clus e . The opic s a emen could be in he o m o a ques ion o se o ela ed ques ions and can include backg ound in o ma ion ha he assesso has conside ed would cla i y his/he in o ma ion need. Fo each opic 4 manual summa ies we e p oduced a NIST. The DUC baseline was a simply sys em ha e u ned all he leading sen- ences (up o 250 wo ds) o he mos ecen documen . All 34 DUC 2006 pa icipa ing sys ems and he baseline we e e alua ed a wo le els: manu- ally (Linguis ic quali y and Responsi eness) and au oma ically (me ics om he package ROUGE, he Recall-O ien ed Unde s udy o Gis ing E alua ion [9]). Manual e alua ion sco ed each aspec o a gi en summa y as 1: e y poo , 2:poo , 3:accep able, 4:good, o 5: e y good. In addi ion, ou un is one o he 21 pa icipan sys ems ha we e also manually e alua ed by means o he py amid me hod [11]. 4.1.2 FEMsum se ings in DUC 2006 Ou goal in pa icipa ing a DUC was o e alua e a numbe o aspec s o ou sys em. We he e o e submi ed h ee di e en kinds o au oma ic sum- ma ies in a single un: one lexically based (LEX), and wo seman ically based (SEM150, SEM250). Ou o he 50 summa ies we we e expec ed o submi , 7 we e p oduced using he LEX app oach, 13 by means o he SEM150 s a egy, and 30 by using he SEM250 one. Ou sys em was assigned he iden i ica ion numbe 19. Gi en he que y, a common c ucial s ep in all he app oaches is o de ec he mos ele an TUs (sen ences in his expe imen s). We decided o ix a maximum numbe o sen ences de ec ed as ele an by RID. Fo ha eason we use he co pus o sen ences de ec ed as pa o a manual summa y in DUC 2005 p oposed by Copeck and Spakowicz [3]. Analyzing P ecision and Recall Nwas empi ically ixed in a maximum o 250. In he LEX app oach, ele an sen ences a e de ec ed by RID and hen SC is applied o ob ain he summa ies. On he o he hand, in bo h SEM app oaches he ini ial c i e ia o sen ence ele ance is ha sen ences om a same documen a e conside ed o ha e a simila ele ance, independen ly o he RID sco e. In he SEM250 s a egy, all he sen ences om he RID ou pu a e aken as CE inpu , whe eas in SEM150 he inpu o CE is he clus e o 150 sen ences om he i s documen s in he se . SEM150 ends o educe he numbe o documen s whose con en is candida e o appea in he summa y. 9 Table 1. Sample o en i onmen buil om a sen ence. Table 2. DUC FEMsum manually e alua ion o linguis ic quali y sco es by app oach. Table 3. DUC FEMsum manual con en esponsi eness sco e. Table 4. DUC Con en esponsi eness sco es by app oach. Table 5. DUC Con en esponsi eness sco es dis ibu ion by app oach Table 6. DUC ROUGE measu es when conside ing 4 manual summa ies as e e ences. Table 7. CHIL ROUGE measu es when conside ed 3 manual summa ies as e e ences. Table 8. CHIL esponsi eness conside ing 3 human models when e alua ing au oma ic summa ies and 2 when e alua ing human summa ies. Table 9. CHIL esponsi eness sco es dis ibu ion by au oma ic sys em. Figu e 1. FEMsum Global A chi ec u e Figu e 2. Con en Ex ac o modul Figu e 3. Linguis ic P ocesso cons i uen s Figu e 4. Sample o a sen ence analysis and he co esponden en i onmen ep eseen a ion Figu e 5. Sample o he g aph ep esen a ion o an En i onmen 16 Table 1 Sample o en i onmen buil om a sen ence ”Romano P odi 1is 2 he 3p ime 4minis e 5o 6I aly 7” i en p ope pe son(1), en i y has quali y(2), en i y(5), i en coun y(7), quali y(4), which en i y(2,1), which quali y(2,5), mod(5,7), mod(5,4) Table 2 DUC FEMsum manually e alua ion o linguis ic quali y sco es by app oach. G amma icali y Non- edundancy Re e en ial cla i y Focus FEMsum mean FEMsum mean FEMsum mean FEMsum mean LEX 3,14 3,45 2,43 4,02 2,43 2,83 3,29 3,73 S150 3,00 3,60 4,15 4,19 3,08 3,09 3,77 3,84 S250 3,33 3,59 4,23 4,27 2,77 3,12 3,20 3,42 Table 3 DUC FEMsum manual con en esponsi eness sco e. Sys em(ID) Sco e Mean Dis ance Human(A-J) 4,75 2,19 Bes (27) 3,08 1,83 FEMsum(19) 2,60 0,04 +www(30) 2,58 0,02 Baseline(1) 2,04 -0,52 Mean(2-35) 2,56 S de 0,28 Table 4 DUC Con en esponsi eness sco es by app oach. Mean(1-35) FEMsum Mean Dis ance LEX 2,36 2,29 -0,07 SEM150 2,55 2,92 0,37 SEM250 2,58 2,53 -0,05 Table 5 DUC Con en esponsi eness sco es dis ibu ion by app oach 1: Ve y Poo 2: Poo 3: Accep able 4: Good 5: Ve y Good LEX 14% 43% 43% 0% 0% SEM150 7,5% 31% 31% 23% 7,5% SEM250 6,7% 50% 26,7% 16,7% 0% 17 Table 6 DUC ROUGE measu es when conside ing 4 manual summa ies as e e ences. Bes (24) FEMsum(19) +www(30) Baseline(1) R-2 0,095 0,076 0,067 0,050 R-SU4 0,155 0,131 0,122 0,098 Table 7 CHIL ROUGE measu es when conside ed 3 manual summa ies as e e ences. SDS LEX LEXnoT +www SEM ROUGE-1 0,293 0,309 0,312 0,333 0,323 ROUGE-2 0,060 0,092 0,102 0,089 0,073 ROUGE-3 0,029 0,056 0,064 0,052 0,032 ROUGE-4 0,019 0,043 0,050 0,043 0,021 ROUGE-L 0,256 0,272 0,279 0,289 0,280 ROUGE-W1.2 0,089 0,098 0,100 0,104 0,098 ROUGE-S1 0,057 0,088 0,097 0,087 0,067 ROUGE-S4 0,064 0,089 0,095 0,094 0,073 ROUGE-S9 0,069 0,095 0,102 0,103 0,083 ROUGE-SU1 0,136 0,162 0,169 0,168 0,152 ROUGE-SU4 0,102 0,126 0,132 0,134 0,115 ROUGE-SU9 0,090 0,116 0,122 0,124 0,105 Table 8 CHIL esponsi eness conside ing 3 human models when e alua ing au oma ic sum- ma ies and 2 when e alua ing human summa ies. M1 M2 M3 SDS LEX LEXnoT +www SEM 3,625 3,400 3,375 1,250 1,775 2,025 1,800 1,800 Table 9 CHIL esponsi eness sco es dis ibu ion by au oma ic sys em. M1 M2 M3 SDS LEX LEXnoT +www SEM 1: Ve y Poo 0% 0% 0% 70% 40% 15% 30% 35% 2: Poo 10% 5% 10% 25% 25% 50% 50% 40% 3: Accep able 20% 35% 30% 5% 35% 30% 15% 20% 4: Good 40% 40% 45% 0% 0% 5% 5% 5% 5: Ve y Good 30% 20% 15% 0% 0% 0% 0% 0% 18 pape s slides no es aw ex Que y Con en Ex ac o Summa y Summa y Compose ASR ansc ip Use Need OUTPUT Se ings CONTENT DOC_RESTRICT SIZE FORMAT No ansc ip s Tex Num. wo ds Keywo ds Gene ic Speech/ ideo In o ma ion Rele an De ec o NLs a emen All Sin hesizedTex MDS Independen P esen a ion Tex English No used Used SDS Scien i ic DOC_STRUCTURE LANGUAGE MEDIA UNIT Spanish GENRE Jou nalis ic INPUT Se ings DOMAIN Independen Scien i ic News Fig. 1. FEMsum Global A chi ec u e Linguis ic P ocesso sen en sin Tex Uni s Selec o Candida es Rele an Tex Uni s sim CONTENT EXTRACTOR Candida es Simila i y Ma ix Gene a o Fig. 2. Con en Ex ac o modul Tokenize POS Tagge Lemma ize NERC Syn ac ic Chunke Seman ic Tagge Anno a o DM Seman ic Analize Discou se Ma ke s Wo dNe Use Need Tex Uni s sen sin en LINGUISTIC PROCESSOR Fig. 3. Linguis ic P ocesso cons i uen s 19 Fig. 4. Sample o a sen ence analysis and he co esponden en i onmen ep eseen- a ion Romano P odi 1 is 2 p ime 4 minis e 57 I aly 0 which_en i y which_quali y mod mod en i y_has_quali y i_en_p ope _pe son quali y en i y i_en_coun y Fig. 5. Sample o he g aph ep esen a ion o an En i onmen 20