Full text
Author: Wilcke,W.X. VU University Amsterdam Ariadne is funded by the European Commission’s 7th Framework Programme. D16.1: First Report on Data Mining
ARIADNE'D16.1'(Public)' ! ! The'research'leading'to'these'results'has'received'funding'from'the'European'Community's' Seventh'Framework'Programme'(FP7JINFRASTRUCTURESJ2012J1)'under'grant'agreement'n°' 313193.'' Version:'1.2'(final)' March'2015' Author:' Wilcke,'W.X,'VU'University'Amsterdam' Contributing'partners:' De'Boer,V.'AVU'University'Amsterdam' Van'Harmelen,'F.A.H.'AVU'University'Amsterdam' De'Kleijn,'M.T.M.'AVU'University'Amsterdam' Wansleeben,'M.A'Leiden'University' ' Quality'Control'Review:' Wright,'H.'E.'–'Archaeology'Data'Service,'University'of'York' ' ' ' ' ' ARIADNE is a project funded by the European Commission under the Community’s Seventh Framework Programme, contract no. FP7 The views and opinions expressed in this presentation are the sole responsibility of the authors and do not necessarily reflect the views of the European Commission.
ARIADNE"D16.1"Public" III" " Table&of&Contents& Document(History(....................................................................................................................(I" List(of(Abbreviations(...............................................................................................................(II" Executive&Summary(............................................................................................................(IV" Recommendations"................................................................................................................................."V" Roadmap"................................................................................................................................................"VI" 1"Introduction(and(Objectives(.............................................................................................(1" 1.1"Structure"of"Report"........................................................................................................................."1" 2"Introduction(to(Linked(Data(.............................................................................................(2" 2.1"The"RDF"Data"Model"......................................................................................................................."2" 2.2"Ontologies"......................................................................................................................................"5" 2.3"The"Semantic"Web".........................................................................................................................."6" 2.4"Linked"Archaeological"Data"............................................................................................................"8" 3"Introduction(to(Data(Mining(..........................................................................................(10" 3.1"Learning"from"Archaeological"Data".............................................................................................."10" 3.2"Knowledge"Discovery"and"Data"Mining"........................................................................................"12" 3.3"Data"Mining"Tasks"........................................................................................................................."13" 3.4"Towards"Mining"the"Semantic"Web".............................................................................................."15" 4"SemanticOWeb(Mining(...................................................................................................(17" 4.1"DataIMining"Tasks"........................................................................................................................"17" 4.2"Applicable"Solutions"....................................................................................................................."22" 5"Domain(Understanding(..................................................................................................(35" 5.1"Relevant"Studies"..........................................................................................................................."36" 5.2"Wishes"of"Domain"Experts"............................................................................................................"38" 5.3"Summary"......................................................................................................................................."39" 6"Data(Understanding(.......................................................................................................(40" 6.1"Data"Produced"using"NaturalILanguage"Processing"....................................................................."40" 6.2"Case"Study"on"Data"Repositories".................................................................................................."41" 6.3"Summary"......................................................................................................................................."43"
ARIADNE"D16.1"Public" IV" " 7"Data(Mining(on(Linked(Archaeological(Data(...................................................................(45" 7.1"Hypothesis"Generation"................................................................................................................."45" 7.2"Assisted"Query"Formulation"........................................................................................................."46" 7.3"Ranking"of"Query"Results".............................................................................................................."46" 7.4"Resource"Recommender"System".................................................................................................."47" 7.5"Data"Quality"Analysis"...................................................................................................................."48" 7.6" Trust"Analysis"................................................................................................................................"49" 8"Conclusions(...................................................................................................................(51" 8.1"Domain"Understanding"................................................................................................................."51" 8.2"Data"Understanding"......................................................................................................................"52" 8.3"Recommendations"........................................................................................................................"53" 8.4"Roadmap"......................................................................................................................................."54" Bibliography(.........................................................................................................................(55" Appendix(A"Reasoning(with(Logic(.........................................................................................(i" A.1"Reasoning"by"Deduction"..............................................................................................................."i" A.2" Reasoning"by"Induction"..............................................................................................................."ii" A.3"Logic"Reasoning"within"the"Semantic"Web"................................................................................."ii" Appendix(B"Vector(Space(Models(.......................................................................................(iv" Appendix(C"Learning(Methods(for(Semantic(Web(Mining(....................................................(v" C.1"Propositional"Learning"................................................................................................................"v" C.2"Statistical"Relational"Learning"..................................................................................................."vii" C.3"Kernel"Methods".........................................................................................................................."xi" Appendix(D"Sample(of(Archaeological(Scenarios(...............................................................(xiv" "
I" " Document&History& 1. 13th"February,"2015"" –"Full"Draft"Version"1.0" 2. 19th"February,"2015"" –"QC"Review"1.0" 3. ""4rd""""""March,"2015"" –"Full"Draft"Version"1.1" 4. ""6th""""""March,"2015"" –"QC"Review"1.1" 5. ""6rd""""""March,"2015"" –"Full"Final"Version"1.2" " "
ARIADNE"D16.1"Public" II" " List&of&Abbreviations& The"following"abbreviations"will"be"used"in"this"report." Abbreviation! Full&Term& ACDM" ARIADNE"Catalogue"Data"Model" ADS" Archaeological"Data"Service" API" Application"Programming"Interface" ARIADNE" Advanced" Research" Infrastructure" for" Archaeological" Dataset" Networking"in"Europe" BGV" Basic"Geo"Vocabulary" CAA" Computer"Applications"&"Quantitative"Methods"in"Archaeology" DCAT" Data"Catalogue"Vocabulary" DINAA" Digital"Index"of"NorthIAmerican"Archaeology" DM" Data"Mining" GIS" Geographic"Information"System" IG" Information"Gain" ILP" Inductive"Logic"Programming" IT" Information"Task" KDD" Knowledge"Discovery"and"Data"Mining" LAD" Linked"Archaeological"Data" LD" Linked"Data" LOD" Linked"Open"Data" ML" Machine"Learning" MRDM" MultiIRelational"Data"Mining" NLP" Natural"Language"Processing" OGC" Open"Geospatial"Consortium" "
ARIADNE"D16.1"Public" III" " Abbreviation! Full&Term& OLAP" Online"Analytical"Processing" OWL" Web"Ontology"Language" PCA" Principal"Component"Analysis" PSL" Probabilistic"Soft"Logic" RDF" Resource"Description"Framework" RDFS" Resource"Description"Framework"Schema" SKOS" Simple"Knowledge"Organization"System" SRL" Statistical"Relational"Learning" SVM" Support"Vector"Machine" SW" Semantic"Web" SWM" Semantic"Web"Mining" TL" Trust"Level" UI" User"Interface" URI" Universal"Resource"Indicator" VSM" Vector"Space"Models" W3C" World"Wide"Web"Consortium" WGS" World"Geodetic"System" WP" Work"Package" " "
ARIADNE"D16.1"Public" IV" " Executive&Summary& ARIADNE,"the"Advanced"Research"Infrastructure"for"Archaeological"Dataset"Networking"in"Europe,"will" facilitate" a" central" web" portal" that" provides" access" to" archaeological" data" from" various" sources" in" a" standardized"and"open"format."This"format"will"likely"adhere"to"the"Linked"Data"paradigm"either"fully"or" partially," with" the" former" being" the" option" that" we" believe" is" needed" to" propel" ARIADNE" towards" a" higher"level"of"interoperability."By"this"assumption,"users"will"be"able"to"use"the"portal"to"browse"and" search"the"data,"thereby"making"use"of"all"the"advantages"that"Linked"Data"has"to"offer."Among"these" advantages" are" advanced" search" abilities," the" inherent"capability" of" drawing" inferences," and" the" enrichment"of"data"by"linkage"to"and"from"external"sources."We"expect"these"features"to"have"a"positive" impact" on" the" archaeological" research" community." However," this" does" not" necessarily" add" to" the" knowledge"contained"within"the"aggregated"data"sets."Ideally,"we"would"like"to"expand"this"knowledge" as"well."One"field"of"expertise"that"specializes"in"exactly"that"is"data"mining."Hereto,"this"field"provides" tools"and"techniques"to"identify"valid,&novel,&potentially&useful,&and&ultimately&understandable&patterns& in&data&(U."M."Fayyad"1996)." This"report"examines"the"applicability"and"feasibility"of"integrating"data"mining"solutions"into"ARIADNE." To"this"end,"we"explored"various"stateIofItheIart"theories,"methods,"and"solutions"to"detect"patterns"in," and"establish"relations"between,"data"from"the"archaeological"domain."Throughout"this"report,"we"made" the"assumption"that"this"data"will"adhere,"either"fully"or"partially,"to"the"principles"of"the"Linked"Data" paradigm."The"subfield"of"data"mining"dedicated"to"this"form"of"data,"known"as"semantic"web"mining," was"deliberately"chosen"over"theImoreItraditional"tabular"data"mining,"for"its"ability"to"fully"exploit"the" graphIlike"structure"of"Linked"Data"without"the"loss"of"knowledge."In"addition"to"data"mining,"our"study" focussed"on"usageIpattern"analysis"and"content"linking,"as"well"as"on"information"retrieval."To"this"end,"a" thorough"analysis"of"users’"needs"and"wishes"was"conducted,"as"well"as"an"exploration"of"the"expected" data’s"characteristics."Furthermore,"recent"and"relevant"literature"and"experience"on"the"topics"involved" were"examined"in"depth."" The" userIrequirements"study"involved" an" analysis" of" the" questionnaires" and" interviews" that" were" conducted" by" work" package"2.1" and" 13.1," respectively."While" providing" valuable" insight" into" the" stakeholders" of" ARIADNE," both" work" packages" only" touched" on" the" possibility" of" data" mining." As" a" result," very" little" could" be" ascertained" as" to" what" path" any" data"mining" solution" should" follow." Moreover,"the"large"majority"of"the"stakeholders"had"little"to"no"experience"with"data"mining"and"were" unaware" of"what" it" actually" entailed." To" mitigate" this" lack" of" direction," several"additional"interview" sessions" with" stakeholders" were" held," during" which" the" possibility" of" data" mining" was" moreIactively" explored."Regardless,"of"all"the"topics"discussed,"only"few"were"relevant"with"respect"to"data"mining." In" its" entirety," the" requirements" study" seemed" to" indicate" that" the" large" majority" of" the" difficulties" experienced"by"stakeholders"could"be"mitigated"by"the"use"of"Linked"Data"alone."Several"of"these"issues" could"additionally"be"improved"upon"even" more" with"the"help" of"data"mining."These"issues" involved" knowing"which"data"is"available,"how"to"locate"relevant"data,"and"how"to"distil"the"relevant"results"from"
ARIADNE"D16.1"Public" V" "" those"that"are"not."In"addition,"the"quality"of"the"data"was"mentioned"prominently,"thereby"emphasizing" their"(lack"of)"completeness"and"the"(lack"of)"trust"bestowed"on"them."Together,"these"were"amongst" the"prime"areas"considered"to"which"a"dataImining"solution"could"be"applied." Exploring"the"data"is"an"important"early"step"within"any"dataImining"process,"during"which"the"data’s" characteristics,"their" quality," and"their" abnormalities"are" inspected." Generally,"a"generous"amount"of" data" is" provided" from" which" conclusions" can" be" drawn" that" influence" choices" made" during" the" development"of"the"eventual"dataImining"solution."Unfortunately,"the"minimal"amount"of"data"currently" available"through"ARIADNE"prevents"such"a"sequence"of"events"to"take"place."Therefore,"Linked"Data" from"several"different"archaeological"repositories"around"the"globe"was"inspected"instead."These"data" were"chosen"for"their"almostIdisjoint"characteristics,"thus"hopefully"providing"good"representations"of" the"different"facets"that"ARIADNE"might"bring"forth."Assuming"they"do,"several"observations"could"be" made:"save"for"the"generally"expected"differences"in"used"ontologies"and"structure,"the"examined"linked" archaeological"data"were"found"to"strongly"depend"on"descriptive"values,"as"well"as"consisting"largely"of" relatively"flat"data"structures."These"aspects"of"the"data"should"be"considered"during"the"development" of"the"forthcoming"dataImining"solution." Understanding" the" domain" and" its" data" are" two" early" but" crucial" steps" in" any" dataImining" process." Together"with"our"updated"knowledge"on"Linked"Data"and"data"mining,"these"all"come"together"to"form" the"field"of"Semantic"Web"mining."This"field"represents"a"young"area"of"research"of"which"many"aspects" are"still"uncertain"or"left"unexplored,"from"both"a"technical"and"practical"perspective."In"fact,"many"of" methods"are"still"under"heavy"development,"with"few"of"them"having"progressed"outside"the"confines"of" academic"research."Therefore,"instead"of"considering"all"possible"approaches,"we"have"solely"focussed" on"the"moreIprominent"movements"as"seen"in"the"literature."" Recommendations& Generally,"the"developer"of"a"typical"data"mining"solution"will"explore"a"large"amount"of"data"with"the" goal"of"revealing"potentially"relevant"patterns."After"careful"inspection,"the"truly"relevant"patterns"will" subsequently" be" generalized" to" the" entirety" of" the" data." Unfortunately," the" small" amount" of" data" currently"available"through"ARIADNE"would"make"it"rather"unlikely"to"successfully"generalize"about"any" discovered" pattern" for"the" large" amount" of" data" that," one" day," will" be" accessible." Instead," a" more" generic"approach"is"suggested,"such"that"its"workings"are"ensured"regardless"of"the"exact"characteristics" of"the"future"data." Based"on"the"study"of"both"domain"and"data,"as"well"as"on"practical"constraints"with"respect"to"time"and" resources,"two"dataImining"solutions"were"chosen"which"were"deemed"the"most"feasible"and"suitable" for" implementation" within" the" ARIADNE" infrastructure." These" constitute" 1)" the" ability" for" users" to" generate" potentiallyIrelevant" hypotheses," and" 2)" analysing" the" quality" of" data" as" well" as" helping" to" improve"it."We"will"briefly"touch"on"these"two"solutions"next." " "
ARIADNE"D16.1"Public" 5" " 2.2 Ontologies& As"discussed"previously","RDF"is"a"simple"data"model"with"which"statements"can"be"specified."In"fact,"RDF" alone"can"be"used"to"make"any"statement,"even"ludicrous"ones"such"as"that"the"Pyramids"of"Giza"are"in" Rome."The"reason"behind"this"is"that"it"does"not"make"any"assumptions"about"the"domain"it"describes," nor"the"semantics"used"to"describe"that"domain."That"is,"RDF"itself"is"“unaware”"of"the"information"it" states."Instead,"this"“awareness”"is"provided"by"ontologies"(Shadbolt,"Hall"and"BernersILee"2006,"van" Harmelen,"et"al."2012,"Heath"and"Bizer"2011)." Within"the"field"of"Computer"Science,"ontologies"describe"a"domain"by"types,"properties,"and"relation" types."They"are"the"culmination"of"a"progression"from"simple"vocabularies"with"a"fixed"list"of"terms"to" fullIfledged"languages"with"a"powerful"expressiveness"(Garshol"2004,"van"Harmelen,"et"al."2012)."To"this" end,"ontologies"generally"provide"a"hierarchy"of"classes"and"properties,"as"well"as"allowing"some"form"of" reasoning" over" them." These"classes" may" encompass" subjects," predicates," and" objects," including" both" resources"and"literals."For"instance,"an"archaeological"ontology"might"define"the"exact"meaning"of"the" concept"excavation,"and"that"this"concept"should"always"contain"an"archaeological"context"to"be"valid." Numerous"ontologies,"both"simple"and"powerful,"modelling"various"domains"have"already"been"made" available"in"an"RDFIcompliant"format."Within"ARIADNE,"the"more"relevant"of"these"is,"unsurprisingly,"the" archaeological"domain."As"this"domain"has"a"strong"geospatial"component"(Wagtendonk,"et"al."2009,"De" Kleijn,"et"al."2014,"Conolly"and"Lake"2006),"this"latter"domain"can"be"regarded"as"quite"relevant"as"well." Therefore,"several"ontologies"describing"these"two"domains"will"be"briefly"discussed"next."Please"note" that,"for"simplicity,"the"technical"aspects"behind"ontologies"have"been"omitted." 2.2.1 Geospatial&Ontologies& There"exist"several"ontologies"holding"geospatial"knowledge,"of"which"the"most"elementary"is"the"Basic" Geo" Vocabulary"(BGV)" (Brickley" 2006)." BGV" allows" for" the" specification" of" points" within" the" World" Geodetic"System"(WGS)"standard."To"this"end,"it"accepts"latitude,"longitude,"and"altitude"declarations." A" geospatial"ontology" with" a" moreIpowerful" level" of" expressiveness" than" with" BGV" is" GeoSPARQL;" a" geographical"query"language"developed"by"the"Open"Geospatial"Consortium3"(OGC)"(OGC"GeoSPARQL"I" A"Geographic"Query"Language"for"RDF"Data"2012)."While"it"emphasizes"the"retrieval"of"knowledge,"i.e." querying,"more"than"describing,"GeoSPARQL"still"offers"a"flexible"ontology"to"describe"topologies." To" accomplish" this," it" incorporates" specifications" from" other" geospatial" standards" designed" by" the" OGC," among"which"are"the"Geographic"Markup"Language4"and"Simple"Feature"Access5."Therefore,"GeoSPARQL" accepts"declarations"of"points,"(multi)"lines,"and"(multi)"polygons,"as"well"as"describing"their"properties" """"""""""""""""""""""""""""""""""""""""""""""""""""""""""""" 3"Open"Geospatial"Consortium,"see"www.opengeospatial.org"" 4"Geospatial"Markup"Language,"see"www.opengeospatial.org/standards/gml"" 5"Simple"Feature"Access,"see"www.opengeospatial.org/standards/sfa"
ARIADNE"D16.1"Public" 6" " by"Region"Connection"Calculus."However,"some"of"the"more"exotic"features"require"that"the"data"store" supports"the"GeoSPARQL"protocol." 2.2.2 Archaeological&Ontologies& An" ontology" within" the" archaeological" domain" will" be" the" ARIADNE" Catalogue" Data" Model" (ACDM)" (Aloia,"et"al."2014)."The"ACDM"attempts"to"provide"a"data"model"to"describe"archaeological"resources," such"as"collections,"data"sets"and"services,"as"well"as"metadata"and"vocabularies."Due"to"this"area"of" focus,"it"is"being"built"upon"the"Data"Catalogue"Vocabulary"(DCAT)"(Maali,"Erickson"and"Archer"2014);"a" vocabulary"commended"by"the"W3C"for"its"ability"to"represent"government"data"catalogues."Instead"of" catalogues" however," ACDM" emphasizes" collections" and" data" sets," with" the" former" being" a" set" of" heterogeneous"items"without"a"formal"structure"and"the"latter"being"a"set"of"structured"records."These" structured"records"are"assumed"to"originate"from"either"a"database"or"from"a"Geographic"Information" System"(GIS)." Two" other" archaeological" ontologies" are" the" CRMarchaeo" and" CRMIEH" extension" of" the" CIDOC" Conceptual" Reference" Model;" an" ontology" for" describing" knowledge" from" the" domain" of" cultural" heritage"(The"CIDOC"Conceptual"Reference"Model"n.d.,"Doerr"and"Schaller"2008)."The"CIDOC"CRM"was" developed" by" in"collaboration" with" the" International" Council" of" Museums," with" the" aim" of" allowing" diverse"perspectives"that"incorporate"different"institutional"histories,"disciplines,"and"objectives."To"this" end,"it"provides"a"solid"core"with"the"ability"to"add"functionality"by"use"of"extensions." " The" CRMarchaeo" constitutes" a" generic" archaeological" extension," developed" within" the" ARIADNE" framework,"which"aims"at"encoding"metadata"on"the"excavation"process"(Cripps,"et"al."2014)."By"offering" this" metadata," CRMarchaeo" endeavors" to" optimize" the" interpretability" of" a" documented" excavation," thereby" providing" the" rational" for" conducting" that" excavation," as" well" as" knowledge" on" previous" excavations"and"studies"on"the"same"site."" " The"second"archaeological"CRM"extension,"the"CIDOC"CRMIEH"(May"n.d.),"was"developed"to"include"the" archaeological"concepts"and"processes"in"use"by"the"English"Heritage6;"a"national"heritage"body"in"the" UK"charged"with"safeguarding"cultural"heritage."To"this"end,"it"offers"numerous"classes"and"(inverse)" properties"divided"amongst"several"modules." 2.3 The&Semantic&Web& The" LOD" cloud" constitutes" a" large" number" of" interconnected" RDF" repositories." These" repositories," commonly"referred"to"as"a"triple(stores,"allow"queries"to"be"processed"on"their"data"(Shadbolt,"Hall"and" BernersILee"2006,"van"Harmelen,"et"al."2012)." """"""""""""""""""""""""""""""""""""""""""""""""""""""""""""" 6"English"Heritage,"see"www.englishIheritage.org.uk""
ARIADNE"D16.1"Public" 7" " A" single," isolated" triple" store"already"provides" a" data" structure" that" can" allow" powerful" methods" of" searching" through" and" reasoning" with" the" data"to" be" used." The" real" advantage" of" LD"however,"only" surfaces"when"multiple"triple"stores"are"available"on"the"web"and"are"linked"to"each"other."Recall"that" this"distinction"is"equivalent"to"the"difference"between"the"four"and"five"star"rating"of"the"LOD"project"as" was" discussed" earlier." Irrespective," these" interlinked"triple" stores"together"form" a" web( of( data;" the" Semantic&Web"(Bizer,"Heath"and"BernersILee"2009,"Heath"and"Bizer"2011,"van"Harmelen,"et"al."2012).""" 2.3.1 Querying&the&Semantic&Web& Nearly"all"search"engines"for"the"currentIday"WorldIWide"Web"are"keywordIbased"(Freitas,"et"al."2012);" given"a"provided"set"of"words"the"most"relevant"result"is"sought."Generally,"this"is"accomplished"by"some" variation"on"Vector(Space(Models((Appendix"B),"which"constitute"a"method"with"which"documents"can" be" identified." Unfortunately," such" a" solution" would" be" unsuited"for"searching"through" the"LOD" that" makes"up"the"Semantic"Web"(SW)"or"web"of"data,"as"it"lacks"the"ability"to"represent"structured"data"as" well" as" their"semantics" (Figure" 2I3)." Instead," the" SW" is" generally" searched" through" by" queryIbased" languages"that"were"specifically"developed"for"this"purpose." In"the"case"of"a"triple"store,"the"commonly"used"query"language"is"that"of"SPARQL"(Prud'hommeaux"and" Seaborne"2008,"van"Harmelen,"et"al."2012,"Heath"and"Bizer"2011,"Shadbolt,"Hall"and"BernersILee"2006," Bizer,"Heath"and"BernersILee"2009)." SPARQL" is" a" W3C" recommended"protocol"and"a"query"language" developed"to"access,"retrieve,"and"modify"RDF"data."Similar"to"other"query"languages,"such"as"SQL,"it" offers" a" wide" range" of" capabilities" ranging" from" simple" pattern" matching" to" complex" queries" with" restrictions"in"range,"time,"and"domain."Moreover,"SPARQL’s"queryIprocessing"engine"enables"the"ability" to"reason"deductively"(Appendix"A)"over"the"data"to"which"it"provides"access."This"ability"can"be"either" relatively"powerful"or"rather"limited,"depending"on"the"expressive"strength"of"the"ontologies"used"to" specify"these"data"with." Querying"a"triple"store"through"a"queryIlanguage"like"SPARQL"is"accomplished"by"connecting"to"a"soI called"endpoint."Endpoints"are"provided"by"the"triple"store"and"form"the"bridge"between"a"user"and"the" data"contained"in"that"triple"store"(van"Harmelen,"et"al."2012,"Heath"and"Bizer"2011)."Once"a"query"has" been"submitted"to"such"a"publically"accessible"endpoint,"it"will"be"processed"and"the"results"returned." For"instance,"a"query"might"request"all"pottery"fragments"with"burn"marks"and"larger"than"20"cm2"that" were"found"between"1950"and"1960"at"a"specific"site"in"Italy."
ARIADNE"D16.1"Public" 8" " " Figure&2D3:&ExpressivityDUsability&tradeDoff&for&querying&with&either&a&word&or&ontologyDbased&search& engine&(Freitas,&et&al.&2012).& 2.4 Linked&Archaeological&Data& While"fairly"new,"the"concept"of"LD"is"not"unheard"of"within"the"archaeological"domain."In"fact,"some" think"of"it"as"being"the"next"logical"step"in"sharing"archaeological"knowledge"(Signore"2009,"Richards" 2006)."Others"believe"however"its"semantic"complexity"lacks"the"ability"to"describe"the"uncertainty"in" archaeological" data" (Isaksen," Martinez," et" al." 2010," Martinez" and" Isaksen" 2010)." Regardless," several" endeavors"regarding"Linked"Archaeological"Data"(LAD)"repositories"have"already"been"undertaken." One" of" the" larger" LAD" repositories"within" Europe"is" disseminated" by" the" Archaeology" Data" Service7" (ADS),"which"provides"a"SPARQL"endpoint"to"a"triple"store"hosted"by"the"University"of"York8"(Charno,"et" al." 2012)." This" triple" store"was" developed" as" part"of" the" STELLAR9"project"(Tudhope," et" al." 2011);" a" collaboration,"funded"by"the"UK"Arts"&"Humanities"Research"Council10,"in"partnership"with"the"University" of"Glamorgan"(now"South"Wales)11,"and"English"Heritage,"with"the"aim"of"improving"the"integration"of"LD" into"the"digital"archaeological"domain."The"data"currently"in"the"ADS"triple"store"were"converted"from" databases"and"spreadsheets"to"RDF,"using"the"CRMIEH"ontology."The"resulting"triples"are"stored"in"an" AllegroGraph12"triple"store,"which"points"to"a"SPARQL"endpoint"and"allows"the"results"to"be"provided"in" one"of"several"RDF"serializations." On"the"other"side"of"the"Atlantic"ocean,"the"Digital"Index"of"North"American"Archaeology"(DINAA)"aims" to"integrate" government"curated" public" data" from" both" offline" and" online" digital" archaeological" """"""""""""""""""""""""""""""""""""""""""""""""""""""""""""" 7"Archaeological"Data"Service,"see"data.archaeologicaldataservice.ac.uk"" 8"University"of"York,"see"www.york.ac.uk"" 9"STELLAR"Project,"see"www.archaeologicaldataservice.ac.uk/research/stellar"" 10"Arts"&"Humanities"Research"Counsel,"see"www.ahrc.ac.uk"" 11"University"of"Glamorgan,"see"www.southwales.ac.uk"" 12"AllegroGraph,"see"www.franz.com/agraph/allegrograph""
ARIADNE"D16.1"Public" 9" " repositories" (Wells," et" al." 2014)." Supported" by" the" National" Science" Foundation13," its" primary" focus" concerns"aiding"the"researcher"in"data"discovery,"as"well"as"filling"the"gap"in"archaeological"information" infrastructures." Moreover," DINAA" emphasizes" the" reuse" of" both" technologies" and" data." These" data," which"are"stored"in"a"MySQL14"database,"are"expressed"with"the"help"of"the"DINAA15"vocabulary;"an" ontology"built"upon"OWL"and"stronglyIinfluenced"by"CIDOC"CRM."Furthermore,"as"these"data"are"open," DINAA"has"been"welcomed"into"the"LOD"cloud." Where"both"ADS"and"DINAA"provide"national"data"for"the"most"part,"ARIADNE"will"attempt"to"extend" this"to"the"whole"of"Europe."Not"all"LAD"projects"aim"at"such"scale"however."For"instance,"(Gruber,"et"al." 2012)"explore" the" feasibility" and" usefulness" of" introducing" LD" into" the" field" of" numismatics," thereby" providing"enhanced"searching"abilities"to"a"database"of"Roman"coins."Another"example"is"from"(Isaksen," Martinez,"et"al."2009),"who"tried"to"improve"our"understanding"of"ancient"trade"networks"by"analyzing" LAD" concerning" the" distribution" of" amphorae" and" marble." As" a" final" example," consider" the" research" done"in"de"Boer,"et"al."(2014)"where"they"paired"LAD"concerning"Dutch"ship"wrecks"with"that"of"Dutch" sailors,"resulting"in"new"insights"on"the"socioIeconomic"realities"of"the"18th"Century."" """"""""""""""""""""""""""""""""""""""""""""""""""""""""""""" 13"National"Science"Foundation,"see"www.nsf.gov"" 14"MySQL,"see"www.mysql.com"" 15"DINNA"Ontology,"see"opencontext.org/vocabularies/dinaa""
ARIADNE"D16.1"Public" 10" " 3 Introduction&to&Data&Mining& Data(mining"(DM)"is"a"fairly"new"and"multiIdisciplinary"field"which"intersects"with"Artificial"Intelligence," Data"Science,"and"Statistics,"as"well"as"partially"overlapping"with"Machine"Learning"(ML)"from"which"it" draws"its"technical"basis"(Kantardzic"2011,"Hastie,"et"al."2009,"Friedman"1998,"Witten,"Frank"and"Hall" 2011)." Therefore," the" people" who" specialize" in" this" field" stem" from" various" backgrounds" and" hold" different"views,"making"it"difficult"to"provide"a"definition"agreed"upon"by"all"those"involved."Consider"the" following"three"definitions,"as"found"in"the"literature:" Data& Mining& it& the& nontrivial& process& of& identifying& valid,& novel,& potentially& useful,& and& ultimately&understandable&patterns&in&data"(U."M."Fayyad"1996)." Data& Mining& is& the& extraction& of& implicit,& previously& unknown,& and& potentially& useful& information&from&data&(Witten,"Frank"and"Hall"2011).& Data&Mining&is&a&decision&support&process&where&we&look&in&large&data&bases&for&unknown& and&unexpected&patterns&of&information&(Parsaye"1996)." A" common" notion" throughout" the" different" definitions" is" that" of" identifying," extracting," and" using" information" from" data" that" was" previously" unknown." In" order" words;" Data" Mining" concerns" learning( from(data((Hastie,"et"al."2009)."" 3.1 Learning&from&Archaeological&Data& Long"before"the"field"of"DM"came"to"be,"statistics"was"the"only"area"that"specialized"in"learning"from" data."Despite"its"potential"usefulness"however,"many"archaeologists"used"to"refrain"from"familiarizing" themselves"with"these"methods"(Baxter"2003)."Hence,"their"use"in"the"archaeological"domain"progressed" rather"slowly."Instead,"most"grew"out"of"necessity,"i.e."to"solve"a"problem,"and"were"originally"developed" in"other"fields"such"as"geography"and"ecology."Only"with"the"emergence"of"New&Archaeology,"did"the" interest" in" statistical" methods"grow." Still," early" uses" were" largely" without"the" thenIcalled"‘complex" statistics’,"amongst"those"listed"were"principal"component"(PCA),"factor,"and"cluster"analysis"(Whallon" 1987,"Kintigh"1987)."This"changed"with"the"increased"availability"of"statistical"applications." " "
ARIADNE"D16.1"Public" 11" " Over"the"last"two"decades,"a"rise"in"computational"power"gave"birth"to"various"statistical"applications" (Baxter"2003)."Some"of"these,"such"as"SPSS16"and,"in"a"lesser"degree,"R17,"were"aimed"at"individuals"who" were" not" really" familiar" with" the" theory" behind" statistical" methods." These" tools" allowed" the" archaeologists" to" apply" and" get" acquainted" with" basic" approaches" such" as" regression" analysis" and" Bayesian"statistics,"as"well"as"providing"a"simple"frontend"to"the"aforementioned"‘complex"statistics’." Recent" years" saw" the" integration" of" statistical" methods" into" various" nonIdedicated" applications."For" example,"consider"a"Geographical"Information"System"(GIS);"a"tool"commonly"used"by"archaeologists"to" perform"some"form"of"spatial"analysis"(Selhofer"and"Geser"2014,"Baxter"2003)."Hereto,"most"modern"GIS" frameworks"provide" a" simplified" frontend" to" several" adapted" statistical" methods," thereby" including" location"and"predictive"modelling,"as"well"as"a"subset"of"the"‘complex"statistics’"mentioned"earlier."" " Figure&3D1&:&Rough&segmentation&of&an&archaeological&aerial&photograph&as&determined&by&a&DM& algorithm&(Kobylinski&and&Walczak&2006).& The"acceptance"of"DM"by"archaeologists"appears"to"follow"a"line"similar"to"that"of"statistics,"with"the" term" “data" mining”" having" been" mentioned" only" sporadically" in" archaeologicallyIrelated" literature." Nevertheless," certain" topics" related" to" DM" appear" to" be" quite" well" represented," especially"those" involving"some"form"of"classification."Amongst"these,"the"oftenIencountered"artefacts"are"coins,"glass," and"ceramics,"which"are"classified"based"on"the"similarity"between"their"visual"characteristics"(van"der" Maaten,"et"al."2006,"Huber,"et"al."2005,"Nolle,"et"al."2003,"Karasik,"et"al."2004)."An"example"of"a"more" specialized"study"involving"classification"is"that"of"Bi,"et"al."(2008),"who"created"a"method"to"spatially" classify" and" partition" archaeological" settlements" based" on" the" discovery" of" nearby" hearths," pits," urn" tombs,"and"pit"tombs."Another"specialized"example"is"the"research"conducted"by"Linderholm"&"Geladi" (2012)," who" developed" an" approach"to" classify" archaeological" soil" and" sediment" samples" based" on" infrared"readings"of"those"samples."As"yet"another"example,"consider"Di"Ludovico"and"Pieri"(2011),"who" explored" various" means" to" classify" entries" within" large" corpora" of" decorations" on" Mesopotamian" cylinder" seals." As" a" final" example," consider" the" research" done" by" Kobylinski" &" Walczak" (2006)," who" """"""""""""""""""""""""""""""""""""""""""""""""""""""""""""" 16"SPSS,"see"www.ibm.com/software/nl/analytics/spss"" 17"The"RIproject,"see"www.rIproject.org""
ARIADNE"D16.1"Public" 12" " developed"a"method"to"automatically"determine"potentially"interesting"features"in"archaeological"aerial" photos"(Figure"3I1)." 3.2 Knowledge&Discovery&and&Data&Mining& Until"now,"we"have"used"the"term"“data"mining”"to"refer"to"the"whole"process"of"discovering"useful" patterns"from"any"form"of"data."For"simplicity"sake,"we"will"continue"to"do"so."Strictly"speaking"however," the"term"merely"denotes"the"act"of"running"an"algorithm"on"a"data"set."This"is"only"one"stage"in"a"larger" knowledge"discovery"process,"generally"known"as"Knowledge"Discovery"and"Data"Mining"(KDD)."During" the"course"of"ARIADNE,"such"a"process"will"be"undertaken"by"the"knowledge"engineers"involved"with"WP" 16."In"fact,"the"research"being"conducted"for"this"report"already"constitutes"a"partial"KDD"process." " Figure&3D2:&Schematic&depiction&of&a&generic&KDD&process.& A"generic"KDD"process"(Figure"3I2)"typically"consists"of"six"different"stages"(Kurgan"and"Musilek"2006):" Domain"Understanding,"Data"Understanding,"Data"Preparation,"Data"Mining,"Evaluation,"and"Knowledge" Consolidation."For"each"of"these,"a"general"description"will"be"given." Domain( Understanding"concerns" familiarizing" oneself" with" the" domain" at" hand" (Fayyad," PiatetskyIshapiro" and" Smyth" 1996," Kantardzic" 2011," Kurgan" and" Musilek" 2006," Maimon" and" Rokach" 2005)." This" entails" a" sufficient" comprehension" of" the" current" state" of" affairs," on" the" problems"therein,"and"on"the"goals"that"have"to"be"reached."Furthermore,"key"figures"and"their" terminology"should" be"identified."Within"ARIADNE,"this"boils"down"to"an"understanding"of"the" archaeological"domain,"as"well"as"of"the"archaeologists"themselves." Data(Understanding(concerns"analysing"the"data,"thereby"inspecting"its"quality"(Witten,"Frank"and" Hall"2011,"Fayyad,"PiatetskyIshapiro"and"Smyth"1996,"Kantardzic"2011,"Kurgan"and"Musilek"2006)." This"involves"the"identification"of"anomalies,"such"as"noise,"outliers,"and"missing"values,"as"well"as" the" selection" of" interesting" subsets" or" features." In" the" case" of" ARIADNE," this" involves" an" understanding"of"LAD"and"its"anomalies." " "
ARIADNE"D16.1"Public" 13" " Data(Preparation"concerns"the"transformation"of"the"data"as"to"make"it"suitable"for"DM"(Witten," Frank" and" Hall" 2011," Fayyad," PiatetskyIshapiro" and" Smyth" 1996," Kantardzic" 2011," Kurgan" and" Musilek"2006,"Maimon"and"Rokach"2005)."This"entails"resolving"problems"in"data"quality,"which" were"discovered"previously,"as"well"as"scaling"and"normalizing"values"if"needed."Furthermore,"a" final"selection"of"interesting"features"is"made."" Data( Mining"concerns" applying" a" suitable"inductiveIreasoning" method" (Appendix" A)" to" the" prepared"data"set,"resulting"in"the"automated"discovery"of"potentially"relevant"patterns"(Witten," Frank" and" Hall" 2011," Fayyad," PiatetskyIshapiro" and" Smyth" 1996," Kantardzic" 2011," Kurgan" and" Musilek"2006,"Maimon"and"Rokach"2005)."These"patterns"are"described"in"a"mathematical"model" that"approximates"the"data." Evaluation"is"the"phase"during"which"the"previously"generated"knowledge"is"interpreted,"as"well" as"being"inspected"for"its"usefulness"(Witten,"Frank"and"Hall"2011,"Fayyad,"PiatetskyIshapiro"and" Smyth"1996,"Kantardzic"2011,"Kurgan"and"Musilek"2006,"Maimon"and"Rokach"2005)."This"typically" involves"a"visualization"of"the"corresponding"patterns."In"the"case"of"ARIADNE,"this"will"involve"an" iterative"review"process"attended"by"both"knowledge"engineers"and"archaeological"researchers." Knowledge( Consolidation"concerns" the" presentation" of" the" new" knowledge" in" a" userIoriented" fashion," followed" by" the" possible"incorporation" of" that" knowledge" into" a" final" system" (Witten," Frank" and" Hall" 2011," Fayyad," PiatetskyIshapiro" and" Smyth" 1996," Kantardzic" 2011," Kurgan" and" Musilek" 2006," Maimon" and" Rokach" 2005)." Within" ARIADNE," this" would" correspond" to" either" showing"the"DM"results"to"the"user"or"adding"these"results"to"the"corresponding"triple"store." The"generic"sequence"of"stages"as"outlined"above"is"only"one"of"many"proposed"KDD"models"(Kurgan" and"Musilek"2006)."Due"to"its"domainIindependent"properties,"we"believe"it"to"be"a"suitable"model"to" follow"during"our"research"within"this"WP."In"fact,"the"steps"of"Domain"and"Data"Understanding"will"be" covered"largely"during"the"course"of"this"report." 3.3 Data&Mining&Tasks& The" domainIunderstanding" stage" of" a" KDD"process" typically" provides" the" knowledge" engineer" with" insight"into"the"desired"goals"of"that"process."These"goals"form"the"main"criteria"when"deciding"on"which" DM"task"to"implement."While"many"variants"exist,"these"tasks"generally"fall"into"one"of"the"following" higherIlevel" tasks"(Fayyad," PiatetskyIshapiro" and" Smyth" 1996," Kantardzic" 2011," Lavrac" and" Dzeroski" 2001)":" Classification," Regression," Clustering," Summarization," Change" and" Deviation" Detection," and" Dependency"Modelling."For"each"of"these,"a"general"description"will"be"given." " "
ARIADNE"D16.1"Public" 14" " Classification"focusses"on"learning"a"predictive"model"that"is"capable"of"correctly"assigning"new" instances"of"unknown"classes"to"one"of"several"predefined"classes"(Fayyad,"PiatetskyIshapiro"and" Smyth"1996,"Witten,"Frank"and"Hall"2011,"Kantardzic"2011,"Lavrac"and"Dzeroski"2001,"Berendt,"et" al." 2004," Hagood" 2012)." Typically," these" classes" represent"related" categories" within" a" certain" domain,"for"example,"red"and"green"denote"classes"within"a"finite"set"of"colours."" Regression(analysis"involves"learning"a"model,"which"can"predict"unknown"attribute"values"based" on" known" values" of" the" other" attributes" belonging" to" the" corresponding" instances" (Fayyad," PiatetskyIshapiro" and" Smyth" 1996," Witten," Frank" and" Hall" 2011," Kantardzic" 2011)." Here," both" known"and"unknown"values"should"have"a"numerical"internal"representation."In"the"case"of"binary" or" categorical" values," e.g." labels," specific" numerical" ranges" are" used." For" instance," consider" predicting" either" true"or" false"by" letting" a" positive" and" negative" value" denote" the" former" and" latter,"respectively."" Cluster( analysis"tries" to" describe"a"finite"set" of" groups"or"clusters" composed"of"instances" with" similar"attribute"values"(Fayyad,"PiatetskyIshapiro"and"Smyth"1996,"Witten,"Frank"and"Hall"2011," Kantardzic"2011,"Lavrac"and"Dzeroski"2001,"Berendt,"et"al."2004,"Hagood"2012)."These"clusters"are" determined" without" prior" knowledge" on" the" underlying" structure" of" the" data," and" may" be" regarded"as"nameless"classes."Hence,"cluster"analysis"can"be"seen"as"a"variant"of"classification." Summarization"concerns"the"methods"that"are"capable"of"compressing"data"into"more"compact" forms"without"losing"too"much"of"the"original"knowledge"(Fayyad,"PiatetskyIshapiro"and"Smyth" 1996," Kantardzic" 2011," Lavrac" and" Dzeroski" 2001)." This" problem" can" be" tackled" through" two" distinct"paths;"either"through"extraction"or"through"abstraction"(Mani"and"Maybury"1999)."Here," the"former"entails"the"automatic"extraction"of"existing"fragments"of"the"data"that"are"deemed" relevant,"while"the"latter"method"generates"new"data"that"describes"these"relevant"aspects"in"a" concise"way."The"alert"reader"might"recognize"the"similarities"between"the"method"of"abstraction" and"the"technique"behind"VSM." Change( and( Deviation( Detection"aims" at" discovering" significant" changes" or" deviations," e.g." outliers,"from"previously"measured"or"normative"values,"respectively"(Fayyad,"PiatetskyIshapiro" and" Smyth" 1996," Kantardzic" 2011," Lavrac" and" Dzeroski" 2001)." In" both" cases," the" key" is" to" determine" whether" the" probability" of" such" an" anomaly" occurring" is" too" low" to" warrant" it" to" actually" happen." For" instance," given" an" average" human" height" of" 1.70"meters," the" arrival" of" someone"with"a"height"of"2.50"meter"would"almost"certainly"stand"out." Dependency(Modelling"consists"of"learning"one"or"more"models"that"are"capable"of"describing" significant" dependencies" between" the" different" variables" found" in" the" data"(Fayyad," PiatetskyI shapiro"and"Smyth"1996,"Kantardzic"2011)."Typically,"these"models"focus"on"a"particular"subset"of" the"data,"thereby"describing"different"dependencies,"and"thus"do"not"cover"the"entire"data"set."" None" of" these" six" highIlevel" tasks" make" assumptions" on" the" domain" at" hand," nor" do" any" of" the" algorithms"used"to"perform"these"tasks."They"do"however,"make"assumptions"on"the"data"to"which"they"
ARIADNE"D16.1"Public" 21" " conceptual( hierarchicalOclustering"algorithm," each" of" the" clusters" may" be" provided" by"a" humanI interpretable"label"(Fisher"1987,"Fanizzi,"d’Amato"and"Esposito"2008)." " Figure&4D5:&Example&of&generating&a&taxonomy&(right)&from&an&hierarchicallyDclustered&data&set&(left).& Here,&assume&this&is&a&dataset&about&Greek&horae&that&is&clustered&based&first&on&civilization,&second& on&shape,&and&third&on&location.&Note&that&the&labels&would&generally&be&more&descriptive&instead&of&the& ones&used&here.& Thus"far,"only"the"generation"of"a"taxonomy"from"new"data"has"been"considered."The"same"method" however,"may"also"be"applied"to"existing"data"sets"that"already"adhere"to"a"certain"ontology."In"those" cases," an" alternative" taxonomy" may" be" offered" which" is" based" on" similarities" between" the"entities" instead" of"on" a" predefined" taxonomy," which" was" made" by" domain" experts." For" instance," instead" of" finding" Nero"listed" under" Roman& Emperor," he" might" alternatively" be" found" in" the" cluster" containing" resources"concerning"Roman&Cities"and"Disaster."Due"to"the"latter"being"based"on"the"data"itself,"it"will" be"less"likely"to"possess"a"userIinduced"bias"with"respect"to"the"hierarchy."However,"the"occurrence"of" erroneous"data"and"variation"in"the"data’s"quality"might"limit"such"effect." Instead" of" clustering" a" whole" data" set" it"could"alternatively"be" limited"to" a" specific" (combination" of)" subgraph(s)."As"such,"this"approach"may"be"applied"to"the"results"returned"by"a"query."An"advantage"of" this"is"that"the"clusters"and"their"hierarchy"are"based"on"the"local"neighbourhood"of"ones’"query,"thus" offering"only"information"on"relevant"data."In"addition,"each"action"that"moves,"narrows,"or"broadens" the" scope" of" the" search" will" trigger" a" recalculation" of" the" clusters," thus," once" again," providing" information" only" on" the" mostIrelevant" data." Furthermore," by" clustering" similar" results," the" user" is" presented"with"a"lessIexhaustive"list,"thereby"allowing"for"a"more"userIfriendly"and"intuitive"browsing" climate." " "
ARIADNE"D16.1"Public" 22" " 4.2 Applicable&Solutions& The"majority"of"the"developments"in"SWM"are"still"fairly"academic."Even"so,"there"have"also"been"a"fair" number"of"projects"that"have"resulted"in"a"moreIrefined"and"usable"product."Some"of"these"might"even" be"suited"for"ARIADNE,"albeit"with"a"number"of"adjustments."Therefore,"a"cross"section"of"the"existing" solutions"will"be"discussed"next."Note"that,"for"readability,"most"technical"details"will"be"omitted."" 4.2.1 SPARQL&extensions& Recall"that"SPARQL"is"the"recommended"(queryIbased)"interface"for"a"triple"store."As"a"result,"nearly"all" triple"stores"support"this"standard."Hence,"the"integration"of"DM"with"SPARQL"might"provide"a"clean"and" natural"solution."Two"different"types"of"extensions"that"implement"such"capabilities"will"be"discussed" next." 4.2.1.1 Assisted+Query+Forming+ While"SPARQL"has"long"been"the"recommended"standard"for"querying"a"triple"store,"its"complexity"still" forms"a"barrier"for"many"users."A"large"part"of"this"complexity"originates"from"the"heterogeneity"of"the" data,"which"may"be"mapped"to"very"different"ontologies."It"is"unlikely"that"the"average"user"is"familiar" with"all"of"these"variations,"thus"resulting"in"an"inability"to"construct"queries"that"would"optimally"exploit" the"data."Therefore,"several"projects"have"focussed"on"extending"SPARQL"with"the"ability"to"assist"in"the" forming"of"queries"(Figure"4I6)."Two"different"approaches"will"be"briefly"discussed"next." SPACE"is"a"queryIdriven"autocompletion"extension"to"SPARQL"(Kramer,"Dividino"and"Gröner"2013)."At"its" heart"lies"an"index"build"from"past"queries"fired"at"endpoints."The"rationale"for"this"approach"stems"from" the"idea"that"the"query"logs"of"a"specific"endpoint"provide"a"good"representation"of"the"data"to"which" that" endpoint" provides" access." Therefore," while"a" query" is" being" written," SPACE" compares" the" progressed"query"to"those"in"its"index"for"the"corresponding"endpoint"and"subsequently"suggests"the" most"similar"past"query." An"alternative"to"a"queryIdriven"approach"is"that"of"a"dataIdriven"approach."(Gombos"and"Kiss"2014," Campinas,"et"al."2012)." That" is," instead" of" predicting" a" query" based" on" previous"queries," restrict"the" possible"queries"to"what"the"data"can"offer."One"such"method"consists"of"generating"a"graph"summary" of"the"RDF"graphs."Once"built,"it"represents"a"generalization"of"the" original"graph"from"which"oftenI paired"RDF"elements"can"be"queried."For"instance,"it"might"link"the"predicate"written&by"to"the"class" author."Consequently,"if"the"latestIwritten"term"of"a"query"would"consist"of"the"predicate"written&by,"a" suggestion"of"the"class"author,"e.g."D.&Wheatley,"might"be"given."
ARIADNE"D16.1"Public" 23" " " Figure&4D6:&Implemented&example&of&a&SPARQL&assisted&query&formulation&(Campinas,&et&al.&2012).& 4.2.1.2 SPARQL;ML+ With" SPARQLIML," a" wide" range" of" prediction" and" classification" methods" are" added" to" the" SPARQL" interface" (Kiefer," Bernstein" and" Locher" 2008," Locher" 2007)." These" methods" stem" from" SRL" and" have" been" modified" to" work" directly" to" graph" data." In" addition," the" methods" can" be" accessed" through" statements"that"follow"the"SPARQL"grammar"and"which"are"similar"to"those"used"by"Microsoft’s"DataI Mining"Extension18."" Under"the"hood,"SPARQLIML"requires"the"data"is"stored"in"the"MonetDB19"database,"which"supports" both"relational"and"graph"data."In"addition,"it"needs"the"Weka20"and"Proximity21"DM"APIs,"through"which" all"DM"operations"are"processed."These"software"packages"are"run"on"the"server"side,"i.e."on"the"server" that"provides"a"SPARQLIML"interface."Hence,"it"spares"the"users"the"burden"of"installing"additional"tools." Furthermore,"all"normal"SPARQL"operations"remain"unaffected." " Figure&4D7:&An&example&query&in&SPARQLDML&to&learn&a&predictive&model." """"""""""""""""""""""""""""""""""""""""""""""""""""""""""""" 18"Microsoft’s"DataIMining"Extension,"see"msdn.microsoft.com/enIus/library/ms132058.aspx"" 19"MonetDB,"see"www.monetdb.org"" 20"Weka"is"an"openIsource"data"mining"tool"for"propositional"data,"see"www.cs.waikato.ac.nz/ml/weka/"" 21"Proximity"is"an"openIsource"data"mining"tool"for"relational"data,"see"kdl.cs.umass.edu/display/public/Proximity""
ARIADNE"D16.1"Public" 24" " 4.2.2 Ranking&Methods& Traditional" ranking" methods," such" as" Google’s" PageRank22," generally" apply" a" form" of" an" authorityI ranking"algorithm."These"algorithms"order"pages"on"the"web"based"on"how"often"they"are"referred"to" from"authoritative"sites."Here,"authoritative"sites"are"defined"as"important"and"trusted"hubs"on"the"web," e.g."due"to"their"influence"within"the"online"community." While"traditional"ranking"methods"work"quite"well"on"the"regular"web,"they"do"not"contain"a"mechanism" for"handling"and"exploiting"semantic"relationships." Hence,"these" are" ill"suited" in"the"case"of"the"SW." Fortunately,"a"handful"of"ranking"schemes"exist"that"specifically"target"the"SW."Five"of"those"will"briefly" be"discussed"next,"thereby"omitting"those"less"relevant"to"ARIADNE." 4.2.2.1 TripleRank+ TripleRank" is" an" authorityIranking" method" that" takes" the" semantics" of" LD" into" account"(Franz," et" al." 2009)."This"is"accomplished"by"representing"the"graph"as"a"thirdIorder"tensor"(C.1.2),"which"is"capable"of" exploiting"these"semantics"in"a"natural"fashion."By"applying"a"specific"factorization"method,"authoritative" sources"can"be"determined."By"subsequently"calculating"the"contributions"of"these"sources"to"a"group"of" triples,"an"ordering"can"be"found."In"addition,"groups"of"semanticallyIsimilar"predicates"and"resources" may"be"identified." 4.2.2.2 ReConRank+ The"ReConRank"method"is"a"fusion"of"two"ranking"algorithms;"ResourceRank"and"ContextRank."(Hogan," Harth"and"Decker"2006)."The"first"of"these"has"been"adapted"from"PageRank"and"thus"applies"a"form"of" authority"ranking."This"is"accomplished"by"iteratively"going"through"the"graph"whilst"ignoring"the" semantics"of"the"connecting"links."ContextRank,"on"the"other"hand,"takes"the"context"graph"into" account."This"graph"consists"of"contextIspecific"resources"and"predicates"which"are"trusted"to"be"valid." Finally,"both"ranking"algorithms"are"combined"to"compute"the"ReConRank"order"of"relevance." 4.2.2.3 xhRank+ xhRank"is"a"ranking"approach"that"endeavours"to"implement"multiple"different"metrics"into"one"single" package"in"the"hope"of"achieving"the"best"of"several"worlds"(He"and"Baker"2011)."Hereto,"it"calculates" the" ranking" based" on" relevance," on" importance," and" on" query" length." Of"these," the" relevance"is" determined"based"on"the"context"graph"of"a"query,"as"well"as"on"the"contextual"similarity"of"phrases"and" terms" within" that" query," and" those" contained" in" the" RDF" graph."The" importance" is" computed" by" considering"authority"nodes,"as"well"as"the"popularity"of"all"relevant"resources."Finally,"the"query"length" is"calculated"by"evaluating"a"(weighted)"context"graph"with"respect"to"the"input"query."Once"all"metrics" have"resulted"in"a"(raw)"rank,"these"are"combined"to"form"the"overall"rank." """"""""""""""""""""""""""""""""""""""""""""""""""""""""""""" 22"PageRank,"see"www.google.com/about/company/products/""
ARIADNE"D16.1"Public" 25" " 4.2.2.4 SemRank+ The"SemRank"relevance"model"constitutes"a"fusion"of"semantic"and"information"theoretic"methods,"as" well" as" heuristics"(Etter" and" Domeniconi" 2014," Anyanwu," Maduko" and" Sheth" 2005)." Together," these" techniques" result" in" a" unified" model" with" which" all" types" of" complex" semantic" relations," known" as" Semantic" Associations" (SA)" (Figure" 4I8)," can" be" ranked"by" relevance." Instead" of" applying" popular" relevance"measures,"such"as"shortest"path"or"leastIfrequently"occurring"path,"SemRank"calculates"the" Information"Gain"(IG)"per"relation."This"metric"conveys"how"much"information"a"user"would"gain"when" presented"with"the"SA"to"which"the"IG"belongs."However,"as"the"developers"acknowledge"that"different" domains"require"different"measures,"they"provide"the"option"of"easily"switching"to"another"relevance" metric." " Figure&4D8&:&Three&types&of&Semantic&Associations&(Anyanwu,&Maduko&and&Sheth&2005).& 4.2.2.5 Vector+Space+Models+ Vector"Space"Models"(VSM)"constitute"an"approach"by"which"(unstructured)"data"can"easily"be"searched" through"to"find"potentially"relevant"answers"that"fit"a"descriptive"query"(Appendix"B)."Such"an"approach" can" be" added" on" top" of" regular" SW" queries," thereby" providing" more"versatile" and" more"scalable" searching"abilities."Moreover,"as"VSMs"represent"the"potential"relevance"on"a"continuous"scale,"these" values"may"be"used"to"rank"the"corresponding"results"as"well." Several" studies" have" focused" on" incorporating" VSM" into" the" SW" (Freitas," et" al." 2012," Mendes," et" al." 2011,"Castells,"Fernandez"and"Vallet"2007,"Tous"and"Delgado"2006)."Hereto,"they"indexed"resources"by" their"vectors."That"is,"each"nonIzero"term"in"a"vector"denoted"a"predicateIresource"pair"belonging"to" that"resource."In"addition,"each"of"these"pairs"was"weighted"to"reflect"how"well"they"represented"their" corresponding" resource." For" instance," given" an"item" in" an"archaeological" data" set," the"predicateI
ARIADNE"D16.1"Public" 26" " resource"pair"instance&of&Dragendorff&33"would"be"far"better"at"describing"the"item"that"the"pair"is&an& artefact." The"implementation"of"VSM"into"the"SW"can"be"approached"in"two"different"ways."Either"a"keywordI based"query"is"used"to"generate"a"SPARQL"query,"or"a"SPARQL"query"is"used"from"which"keywords"are" extracted."Either"way,"both"query"and"keyword"are,"at"some"point,"available"for"further"processing."This" process"continues"by"the"execution"of"the"query"by"the"query"engine,"after"which"the"results"are"ranked" based"on"their"similarity"to"the"keywords"(Figure"4I9)." " Figure&4D9:&Workflow&as&to&how&VSM&may&aid&in&ranking&query&results&(Castells,&Fernandez&and&Vallet& 2007).& 4.2.3 Frameworks& Several"frameworks"exist"that"focus"on"adding"DM"capabilities"to"graph"data"or"data"composed"of"formal" ontologies."Of"these,"four"relevant"examples"will"be"discussed"next." 4.2.3.1 SUNS+ Statistical"Unit"Node"Set"(SUNS)"is"a"ML"plugin"for"the"Large"Knowledge"Collider23"(LarKC");"a"largeIscale" integration"project"aimed"at"developing"a"platform"for"massive"distributed"incomplete"reasoning"on"the" SW."It"has"since"been"ported"to"work"on"relational"data"as"well"(Huang,"Tresp"and"Kriegel,"et"al."2009)." """"""""""""""""""""""""""""""""""""""""""""""""""""""""""""" 23"Large"Knowledge"Collider;"a"FP7"project,"see"www.larkc.eu""
ARIADNE"D16.1"Public" 27" " The"SUNS"framework"(Huang,"Tresp"and"Bundschus,"et"al."2011,"Huang"and"Tresp"2010)"centres"on"the" concepts" of" statistical& unit"and" population," which" it" defines" as" an" instance" of" a" certain" class" and" all" instances"under"consideration,"respectively."In"addition,"it"defines"each"potential"triple"as"a"binary"triple& node"of"which"the"value"is"true"if"the"triple"is"known"to"exist"and"false"if"the"triple"is"known"not"to"exist." Moreover,"the"entirety"of"triple&nodes"that"belong"to"a"statistical&unit"is"defined"as"the"statistics&unit& node&set."For"instance,"an"arbitrary"artefact"of"the"class"Dragendorff&33"would"be"a"statistical&unit"of"that" class."Each"triple"that"states"a"fact"about"that"artefact,"i.e."that"it"has"the"colour"red,"is"a"positive"triple& node."" At"its"core,"SUNS"applies"a"multivariate24"prediction"algorithm,"which"is"advocated"by"SUNS’"developers" as"providing"an"improved"predictive"performance"when"compared"to"traditional"algorithms."Regardless," it"constitutes"a"propositional"approach,"thus"requiring"the"graph"data"to"be"translated"into"a"relational" matrix." The" unknown" triples" are" subsequently" predicted" by" factorization" (C.1.1)," with" the" option" of" integrating"this"new"knowledge"into"the"triple"store." " Figure&4D10&:&A&schematic&overview&of&the&SUNS&framework&(Huang&and&Tresp&2012).& 4.2.3.2 LiDDM+ The"LinkedIData"Data"Miner"(LiDDM)"aims"at"providing"a"framework"that"offers"a"LDIspecific"alternative" to" the" moreIregular" KDD" schema" (Figure" 4I11)" (Narasimha," et" al." 2011," Ramezani," Saraee" and" Nematbakhsh"2013)."Moreover,"it"emphasizes"the"notion"of"simplicity"with"the"developers"advocating" the"use"of"regular"SPARQL"queries"in"order"to"prevent"alienating"the"user"with"a"complex"learning"curve." """"""""""""""""""""""""""""""""""""""""""""""""""""""""""""" 24"Multivariate"methods"are"a"fusion"of"supervised"and"unsupervised"methods,"which"use"known"input"features"to" predict"several"variables"jointly,"thereby"generally"increasing"their"predictive"strength."
ARIADNE"D16.1"Public" 28" " The"first"step"in"the"LiDDM"framework"is"the"import"of"data"as"formulated"as"a"SPARQL"query."By"going" through"several"preIprocessing"steps"the"data"is"then,"among"other"things,"translated"to"a"propositional" format." Once" completed," propositional" DM" methods" can" be" applied" as" provided" by" an" external" DM" processing"engine."The"results"of"this"can"subsequently"be"visualized."" In"order"to"test"their"framework,"the"developers"created"the"LiDDM"Tool"(LiDDMT)."To"facilitate"the"DM" methods"they"implemented"the"Weka"API."Based"on"the"results"gained,"they"conclude"that"the"strength" in"the"proposed"framework"lies"in"mining"several"aggregated"data"sets"simultaneously,"thereby"offering" a"flexibility"with"respect"to"the"data"format."In"addition,"the"developers"hope"to"automate"many"of"its" features"in"the"near"future." " Figure&4D11&:&Schematic&depiction&of&the&LiDDM&architecture&(Narasimha,&et&al.&2011).& 4.2.3.3 OLAP+ In"the"field"of"Business"Intelligence,"Online"Analytical"Processing"(OLAP)"involves"a"framework"for"the" analysis"of"multidimensional"relational"data"(Codd,"Codd"and"Salley"1993)."Typically,"this"concerns"an" interactive"DM"process"from"which"the"results"may"lead"to"business"and"financial"reports." The"core"concept"of"any"OLAP"application"is"the"OLAP"Cube"(Figure"4I12)."Simply"put,"an"OLAP"Cube" represents"a"generalization"of"tabular"data,"thereby"placing"certain"aspects"of"a"multidimensional"data" set"on"the"axes"of"the"cube."When"requiring"more"than"the"three"dimensions"that"a"cube"provides,"it"is" customary"to"speak"of"an"OLAP"Hypercube."New"insights"may"next"be"obtained"by"interactively"selecting"
ARIADNE"D16.1"Public" 29" " and"analysing"slices"of"this"(hyper)"cube."For"instance,"data"on"artefacts,"location"of"their"discovery,"and" their"carbon"dating"might"be"brought"together"in"a"threeIdimensional"OLAP"Cube"to"analyse"the"possible" relations"between"them." Analysing"data"on" the"SW" by"OLAP"has"slowly"been"gaining"momentum."However,"the"current"focus" leans" more" towards" the" translation" of" relational" data" to" an" OLAP" representation" on" the" SW." To" accomplish" this," the" W3C" recommends" the" use" of" the" RDF" Data" Cube" ontology" (QB)" (Cyganiak" and" Reynolds" 2014)." Others" however," deem" this" ontology" too" limited" (Etcheverry" and" Vaisman" 2012," ragimov,"et"al."2014),"and"have"extended"it"with"QB4OLAP"to"enable"all"analytical"abilities"of"OLAP."As"a" result,"relational"data"published"on"the"SW"using"the"QB"ontology"can"be"analysed"with"OLAP"techniques" by" using" QB4OLAP." A" more"native" alternative" is" to" use" the" Open" Cube" vocabulary"(Etcheverry" and" Vaisman" 2012)," which" combines" QB" and" QB4OLAP" into" a" single" ontology." Moreover," it" allows" for" performing"OLAP"operations"via"SPARQL"queries." " Figure&4D12:&Schematic&representation&of&an&OLAP&workflow.& 4.2.3.4 AITION+ AITION"is"an"interactive"DM"solution"for"the"biomedical"domain"(Dimitropoulos,"et"al."2012,"Metaxas,"et" al."2014)." Developed" by"the" University" of" Athens25"for" the" FP6"HealthIeIChild26"project,"it"specifically" aims"at"discovering"knowledge"in"a"medical"processing"environment."More"specific,"it"provides"a"full"KDD" solution"by"which"(biomedical)"researchers"can"preIprocess,"simulate,"and"visualize"relational"data,"as" """"""""""""""""""""""""""""""""""""""""""""""""""""""""""""" 25"University"of"Athens,"see"www.uoa.gr"" 26"HealthIe"Child,"see"www.healthIeIchild.org""
ARIADNE"D16.1"Public" 30" " well" as" construct" statistical" models" and" subsequently" infer" from" them" (Figure& 4D13)." Hereto," AITION" offers"a"userIfriendly"graphical"interface"similar"to"those"of"statistical"software"packages."This"interface" allows"one"to"tweak"the"KDD"process"to"his"or"her"needs,"amongst"which"are"the"selection"of"algorithm" and"the"optional"specification"of"prior"knowledge"and"medical"ontologies." Originally,"AITION"was"developed"as"a"standIalone"desktop"application."However,"due"to"limitations"in" processing" capabilities" it" was" later" extended" to" a" serverIoriented" design." This" allows" it" to" run" on" distributed" architectures" such" as" clusters," grids," and" clouds." These" architectures" however," should" provide"access"to"a"relational"(big"data)"database"for"AITION"to"work"properly."Hence,"AITION"expects" the"data"to"conform"to"the"specifications"of"MRDM."That"is,"graphIbased"data"such"as"LD"is"unsupported" at"this"time." " Figure&4D13:&Schematic&depiction&of&the&AITION&framework&(Dimitropoulos,"et"al."2012).& 4.2.4 Platforms& Whereby"previous"solutions"were"either"mostly"theoretical,"or"extensions"to"triple"stores,"the"following" two" platforms" constitute" relatively" complete" products" with" a" refined" user" interface." These" will" be" touched"on"next." 4.2.4.1 Rapidminer+ Rapidminer" is" a" popular" DM" and" Business" Analytics" platform" that" aims" at" providing" the" whole" KDD" process"to"business"users."To"this"end,"it"combines"a"wide"range"of"DM"and"DMIrelated"techniques"with" an"intuitive"interface"(Figure"4I14)."Moreover,"its"developers" try"to"stay"at"the"front"of"technological" innovation," thereby" offering" versions" capable" of" running" on" highIperformance" and" distributed" architectures,"as"well"as"running"directly"from"the"cloud."
ARIADNE"D16.1"Public" 37" " apparent" lack" of" currentlyIavailable" DM" applications" within" the" domain." That" is," of" the" users" that" participated"in"the"survey,"less"than"a"fifth"rated"their"availability"good"(12%)"or"very"good"(5%).""This" state"of"affairs"prevents"users"from"gaining"any"practical"experience,"and"thus"impeding"them"in"forming" a"realistic"view"of"the"added"value"that"DM"could"offer."" In"the"case"of"data"enhancement,"only"few"stakeholders"of"the"sample"group"–"all"affiliated"with"data" repositories"–"appear"to"hold"the"belief"that"DM"might"be"beneficial."In"fact,"only"one"of"them"strongly" beliefs" that" the" integration" of" DM" in" repositories" should" be" regarded" as" a" very" important" aspect." Additional"interviews"(n&=&3)"were"conducted"by"(Hollander"and"Hoogerwerf"2014)"as"part"of"ARIADNE" Deliverable"13.1,"from"which"a"desire"to"enhance&metadata"was"concluded."This"enhancement"would" entail,"among"other"thing,"automated"metadata"extraction,"duplication"detection,"link"prediction,"and" overall"semantic"enrichment." During"the"survey,"the"sample"group"appeared"to"hold"mixed"beliefs"about"whether"DM"is"an"important" aspect"of"their"research"(Selhofer"and"Geser"2014)."More"specifically,"half"of"them"deemed"it"either"very" important"(16%)"or"rather"important"(34%),"while"the"remainder"thought"it"to"be"rather"unimportant" (26%)"or"very"unimportant"(23%)."Even"less"of"the"survey’s"participants"had"ever"used"a"DM"solution" during"their"research,"with"less"than"a"tenth"(8%)"having"used"it"very"often,"and"slightly"more"(16%)" having"applied"it"frequently."If"fact,"near"half"(41%)"had"(almost)"never"used"DM"at"all."These"numbers" may" partially" be" inaccurate" given" that," as" indicated" earlier," many" participants" were" unaware" of" the" possibilities"that"the"integration"of"DM"tools"into"their"research"may"provide,"and"may,"in"fact,"have"used" such" tools" unwittingly."Examples" of" such"tools" are" those" integrated"into" a" GIS," with" which" advanced" analysis"of"archaeological"data"is"already"possible"(Conolly"and"Lake"2006)."" According" to" the" survey’s" participants," the" most" important" type" of" data" they" are" using" during" their" research"is,"unsurprisingly,"that"of"excavation"data"(Selhofer"and"Geser"2014)."Of"the"other"important" sources,"many" appear" to"involve"a"geographical"component,"such" as" GIS"or"satellite"data."Therefore," augmenting"those"types"of"data"with"the"help"of"DM"would"likely"offer"the"most"effective"improvement" during" the" knowledgeIdiscovery" stage" as" perceived" by" the" user." However," it" may" also" be" true" that" researchers"are" using" the"other"types"of"data"less" often" due" to"the"poor"transparency"that" currently" exists." Irrespective," Hollander" and" Hoogerwerf" (2014)"emphasize" the" need" to" present" the" user" with," among" others," a" geoIintegrated" search," thereby" offering" numerous" options" such" as" an" interactive" timeline"and"various"background"layers." Another" important" part" of" data" transparency" is" the" ability" to" easily" distil" relevant" information" from" exhausting" quantities" of" potentiallyDrelevant"information." One" approach" to" achieve" this" might" be" to" apply" some" form" of" ranking" to" the" data" (Hollander" and" Hoogerwerf" 2014)," either" automatically" or" manually"by"the"user."Alternatively,"a"combination"of"the"two"may"be"possible,"whereby"the"users"can" correct"or"otherwise"influence"the"automatically"generated"rankings."
ARIADNE"D16.1"Public" 38" " 5.2 Wishes&of&Domain&Experts& In"an"effort"to"gain"additional"insight"into"the"needs"of"archaeological"researchers"we"organized"five"oneI hour"brainstorming"sessions."Four"of"these"were"held"at"Leiden"University"and"one"at"the"VU"University" Amsterdam." Each" of" these" sessions" consisted" of" an" open" discussion" with" a" different" archaeological" specialist,"among"which"were"junior"and"senior"researchers"as"well"as"Ph.D."candidates."Furthermore," their" fields" of" expertise" ranged" from"Prehistoric" hunter" gatherers" and"farming" communities," to" longI term"developments"of"settlements,"landIuse"dynamics,"and"spatial"dynamic"modelling."" The" participants" were" asked" to" write" down" several" archaeological" research" scenarios," which" they" deemed"difficult"with"the"current"information"infrastructure."Moreover,"the"assumption"was"made"that" all"archaeological"publications"and"data"would"have"been"made"available"in"one"large"database."In"total," the"sessions"resulted"in"26"scenarios,"of"which"an"English"translation"is"given"in"Appendix"D." An" analysis" of" the" scenarios"posed" revealed"that" the" large" majority" (77%)" of" the" difficulties" that" researchers" experience" concern" an" informationIgathering" task." Hence," the" aggregation" of" data" is" desired."In"several"occasions,"these"data"are"likely"to"originate"from"multiple,"distinct"sources"and"stem" from"different"domains,"amongst"which"are"plant"and"dental"records."Moreover,"nearly"half"of"these" informationIgathering" scenarios"(40%)" relate" some" archaeological" entity" to" a" geospatial" component," thereby" making" this" latter" domain" the" second" most" prominent." Finally," three" scenarios"mention"" difficulty"in"finding"relevant"data"using"a"keywordIbased"search"engine." The" remaining" research" scenarios"(23%)" involve" a" factIfinding" task," among" which" are" the" search" for" contact"information,"specific"GIS"files,"and"the"number"of"a"physical"storage"container."Two"scenarios" that"stand"out"concern"the"trustworthiness"of"certain"data,"which"is"dependent"on"the"corresponding" archaeological"context." InformationDSeeking&Task& Number&of&Scenarios& Fact"Finding" 6" Information"Gathering" 20" Keeping"UpItoIdate" 0" Table&5D1:&Dissemination&of&the&archaeological&research&scenarios&into&the&three&distinct&InformationD Seeking&tasks.& When"assuming"an"ideal"situation"in"which"all"required"data"has"been"made"available"as"LD,"most,"if"not" all"of" these" informationIgathering" tasks" would" be" accomplishable" without" much"trouble."In"contrast," none"of"the"factIfinding"tasks"would"require"LD;"for"these,"a"simple"relational"database"would"suffice." These"results"are"similar"to"those"of"the"survey"and"the"interviews."That"is,"the"abilities"that"are"provided" by"LD"are"likely"to"greatly"improve"the"knowledgeIdiscovery"process"of"an"archaeological"researcher." Also"similar,"unfortunately,"is"the"difficulty"of"distilling"tasks"to"which"DM"might"prove"beneficial.""
ARIADNE"D16.1"Public" 39" " 5.3 Summary& Based"on"the"previouslyIdiscussed"survey"and"interviews,"the"following"points"can"be"distilled"which"are" deemed"relevant"with"respect"to"DM."Please"note"however,"that"a"solution"to"the"last"point"would"more" likely"be"considered"an"issue"treatable"by"ML"rather"than"DM." Unawareness(of(the(available(data"reflects"the"lack"of"knowledge"users"possess"on"what"data"is" available" to" them." This" hinders" them" during" the" early" stages" of" their" research," as" they" fail" to" explore"data"outside"the"scope"of"their"current"search." Uncertainty( on( how( to( locate( relevant( data"concerns" the" difficulty" of" users" to" get" an" understanding"of"the"actions"required"to"find"and"access"the"data"they"are"looking"for."Instead," they"run"into"ambiguity"issues"with"the"terms"they"use"to"search." Inability(to(effectively(distil(relevant(data"involves"the"issues"that"users"experience"when"faced" with"an"exhaustive"list"of"potentially&relevant"data."To"determine"what"data"is"relevant"and"what"is" not,"they"need"to"thoroughly"examine"each"entry"in"the"list." Incompleteness(of(Data"concerns"the"perceived"gaps"in"information,"due"to"having"been"omitted" during"either"the"study"or"the"subsequent"digitization."As"researchers"cannot"ascertain"whether" the"missing"data"is"of"importance,"this"hinders"them"from"trusting"the"data"set"as"a"whole."
ARIADNE"D16.1"Public" 40" " 6 Data&Understanding& Data"Understanding"is"an"important"step"within"any"KDD"and"concerns"inspecting"the"data,"their"quality," and"their"abnormalities."Obtaining"a"good"overview"of"these"aspects"contributes"to"the"performance"of" any"future"DM"applications"within"the"ARIADNE"infrastructure."After"completion"of"this"project"at"the" end" of" February" 2017," it" is" hoped"that"Europe's" archaeological" communities"will" adopt" it." Until" that" moment"however,"the"available"data"will"be"comprised"of"what"already"exists,"as"well"as"small"amounts" of"new"data"which"will"be"produced"by"those"involved"with"Natural"Language"Processing."Therefore,"the" conclusions"at"the"end"of"this"report"will"apply"only"to"these"data."It"is"expected"however,"that"these" data"will"constitute"a"sufficient"representation"of"the"future"data,"and"thus"allow"for"generalization,"such" that"the"conclusions"will"hold." We"expect"that"the"large"majority"of"the"data"that"ARIADNE"will"offer"during"its"first"couple"of"years"will" be"provided"by"the"currently"existing"digitallyIaccessible"data"infrastructures."Only"a"few"of"those"have" already"explored"the"possibility"of"publishing"their"data"as"LD,"and"even"fewer"have"embraced"it"entirely." A"good"example"of"a"repository"that"does"embrace"LD"is"the"ADS;"the"Archaeological"Data"Service"based" at"the"University"of"York,"and"which"has"adopted"all"facets"of"the"LD"paradigm"for"a"section"of"their" archive"data,"and"for"all"their"resource"discovery"metadata."An"example"of"a"repository"that"is"well"on"its" way" might" be" EASY29;" the" digital" warehouse"on" Digital" Humanities" hosted" by" DANS30," which" offers" unstructured"data"together"with"their"corresponding"LD"counterpart."In"some"cases"however,"data"sets" consisting"solely"of"LD"are"published31." 6.1 Data&Produced&using&NaturalQLanguage&Processing& In"addition"to"the"Linked"Archaeological"Data"(LAD)"from"existing"infrastructures,"ARIADNE"will"provide" LAD"that"has"been"generated"semiIautomatically"from"unpublished"archaeological"reports"by"means"of" NaturalILanguage"Processing"(NLP)."These"reports,"the"soIcalled"grey&literature,"are"increasingly"‘born" digital’."Those"that"are"not"are"scanned"to"create"a"copy,"which"is"usually"made"available"in"PDF"format." In" either" case," this"may"be"followed" by" processing"the"text"through"specialised"tools," to" convert" the" report"to"LD."" Within"ARIADNE,"the"task"of"exploring"the"applicability"of"NLP"falls"under"WP16"with"task"number"16.2." The"main"contributors"to"this"task"are"the"University"of"South"Wales,"the"ADS,"and"Leiden"University."As" it"may"be"assumed"that"other"NLP"endeavours"will"proceed"in"a"similar"fashion,"there"will"thus"be"no" """"""""""""""""""""""""""""""""""""""""""""""""""""""""""""" 29"Electronic"Archiving"System,"see"easy.dans.knaw.nl/"" 30"Data"Archiving"and"Networked"Services,"see"dans.knaw.nl/en"" 31"For"example,"consider"the"recent"CLARIN"Dutch"Ships"and"Sailors"data"set""
ARIADNE"D16.1"Public" 41" " difference"between"data"from"either"origin"with"respect"to"accessing"and"retrieving"it."That"is,"both"will" be"stored"on"the"same"(or"similar)"triple"store"using"the"same"(or"similar)"ontologies." The"main"challenge"with"which"the"field"of"NLP"struggles"is"very"similar"to"that"which"the"Semantic"Web" tries" to" solve," namely" the" problem" of" making" knowledge" interpretable" by" software" agents" (Briscoe" 1991)."However,"the"Semantic"Web"tries"to"attack"this"problem"by"structuring"the"knowledge,"whereas" the" field" of" NLP" generally" tries" to" create" semanticallyIaware" agents." Unfortunately," due" to" its" complexity,"this"latter"approach"is"still"far"from"perfect."Therefore,"any"knowledge"engineer"should"be" wary"of"possible"flaws"in"the"data"converted"by"NLP,"which"were"not"detected"during"manual"or"(semiI)" automatic"checks." 6.2 Case&Study&on&Data&Repositories& Of"the"existing"LAD"infrastructures"(Fentress"2014),"four"have"been"chosen"from"which"the"RDF"data"was" deemed"a"suitable"representative"of"their"specific"area"of"expertise."In"addition,"this"selection"favoured" minimal"overlap"between"the"corresponding"data"sets,"as"to"provide"as"high"a"degree"of"variability"as" possible."These"four"are"the"ADS,"EASY,"Open"Context,"and"Pleiades." 6.2.1 ADS& Recall"that"the"Archaeological"Data"Service"(ADS)"provides"a"web"portal"to"a"triple"store."At"the"moment" of"writing,"it"holds"the"aggregated"data"of"nearly" 500"distinct"collections."Together,"these"collections" provide"472,172"triples"on"103,148"resources32." The" data" appears" to" be" structured" in" three" sections," each" with" a" different"purpose." The" smallest" of" these,"comprising"less"than"a"few"percent"of"the"whole"graph,"describes"several"types"of"amphorae"by" the"distinct"physical"features"that"make"up"their"shape."The"corresponding"statements"consist"largely" out"of"SKOS"statements."" With"a"slightly"higher"percentage"follow"numerous"descriptions"of"the"sources"and"studies"from"which" the"data"originated,"all"of"which"mainly"make"use"of"Dublin"Core"(DC)."It"should"be"noted"however,"that" most"of"the"DC"triples"have"literals"as"their"object."Furthermore,"it"appears"as"if"the"range"of"certain"DC" attributes,"e.g."dc:coverage,"covers" disjointed" sets" such" as" the" disjoint"union" of" temporal" period" and" geographical"location." The" large" majority" of" the" graph"describes" the" artefacts" provided" by" the" sources"and" studies." The" corresponding"statements"are"specified"using"the"CIDOC"CRMIEH"ontology,"of"which"the"properties"start" with"a"unique"identifier."In"addition,"descriptions"of"the"same"resource"frequently"occur"more"than"once," with"each"occurrence"specifying"a"different"set"of"properties." """"""""""""""""""""""""""""""""""""""""""""""""""""""""""""" 32"These"figures"are"based"on"the"graphs"retrieved"in"December"2014."
ARIADNE"D16.1"Public" 42" " 6.2.2 EASY& Recall" that" EASY" concerns" a" digital" repository" on" Digital" Humanities" hosted" by" DANS," and" which" experiments"with"offering"LD"alongside"their"corresponding"unstructured"counterpart,"among"which"are" databases,"images,"and"reports."On"the"whole,"the"data"within"EASY’s"test"triple"store"constitutes"more" than"25,000"data"sets"which,"together,"results"in"roughly"28,000"resources"consisting"of"836,447"triples" in"total33." Little"variation"between"concept"descriptions"exist"within"the"graph,"with"almost"every"one"consisting"of" a"single"data"sets"with"similar"properties."These"properties"mostly"concern"metadata"and"are"described" in" Dublin" Core34"(DC)"with" literal" values."Therefore,"while" the" graph" is"relatively" uniformly" divided,"it" does"contain"little"crossIlinking." While"not"serialised"in"RDF,"EASY"additionally"provides"each"data"set"with"more"contextual"information" in"the"form"of"a"XML"file;"the"“pakbon”,"or"packing"slip"in"English."This"file"consists"of"numerous"fixed" attributes,"developed"in"cooperation"with"the"SIKB35,"that"are"to"be"supplied"by"the"submitter"of"the" data" set." If" converted" to" RDF," the" information" herein" might" prove" quite" beneficial" to" archaeological" researchers." 6.2.3 Open&Context& Open"Context36"is"a"web"portal"that"allows"researchers"to"publish"and"access"scientific"data"from"various" domains," such" as" zooarchaeological"and" spatial" archaeology," as" well" as" numismatics." At" the" time" of" writing," the" portal" offered" access" to" data" from" 20" projects," with" 34" others" forthcoming37." Of" the" corresponding"data"sets,"only"a"fragment"have"been"converted"to"RDF."Instead,"most"data"resides"in"a" combination"of"JSON38,"KML39,"and"ArchaeoML40."Here,"the"latter"two"are"XML"formats,"which"aim"at" expressing"geospatial"and"archaeological"data,"respectively."However,"OpenContext"is"currently"working" on"aligning"its"metadata"to"the"CIDOCICRM." The"currently"available"RDF"data"consists"of"nearly"5,000"triples"which,"together,"describe"roughly"1,250" resources41."Of"these"resources,"the"majority"consist"of"coin,"region,"and"site"descriptions."Little"variation" between" concept" descriptions" exist" within" or" even" between" these" domains," which" all" apply" a" static" """"""""""""""""""""""""""""""""""""""""""""""""""""""""""""" 33"These"figures"are"based"on"the"graphs"retrieved"in"December"2014." 34"Dublin"Core,"see"www.dublincore.org"" 35"Stichting"Infrastructuur"Kwaliteitsborging"Bodemonderzoek,"see"http://www.sikb.nl"" 36"Open"Context,"see"www.opencontext.org"" 37"These"figures"are"based"on"the"graphs"retrieved"in"December"2014." 38"JavaScript"Object"Notification,"see"www.json.org"" 39"Keyhole"Markup"Language"(KML),"see"www.opengeospatial.org/standards/kml/"" 40"Archaeo"Markup"Language"(ArchaeoML),"see"www.opencontext.org/about/concepts/"" 41"These"figures"are"based"on"the"graphs"retrieved"in"December"2014."
ARIADNE"D16.1"Public" 43" " generic"set"of"properties"whereby"only"the"values"differ."These"values"often"involve"custom"resources,"as" well"as"those"provided"by"Geonames"42." 6.2.4 Pleiades& Pleiades43"is"a"web"portal"that"provides"historical"geographical"information"about"place"from"the"ancient" world."Hereto,"they"use"their"own"definition"of"place,"which"constitutes"a"geographical"location"with"an" ancient" name," which" may" vary" over" time." Currently," the" database" contains" close" to" 3,500" places," resulting"in"2,258,807"triples44."Furthermore,"the"corresponding"resources"on"authors,"place"types,"and" time"periods"consist"of"an"additional"5,000"triples." Within"the"triple"store"each"resource"type"has"its"own"graph,"with"the"place"descriptions"encompassing" nine" graphs." These" latter" graphs" constitute" the" majority" of" the" data." While" moderate" variation" in" concept"descriptions"exist"between"the"typeIspecific"graphs,"there"is"little"variation"within"them." Particularly"interesting"is"the"choice"to"include"an"errata"graph."The"goal"of"this"graph"is"to"hold"the" falsified" statements" from" other" graphs" instead" of" correcting" the" errors" directly." Therefore," both" the" correct"and"the"incorrect"version"are"available."" 6.3 Summary& Based"on"the"case"study"of"four"different"data"sets,"as"well"as"on"the"differences"between"them,"the" following"notes"can"be"distilled"which"might"require"special"attention"when"designing"a"DM"solution." Differences(in(ontologies"used"exist"between"the"data"sets"from"different"sources,"whereas"this" occurs"less"so"within"those"from"the"same"source."Additionally,"the"usage"of"different"versions"of" the"same" ontology"can"be"observed,"which"may" cause"such"an"ontology"to"be"either"under"or" overrepresented"if"not"dealt"with"accordingly."Moreover,"attention"should"be"given"to"the"use"of" custom"ontologies."Note"that,"ideally,"all"data"will"be"translated"using"a"single"ontology"such"as" CIDOC"CRMarchaeo,"CRMIEH,"or"the"ACDM." Structural(variation"within"data"sets"from"the"same"source"appears"relatively"little,"with"most"of"a" dataset’s"concept"descriptions"following"roughly"the"same"structural"schema."More"specifically," nearly"all"descriptions"in"a"data"set"use"roughly"the"same"set"of"properties,"which"are"specified" using" the" same" ontologies." The" inverse" appears" true" between" the" data" sets" from" different" sources,"which"all"have"their"own"distinct"schema"structure." """"""""""""""""""""""""""""""""""""""""""""""""""""""""""""" 42"Geonames,"see"www.geonames.org"" 43"Pleiades,"see"pleiades.stoa.org/"" 44"These"figures"are"based"on"the"graphs"retrieved"in"December"2014."
ARIADNE"D16.1"Public" 44" " StructurallyOflat( graphs"appear" to" be" quite" common." This" phenomena" occurs" due" to" the" scarceness"of"URIs"per"resource,"either"because"a"resource"only"has"a"small"number"of"properties," or"because"of"an"extensive"use"of"literals."In"the"latter"case,"numerous"occurrences"of"unnecessary" use"have"been"observed."That"is,"literals"were"used"to"denote"property"values"for"which"URIs"were" available." Strong( Dependency( on( descriptive( values"tends" to" occur" frequently" throughout" all" data" sets." Often," these" descriptions" provide" crucial" information" and" thus" cannot" be" ignored" without" a" significant"loss"of"knowledge."Therefore,"such"values"should"be"given"additional"thought"during" the"development"of"a"DM"solution." Concurrent(RDF(statements"were"observed"within"one"data"set"in"which"both"the"correct"and"the" corresponding" falsified" statements" on" the" same" concept" were" kept" available." Due" to" this" construction,"duplicate"and"possibly"conflicting"entries"might"surface."Hence,"methods"for"conflict" resolution"should"be"explored."
ARIADNE"D16.1"Public" 45" " 7 Data&Mining&on&Linked&Archaeological&Data& As"discussed"earlier,"a"typical"KDD"process"starts"off"with"a"deep"understanding"of"the"domain"and"the" data."At"this"moment,"expertise"within"the"former"is"readily"available"within"ARIADNE."Unfortunately," the"same"cannot"be"said"about"the"data"aspect."This"is"understandable,"as"ARIADNE"has"reached"half"of" its" fourIyear" run." However," as" the" exploration" of" data" is" an" important" step" in" any" KDD" process," this" means"it"is"difficult"to"predict"what"kind"of"new"knowledge"may"be"brought"to"light"as"a"result"of"this" process."Therefore,"this"section"will"focus"on"the"moreIgeneric"options"expected"to"function"properly"on" any"form"of"Linked"Archaeological"Data"(LAD)." Based"on"both"the"Domain"and"Data"Understanding"step"within"our"selected"KDD"process,"several"DM" solutions"were"selected"which"we"deem"suitable"for"a"LAD"framework"such"as"ARIADNE"will"likely"be."We" will"next"discuss"these"solutions"in"more"depth." 7.1 Hypothesis&Generation& DM"methods"are"capable"of"detecting"patterns"in"data."Interesting"and"potentially"relevant"subsets"of" these"patterns"can"then"be"presented"to"users"as"starting"points"for"forming"new"research"hypotheses." For"instance,"the"system"might"detect"that"specific"types"of"pottery"are"most"often"found"near"coastal" areas." This"might" already" be" known" to" the"researcher," or" it" might" be" something" the" researcher" is" interested" in"exploring"further."The"interestingness"of"patterns"will"be"derived"algorithmically"on"the" basis"of"predefined"criteria"and"user"feedback."To"facilitate"this,"any"LAD"repository"should"ideally"offer"a" KDD"interface"that"provides"its"users"with"such"capabilities."These"capabilities"would"then"be"applied" directly"to"the"data"from"one"or"more"repositories." The"integration"of"KDD"capabilities"into"a"structured"query"interface"such"as"SPARQL"or"similar"would" provide"a"solution"that"largely"satisfies"the"earlierImentioned"criteria."As"a"result,"any"query"can"easily" be" extended" with" dataImining" operations" capable" of" generating" potential" hypotheses."Moreover," as" using" these" additional" operations" is" purely" optional," their" presence" would" not" hinder" users" who" are" solely"interested"in"regular"queries."Furthermore,"it"might"prove"useful"to"assist"with"the"formulation"of" queries"for" those"users"who" are" unfamiliar" with" the"syntax"of" a" structured" query" language"such" as" SPARQL." Alternatively," the" capabilities" or" several" higherIlevel" abstractions" could" be" integrated" into" a" graphical"UI,"thus" lowering" their"learning"curve."Moreover,"multiple"hypotheses"could"then"easily"be" presented"to"users"in"nonIintrusive"ways"to"allow"for"quick"scanning"for"potentially"valuable"directions." " "
ARIADNE"D16.1"Public" 46" " 7.2 Assisted&Query&Formulation& The"complexity"of"accessing"DM"capabilities"through"a"query"language"might"pose"a"hurdle"to"interested" users,"i.e."those"who"want"to"use"these"capabilities"but"who"are"unfamiliar"with"their"syntax."To"lower" this"barrier,"it"might"prove"useful"to"assist"the"user"in"its"formulation"of"queries."Instead"of"limiting"this" assistance"to"solely"the"DM"capabilities,"it"may"be"extended"to"aid"with"formulating"regular"queries"as" well." This" may" manifest" itself" as" either" predicting" or" autocompleting" the" query." Alternatively," a" combination"of"the"two"may"be"applied." Predicting"a"query"involves"learning"from"past"queries."Based"on"a"partiallyIwritten"query,"the"remainder" is" predicted" by" comparing" the" similarity" between" the" written" part" and" (parts" of)"past" queries." A" distinction"herein"is"whether"a"more"local"or"more"global"view"should"be"maintained."Here,"a"more"local" view"would"concern"queries"from"users"who"share"a"similar"background"or"interest,"as"well"as"users"who" have"accessed"much"of"the"same"data."This"however,"would"require"user"profiles."In"contrast,"a"global" view"would"consider"the"queries"of"all"users."A"combination"is"possible"as"well,"thereby"favouring"more" local"queries"over"global"ones."" Whereas"the"previously"discussed"approach"is"queryIdriven,"a"dataIdriven"approach"can"be"used"as"well." This" would" entail" learning" from" the" available" data," thereby" determining" frequently"occurring" combinations"of"relations"which"may"be"offered"as"suggestions."This"could"additionally"be"coupled"with" relevant"ontologies"and"knowledge"of"the"query"language"and"its"DM"extension."Such"a"coupling"would" prevent"offering"invalid"suggestions"as"well."Hence,"it"could"be"considered"as"a"form"of"autocompletion." 7.3 Ranking&of&Query&Results& Integrating"KDD"capabilities"into"a"query"interface"allows"the"ability"to"rank"the"results"of"such"a"query." This"ability"concerns"the"ordering"of"results"based"upon"certain"criteria."Within"a"LAD,"a"strongly"desired" criterion"is"that"of"relevancy."This"criterion"is"commonly"regarded"as"the"most"challenging"to"determine," as"it"is"intertwined"with"the"researchers"flow"of"thought"(Franz,"et"al."2009)."However,"when"using"a" query" language" capable" of" representing" structure" and" semantics," e.g." SPARQL," determining" this" relevancy"becomes"more"manageable." While"SPARQL"or"a"similar"query"language"would"always"form"the"bridge"between"a"LAD’s"frontend"and" backend,"it"might"not"necessarily"be"the"interface"that"is"used"to"search"by"the"researchers."That"is,"the" frontend"might"facilitate"faceted"searching,"thereby"effectively"posing"as"a"wrapper"to"the"underlying" query" language" with" its" DM" extension." In" addition," it" may" provide" a" keywordIbased" search" as" well," thereby"requiring" a" NLP" solution" which" translates" the" query" to" a" SPARQLIbased" language" or" similar." While"such"a"translation"would"likely"result"in"some"loss"in"precision,"it"does"allow"for"a"more"userI friendly" search." Moreover," it" opens" up" the" possibility" of" incorporating" VSM," thereby" enabling" the" traditional" ranking" of" documents" based" on" their" relevance" to" the" provided" keywords." Note" that" the"
ARIADNE"D16.1"Public" 53" " Strong(Dependency(on(descriptive(values" All"studied"data"sets"appear"to"depend"strongly"on"descriptions."In"most"cases,"these"descriptions" provide"crucial"information"and"thus"cannot"be"omitted"without"a"significant"loss"of"knowledge." However,"such"descriptions"are"often"represented"as"a"single"literal,"thus"being"little"more"than" unstructured"text."" Concurrent(RDF(statements" Data"sets"were"observed"in"which"both"the"correct"and"incorrect"statements"for"the"same"concept" were"kept"available."That"is,"instead"of"updating"falsified"statements,"an"erratum"was"supplied." Due"to"this"construction,"duplicate"and"possibly"conflicting"entries"might"surface."" 8.3 Recommendations& Generally,"the"developer"of"a"typical"Data"Mining"solution"is"supplied"with"a"generous"amount"of"data" from"which"the"exploration"might"reveal"potentially"relevant"patterns."After"careful"inspection"of"the" data"currently"available," relevant" patterns" can" likely"be" generalized" to" the" entirety" of" the" data." Unfortunately,"the"minimal"amount"of"data"current"available"through"ARIADNE"prevents"such"sequence" of"events"to"take"place."Therefore,"it"would"be"rather"unlikely"to"successfully"generalize"any"discovered" pattern"to"the"large"amount"of"data"that,"one"day,"will"be"accessible"through"ARIADNE."Instead,"a"moreI generic"approach"is"suggested,"such"that"its"workings"are"ensured"regardless"of"the"exact"characteristics" of"the"future"data." Based" on" the" study" of" both" the"domain" and" data"generated" by" the" domain," as" well" as" on" practical" constraints" with" respect" to" time" and" resources,"two" data" mining" solutions" were" chosen" which" were" deem"the"mostIfeasible"and"suitable"for"implementation"within"the"ARIADNE"framework."These"are:" Hypothesis(Generation((7.1)" The"official"project"proposal"of"ARIADNE"mentions"the"ability"to"detect"patterns"in"archaeological" data"or"related"data"and"applications"within"the"ARIADNE"infrastructure."DataImining"methods" are" capable" of" detecting" such" patterns." Interesting" and" potentially"relevant" subsets" of" these" patterns"can"then"be"presented"to"users"as"starting"points"for"forming"new"research"hypotheses." This"may"already"be"known"to"the"researcher,"or"it"might"be"something"they"are"interested"in" exploring"further."The"interestingness"of"patterns"will"be"determined"algorithmically"on"the"basis" of"predefined"criteria"and"user"feedback."To"facilitate"this,"a"user"interface"should"provide"access" to" a" data"mining" backend." Initially," this" might" best" be" integrated" within" a" textIbased" query" interface"such"as"SPARQL"or"similar."At"a"later"stage,"a"wrapper"for"the"graphical"user"interface" should"be" made." This" will" allow" multiple" hypotheses" to" be" presented" in" nonIintrusive" ways," thereby"allowing"users"to"quickly"scan"in"potentially"valuable"directions." " "
ARIADNE"D16.1"Public" 54" " Data(Quality(Analysis((7.5)( Two"aspects"that"reflect"poorly"on"the"quality"of"data"are"the"occurrence"of"gaps"and"errors"in"the" knowledge" contained" therein."In" case" of" the" former," filling" these" voids" involves" predicting" the" mostIlikely" resource," link," or" literal."In" the" latter" case," these" errors" typically" include"anomalies" within"the"data"which"cannot"be"explained"by"any"of"the"discovered"patterns"alone."Depending"on" the"likelihood"of"them"being"erroneous,"the"detected"errors"could"be"suggested"for"removal"or" tagged"as"dubious."Alternatively,"they"could"be"replaced"by"a"prediction"of"the"correct"value."This" could"reuse"the"earlierImentioned"prediction"method"and"data"mining"backend." 8.4 Roadmap& The" sequel" to" this" report," i.e." Deliverable" 16.3," will" present" the" final" results" of" the" applicability" and" feasibility" of" data" mining" within" the" ARIADNE" framework." To" this" end," the" aforementioned" recommendations"will"be"explored"and"experimented"with"further."This"process"will"consist"of"several" phases." Following"this"report,"a"more"extensive"study"into"the"recommended"topics"will"first"be"performed."This" will"assist"us"in"narrowing"down"the"list"of"possible"options"to"only"those"that"we"believe"possess"the" most" potential." The" remaining" options" will" subsequently" be" implemented" into" our" experimental" environment"on"site"where"they"will"thoroughly"be"tested"on"various"linked"archaeological"data."The" result" of" these" tests" will" determine" whether" the" selected" options" are" suitable" for" integration" within" ARIADNE."Based"on"the"experience"gained"during"the"current"study,"we"expect"that"most"of"the"possible" options"will"need"to"be"adapted"to"suit"both"the"data"and"the"user’s"needs."If,"for"some"reason,"none"of" these"options"are"found"to"be"suitable,"a"custom"solution"will"be"developed"instead." Once"the"selected"options"have"been"successfully"implemented"within"the"experimental"environment," internal"evaluation"rounds"will"be"organized"during"which" domain"enthusiasts"and" experts"of" varying" levels"of"expertise"will"be"asked"to"experiment"with"the"implementation."Here,"we"expect"the"groups"to" be"comprised"of"students,"junior"and"senior"archaeological"researchers,"and"local"data"and"repository" managers." This" will" additionally" provide" the" input" needed" for" the" development" of" (elements" of)" a" graphical"user"interface."Similarly"as"before,"an"iterative"scheme"will"be"followed."" The" final" phase" will" consist" of" potentially" implementing" the" data"mining" solutions" into" the" ARIADNE" infrastructure." This" would" be"followed" by" extensively" experimenting" on"the"various" data" accessible" through"ARIADNE."This"implementation"will"be"improved"upon"further"during"a"series"of"iterative"and" open"evaluation"sessions"for"the"remainder"of"the"ARIADNE"project."
ARIADNE"D16.1"Public" 55" " Bibliography& Aloia," N," C" Meghini," D" Gavrilis," and" C" Papatheodorou." ARIADNE& Catalogue& Data& Model."Deliverable," ARIADNE,"2014." Amin,"A,"J"Van"Ossenbruggen,"L"Hardman,"and"A"van"Nispen.""Understanding"cultural"heritage"experts'" information" seeking" needs."" The& 8th& ACM/IEEEDCS& joint& conference& on& Digital& libraries."ACM," 2008."39I47." Anyanwu,"K," A" Maduko," and" A" Sheth.""SemRank:"ranking"complex" relationship"search" results" on" the" semantic"web.""14th&international&conference&on&World&Wide&Web&."ACM,"2005."117I127." Artz," D," and" Y" Gil." "A" survey" of" trust" in" computer" science" and" the" semantic" web."" Web& Semantics:& Science,&Services&and&Agents&on&the&World&Wide&Web,"2007:"58I71." Balmin," A," V" Hristidis," and" Y" Papakonstantin." "Objectrank:" authorityIbased" keyword" search" in" databases.""VLDB,"2004:"564I575." Baxter,"M."Statistics&in&archaeology."London:"Arnold,"2003." Berendt," B," A" Hotho," and" G" Stumme." "Towards" semantic" web" mining."" The& Semantic& Web—ISWC." Springer"Berlin"Heidelberg,"2002."264I278." Berendt,"B.,"Hotho,"A,"D"Mladenic,"M"Van"Someren,"M"Spiliopoulou,"and"G"Stumme.""A"roadmap"for" web"mining.""Web&to&semantic&web,"2004:"1I22." Bi,"S,"S"Xue,"Y"Xu,"Pei,"and"A.""Spatial"Data"Mining"in"Settlement"Archaeological"Databases"Based"on" Vector"Features.""Fuzzy&Systems&and&Knowledge&Discovery."Jinan"Shandong:"IEEE,"2008."277I281." Bicer," V," T" Tran," and" A" Gossen." "Relational" kernel" machines" for" learning" from" graphIstructured" RDF" data.""The&Semantic&Web:&Research&and&Applications,"2011:"47I62." Bizer,"C,"and"R"Oldakowski.""Using"ContextI"and"ContentIBased"Trust"Policies"on"the"Semantic"Web."" 13th&International&World&Wide&Web&Conference."New"York,"NY:"ACM"Press,"2004."228I229." Bizer,"C,"T"Heath,"and"T"BernersILee.""Linked"dataIthe"story"so"far.""International&journal&on&semantic& web&and&information&system"3,"no."5"(2009)." Bloehdorn,"S,"and"Y"Sure.""Kernel"methods"for"mining"instance"data"in"ontologies.""2007:"58I71." Borgwardt,"K"M,"N"N"Schraudolph,"and"S"Vishwanathan.""Fast"computation"of"graph"kernels.""Advances& in&neural&information&processing&systems,"2006:"1449I1456."
ARIADNE"D16.1"Public" 56" " Bray,"T,"J"Paoli,"C"M"SperbergIMcQueen,"E"Maler,"and"F"Yergeau.""Extensible"Markup"Language"(XML)"1."" W3C."November"26,"2008."www.w3.org/TR/xml." Brickley,"D.""Basic"Geo"Vocabulary.""W3C."February"1,"2006."www.w3.org/2003/01/geo." Briscoe,"T.""Lexical"issues"in"natural"language"processing.""Natural&language&and&speech,"1991:"39I68." Buneman,"P.""Semistructured"data.""The&sixteenth&ACM&SIGACTDSIGMODDSIGART&Principles&of&database& systems."ACM,"1997."117I121." Campinas,"S,"T"E"Perry,"D"Ceccarelli,"R"Delbru,"and"G"Tummarello.""Introducing"rdf"graph"summary"with" application" to" assisted" sparql" formulation."" 23rd& International& Workshop& on& Database& and& Expert&Systems&Applications."2012." Castells,"P,"M"Fernandez,"and"D"Vallet.""An"adaptation"of"the"vectorIspace"model"for"ontologyIbased" information"retrieval.""Knowledge&and&Data&Engineering."IEEE"Transactions,"2007."261I272." Charno,"M,"S"Jeffrey,"C"Binding,"D"Tudhope,"and"K"May.""From"the"Slope"of"Enlightenment"to"the"Plateau" of" Productivity:" Developing" Linked" Data" at" the" ADS."" 40th& Annual& Conference& of& Computer& Applications& and& Quantitative& Methods& in& Archaeology."Southampton:" Amsterdam" University" Press,"2012."216I223." Chen,"M"S,"J"Han,"and"P"S"Yu.""Data"mining:"an"overview"from"a"database"perspective.""Knowledge&and& data&Engineering,"1996:"866I883." Cleave,"J"P."A&study&of&logics."Oxford"University"Press,"1991." Codd,"E"F.""A"relational"model"of"data"for"large"shared"data"banks.".""Communications&of&the&ACM,"1970:" 377I387." Codd,"E"F,"S"B"Codd,"and"C"T"Salley."Providing&OLAP&(onDline&analytical&processing)&to&userDanalysts:&An&IT& mandate."Codd"and"Date,"1993." Cohen,"W"W,"and"Z"Kou.""Stacked"graphical"learning:"approximating"learning"in"markov"random"fields" using"very"short"inhomogeneous"markov"chains.""Technical"report,"2006." Conolly," J," and" M" Lake." Geographical& Information& Systems& in& Archaeology."Cambridge:" Cambridge" University"Press,"2006." Cripps,"P,"et"al."CRMarchaeo:&the&Excavation&Model."CIDOC"CRM,"2014." Cyganiak," R," and" D" Reynolds." "The" RDF" Data" Cube" Vocabulary."" W3C."January" 2014," 2014." www.w3.org/TR/vocabIdataIcube." d'Amato,"C,"N"Fanizzi,"and"F"Esposito.""Inductive"learning"for"the"Semantic"Web:"What"does"it"buy?"" Semantic&Web,"2010:"53I59."
ARIADNE"D16.1"Public" 57" " De" Kleijn," M," N" van" Manen," J" Kolen," and" H" Scholten." "Towards" a" UserIcentric" SDI" Framework" for" Historical" and" Heritage" European" Landscape" Research."" International& Journal& of& Spatial& Data& Infrastructures&Research,"2014:"1I35." Di" Ludovico," A," and" G" Pieri." "Artificial" Neural" Networks" and" ancient" artefacts:" Justifications" for" a" multiform"integrated"approach"using"PST"and"AutoICM"models.""Archeologia&e&calcolatori,"2011:" 91I128." Dimitropoulos," H," et" al." "AITION:" a" scalable" platform"for" interactive" data" mining."" Scientific& and& Statistical&Database&Management,"2012:"646I651." Doerr,"M,"and"K"Schaller.""The"Dream"of"a"Global"Knowledge"Network"I"A"new"Approach.""ACM&Journal& on&Computers&and&Cultural&Heritage,"2008." Doerr," M," K" Schaller," and" M" Theodoridou." "Itegration" of" complementary" archaeological" sources."" Computer&Applications&and&Quantitative&Methods&in&Archaeology."Prato,"Itally","2004." Dubin,"D.""The"most"influential"paper"Gerard"Salton"never"wrote.""Library&Trends,"2004:"748I764." Earl," G," T." T" Sly," and" D" D" Wheatley." "Archaeology" in" the" Digital" Era."" Computer& Applications& and& Quantitative&Methods&in&Archaeology."Southampton:"Amsterdam"University"Press,"2014." Etcheverry,"L,"and"A"A"Vaisman.""Enhancing"OLAP"analysis"with"web"cubes.""Semantic&Web:&Research& and&Applications,"2012:"469I483." Etcheverry,"L,"and"A"A"Vaisman."QB4OLAP:&A&New& Vocabulary&for&OLAP&Cubes&on&the&Semantic&Web." R1210LAC004,"2012." Etter,"D,"and"C"Domeniconi.""SemRank:"Semantic"Rank"Learning"for"Multimedia"Retrieval.""2014." Fanizzi,"N,"C"d’Amato,"and"F"Esposito.""Conceptual"clustering"and"its"application"to"concept"drift"and" novelty"detection.""Munich:"Springer"Berlin"Heidelberg,"2008." Fayyad,"U"M.""Data"mining"and"knowledge"discovery:"Making"sense"out"of"data.""IEEE&Intelligent&Systems" 11,"no."5"(1996):"20I25." Fayyad," Usama," Gregory" PiatetskyIshapiro," and" Padhraic" Smyth." "From" Data" Mining" to" Knowledge" Discovery"in"Databases.""AI&Magazine"17"(1996):"37I54." Fentress,"E."Register&of&Online&Archaeological&Databases."Deliverable,"ARIADNE,"2014." Fisher,"D"H.""Knowledge"acquisition"via"incremental"conceptual"clustering.""Machine&learning"2,"no."2" (1987)." Franz," T," A" Schultz," S" Sizov," and" S" Staab." "Triplerank:" Ranking" semantic" web" data" by" tensor" decomposition.""2009:"213I228."
ARIADNE"D16.1"Public" 58" " Freitas,"A,"E"Curry,"J"G"Oliveira,"and"S"O'Riain.""Querying"heterogeneous"datasets"on"the"linked"data"web:" Challenges,"approaches,"and"trends.""Internet&Computing,&IEEE,"2012:"24I33." Friedman,"J"H.""Data"Mining"and"Statistics:"What's"the"connection?""Computing&Science&and&Statistics"29," no."1"(1998):"3I9." Garshol," L" M." "" Metadata?" Thesauri?" Taxonomies?" Topic" maps!" Making" sense" of" it" all."" Journal& of& information&science"4,"no."30"(2004):"378I391." Gärtner,"T.""A"survey"of"kernels"for"structured"data.""ACM&SIGKDD&Explorations&Newsletter,"2003,"5"ed.:" 49I58." Getoor,"L,"and"B"Taskar."Introduction&to&statistical&relational&learning."MIT"press,"2007." Golbeck,"J,"B"Parsia,"and"J"Hendler.""Trust"networks"on"the"semantic"web.""Journal&of&Web&Semantics," 2003:"238I249." Gombos,"G,"and"A"Kiss.""SPARQL"query"writing"with"recommendations"based"on"datasets.""Information& and&Knowledge&Design&and&Evaluation,"2014:"310I319." Gruber," E," G" Bransbourg," S" Heath," and" A" Meadows." "Linking" Roman" Coins:" Current" Work" at" the" American" Numismatic" Society."" 40th& Annual& Conference& of& Computer& Applications& and& Quantitative& Methods& in& Archaeology."Southampton:"Amsterdam"University"Press,"2012."249I 258." Gupta,"M,"Y"Sun,"and"J"Han.""Trust"analysis"with"clustering.""20th&international&conference&companion&on& World&Wide&Web&."ACM,"2011."53I54." Hagood,"J.""A"brief"introduction"to"data"mining"projects"in"the"humanities.""Bulletin& of& the& American& Society&for&Information&Science&and&Technology"38,"no."4"(2012):"20I23." Hastie,"T,"R"Tibshirani,"J"Friedman,"T"Hastie,"J"Friedman,"and"R"Tibshirani."The& elements& of& statistical& learning."New"York:"Springer,"2009." He," X," and" M" Baker." "xhRank:" Ranking" Entities" for" Semantic" Web" Searching."" Fifth& International& Conference&on&Advances&in&Semantic&Processing."2011."62I68." Heath," T," and" Christian" Bizer." Linked& Data:& Evolving& the& Web& into& a& Global& Data& Space."Morgan" &" Claypool"Publishers,"2011." Hogan,"A,"A"Harth,"and"S"Decker.""Reconrank:"A"scalable"ranking"method"for"semantic"web"data"with" context.""2006." Hollander,"Hella,"and"Maarten"Hoogerwerf."D13.1:&Service&Design."ARIADNE,"2014."
ARIADNE"D16.1"Public" 59" " Huang," B," A" Kimmig," L" Getoor," and" J" Golbeck." "Probabilistic" soft" logic" for" trust" analysis" in" social" networks.""International&Workshop&on&Statistical&Relational&AI."2012."1I8." Huang," Y," and" V" Tresp." "Accessing" the" Semantic" Web" via" Statistical" Machine" Learning."" ESWC& 2012& Tutorial."May" 2012," 2012." http://www.dbs.ifi.lmu.de/~huang/eswc2012tutorial/ESWC2012I TutorialV10.pdf." Huang,"Y,"and"V"Tresp."Relation&Prediction&in&Semantic&Domains&using&Multivariate&Prediction."Munich," Germany:"Siemens"AG,"2010." Huang,"Y,"V"Tresp,"H"Kriegel,"and"P.""Multivariate"prediction"for"learning"in"relational"graphs.""Workshop:& Analyzing&Networks&and&Learning&With&Graphs."2009." Huang,"Y,"V"Tresp,"M"Bundschus,"A"Rettinger,"and"H"P"Kriegel.""Multivariate"prediction"for"learning"on" the"semantic"web.""Inductive&Logic&Programming,"2011:"92I104." Huber," R," H" Ramoser," K" Mayer," H" Penz," and" M" Rubik." "Classification" of" coins" using" an" eigenspace" approach.""Pattern&Recognition&Letters,"2005:"61I75." Huggett," J." "Disciplinary" Issues:" Challenging" the" Research" and" Practice" of" Computer" Applications" in" Archaeology.""40th&Annual&Conference&of&Computer&Applications&and&Quantitative&Methods&in& Archaeology."Southampton:"Amsterdam"University"Press,"2012."14I24." Isaksen," L," G" Earl," K" Martinez," S" Keay," and" N" Gibbins." "Linking" archaeological" data."" International& Conference&on&Computer&Applications&and&Quantitative&Methods&in&Archaeology."2009." Isaksen,"L,"K"Martinez,"N"Gibbins,"G"Earl,"and"S"Keay.""Linking"archaeological"data.""CAA,"2009." Isaksen," L," K" Martinez," N" Gibbins," Graeme" Earl," and" S" Keay." "Interoperate" with" whom?" Fomality," Archaeology"and"the"Semantic"Web.""WebScience."Raleigh,"NC,"2010." Kantardzic,"M."Data&mining:&concepts,&models,&methods,&and&algorithms."John"Wiley"&"Sons,"2011." Karasik," A," I" Sharon," U" Smilansky," and" A" Gilboa." "Typology" and" classification" of" ceramics" based" on" curvature" analysis."" Computer& Applications& and& Quantitative& Methods& in& Archaeology."2004." 472I475." Kellar," M," C" Watters," and" K" M" Inkpen." "An" exploration" of" webIbased" monitoring:" implications" for" design.""The&SIGCHI&conference&on&Human&factors&in&computing&systems."ACM,"2007."377I386." Khan," M" A," G" A" Grimnes," and" A" Dengel." "Two" preIprocess" operators" for" improved" learning" from" semanticweb"data.""First&Rapidminer&Community&Meeting&and&Conference."2010." Kiefer,"C,"A"Bernstein,"and"A"Locher.""Adding"data"mining"support"to"SPARQL"via"statistical"relational" learning"methods.""The&Semantic&Web:&Research&and&Applications,"2008:"478I492."
ARIADNE"D16.1"Public" 60" " Kintigh,"K.""Quantitative"methods" designed"for" archaeological"problems.""In"Quantitative& Research& in& Archaeology:& Progress& and& Prospects," by" M" S" Aldenderfer," 135I150." Newbury" Park," NJ:" Sage," 1987." Knobbe,"A"J."MultiDrelational&data&mining."Ios"Press,"2006." Kobylinski," L," and" K" Walczak." "Data" mining" approach" to" classification" of" archaeological" aerial" photographs.""Intelligent&Information&Processing&and&Web&Mining,"2006:"479I487." Kolda,"T"G,"and"B"W"Bader.""Tensor"decompositions"and"applications.""SIAM&review"51,"no."3"(2009):" 455I500." Koller,"D,"and"N"Friedman."Probabilistic&graphical&models:&principles&and&techniques."MIT"press,"2009." Kramer," K," R" Q" Dividino," and" G" Gröner." "SPACE:" SPARQL" Index" for" Efficient" Autocompletion."" International&Semantic&Web&Conference."2013."157I160." Krogel," M" A," S" Rawles," F" Železný," P" A" Flach," N" Lavrač," and" S" Wrobel." "Comparative" evaluation" of" approaches" to" propositionalization."" 13th& International& Conference& on& ILP."Szeged:" Springer" Berlin"Heidelberg,"2003."197I214." Kurgan,"L"A,"and"P"Musilek.""A"survey"of"Knowledge"Discovery"and"Data"Mining"process"models.""The& Knowledge&Engineering&Review,"2006:"1I24." Lavrač,"N,"A"Vavpetič,"L"Soldatova,"I"Trajkovski,"and"P"K"Novak.""Using"ontologies"in"semantic"data"mining" with"segs"and"gIsegs.""Discovery&Science,"2011:"165I178." Lavrac,"N,"and"S"Dzeroski."Relational&Data&Mining."Springer,"2001." Linderholm," J," and" P" Geladi." "Classification" of" archaeological" soil" and" sediment" samples" using" near" infrared"techniques.""NIR&news,"2012:"6." Locher," A." SPARQLDML:& Knowledge& Discovery& for& the& Semantic& Web."Thesis" University" of" Zurich," University"of"Zurich,"2007." Maali," F," J" Erickson," and" P" Archer." "Data" Catalogue" Vocabulary" (DCAT)."" W3C."January" 16," 2014." www.w3.org/TR/vocabIdcat." Maedche,"A,"and"V"Zacharias.""Clustering"ontologyIbased"metadata"in"the"semantic"web."In.""Principles& of&Data&Mining&and&Knowledge&Discovery."Springer"Berlin"Heidelberg,"2002."348I360." Maimon,"O"Z,"and"L"Rokach."Data&mining&and&knowledge&discovery&handbook."New"York:"Springer,"2005." Mani," Inderjeet," and" Mark" T" Maybury." Advances& in& Automatic& Text& Summarization."Cambridge:" MIT" Press,"1999."
ARIADNE"D16.1"Public" 61" " Martinez,"K,"and"L"Isaksen.""The"Semantic"Web"Approach"to"Increasing"Access"to"Cultural"Heritage."" Revisualizing&Visual&Culture."Varnham,"2010."29I44." May," K." Hypermedia& Research& Unit& D& CIDOC& CRMDEH& Ontology."n.d." http://hypermedia.research.southwales.ac.uk/kos/CRM/"(accessed"November"02,"2014)." Mendes,"P"N,"M"Jakob,"A"GarcíaISilva,"and"C"Bizer.""DBpedia"spotlight:"shedding"light"on"the"web" of" documents.""7th&International&Conference&on&Semantic&Systems."ACM,"2011."1I8." Metaxas,"O,"H"Dimitropoulos,"Y"Ioannidis,"and"M"Paedigree.""AITION:"A"scalable"KDD"platform"for"Big" Data"Healthcare.""Biomedical&and&Health&Informatics."IEEE,"2014."601I604." Middleton,"S"E,"N"R"Shadbolt,"and"D"C"De"Roure.""Ontological"user"profiling"in"recommender"systems."" ACM&Transactions&on&Information&Systems&(TOIS)&22.1,"2004:"54I88." Narasimha,"V,"P"Kappara,"R"Ichise,"and"O"P"Vyas.""LiDDM:"A"Data"Mining"System"for"Linked"Data.""CEUR& Workshop&Proceedings:&Linked&Data&on&the&Web."2011." Nickel,"M,"V"Tresp,"and"H"Kriegel.""A"threeIway"model"for"collective"learning"on"multiIrelational"data."" 28th&international&conference&on&machine&learning."2011."809I816." Nolle,"M,"H"Penz,"M"Rubik,"K"Mayer,"I"Hollander,"and"R"Granec.""DagobertIa"new"coin"recognition"and" sorting" system."" International& Conference& on& Digital& Image& Computing,& Techniques,& and& Applications."Sydney:"CSIRO"Publishing,"2003."329I338." Novak,"P"K,"A"Vavpetic,"I"Trajkovski,"and"N"Lavrac.""Towards"semantic"data"mining"with"gIsegs.""11th& International&Multiconference&Information&Society."2009." OGC& GeoSPARQL& D& A& Geographic& Query& Language& for& RDF& Data."Specification," Open" Geospatial" Consortium,"2012." Padawitz,"P."Computing&in&Horn&clause&theories."Springer"Publishing"Company,"2012." Parsaye,"K.""Surveying"Decision"Support:"New"Realms"of"Analysis.""Database&Programming&and&Design," 1996:"26I33." Paulheim,"H,"and"J"Fümkranz.""Unsupervised"generation"of"data"mining"features"from"linked"open"data."" International&conference&on&web&intelligence,&mining&and&semantics."ACM,"2012."31." Potoniec,"J,"and"A"Lawrynowicz.""RMonto:"ontological"extension"to"RapidMiner.""ISWC."2011."1I4." —.""RMontoItowards"KDD"workflows"for"ontologyIbased"data"mining.""eCML&PKDD."2011b."11." Prud'hommeaux," E," and" A" Seaborne." "SPARQL" Query" Language" for" RDF."" W3C."January" 18," 2008." www.w3.org/TR/rdfIsparqlIquery."
ARIADNE"D16.1"Public" 62" " Pu," L," and" B" Faltings." "Understanding" and" improving" relational" matrix" factorization" in" recommender" systems.""7th&ACM&conference&on&Recommender&systems."ACM,"2013."41I48." ragimov,"D,"K"Hose,"T"B"Pedersen,"and"E"Zimányi.""Towards"Exploratory"OLAP"over"Linked"Open"Data–A" Case"Study.""2014:"18." Ramezani,"R,"M"Saraee,"and"M"A"Nematbakhsh.""Finding"association"rules"in"linked"data,"a"centralization" approach.""Iranian&Conference&on&Electrical&Engineering."IEEE,"2013."1I6." Rapidminer& SemWeb."n.d." https://code.google.com/p/rapidminerIsemweb" (accessed" September" 12," 2014)." Rettinger,"A,"U"Lösch,"V"Tresp,"C"d'Amato,"and"N"Fanizzi.""Mining"the"semantic"web.""Data&Mining&and& Knowledge&Discovery,"2012:"613I662." Richards," J" D." "Archaeology," eIpublication" and" the" semantic" web."" ANTIQUITYDOXFORD"80," no." 310" (2006):"970I979." Ristoski,"P,"and"H"Paulheim.""A"Comparison"of"Propositionalization"Strategies"for"Creating"Features"from" Linked"Open"Data.""Linked&Data&for&Knowledge&Discovery,"2014:"6I17." Ristoski," P," C" Bizer," and" H" Paulheim." "Mining" the" web" of" linked" data" with" rapidminer."" International& Semantic&Web&Conference."2014." Rocha,"C,"D"Schwabe,"and"M"P"Aragao.""A"hybrid"approach"for"searching"in"the"semantic"web.""13th& Iinternational&conference&on&World&Wide&Web."ACM,"2004."374I383." Salton,"G,"A"Wong,"and"C"S"Yang.""A"vector"space"model"for"automatic"indexing.""Communications&of&the& ACM,"1975:"613I620." Schölkopf," B," and" A" J" Smola." Learning& with& kernels:& Support& vector& machines,& regularization,& optimization,&and&beyond."MIT"press,"2002." Selhofer,"Hannes,"and"Guntram"Geser."D2.1:&First&Report&on&Users'&Needs."Salzburg:"ARIADNE,"2014." Sen,"P,"G"Namata,"M"Bilgic,"L"Getoor,"B"Galligher,"and"T"EliassiIRad.""Collective"classification"in"network" data.""AI&magazine"29,"no."3"(2008)." Shadbolt,"N,"W"Hall,"and"T"BernersILee.""The"semantic"web"revisited.""Intelligent&Systems"(IEEE)"3,"no."21" (2006)." Signore," O." "Representing" knowledge" in" archaeology:" from" cataloguing" cards" to" semantic" web."" Archeologia&e&calcolatori,"no."20"(2009):"111I128." Singla,"P,"and"P"Domingos.""Entity"resolution"with"markov"logic.""Sixth&International&Conference&on&Data& Mining."IEEE,"2006."572I582."
ARIADNE"D16.1"Public" v" " Appendix&C Learning& Methods& for& Semantic& Web& Mining& SemanticIWeb"Mining"offers"a"wide"range"of"potential"methods,"most"of"which"are"rather"experimental." Of" these" methods"(Knobbe" 2006," Rettinger," et" al." 2012," Berendt," et" al." 2004)," two" are" featured" prominently" in" recent" literature," namely"Propositionalization"and"Statistical"Relational" Learning." Of" these,"the"core"concepts"will"be"discussed"next."In"addition,"we"will"briefly"look"at"the"promising"and" rather"new"Kernel"Methods,"which"have"recently"become"quite"popular"within"ML"communities." C.1 Propositional&Learning& Recall"that"propositional"data"are"assumed"to"be"independent"and"identically"distributed,"hence"allowing" for" statistical" machine" learning" algorithms" to" be" applied." Given" that" these" assumptions" do" not" necessarily" hold" for" LD," applying"them" anyway" would" likely" result" in" false" conclusions." Instead" of" abandoning" these" proven" methods" however," an" alternative" would" be" to" temporarily" convert" LD" to" propositional"data;"a"process"known"as"propositionalization"(Ristoski"and"Paulheim"2014,"Tresp,"et"al." 2008)." A"wellIknown"method"during"propositionalization"is"factorization,"which"involves"the"decomposition"of" an"object"into"a"product"of"smaller"objects;"its"factors."When"put"together"again,"these"factors"return"(an" estimation"of)"the"original"object."Within"the"context"of"LD,"it"entails"the"translation"of"graph"data"to" propositional"data,"and"vice"versa."While"in"this"proposition"format"statistical"ML"can"be"applied." C.1.1 Relational&Matrices& A"matrix"is"a"twoIdimensional"array"arranged"in"𝑛"rows"and"𝑚"columns."Each"point"(𝑖,𝑗)"with"𝑖∈𝑛"and" 𝑗∈𝑚"is"an"element"of"the"matrix."When"two"or"more"matrices"are"needed"to"describe"a"single"data"set," these" are" often"denoted"as"relational" matrices." Note" that"this" strongly" resembles"the" relational"data" model"as"is"used"in"relational"database." Translating"LD"into"relational"matrices"(Figure"CI1)" is" a"fairly"straightforward" procedure"(Tresp," et" al." 2008,"Pu"and"Faltings"2013,"Rettinger,"et"al."2012)."Each"matrix"represents"a"single"predicate"in"the"data" set." Hence," the" number" of" required" matrices" equals" the"number" of" predicates." Next," of" each" triple" containing" a" certain" predicate," its" subject"and" object"are" placed" on" row" 𝑖"and" column" 𝑗"of" the" corresponding"matrix,"respectively."Within"that"matrix,"the"element"(𝑖,𝑗)"now"contains"a"value"of"1.0," thereby"representing"that"the"predicate"holds"with"respect"to"the"corresponding"subject"and"object.& & &
ARIADNE"D16.1"Public" vi" " After"propositionalization,"each"(𝑖,𝑗)"of"a"matrix"contains"either"a"value"of"1.0"(holds)"or"0.0"(otherwise)" (Pu" and" Faltings" 2013," Rettinger," et" al." 2012)." Factorization" is" then" applied" to" determine" the" latent" features"that"are"hidden"between"the"entities."This"is"similar"to"the"wellIknown"statistical"technique"of" Principal"Component"Analysis"(PCA)."Once"split"up,"the"matrices"are"multiplied"again,"thereby"creating" an"estimation"of"the"original"matrix."However,"where"before"some"of"the"entities"had"the"value"0.0,"they" now"have"a"value"between"0.0"and"1.0,"thereby"representing"confidence"values"that"the"corresponding" statement"holds." " Figure&CD1:&Propositionalization&of&graph&data&to&relational&matrices.&Note&that,&for&clarity,&only&three& distinct&RDF&predicates&have&been&depicted&as&relational&matrices,&whereas&the&graph&depicts&a&maximum& of&nine.& An" advantage" of" using" relational" matrices" is" that" the" processes" of" applying" propositionalization" and" factorization"are"fairly"straightforward."While"many"different"approaches"to"factorizing"a"matrix"exist,"a" technical"detail"omitted"here,"it"constitutes"a"proven"method"overall"that"is"in"use"at"many"companies" (Pu"and"Faltings"2013,"Rettinger,"et"al."2012)."Furthermore,"once"an"RDF"graph"has"gone"through"this" process,"the"reconstructed"statements"that"were"previously"unknown"could"be"integrated"as"(weighted)" triples"(Tresp,"et"al."2008)."Care"should"be"taken"however,"as"perhaps"not"every"unknown"statement"is" justified."Another"aspect"of"which"one"should"be"wary"is"that"all"statements"of"a"certain"predicate"are" estimated" in" a" single" step." Extending" this" over" the" entire" graph" provides" a" performance" that" scales" proportionally"to"the"number"of"predicates,"as"well"as"to"their"frequency."That"is,"every"additional"triple" requires"an"additional"entry"in"the"corresponding"(possibly"not"yet"existing)"matrix." C.1.2 Tensors& Tensors"may"be"thought"of"as"the"generalization"of"the"matrix,"thereby"allowing"for"an"arbitrary"number" of"modes,"called"orders."Differently"put;"a"matrix"can"be"regarded"as"a"secondIorder"tensor."Depending" on"the"number"of"orders"𝑁,"a"tensor"consists"of"points"(𝑖,𝑗,…,𝑛)"with"𝑖,𝑗,…,𝑛∈𝑁."Note"however,"that" orders"are"independent"of"the"spatial"dimensions." During"propositionalization"(Figure"CI2)," each"graph"is"translated"to"a"single"tensor" (Kolda"and"Bader" 2009,"Rettinger,"et"al."2012,"Nickel,"Tresp" and"Kriegel" 2011)." Given" that" each" triple"consists"of" three" entities,"those"tensors"will"be"of"the"third"order."Of"these"orders,"the"first"will"contain"all"subjects,"the" second"all"predicates,"and"the"third"all"objects"of"the"graph."Assuming"indices"𝑖,𝑗,𝑘"for"the"first,"second,"
ARIADNE"D16.1"Public" vii" " and"third"order"respectively,"each"element" 𝑖,𝑗,𝑘"will"now"contain"the"value"1.0"if"the"statement"holds," and"0.0"otherwise."" Factorization" is" then" applied" to" determine" the" latent" features" that" are" hidden"between" the" entities." Once"split"up,"the"tensors"are"multiplied"again,"thereby"creating"an"estimation"of"the"original"tensor" (Kolda" and" Bader"2009,"Rettinger,"et"al." 2012)."However,"where"before" some"of"the"entities"had"the" value"0.0,"they"now"have"a"value"between"0.0"and"1.0,"thereby"representing"confidence"values"that"the" corresponding"statement"holds." " Figure&CD2:&Propositionalization&of&graph&data&to&a&thirdDorder&tensor&in&a&threeDdimensional&vector& space.&Note&that,&for&clarity,&only&two&RDF&statements&have&been&depicted&in&the&tensor,&whereas&the& graph&depicts&nine.& An" advantage" of" using" tensors" is" the" simplicity" of" applying" propositionalization" and" factorization." In" addition," tensors" permit" the" inclusion" of" contextual" information"(Rettinger," et" al." 2012)," as" well" as" allowing"for"collective"learning"such"that"it"is"less"dependent"on"the"explicit"aggregation"of"information."" Furthermore,"once"an"RDF"graph"has"gone"through"this"process,"the"reconstructed"statements"that"were" previously"unknown"could"be"integrated"as"(weighted)"triples"(Tresp,"et"al."2008)."Care"should"be"taken" however,"as"perhaps"not"every"unknown"statement"is"justified."Another"aspect"of"which"one"should"be" wary"is"that"all"statements"of"the"whole"graph"are"estimated"in"a"single"step."Nevertheless,"tensors"tend" to"scale"rather"well,"despite"their"large"size."That"is,"each"order"ranges"proportionally"to"the"number"of" corresponding"entities"within"the"graph."" C.2 Statistical&Relational&Learning& Statistical" Relational" Learning(is" an" umbrella" term," which" encompasses" a" large" number" of" various" methods" to" represent," reason," and" learn" in" domains" with" complex" relational" and" rich" probabilistic( structures((Getoor"and"Taskar"2007,"Rettinger,"et"al."2012)."Typically,"these"methods"are"based"on"either" logicI"or"frameIbased"formalisms,"such"as"rules"or"graphical"models"respectively."" C.2.1 Inductive&Logic&Programming& Inductive"Logic"Programming"(ILP)"encompasses"methods"that"attempt"to"learn"logical"(Horn)"clauses" directly" from" relational" data"(Lavrac" and" Dzeroski" 2001," Getoor" and" Taskar" 2007," Tresp," et" al." 2008," Maimon" and" Rokach" 2005)." These" clauses" typically" consist" of" conjunctions" of" positive" and" negative"
ARIADNE"D16.1"Public" viii" " logical"atoms,"which"together,"can"be"seen"as"Logic"Programs."In"addition,"these"Logic"Programs"allow" for"the"integration"of"valid"background"(domain)"knowledge"as"well." Given"a"set"of"positive"and"negative"atoms"of"target"relation"𝑝,"as"well"as"given"a"set"of"background" relations"𝑞!,"the"task"is"to"learn"a"definition"of"relation"𝑝"that"is"consistent"and"complete"(Lavrac"and" Dzeroski"2001,"Getoor"and"Taskar"2007,"Maimon"and"Rokach"2005)."That"is,"this"definition"should"be" able"to"explain"all"specified"positive"and"negative"atoms."Within"the"context"of"LD,"both"types"of"atoms" can"be"thought"of"as"triples"that"either"do"or"do"not"hold."In"that"case,"the"outcome"of"a"single"clause" would"then"constitute"the"truth"value"of"a"hypothesized"triple." As"an"example,"consider"wanting"to"determine"whether"a"colour"difference"in"a"layer"of"excavated"soil" are"the"remains"of"an"ancient"water"well."Given"numerous"examples"of"colour"differences"that"were" found"to"be"such"a"water"well,"and"given"numerous"more"examples"of"those"that"were"not,"a"new"rule" could"be"learned"that"discriminates"between"the"two"cases."A"very"crude"version"of"such"a"rule,"or"Logic" Program,"might"resemble"Equation"DI1,"which"evaluates"the"colour"difference"𝐶𝐷"as"being"a"well"(𝑉= 𝑡)"if"its"radius"is"larger"than"50"units"and"its"depth"larger"than"200"units,"as"well"as"having"a"difference"𝐷" in"colour"with"respect"to"the"surrounding"area"𝑆𝐴"that"exceeds"a"value"of"1.42"units." &𝑤𝑎𝑡𝑒𝑟𝑊𝑒𝑙𝑙 𝐶𝐷,𝑉=𝑡←& & & & & & & &&&&&Equation&CD1" """""𝑟𝑎𝑑𝑖𝑢𝑠 𝐶𝐷 >50 ∧𝑑𝑒𝑝𝑡ℎ𝐶𝐷 >200" ∧𝑐𝑜𝑙𝑜𝑢𝑟𝐷𝑖𝑓𝑓𝑒𝑟𝑒𝑛𝑐𝑒 𝐶𝐷,𝑆𝐴,𝐷∧𝐷>1.42!" A"big"advantage"of"ILP"is"the"experience"it"has"on"(multiI)"relational"DM"(Getoor"and"Taskar"2007,"Lavrac" and"Dzeroski"2001,"Maimon"and"Rokach"2005),"which,"while"not"the"same,"shares"many"of"the"hurdles" that"are"also"present"with"graph"data."Another"advantage"is"its"logical"representation,"which"provides"a" strong"expressive"power"and"which"fits"naturally"to"the"formal"logics"behind"LD."This"fit"allows"for"an" easy"integration"of"background"knowledge,"as"well"as"allowing"for"an"equally"easy"integration"of"inferred" knowledge"into"the"original"graph."However,"while"some"variants"do"exist,"this"new"knowledge"is"mainly" limited" to" statements" that" either" do" or" do"not" hold." Another" disadvantage" is" the" need" for" multiple" positive"and"negative"atoms"per"relation,"which"in"the"case"of"LD,"is"challenging"due"to"the"oftenIexisting" sparseness"in"RDF"graphs."Moreover,"the"need"for"negative"atoms"translated"into"a"need"for"explicitly" specified"triples"that"state"that"a"relation"does"not"hold;"a"practice"that"is"rare"in"a"system"that"adheres" to"the"OpenIWorld"Assumption." C.2.1.1 Propositional+ILP+ In"addition"to"relational"learning,"ILP"also"has"a"strong"relation"with"propositional"learning."In"fact,"it"was" in" the" area" of" ILP" where" the" term" propositionalization"was" originally" used." Within" that" context," it" constitutes"the"translation"of"firstIorder"clauses"into"features"to"which"statistical"ML"algorithms"can"be" applied." However," this" process" may" result" in" the" loss" of" information," and" thus" is" said" to" exchange" accuracy"for"efficiency"(Krogel,"et"al."2003)."
ARIADNE"D16.1"Public" ix" " With"propositionalization"(Figure"CI3),"a"set"of"Boolean"features"is"sought"whereby"each"feature"can"be" defined"as"a"corresponding"clause"(Krogel,"et"al."2003,"Lavrac"and"Dzeroski"2001,"Tresp,"et"al."2008)." Given"𝑛"features,"a"propositionalization"of"a"relationalIlearning"problem"is"a"set"of"𝑛"clauses,"with"each" clause"constituting"one"or"more"logical"literals."These"literals"are"derived"from"the"relational"background" knowledge."Once"all"features"have"been"defined,"they"are"evaluated"with"respect"to"an"instance,"thus" resulting"in"a"sequence"of"𝑛"Boolean"values."" Within"the"field"of"ILP,"there"are"two"approaches"to"propositionalization"(Krogel,"et"al."2003,"Lavrac"and" Dzeroski" 2001);" either" completely" or" partially." With" complete" propositionalization," all" knowledge" contained"within"a"data"set"is"translated"to"feature"definitions,"whereas"with"the"partial"variant"this"is" done" only" for" the" mostIrelevant" subset" of" this" knowledge." Therefore," in" case" of" the" latter," certain" knowledge"is"lost."Moreover,"determining"which"features"are"of"importance"is"typically"accomplished" through"heuristics,"thus"introducing"assumptions." " Figure&CD3:&Propositionalization&of&graph&data&to&a&set&of&ILP&feature&definitions.& An" advantage" of" propositional" ILP" are" that," in" most" cases," it" is" a" very" robust" and" effective" method" (Krogel,"et"al."2003,"Lavrac"and"Dzeroski"2001,"Tresp,"et"al."2008)."Only"a"few"data"sets"are"known"to" exist"for"which"this"does"not"hold."Another"benefit"is"that"the"partial"variant"may"easily"be"tailored"to" one’s"needs"by"specifying"the"number"of"features."This"latter"feature"can"additionally"be"regarded"as"a" disadvantage,"as"it"may"quickly"lead"to"lost"knowledge."Moreover,"it"might"introduce"assumptions"that" may"bias"any"future"reasoning."" C.2.2 Relational&Graphical&Models& Probabilistic"Graphical"Models"(PGM)"are"a"general"framework"(Figure"CI4)"to"represent"complex"realI world"phenomena"over"a"highIdimensional"space"by"the"combination"of"probability"theory"and"logical" structures"(Getoor"and"Taskar"2007,"Cohen"and"Kou"2006,"Koller"and"Friedman"2009)."Therefore,"PGMs" are"able"to"reason"with"uncertainty,"as"well"as"with"dependencies"between"entities."While"numerous" kinds"exist,"the"majority"of"the"PGMs"can"easily"be"depicted"by"either"a"directed"or"undirected"graph," such" as" a" Bayesian" and" Markov" model," respectively." Within" this" graph," the" nodes" map" to" domain" variables" and" the" edges" correspond" to" direct" probabilistic" interactions" between" these" variables." Furthermore,"PGMs"have"a"fixed"graphical"structure,"which"limits"their"ability"to"reason"about"a"varying" number"of"entities"in"a"variety"of"configurations."
ARIADNE"D16.1"Public" x" "" Relational" Graphical" Models" (RGM)" are" a" specific" kind" of" PGMs" (Figure" CI4)" that" extend" the" PGM" framework"with"a"flexible"graphical"structure,"as"well"as"with"concepts"of"objects,"their"properties,"and" relations"between"them"(Getoor"and"Taskar"2007,"Tresp,"et"al."2008,"Lavrac"and"Dzeroski"2001)."This" separation"is"similar"to"that" of"propositional"and"relational"logics."Irrespective,"whether" the"relations" hold"is"encoded"by"their"corresponding"binary"variables,"which"are"represented"as"nodes"in"the"RGM." Within"the"context"of"LD,"these"variables"denote"the"potential"RDF"triples"and"not"the" nodes" of" the" corresponding"RDF"graph,"which"are"either"resources"or"literals."More"specifically,"a"variable"with"the" value"1.0"would"denote"that"the"corresponding"triple"holds,"whereas"the"value"0.0"would"denote"the" opposite."Any"value"in"between"would"indicate"a"certainty"value." " Figure&CD4:&Venn&diagram&of&the&hierarchy&within&the&PGM&framework.&Note&that&the&cursive&terms&within& the&boundaries&denote&that&class’&mostDcommon&model.& " As"with"PGMs,"RGMs"may"be"either"directed"or"undirected"(Figure"CI4)."In"case"of"the"former,"these"are" generally"known"as"Probabilistic"Relational"Models"(PRM)"them"(Getoor"and"Taskar"2007,"Tresp,"et"al." 2008,"Lavrac"and"Dzeroski"2001,"Rettinger,"et"al."2012)."In"addition,"the"type"of"relation"and"its"direction" are"assumed"to"be"known;"an"assumption"intended"to"simplify"the"model."An"extension"to"this" additionally"considers"two"types"of"structural"uncertainty:" Reference(uncertainty"concerns"the"case"in"which"a"relation,"and"only"one"of"its"two"members,"is" known." That" is," it" is" unknown" whether" a" certain" entity" is" dependent" upon" one" or" more" other" entities,"and"if"so,"which"those"other"entities"are."" Existence(uncertainty"concerns"the"case"in"which"the"relation"between"two"or"more"entities"is" unknown."That"is,"given"two"entities,"it"is"unknown"whether"one"depends"on"the"other." " "
ARIADNE"D16.1"Public" xi" " As"an"example,"consider"translating"a"simple"RDF"graph"(Figure"CI5"Left)"to"a"PRM"(Figure"CI5"Right)." Here," the"former"involves" an"anomaly" discovered"within"the"soil," which" was"found"to"be" an" ancient" water"well."This"conclusion"was"based"upon"the"radius"and"depth"of"the"anomaly,"thus"making"these" measurements"dependencies"for"that"conclusion."Therefore,"a"PRM"would"represent"the"corresponding" triples,"i.e."Soil&Anomaly&has&radius&60"and"Soil&Anomaly&has&depth&250"as"parent"nodes"of"the"concluding" triple"Soil&Anomaly&a&Water&Well." " Figure&CD5:&Left)&a&simple&RDF&graph&concerning&an&anomaly&in&the&soil&that&once&constituted&a&water& well.&Right)&a&PRM&of&the&RDF&graph&with&every&triple&being&true.& An" advantage" of" RGMs" is" the" flexible" but" powerful" expressiveness" of" their" graphical" representation" (Getoor"and"Taskar"2007,"Tresp,"et"al." 2008," Lavrac"and"Dzeroski"2001,"Rettinger,"et"al."2012)."Being" grounded"in"a"sound"statistical"framework,"they"allow"for"direct"learning"from"graph"data"without"the" need" for" propositionalization." In" addition," ontology" background" knowledge" can" be" integrated," and" inferred" knowledge" can" be" incorporated" into" the" RDF" graph" as" weighted" triples." This" makes" it" particularly"well"suited"to"exploratory"data"analysis."Furthermore,"it"possesses"the"ability"to"learn"and" perform" inference" in" large" networks." However," both" tasks" tend" to" be" expensive," as" computational" requirements"scale"with"the"number"of"statements"whose"truth"value"is"known."Alternative"approaches" have" been" suggested," which" limit" the" learning" and" reasoning" to" relevant" subgraphs" only." These" approaches"however,"are"still"fairly"new"and"experimental."Another"challenge"still"largely"unresolved"is" how"to"deal"with"missing"data."Finally,"one"should"note"that"RGMs"are"mainly"limited"to"learning"and" predicting"truth"values"of"RDF"statements." C.3 Kernel&Methods& Kernel"methods"are"a"group"of"techniques"that"try"to"solve"mathematically"complex"ML"problems"by" translating" them" into" moreImathematicallyIworkable" ML" problems"(Schölkopf" and" Smola" 2002)." The" function"that"performs"this"translation"is"called"the"kernel."How"such"a"kernel"is"defined"depends"greatly" on"the"characteristics"of"the"problem"at"hand."In"fact,"many"different"kernels"exist,"each"designed"with"a" specific"goal"in"mind."
ARIADNE"D16.1"Public" xii" " Mathematically"speaking45;"given"a"problem"specified"in"input&space"𝒳,"a"mapping"Φ"is"defined"which" translates"points"in"𝒳"to"feature&space"ℋ"(Equation"DI2)."A"kernel"𝑘"is"then"defined"such"that"it"maps" input"vectors"𝑥"and"𝑥′"from"𝒳"to"ℋ"using"mapping"Φ"(Equation"DI3)."" 𝛷∶𝒳→ℋ&&&&&&&Equation&CD2& 𝑘𝑥,𝑥!⟼!!𝛷𝑥,𝛷𝑥!!&&&&&Equation&CD3& Amongst"the"most"popular"methods"that"use"kernels"are"SupportIVector"Machines"(SVM)"(Schölkopf" and"Smola"2002,"Gärtner"2003)."SVMs"are"commonly"used"to"perform"classification"tasks"(Figure"CI6)." However," they" can" additionally" be" used" for" either" clustering" or"regression." By" incorporating" kernels," which"they"refer"to"as"applying"the"kernel!trick,"SVMs"are"able"to"mitigate"several"of"the"issues"that"are" encountered"by"traditional"ML"algorithms."One"prime"example"of"such"an"issue"is"local"minima,"which" cause" an" optimization" process" to" stall."Moreover," the" kernel" trick" allows" SVMs" to" learn" nonIlinear" separation"boundaries." " Figure&CD6:&Example&of&how&SVMs&translate&complex&problems&(left),&here&a&binaryDclassification&problem,& to&moreDworkable&problems&(right)&(Schölkopf&and&Smola&2002).& " " """"""""""""""""""""""""""""""""""""""""""""""""""""""""""""" 45"For" simplicity," the" terminology" chosen" is" the" one" from" commonlyIused" in" SVMs." Regardless," the" principles" remain"the"same."
ARIADNE"D16.1"Public" xiii" " An"advantage"of"using"kernel"methods"is"they"do"not"restrict"their"arguments"to"solely"vectorItype"data." Instead,"they"may"be"defined"on"almost"any"type"of"data"(Bloehdorn"and"Sure"2007,"Gärtner"2003)."As"a" result," they" can" be" applied" directly" to" heterogeneous" and" interconnected" data" without" the" need" for" converting" these" data" to" vectors."A" related" advantage" of"kernels" is" their" ability" to"simplify"complex" learning" problems" by" mapping" them"to" lessIcomplex" learning" problems."Moreover," SVM," the" field’s" most"wellIknown"method,"is"one"of"the"most"successful"recent"developments"in"ML."In"fact,"it"has"been" shown"to"effectively"solve"learning"problems"which"were"left"unsolved"by"ILP;"a"field"with"many"more" years" of" experience." However," most" of" the" success" stories" on" kernel" methods" stem" from"academic" research."So"far,"very"little"experience"has"been"achieved"outside"of"those"confines."Due"to"the"field’s" current"popularity"however,"this"disadvantage"may"eventually"be"nullified." C.3.1 Kernels&for&Structured&Data& There" are" two" common" approaches" when" developing"kernels"for"structured" data." These" approaches" concern" the" decision"to" use"either"modelIdriven"or" syntaxIdriven"kernels" (Gärtner" 2003)." Here," the" former"is"often"applied"when"either"background"knowledge"or"states46"are"of"importance."In"contrast," the"latter"emphasizes"the"semantics"within"a"data"set."These"kernels"typically"involve"rules,"trees,"and" graphs."" Kernels"developed"for"learning"on"the"SW"often"fall"into"the"group"of"Graph(Kernels."Each"of"these"graph" kernels"constitutes"a"function"that"translates"a"graph"to"an"element"of"which"the"similarity"with"other" elements" can" easily" be" computed." Unfortunately," this" translation" step" is" generally" rather" resource" expensive," especially" when" isomorphism47"has" to" be" taken" into" account"(Gärtner" 2003," Borgwardt," Schraudolph"and"Vishwanathan"2006,"Vishwanathan,"et"al."2010)."A"compromise"is"to"use"approximate" kernels"instead."One"wellIknown"and"simple"example"of"these"is"a"RandomIWalk"Graph"kernel,"which" computes"the"translation"of"a"graph"by"randomly"walking"over"its"vertices."Two"graphs"𝐺"and"𝐺′"can" subsequently"be"checked"on"similarity"by"comparing"their"corresponding"random"walks." Another"kind"of"kernel"that"is"appropriate"to"learn"from"the"SW"is"the"Clause(Kernel"(Bicer,"Tran"and" Gossen" 2011)." Each" clause"kernel" contains" an" ILP" clause" which" corresponds" to" a" feature" of" the" corresponding"RDF"graph."This"approach"adds"the"advantage"of"dynamically"defining"kernels"based"on" the"likelihood"of"them"being"able"to"explain"a"set"of"provided"examples."Furthermore,"several"of"these" kernels" may" subsequently" be" combined" into"one" composite" kernel," thereby" improving" efficiency." In" addition,"combining"kernels"provides"resilience"to"sparse"data." """"""""""""""""""""""""""""""""""""""""""""""""""""""""""""" 46"A"state"involves"a"snapshot"of"a"data"set"at"a"specific"time."Applying"an"operation,"e.g."add"or"remove,"to"a"state" results"in"a"new"state." 47"Isomorph"graphs"only"differ"in"the"enumeration"of"their"vertices."
ARIADNE"D16.1"Public" xiv" " Appendix&D Sample&of&Archaeological&Scenarios& Several"junior"and"senior"researchers,"as"well"as"a"couple"of"Ph.D."candidates,"were"invited"to"participate" in" a" brainstorming" session" at" the" Faculty" of" Archaeology" of" Leiden" University" as" well" as" at" the" VU" University" Amsterdam."Those" participating" were" asked" to" write" down" an" archaeological" research" scenario," which" they" deemed" difficult" with" the" current" information" infrastructure."In" addition," participants" were" instructed" to" assume" that" all" required" data" was" readily" available" within" this" infrastructure." The"following"scenarios"were"submitted"(translated"from"Dutch):" 1) The&need&to&retrieve&all&literature&about&individual&finds&from&very&(often&obscure)&local&journals& about&a&particular&excavation&(i.e.&the&late&bronze&age&settlement&in&Bovenkarspel).& 2) The&need&for&information&on&plants&which&are&useful&for&humans,&both&as&food,&oil,&medicinal&or& as&material&(for&creating&ropes,&buckets&etc.)&found&in&Holocene&contexts&dated&to&Late&prehistory& (i.e.&Late&Neolithic,&Bronze&Age&and&Iron&Age).& 3) All&publications&from&excavated&settlements&which&are&dated&to&the&Bronze&Age&and&located&on&a& specific&geomorphological&unit,&in&an&area&from&Denmark&to&North&France.&& 4) All&images&(images&and&drawings)&with&metadata&about&the&dating&and&archaeological&context& from&prehistoric&traps&made&of&willow.&In&addition&a&list&of&persons&who&researched&these&artefact& types.& 5) Overview&of&all&Neolithic&axes&made&of&stone,&and&in&particular&flint&found&in&a&Roman&context.& (i.e.&found&during&excavations,¬&individual&finds).& 6) The&reliability&and&usefulness&of&C14&dating&depends&on&a&variety&of&factors&(i.e.&stratigraphical& position,&soil&disturbances,&type&of&sample&etc.).&This&information&is¬ed&on&separate&sample& forms.&It&would&be&very&useful&if&this&information&were&accessible&on&to&the&level&of&individual&C14& identifiers.&(e.g.&www.lumid.nl&provides&TL&information,&but¬&as&LOD).& 7) Functionality& to&retrieve&all&the& GIS& data&from&a&specific&archaeological& area&(e.g.&all&GIS& data& from&Roman&excavations&on&the&Kops&Plateau&in&Nijmegen).& 8) All& published& radiometric& dated& (in& years)& material& (with& id& numbers)& on& Neanderthal& sites& in& France.& 9) All&types&of&arrowheads&dated&to&the&middle&Neolithic&period.&&& 10) All&characteristics&of&hand&axes&from&the&South&East&of&the&Netherlands.& 11) All&information&on&the&origins&of&the&Levallois&technique.&(site&locations&X,Y&coordinates).& 12) Information&on&the&influence&of&the&pH&value&soils&on&the&conservation&conditions&of&charcoal.& 13) All&archaeological&contexts&in&which&broken&flint&axes&dated&from&the&Middle&Neolithic&period&are& found.&(i.e.&settlements,&funerary,&individual&finds&etc.).& 14) All&sources,&bothðnographic,&historic&and&archaeological,&in& which&land& clearances&by& fire&is& mentioned&with®ard&to&their&influence&the&landscape&vegetation.& 15) Functionality&to&automatically&search&for&keyword&synonyms&(for&example&in&Google&scholar).&&