| Creator | Datahub/dbnary#Nf82de97a0daa4f518eac989b52108162 |
| Description | Extracts of wiktionary data for several languages, structured as an RDF graph, based mainly on the LEMON model. English, Finnish, French, German, Greek, Italian, Japanese, Portuguese, Russian and Turkish. |
| Rights | http://www.opendefinition.org/licenses/cc-by-sa |
| Source | DataHub |
| Title | dbnary |
| Creator | Datahub/dbpedia-abstract-corpus#N5b719f1e67ea4f4683b230acc01c8ad9 |
| Description | This corpus contains a conversion of Wikipedia abstracts in six languages (dutch, english, french, german, italian and spanish) into the I used the NLP Interchange Format (NIF). The corpus contains the abstract texts, as well as the position, surface form and linked article of all links in the text. As such, it contains entity mentions manually disambiguated to Wikipedia/DBpedia resources by native speakers, which predestines it for NER training and evaluation. Furthermore, the abstracts represent a special form of text that lends itself to be used for more sophisticated tasks, like open relation extraction. Their encyclopedic style, following Wikipedia guidelines on opening paragraphs adds further interesting properties. The first sentence puts the article in broader context. Most anaphers will refer to the original topic of the text, making them easier to resolve. Finally, should the same string occur in different meanings, Wikipedia guidelines suggest that the new meaning should again be linked for disambiguation. In short: The type of text is highly interesting. Acknowledgments: The conversion of this corpus was supported by the [FREME H2020 project](http://www.freme-project.eu/). |
| Rights | http://www.opendefinition.org/licenses/cc-by |
| Source | DataHub |
| Title | DBpedia abstract corpus |
| Description | |
| Rights | http://www.opendefinition.org/licenses/cc-by-sa |
| Source | DataHub |
| Title | DBpedia in Czech |
| Creator | Datahub/dbpedia-es#N0a1bb97f4153454f91640b28bda77ef2 |
| Description | These data correspond to the ontology DBpedia version 2014. |
| Rights | http://www.opendefinition.org/licenses/cc-by-sa |
| Source | DataHub |
| Title | DBpedia in Spanish |
| Creator | Datahub/dbpedia-fr#N9364835b5a7c4e62817215a464e4389f |
| Description | DBpedia in French dataset. Part of the DBpedia internationalisation effort. Data are extracted here from French speaking pages of wikipedia. |
| Rights | http://www.opendefinition.org/licenses/cc-by-sa |
| Source | DataHub |
| Title | DBpedia in French |
| Description | |
| Rights | http://www.opendefinition.org/licenses/cc-by-sa |
| Source | DataHub |
| Title | DBpedia in Korean |
| Creator | Datahub/dbpedia-live#N29a21abc36424dd395785d62563b3edc |
| Description | DBpedia.org is a community effort to extract structured information from Wikipedia and to make this information available on the Web. DBpedia allows you to ask sophisticated queries against Wikipedia and to link other datasets on the Web to Wikipedia data. The DBpedia knowledge base currently describes more than 3.4 million things, out of which 1.5 million are classified in a consistent Ontology, including 312,000 persons, 413,000 places, 94,000 music albums, 49,000 films, 15,000 video games, 140,000 organizations, 146,000 species and 4,600 diseases. The DBpedia data set features labels and abstracts for these 3.2 million things in up to 92 different languages; 841,000 links to images and 5,081,000 links to external web pages; 9,393,000 external links into other RDF datasets, 565,000 Wikipedia categories, and 75,000 YAGO categories. The DBpedia knowledge base altogether consists of over 1 billion pieces of information (RDF triples) out of which 257 million were extracted from the English edition of Wikipedia and 766 million were extracted from other language editions. |
| Rights | http://www.opendefinition.org/licenses/cc-by-sa |
| Source | DataHub |
| Title | DBpedia-Live |
| Creator | Datahub/dbpedia-nl#N3fd5515acd1447ec8f295fa84ef40b5c |
| Description | DBpedia is a \"community effort to extract structured information from Wikipedia and to make this information available on the Web. DBpedia allows you to ask sophisticated queries against Wikipedia, and to link other data sets on the Web to Wikipedia data. We hope this will make it easier for the amazing amount of information in Wikipedia to be used in new and interesting ways, and that it might inspire new mechanisms for navigating, linking and improving the encyclopaedia itself.\" |
| Rights | http://www.opendefinition.org/licenses/cc-by-sa |
| Source | DataHub |
| Title | DBpedia in Dutch |
| Creator | Datahub/dbpedia#Nf7838f5dc3c24b88b435398442a0b465 |
| Description | ### Description From the front page: > DBpedia.org is a community effort to extract structured information from Wikipedia and to make this information available on the Web. DBpedia allows you to ask sophisticated queries against Wikipedia and to link other datasets on the Web to Wikipedia data. > > The DBpedia knowledge base currently describes more than 3.64 million things, out of which 1.83 million are classified in a consistent Ontology, including 416,000 persons, 526,000 places, 106,000 music albums, 60,000 films, 17,500 video games, 169,000 organisations, 183,000 species and 5,400 diseases. The DBpedia data set features labels and abstracts for these 3.64 million things in up to 97 different languages; 2,724,000 links to images and 6,300,000 links to external web pages; 6,200,000 external links into other RDF datasets, 740,000 Wikipedia categories, and 2,900,000 YAGO categories. The DBpedia knowledge base altogether consists of over 1.2 billion pieces of information (RDF triples) out of which 335 million were extracted from the English edition of Wikipedia and 865 million were extracted from other language editions. |
| Rights | http://www.opendefinition.org/licenses/cc-by-sa |
| Source | DataHub |
| Title | DBpedia |
| Description | |
| Rights | http://www.opendefinition.org/licenses/cc-by-sa |
| Source | DataHub |
| Title | DBpedia in Russian |
| Creator | Datahub/dbpedia-spotlight-nif-ner-corpus#Na70340f2fbf44ff28ce4e219af2cb7eb |
| Description | Based on P. N. Mendes, M. Jakob, A. García-Silva, and C. Bizer. DBpedia Spotlight: shedding light on the web of documents. In Proc. of the 7th Int. Conf. on Semantic Systems, 2011. It contains 60 natural language sentences from ten different New York Times articles with overall 249 annotated DBpedia entities, i. e. the entities are not explicitely bound to mentions within the texts, which causes a certain lack of clarity. Therefore, we (in all conscience) retroactively have allocated the entities to their positions within the texts. The entities dbp:Markup_Language and dbp:PBC_CSKA_Moscow could not be linked in the texts, since there was also a more specific entity enlisted occupying their solely possible location, e. g. hypertext markup language has been annotated with dbp:HTML rather than dbp:Markup_language. |
| Rights | http://www.opendefinition.org/licenses/cc-by |
| Source | DataHub |
| Title | DBpedia Spotlight NIF NER Corpus |
| Description | DBpedia Spotlight is a tool for annotating mentions of DBpedia resources in text, providing a solution for linking unstructured information sources to the Linked Open Data cloud through DBpedia. DBpedia Spotlight performs named entity extraction, including entity detection and Name Resolution (a.k.a. disambiguation). It can also be used for building your solution for Named Entity Recognition, amongst other information extraction tasks. The datasets you find here were produced by the DBpedia Spotlight team and can be reused in many natural language processing tools, as well as general Web applications that need to connect text to unique URIs from DBpedia. |
| Rights | http://www.opendefinition.org/licenses/cc-by |
| Source | DataHub |
| Title | DBpedia Spotlight |
| Creator | Datahub/emn#Nd03e736197e5424ca4915f46e867d11b |
| Description | The Terminology of the European Migration Network in RDF |
| Rights | http://www.opendefinition.org/licenses/cc-by |
| Source | DataHub |
| Title | EMN |
| Creator | Datahub/environmental-applications-reference-thesaurus#N5bb1ca2f37ec4e879042c0e2426c5cfe |
| Description | The Environmental Applications Reference Thesaurus (EARTh) has been compiled and is maintained by the CNR-IIA-EKOLab to facilitate the indexing, retrieval, harmonizing and integration of human- and machine-readable environmental information from disparate sources, across the cultural and linguistic barriers. Ownership of such material always remains with the CNR-IIA-EKOLab. EARTh has been firstly made available as linked data as an activity within the European Project NatureSDIPlus (ECP-2007-GEO-317007). It is currently maintained in the context of LusTRE, a framework under development within the EU project eENVplus (CIP-ICT-PSP grant No. 325232) that aims at combining existing thesauri to support the management of environmental resources. LusTRE considers the heterogeneity in scopes and levels of abstraction of environmental thesauri as an asset when managing environmental data, it exploits linked data best practices SKOS (Simple Knowledge Organization System) and RDF (Resource Description Framework) in order to provide a multi-thesauri solution for INSPIRE data themes related to the environment. |
| Rights | http://creativecommons.org/licenses/by-nc/2.0/ |
| Source | DataHub |
| Title | EARTh |
| Creator | Datahub/eurosentiment#Ne49e1534eb9b4fa6846d42e3fcf3327d |
| Description | Gabriela Vulcu, Raul Lario Monje, Mario Munoz, Paul Buitelaar and Carlos A. Iglesias (2014), Linked-Data based Domain-Specific Sentiment Lexicons, In: Proceedings of the 3rd Workshop on Linked Data in Linguistics (LDL-2014), Reykjavik, Iceland, May 2014 Resource Type: Lexicon Resource Name: EUROSENTIMENT domain-specific sentiment lexicons Size: 9160 Resource Production Status: Newly created-in progress Language(s): W Modality: Written Use of the Resource: Emotion Recognition/Generation Resource Availability: Freely Avalable Resource URL (if available): http://140.203.155.231:8080/eurosentiment/ Resource Description: Please see the submitted article. It describes this language resource and all other involved. |
| Rights | http://www.opendefinition.org/licenses/gfdl |
| Source | DataHub |
| Title | EuroSentiment |
| Creator | Datahub/fao-geopolitical-ontology#N520a3859e3d8418fbee243296fa13db6 |
| Description | The FAO geopolitical ontology provides a master reference for geopolitical information, as it manages names in multiple languages (English, French, Spanish, Arabic, Chinese, Russian and Italian); maps standard coding systems (UN, ISO, FAOSTAT, AGROVOC, DBPedia, etc); provides relations among territories (land borders, group membership, etc); and tracks historical changes. The ontology contains *number of triples : 22495 triples *links to other data sets: 195 links to DBPEDIA The Food and Agriculture Organization of the United Nations (FAO) leads international efforts to defeat hunger and serves as a knowledge network. |
| Rights | http://www.opendefinition.org/licenses/cc-by-sa |
| Source | DataHub |
| Title | FAO geopolitical ontology |
| Creator | Datahub/fiesta#Nb6b2f4b3ba764141931f4f1a69617789 |
| Description | FiESTA (short for \"Format for extensive spatiotemporal annotations\") is a generic format for linguistic and behavioral annotations. |
| Rights | http://www.opendefinition.org/licenses/cc-by-sa |
| Source | DataHub |
| Title | FiESTA |
| Creator | Datahub/french-timebank#Nd85865e038964dd2a23c7b4ed6bd36cb |
| Description | The French TimeBank consists of a set of 109 journalistic articles from 7 different sub-genres annotated according to the ISO-TimeML standard, adapted for the French language. Eventualities (events and states) and temporal expressions (dates, durations, frequencies, quantified intervals) are marked up with in-line annotation. The temporal relations that hold among these entities are also annotated, as are the aspectual and modal subordination relations between eventualities. The corpus is available under the Lesser General Public Licence for Linguistic Resources (LGPL-LR). |
| Rights | http://www.opendefinition.org/licenses/cc-by |
| Source | DataHub |
| Title | French TimeBank |
| Description | This is the Galician EuroWordNet-Lemon lexicon. The lexicon was created from the Spanish Word-Net-LMF lexicon which is part of the Multilingual Central Repository (MCR http://adimen.si.ehu.es/web/MCR). The lexicon conforms to the 'lemon' specification. Gloss and rgloss relations between synsets are not included. For LexicalEntries and LexicalSenses, original ID's are encoded in dcterms:source and the URIs follow the pattern '../lemma-PoS'. For Synsets and Translations original IDs are used in the URIs (.../ID). Synset rdfs:labels were generated as follows: INSERT {?synset rdfs:label ?labels } WHERE {SELECT ?synset (GROUP_CONCAT(?label; separator = ' ; ') as ?labels) { ?sense lemon:reference ?synset; rdfs:label ?label . } GROUP BY ?synset } http://lodserver.iula.upf.edu/id/WordNetLemon/GL/ |
| Rights | http://www.opendefinition.org/licenses/cc-by |
| Source | DataHub |
| Title | Galician EuroWordNet-lemon lexicon (3.0) |
| Creator | Datahub/gemeenschappelijke-thesaurus-audiovisuele-archieven#N5f86544c9f384fd8a38f8ea842f161dd |
| Description | The Netherlands Institute for Sound and Vision <http://portal.beeldengeluid.nl/> is the Dutch archive for public broadcast television. They employ the GTAA, which is a Dutch acronym for Common Thesaurus [for] Audiovisual Archives, to index and disclose their audiovisaul documents. The GTAA closely follows the ISO-2788 standard for thesaurus structures. The thesaurus consists of several facets for describing TV programs: subjects; people mentioned; named entities (Corporation names, music bands etc); locations; genres; makers and presentators. The GTAA contains approximately 160.000 terms: ~3800 Subjects, ~97.000 Persons, ~27.000 Names, ~14.000 Locations, 113 Genres and ~18.000 Makers, and is continually updated as new concepts emerge on TV. |
| Rights | http://www.opendefinition.org/licenses/odc-odbl |
| Source | DataHub |
| Title | Gemeenschappelijke Thesaurus Audiovisuele Archieven – Common Thesaurus Audiovisual Archives |
| Creator | Datahub/gemet-annotated#Ndada529569ca4b958359f9509831cdd6 |
| Description | Details about how this dataset was built are described in the article: Are SKOS concept schemes ready for multilingual retrieval applications? — Diana Tanase and Epaminondas Kapetanios |
| Rights | http://www.opendefinition.org/licenses/odc-odbl |
| Source | DataHub |
| Title | gemet-annotated |