| Contact Point | Metashare/8c13600ccd0711e1a404080027e73ea2f9cfd28f51d5437b8f5827c516c348fe#contact Person |
| Contributor | Dan Tufis |
| Creator | Amália Mendes |
| Description |
This lexicon includes multiword expressions (MWE) of European Portuguese extracted from a balanced 50,8M word written corpus – a subcorpus of the Reference Corpus of Contemporary Portuguese (CRPC). This corpus covers different genres, being mainly constituted by journalistic texts (59%), but it also includes texts from literature (21%), magazines (15%), miscellaneous, supreme court verdicts, parliament sessions and leaflets (5%). The MWE lexicon covers 1.198 lemmas (composed of single words from different POS categories: nouns, adjectives, verbs and adverbs) and a total of 12.753 MWE lemmas (which include inflectional variants of the MWE lemmas) and 242.233 concordances of those MWE expressions manually verified.
|
| Rights | underNegotiation |
| Source | META-SHARE |
| Title |
LEX-MWE-PT: Word Combination in Portuguese Language
|
| Type | Lexical Conceptual Resource |
| Contact Point | Metashare/12fdc090a35e11e1a404080027e73ea2c200dc17aff642fe980ba7a2da7f5ca1#contact Person |
| Contributor | Dan Tufis |
| Creator | Maria Fernanda Bacelar do Nascimento |
| Description |
This resource includes a spoken Portuguese corpus exemplifying the Portuguese spoken in Portugal, Brazil, Angola, Cape Verde, Guinea-Bissau, Mozambique, Sao Tome and Principe, Macao, Goa and East-Timor - with aligned sound and orthographic transcription - collected among sociolinguistically diverse speakers. It consists of recordings from informal conversations, conferences and media.
|
| Rights | underNegotiation |
| Source | META-SHARE |
| Title |
Spoken Portuguese - Geographical and Social Varieties
|
| Type | Corpus |
| Contact Point | Metashare/f30f4d04486111e2a2aa782bcb07413522a84532f1d443ffb62c5fcf3c59545a#contact Person |
| Contributor | Dan Tufis |
| Creator | Amália Mendes |
| Description |
The EUROPARL Corpus (subpart Portuguese-English of the parallel corpora), available at http://www.statmt.org/europarl/, was extracted from the proceedings of the European Parliament (Koehn, 2005). It contains transcriptions of sessions dating back from 1996 to 2011, in a total of approximately 58,324,562 tokens words of European Portuguese (L1) and 49,216,896 tokens of English (translation).
|
| Language | English |
| Portuguese | |
| Rights | CC-BY-SA |
| Source | META-SHARE |
| Title |
EUROPARL Corpus
Parallel Corpora: Portuguese-English
|
| Type | Corpus |
| Contact Point | Metashare/2d875be6a35a11e1a404080027e73ea2fd1adfb79b3648ae9fff2932e82cf95a#contact Person |
| Contributor | N/A |
| Creator | N/A |
| Description |
This is the Maltese version of the Acquis Communautaire (AC), which is the total body of European Union (EU) law applicable in the EU Member States. It consists of selected texts between the 1950s and today, translated to Maltese.
|
| Rights | other |
| Source | META-SHARE |
| Title |
Maltese Acquis Communautaire
|
| Type | Corpus |
| Contact Point | Metashare/fe32ebf2485511e2a2aa782bcb074135aa0fdcd287ac45e7b67de9c36d8d2890#contact Person |
| Contributor | Dan Tufis |
| Creator | António Branco |
| Amália Mendes | |
| Description |
CINTIL-Corpus Internacional do Português is a linguistically interpreted corpus of Portuguese. At present it is composed of 1 Million annotated tokens, verified by human expert annotators. The annotation comprises information on part-of-speech, open classes lemma and inflection, multi-word expressions pertaining to the class of adverbs and to the closed POS classes, and multi-word proper names (for named entity recognition). The corpus has been developed at the University of Lisbon by the NLX group at the Faculty of Sciences and the Anagrama group at the Cenro de Linguística da Universidade de Lisboa.
|
| Language | Portuguese |
| Rights | ELRA_END_USER |
| Source | META-SHARE |
| Title |
CINTIL-Corpus Internacional do Português
|
| Type | Corpus |
| Contact Point | Metashare/2f2a00e4b92f11e1a404080027e73ea2eccd095ad8b0407989b2adb143ab6095#contact Person |
| Contributor | Dan Tufis |
| Creator | Rosa Del Gaudio |
| Description |
The corpus presented here is a collection of several tutorials and scientific papers in the field of Information Technology with 603 annotated definitions from Portuguese. The texts were collected from the Web at the beginning of the 2006 and they are organised in 32 files of three different sub-domains with 268,064 tokens: Information Society (91,825 tokens), Information Technology (80,483 tokens), and e-Learning (94,756 tokens).
|
| Rights | underNegotiation |
| Source | META-SHARE |
| Title |
CINTIL-Definitions
|
| Type | Corpus |
| Contact Point | Metashare/72cc03d88be311e294080015171445924c5e7298251340b291958854f48d783e#contact Person |
| Metashare/72cc03d88be311e294080015171445924c5e7298251340b291958854f48d783e#contact Person2 | |
| Contributor | Anđelka Zečević |
| Creator | Anđelka Zečević |
| Krstev Cvetana | |
| Description |
NERosetta is a multiuser web application that aims to facilitate retrieval and comparison of named entities in a single or parallel texts. The main named entity categorization is realized according to the Quaero annotation recommendation and provides a user with approximately 50 different search options. Registered users have an extra possibility to share annotated resources (in XML format) and annotation schemas for the set of languages as well as to manage their own resources and schemas. The initial version supports four annotation schemas (Stanford NER 3 and Stanford NER 7 for English, Krstev&Vitas for Serbian and Maurel for French) and three annotated parallel versions of Jules Verne's Around The World in Eighty Days (English-Serbian, French-Serbian and French-English).
|
| Rights | GPL |
| Source | META-SHARE |
| Title |
NERosetta
|
| Type | Tool Service |
| Contact Point | Metashare/362a2020cf5711e1a404080027e73ea28eaaf998e9aa47739841451ea4e16f51#contact Person |
| Contributor | Dan Tufis |
| Creator | Maria Fernanda Bacelar do Nascimento |
| Description |
This resource includes a spoken corpus with approximately 300.000 words, covering both formal (152.755 words) and informal (165.838 words) speech, with aligned sound and orthographic transcription and POS-tag information.
|
| Rights | ELRA_END_USER |
| Source | META-SHARE |
| Title |
C-ORAL-ROM_EXM
|
| Type | Corpus |
| Contact Point | Metashare/27607ab28b2c11e2975a00151714459237c30a10120c409ab292a1ed8f3ec9fc#contact Person2 |
| Metashare/27607ab28b2c11e2975a00151714459237c30a10120c409ab292a1ed8f3ec9fc#contact Person | |
| Contributor | Mirko Spasić |
| Creator | Mirko Spasić |
| Duško Vitas | |
| Description |
Serbian NGrams (SrpNGrams) represent set of N-grams extracted from Serbian Lemmatized and PoS Annotated Corpus (SrpLemKor) for N from 1 to 5. Each unigram is maximum continuous chunk of non-whitespace lower-case characters. The resource contains all unique N-grams preceded by number of occurrencies. It also contains n-gram language models (1-5) in the standard ARPA text and binary format, created by IRST Language Modeling Toolkit. SrpKor texts consist of: fiction written by Serbian authors in 20th and 21th century, various scientific texts from various domains (both humanities and sciences), legislative texts and general texts. General texts represent daily news published in newspaper \"Politika\" 2000-2002 and 2005-2010, texts in journals and magazines 1991-2002 (\"Danica\", \"Ebit\", \"Ekonomist\", \"Glasnik\", \"NIN\", \"Ilustrovana politika\", \"Kalibar\", \"Moje srce\", \"Mostovi\", \"Pravoslavlje\", \"Svet\", \"Teološki pogledi\", \"Trn\", \"Viva\", \"Republika\"), internet portal texts 2011-2012 (Peščanik), TANJUG agency news 1995-96, newspaper feuilletons published in newspapers \"Politika\" (2001-2003), \"Večernje novosti\" (2008-2011) and \"Danas\" (2002-2006).
|
| Rights | MS-NC-NoReD-ND |
| Source | META-SHARE |
| Title |
Serbian NGrams
|
| Type | Corpus |
| Contact Point | Metashare/259b504e614c11e2ad2c842b2b6a04d7f20e5b8fbb5449338534ed92601521dd#contact Person |
| Contributor | Harris Papageorgiou |
| Maria Koutsombogera | |
| Creator | Maria Koutsombogera |
| Language | Greek (modern) |
| Source | META-SHARE |
| Title |
Greek Interview Multimodal Corpus
|
| Type | Corpus |
| Contact Point | Metashare/58b341b48be311e294150015171445929ea8576278db448b817323b587beaf95#contact Person |
| Metashare/58b341b48be311e294150015171445929ea8576278db448b817323b587beaf95#contact Person2 | |
| Contributor | Miljana Mladenović |
| Creator | Cvetana Krstev |
| Miljana Mladenović | |
| Description |
This tool is a web application for ontological based emotions recognition and tagging of Serbian texts. The application uses RDFS which are created by using nine discrete emotion psychological theories. Also, it uses associative dictionary of Serbian with about 11 thousands words and Serbian morphological electronic dictionary which contains approximately 4.4 million different inflectional forms of simple words. The application offers a representation of summary results in a graphical form. Annotation of an uploaded text is possible for XML and textual documents as well as a text from Web.
|
| Rights | GPL |
| Source | META-SHARE |
| Title |
Emotions Annotation Tool
|
| Type | Tool Service |
| Contact Point | Metashare/a794730e359c11e28aab080027f903f2139cf0d61cd949ac93267e364eec28ec#contact Person2 |
| Metashare/a794730e359c11e28aab080027f903f2139cf0d61cd949ac93267e364eec28ec#contact Person | |
| Metashare/a794730e359c11e28aab080027f903f2139cf0d61cd949ac93267e364eec28ec#contact Person3 | |
| Contributor | Rui Lageira |
| Gonçalo Simões | |
| Helena Galhardas | |
| Creator | Gonçalo Simões |
| Description |
Etxt2DB is a framework for specifying and executing Entity Recognition (ER) programs. These programs accept as input a text containing potentially interesting entities to be extracted and produce the input text annotated with the recognized entities.
The Etxt2DB functioning mode involves two distinct phases. First, the training phase consists in creating a model based on a given ER technique and one or more resources that guide the creation of the classification model. Examples of these resources are dictionaries for rule-based ER techniques or training data for statistical learning techniques (e.g., Conditional Random Fields). Second, in the execution phase, a classification model previously created receives as input plain text and produces annotations corresponding to the recognized entities.
The Etxt2DB framework consists of a software layer, built on top of Minorthird and Lingpipe, offering a command-like specification language. Existing Machine Learning Java APIs (such as Minorthird and Lingpipe) provide implementations of Entity Recognition techniques. Some developers of ER applications do not want to get involved in the implementation details of the techniques used. Instead, they are willing to focus on: the choice of the technique to be used; the resources used in the process (e.g., dictionaries); a good set of features that help the ER program to take adequate decisions. The objective of the Etxt2DB specification language is to turn the development and tuning of ER programs easier for developers that are mainly concerned with these topics.
In the context of the METANET project, the goal was to build a component-generator tool that encapsulates Etxt2DB. In the training phase, this tool accepts a training data set as input and produces a classification model and a U-Compare component that is able to interpret that model. In the execution phase, the component produced is loaded into the U-Compare platform and then is ready to be used for recognizing entities from text.
|
| Rights | GPL |
| Source | META-SHARE |
| Title |
U-Compare E-txt2DB: Giving structure to unstructured data
|
| Type | Tool Service |
| Contact Point | Metashare/0cb6205066e111e2bac9525400d761476ef3f57fa58942a4aac54d4206216320#contact Person |
| Contributor | Łukasz Dróżdż |
| Piotr Pęzik | |
| Creator | Łukasz Dróżdż |
| Piotr Pęzik | |
| Description |
A subset of the PELCRA PLEC corpus, containing 15 hours (131 000 transcribed words) of recordings of informal interviews with Polish learners of English, time-aligned on the utterance and annotated manually for mispronounciation errors, provided as TEI P5-conformant XML and EAF (ELAN) files.
|
| Language | English |
| Rights | CC-BY-NC |
| Source | META-SHARE |
| Title |
PELCRA Spoken Learner English Corpus
|
| Type | Corpus |
| Contact Point | Metashare/b6646fb866e011e29895525400d761474918098ee988487eaf619d3da163a80b#contact Person2 |
| Metashare/b6646fb866e011e29895525400d761474918098ee988487eaf619d3da163a80b#contact Person | |
| Contributor | Piotr Pęzik |
| Łukasz Dróżdż | |
| Creator | Łukasz Dróżdż |
| Piotr Pęzik | |
| Description |
A subset of the PELCRA corpus of conversational Polish, time-aligned on the utterance level, licensed under the CC-BY-NC license. This resource contains 386 744 words in 73 transcriptions of over 43 hours of recordings made in the years 2008-2010. The texts are provided as TEI P5-compliant XML files with custom PELCRA extensions and in the XLIFF format.
|
| Language | Polish |
| Rights | CC-BY-NC |
| Source | META-SHARE |
| Title |
PELCRA time-aligned spoken corpus of Polish (CC-BY-NC)
|
| Type | Corpus |
| Contact Point | Metashare/1f82f6866b0011e284b6000423bfd61c95584808dd944c28a354bba3bb390dbd#contact Person |
| Contributor | Alina Wróblewska |
| Creator | Alina Wróblewska |
| Description |
Statistical dependency parsing model is trained on the Polish Dependency Bank (PDB, Pol. Składnica zależnościowa) with the the publicly available parsing system -- MaltParser. MaltParser is a transition-based dependency parser that uses a deterministic parsing algorithm. The deterministic parsing algorithm builds a dependency structure of an input sentence based on transitions (shift-reduce actions) predicted by a classifier. The classifier learns to predict the next transition given training data and the parse history.
|
| Rights | GPL |
| Source | META-SHARE |
| Title |
Dependency Parsing Model for Polish
|
| Type | Tool Service |
| Contact Point | Metashare/f05f0f1c63f111e2bff4525400d76147f9863da9a70143bebd894e55197705a1#contact Person |
| Metashare/f05f0f1c63f111e2bff4525400d76147f9863da9a70143bebd894e55197705a1#contact Person2 | |
| Contributor | Łukasz Dróżdż |
| Piotr Pęzik | |
| Creator | Piotr Pęzik |
| Łukasz Dróżdż | |
| Language | French |
| Spanish | |
| Italian | |
| German | |
| English | |
| Source | META-SHARE |
| Title |
PELCRA mutlilingual parallel corpora (CC-BY)
|
| Type | Corpus |
| Contact Point | Metashare/4afd693e6ba711e2aa7c68b599c26a0651325e2993be4cc2a4950546705978a0#contact Person2 |
| Metashare/4afd693e6ba711e2aa7c68b599c26a0651325e2993be4cc2a4950546705978a0#contact Person | |
| Contributor | Attila Mártonfi |
| Description |
Hungarian historical corpus (further as HHC) is a collection of texts written between 1772 and 1997 in different genres, containing ca. 27 million tokens. During the compilation of HHC, text samples were selected by professionals (literary historians, historians, mathematicians etc.) from printed works. A relative majority (40%) of the texts are dated from the second half of the 20th century. The corpus is the product of the Department of Lexicography and Lexicology at RIL HAS, made between 1986 and 1997, maintained continuously since then. As an innovation, genre labeling was unified. Thus, genres and text types in HHC and HNC are marked similarly, this makes possible to search data of these corpora by using the same query structure.
|
| Rights | MS-NC-NoReD |
| Source | META-SHARE |
| Title |
HHC: Hungarian historical corpus
|
| Type | Corpus |
| Contact Point | Metashare/dad2b9848be011e29ebd001517144592d5a00254a9fd45bb9383caa72801461a#contact Person |
| Contributor | Cvetana Krstev |
| Creator | Cvetana Krstev |
| Description |
Morphological electronic dictionary of Serbian (Ekavian pronunciation) (SrpMD) released in the scope of the EU-funded CESAR project is a version of morphological dictionary of Serbian used in the NooJ corpus processing system and consituting the part of the Serbian Nooj Module (see section 6.8). This version is compliant to MULTEXT-East morphosyntactic specification for Serbian (http://nl.ijs.si/ME/V4/msd/html/msd-sr.html) (with one small deviation form it – see section 6.10). It comprises of 3,630,613 entries for 85,721 lemmas covering 11 PoS: nouns (646,867/40,425), adjectives (2,315,640/25,826), verbs (654,159/15,359), adverbs (3233), numerals(4,794/175), conjunctions (83), interjections (218), prepositions (169), pronouns (5,321/104), particles (103), abbreviations (26).
|
| Rights | MS-NC-NoReD |
| Source | META-SHARE |
| Title |
Serbian Morphological Dictionary (Multext-East)
|
| Type | Lexical Conceptual Resource |
| Contact Point | Metashare/936f54fe8bdf11e2bffb0015171445921ba762cd361b47939ea5c40cb0b79fbe#contact Person |
| Contributor | Miloš Utvić |
| Creator | Miloš Utvić |
| Ivan Obradović | |
| Duško Vitas | |
| Description |
This corpus consists of English source texts translated into Serbian, and Serbian source texts translated into English, and several aligned English and Serbian translations of literary texts originally in French. The texts belong to various domains: fiction, general news, scientific journals, web journalism, health, law, education, movie sub-titles. The corpus also contains several Serbian translations of texts from the ‘Acquis communautaire’ corpus and from the ‘Intera’ corpus aligned with their originals. The alignment was performed on the subsentencial level. The texts were segmented and aligned automatically and then manually checked. In most cases the alignment is one-to-one. The size of the corpus is 5,078,280 words (2,672,911 in the English part, 2,405,369 in the Serbian part). More about the content of this corpus can be found at: http://www.korpus.matf.bg.ac.rs/SrpEngKor/SrpEngKor_2013_01.pdf
|
| Language | English |
| Rights | CC-BY-NC |
| Source | META-SHARE |
| Title |
English-Serbian Aligned Corpus
|
| Type | Corpus |
| Contact Point | Metashare/0d68b2f28b3411e2ab9f001517144592e9978ff1de0d4abebd4d6c8935fcb9af#contact Person |
| Contributor | Miloš Utvić |
| Ranka Stanović | |
| Creator | Miloš Utvić |
| Duško Vitas | |
| Ranka Stanković | |
| Description |
Serbian NooJ module (SrpNooJ) was produced in the scope of the EU-funded CESAR project. It consists of a set of resources in both alphabets that are in use for Serbian: Cyrillic and Latin. Each set consists of: the dictionary properties’ definition file (metadata), one text – a novel “Dva carstva” (Two empires) from a Serbian author Branimir Ćosić comprising of 106684 tokens, a sample dictionary in readable form with 35 lemma that belong to 9 grammatical classes, with examples of multiword units and derivational morphology, a sample of morphological grammars used for lemmas from a sample dictionary – three for simple nouns, two for adjectives, two for verbs, and one for a multiunit noun, a readable sample dictionary of inflected forms automatically produced from a sample dictionary of lemmas and a sample morphological grammars, a syntactic grammar for recognition of one class of named entities – full personal names with their roles or functions, a full compiled dictionary (divided in three files: nouns, verbs, and other). It comprises of 85868 entries: nouns (40886), adjectives (25558), verbs (15366), and other (4058).
|
| Rights | CC-BY |
| Source | META-SHARE |
| Title |
Serbian NooJ module
|
| Type | Lexical Conceptual Resource |
| Contact Point | Metashare/fc91787a6b7f11e29f6e000423bfd61cad17bb05bcbd470da8cec4ebdda3481e#contact Person |
| Contributor | Max Silberztein |
| Creator | Mladen Stanojević |
| Description |
NooJ is a linguistic development environment that allows linguists to formalize several levels of linguistic phenomena: typography and spelling; lexicons of simple words, multiword units and discontinuous expressions; inflectional, derivational and productive morphology; local and structural syntax, transformational and semantic analysis and generation. For each of these levels NooJ provides linguists with one formal framework specifically designed to facilitate the description of each phenomenon, as well as parsing/development/debugging tools designed to be as computationally efficient as possible, from Finite-State machines to Turing machines. This approach distinguishes NooJ from other computational linguistic frameworks which provide a unique formalism based on a compromise between power and efficiency. As a corpus processing tool, NooJ allows all researchers and professional to extract information from general or technical corpora by applying sophisticated
queries based on concepts rather than word forms and build indices, add semantic annotations, perform statistical analyses, etc. MONO version of NooJ is operative on all platforms that support MONO.
|
| Rights | MS-NC-NoReD-ND |
| Source | META-SHARE |
| Title |
MONO version of NooJ
|
| Type | Tool Service |