Polish Named Entity Recognition Tool

Instance of: Resource Info
Contributor Jakub Waszczuk
Creator Jakub Waszczuk
Michał Lenart
Description Nerf is a statistical tool for Named Entity Recognition (NER) based on the Conditional Random Fields (CRF) modelling method. The tool has been constructed as a part of the National Corpus of Polish project. It has been adapted to recognize tree-like structures of NEs (i.e., with recursivelly embeded NEs) using the Joined Label Tagging (JLT) method. The JLT method is a simple method of encoding NE structures as a sequence of labels. With this method various additional informations about NEs of categorical nature – type, subtype, type of derivation – can be encoded on the level of labels and subsequently recognized using the resultant CRF model. The tool can be configured to use various types of observations during the training and recognition process, for example: lexical informations from textual level, or grammatical informations from morphosyntactic level.
Rights GPL
See Also http://metashare.elda.org/repository/browse/bbe8ee646aff11e284b6000423bfd61cf73819a718ef48b9a9608b54de8bba8e/
Source META-SHARE
Title Polish Named Entity Recognition Tool
Type Dataset
Type Tool Service

Contact Point

Communication Info Metashare/bbe8ee646aff11e284b6000423bfd61cf73819a718ef48b9a9608b54de8bba8e#communication Info
Given Name Jakub
Surname Waszczuk
Type Contact Person
Person
Person Info Type

Distribution Info

Availability Available-restricted Use
License
Delivery Channel Downloadable
Fee free of charge
Permission
Action http://creativecommons.org/ns/ShareALike
Duty Metashare/bbe8ee646aff11e284b6000423bfd61cf73819a718ef48b9a9608b54de8bba8e#permission
Type Duty
Permission
Restrictions Of Use
Same As http://www.gnu.org/copyleft/gpl.html
Type Licence Info
Type Distribution
Distribution Info

Identification Info

Description Nerf is a statistical tool for Named Entity Recognition (NER) based on the Conditional Random Fields (CRF) modelling method. The tool has been constructed as a part of the National Corpus of Polish project. It has been adapted to recognize tree-like structures of NEs (i.e., with recursivelly embeded NEs) using the Joined Label Tagging (JLT) method. The JLT method is a simple method of encoding NE structures as a sequence of labels. With this method various additional informations about NEs of categorical nature – type, subtype, type of derivation – can be encoded on the level of labels and subsequently recognized using the resultant CRF model. The tool can be configured to use various types of observations during the training and recognition process, for example: lexical informations from textual level, or grammatical informations from morphosyntactic level.
Distribution
Access URL http://zil.ipipan.waw.pl/Nerf
Type Distribution
URL
Identifier 404
Meta Share Id NOT_DEFINED_FOR_V2
Resource Short Name Nerf
Title Polish Named Entity Recognition Tool
Type Identification Info

Resource Creation Info

Creation Start Date 2010-03-01 Date
Creator
Person Info
Communication Info
Address Jana Kazimierza 5
City Warsaw
Distribution
Access URL http://zil.ipipan.waw.pl/MichalLenart
Type Distribution
URL
Email michal.lenart@gmail.com
Fax Number +48 22 38 00 510
Telephone Number +48 22 38 00 560
Type Communication Info
Zip Code 01-248
Given Name Michał
Position Programmer
Surname Lenart
Type Person
Person Info Type
Type Actor
Person Info
Communication Info
Address Jana Kazimierza 5
City Warsaw
Email waszczuk.kuba@gmail.com
Fax Number +48 22 38 00 510
Telephone Number +48 22 38 00 566
Type Communication Info
Zip Code 01-248
Given Name Jakub
Surname Waszczuk
Type Person
Person Info Type
Type Actor
Funding Project
Distribution
Access URL http://nkjp.pl/
Type Distribution
URL
Funder Polish Ministry of Science and Higher Education
Funding Country Poland
Funding Type National Funds
Project End Date 2011-06-30 Date
Project Name National Corpus of Polish
Project Short Name NKJP
Project Start Date 2007-01-01 Date
Type Project Info Type
Type Resource Creation Info

Resource Documentation Info

Documentation
Document Unstructured Jakub Waszczuk, Katarzyna Głowińska, Agata Savary, Adam Przepiórkowski, \"Tools and methodologies for annotating syntax and named entities in the National Corpus of Polish\", Proceedings of the 2010 International Multiconference on Computer Science and Information Technology (IMCSIT).
Type Documentation Info Type
Tool Documentation Type Help Functions
Type Resource Documentation Info

Tool Service Info

Input Info
Media Type Text
Modality Type Written Language
Resource Type Language Description
Type Input Info
Language Dependent false Boolean
Output Info
Media Type Text
Modality Type Written Language
Resource Type Lexical Conceptual Resource
Type Output Info
Resource Type Tool Service
Tool Service Creation Info
Formalism Joined Label Tagging
Conditional Random Fields
Implementation Language C
Cython
Python
Type Tool Service Creation Info
Tool Service Evaluation Info
Evaluated true Boolean
Evaluation Criteria Intrinsic
Evaluation Details Cross validation of the Nerf tool on the NKJP corpus, with respect to the NKJP Named Entities hierarchy, yielded F-measure of 79%.
Evaluation Level Diagnostic
Evaluation Measure Automatic
Evaluation Report
Document Unstructured Results (which will be described in the NKJP book; earlier evaluation results can be found in the \"Tools and methodologies for annotating syntax and named entities in the National Corpus of Polish\", Proceedings of the 2010 International Multiconference on Computer Science and Information Technology (IMCSIT) can be acquired by running cross_validate.py script distributed with the Nerf tool. The process is described in details in the README file, also distributed with Nerf.
Type Documentation Info Type
Evaluation Type Black Box
Evaluator
Person Info
Communication Info
Address Jana Kazimierza 5
City Warsaw
Email waszczuk.kuba@gmail.com
Fax Number +48 22 38 00 510
Telephone Number +48 22 38 00 566
Type Communication Info
Zip Code 01-248
Given Name Jakub
Surname Waszczuk
Type Person
Person Info Type
Type Actor
Type Tool Service Evaluation Info
Tool Service Operation Info
Operating System Linux
Running Environment Info
Required Software Cython (version 0.14 or higher)
Python (version 2.6 or higher)
Type Running Environment Info
Type Tool Service Operation Info
Tool Service Type Tool
Type Tool Service Info

Usage Info

Actual Use Info
Actual Use Nlp Applications
Actual Use Details Named entity-based mention detection for the Polish coreference resolution module.
Type Actual Use Info
Usage Project
Distribution
Access URL http://zil.ipipan.waw.pl/CORE
Type Distribution
URL
Funder National Science Centre (100%)
Funding Country Poland
Funding Type National Funds
Project End Date 2014-04-17 Date
Project Name Computer-based methods for coreference resolution in Polish texts
Project Short Name CORE
Project Start Date 2011-04-18 Date
Type Project Info Type
Usage Report
Document Unstructured Ogrodniczuk M., Kopeć M. Rule-based coreference resolution module for Polish. In Proceedings of the 8th Discourse Anaphora and Anaphor Resolution Colloquium (DAARC 2011), pp. 191–200. Faro, Portugal.
Type Documentation Info Type
Use NLPSpecific Parsing
Actual Use Nlp Applications
Actual Use Details Recognition of Named Entities for the UIMA Language Processing Chain in ATLAS CMS.
Type Actual Use Info
Usage Project
Distribution
Access URL http://www.atlasproject.eu
Type Distribution
URL
Funder European Commission (50%)
The Polish Ministry of Science and Higher Education (50%)
Funding Country Poland
Funding Type Eu Funds
National Funds
Project End Date 2013-02-28 Date
Project Name Applied Technology for Language-Aided CMS
Project Short Name ATLAS
Project Start Date 2010-03-01 Date
Type Project Info Type
Usage Report
Document Unstructured Ogrodniczuk M., Przepiórkowski A. Polish Language Processing Chains for Multilingual Information Systems. G. Bouma et al. (ed.): NLDB 2012, LNCS 7337, pp. 152–157. Springer, Heidelberg.
Type Documentation Info Type
Use NLPSpecific Named Entity Recognition
Actual Use Nlp Applications
Actual Use Details Recognition of Named Entities in the National Corpus of Polish. Tool trained on the manually annotated million-word subcorpus has been used to annotate the entire NKJP corpus.
Type Actual Use Info
Usage Project
Distribution
Access URL http://www.nkjp.pl
Type Distribution
URL
Funder The Polish Ministry of Science and Higher Education (100%)
Funding Country Poland
Funding Type National Funds
Project End Date 2011-06-12 Date
Project Name National Corpus of Polish
Project Short Name NKJP
Project Start Date 2007-12-13 Date
Type Project Info Type
Usage Report
Document Unstructured Agata Savary, Jakub Waszczuk and Adam Przepiórkowski. 2010. Towards the Annotation of Named Entities in the National Corpus of Polish. In Proceedings of the Seventh International Conference on Language Resources and Evaluation, LREC 2010, Valletta, Malta. ELRA.
Type Documentation Info Type
Document Unstructured Jakub Waszczuk, Katarzyna Głowińska, Agata Savary and Adam Przepiórkowski. 2010. Tools and Methodologies for Annotating Syntax and Named Entities in the National Corpus of Polish. In Proceedings of the International Multiconference on Computer Science and Information Technology (IMCSIT 2010): Computational Linguistics – Applications (CLA’10), pages 531–539, Wisła, Poland. PTI.
Type Documentation Info Type
Use NLPSpecific Named Entity Recognition
Foreseen Use Info
Foreseen Use Nlp Applications
Type Foreseen Use Info
Use NLPSpecific Named Entity Recognition
Type Usage Info

Version Info

Has Version 0.2
Modified 2011-10-11 Date
Type Version Info

Metashare/bbe8ee646aff11e284b6000423bfd61cf73819a718ef48b9a9608b54de8bba8e#metadata Creator

Instance of: Actor
Is Creator of Metashare/bbe8ee646aff11e284b6000423bfd61cf73819a718ef48b9a9608b54de8bba8e#metadata Info

Metashare/bbe8ee646aff11e284b6000423bfd61cf73819a718ef48b9a9608b54de8bba8e#Header

Instance of: Catalog Record
Issued 2014-09-23T00:39:49Z Date
Primary Topic Polish Named Entity Recognition Tool
Set Spec toolService:tool
toolService