Matches in ScholarlyData for { <https://w3id.org/scholarlydata/inproceedings/lrec2008/papers/856> ?p ?o. }
Showing items 1 to 13 of
13
with 100 items per page.
- 856 creator alexandre-allauzen.
- 856 creator helene-bonneau-maynard.
- 856 type InProceedings.
- 856 label "Training and Evaluation of POS Taggers on the French MULTITAG Corpus".
- 856 sameAs 856.
- 856 abstract "The explicit introduction of morphosyntactic information into statistical machine translation approaches is receiving an important focus of attention. The current freely available Part of Speech (POS) taggers for the French language are based on a limited tagset which does not account for some flectional particularities. Moreover, there is a lack of a unified framework of training and evaluation for these kinds of linguistic resources. Therefore in this paper, three standard POS taggers (Treetagger, Brill s tagger and the standard HMM POS tagger) are trained and evaluated in the same conditions on the French MULTITAG corpus. This POS-tagged corpus provides a tagset richer than the usual ones, including gender and number distinctions, for example. Experimental results show significant differences of performance between the taggers. According to the tagging accuracy estimated with a tagset of 300 items, taggers may be ranked as follows: Treetagger (95.7%), Brill s tagger (94.6%), HMM tagger (93.4%). Examples of translation outputs illustrate how considering gender and number distinctions in the POS tagset can be relevant.".
- 856 hasAuthorList authorList.
- 856 hasTopic Linguistics.
- 856 isPartOf proceedings.
- 856 keyword "Machine Translation, SpeechToSpeech Translation".
- 856 keyword "Tagging".
- 856 keyword "Validation of LRs".
- 856 title "Training and Evaluation of POS Taggers on the French MULTITAG Corpus".