Matches in ScholarlyData for { <https://w3id.org/scholarlydata/inproceedings/lrec2008/papers/313> ?p ?o. }
Showing items 1 to 13 of
13
with 100 items per page.
- 313 creator michael-mohler.
- 313 creator rada-mihalcea.
- 313 type InProceedings.
- 313 label "Babylon Parallel Text Builder: Gathering Parallel Texts for Low-Density Languages".
- 313 sameAs 313.
- 313 abstract "This paper describes Babylon, a system that attempts to overcome the shortage of parallel texts in low-density languages by supplementing existing parallel texts with texts gathered automatically from the Web. In addition to the identification of entire Web pages, we also propose a new feature specifically designed to find parallel text chunks within a single document. Experiments carried out on the Quechua-Spanish language pair show that the system is successful in automatically identifying a significant amount of parallel texts on the Web. Evaluations of a machine translation system trained on this corpus indicate that the Web-gathered parallel texts can supplement manually compiled parallel texts and perform significantly better than the manually compiled texts when tested on other Web-gathered data.".
- 313 hasAuthorList authorList.
- 313 hasTopic Linguistics.
- 313 isPartOf proceedings.
- 313 keyword "Endangered languages".
- 313 keyword "LR Infrastructures and Architectures".
- 313 keyword "Multilinguality".
- 313 title "Babylon Parallel Text Builder: Gathering Parallel Texts for Low-Density Languages".