A Rule-Based Approach for Aligning Japanese-Spanish Sentences from A Comparable Corpora
Authors/Creators
Description
The performance of a Statistical Machine Translation System (SMT) system is proportionally directed to the quality and length of the parallel corpus it uses. However for some pair of languages there is a considerable lack of them. The long term goal is to construct a Japanese-Spanish parallel corpus to be used for SMT, whereas, there are a lack of useful Japanese-Spanish parallel Corpus. To address this problem, In this study we proposed a method for extracting Japanese-Spanish Parallel Sentences from Wikipedia using POS tagging and Rule-Based approach. The main focus of this approach is the syntactic features of both languages. Human evaluation was performed over a sample and shows promising results, in comparison with the baseline.
Files
1312ijnlc01.pdf
Files
(138.7 kB)
| Name | Size | Download all |
|---|---|---|
|
md5:058d90866c00d193fc01a13790ec1e0c
|
138.7 kB | Preview Download |
Additional details
Identifiers
Related works
- Is metadata for
- 10.5121/ijnlc.2012.1301 (DOI)
Dates
- Available
-
2012
References
- [1] Adafre, Sisay F. & De Rijke, Maarten, (2006) "Finding Similar Sentences across Multiple Languages in Wikipedia", In Proceeding of EACL-06, pages 62-69. [2] Bunescu, Razvan & Pasca, Marius (2006) "Using Encyclopedic Knowledge for Named Entity Disambiguation", In Proceeding of EACL-06, pages 9-16. [3] Fung, Pascale & Cheung Percy, (2004) " Multi-level Bootstrapping for extracting Parallel Sentences from a quasi-Comparable Corpus", In Proceeding of the 20th International Conference on Computational Linguistics. Pages 350 [4] RamÃrez, Jessica, Asahara, Masayuki & Matsumoto, Yuji , (2008) "Japanese-Spanish Thesaurus Construction Using English as a Pivot", In Proceeding of The Third International Joint Conference on Natural Language Processing (IJCNLP), Hyderabad, India. pages 473-480. [5] Rigau, German, Magnni, Bernardo, Aguirre, Eneko & Carroll, John, (2002) " A Roadmap to Knowledge Technologies", In Proceeding of COLING Workshop on A Roadmap for Computational Linguistics. Taipei, Taiwan. [6] Tillman, Christoph, (2009) . " A Bean-Search Extraction Algorithm for Comparable Data", In Proceeding of ACL, pages 225-228 [7] Tillman, Christoph & Xu, Jian-Ming (2009) "A Simple Sentence-Level Extraction Algorithm for Comparable Data", In Proceeding of HLT/NAACL, pages 93-96.