Spell-checking in Spanish: The case of diacritic accents

Jordi Atserias, Maria Fuentes, Rogelio Nazar, IRENE RENAU ARAQUE

Resultado de la investigación: Capítulo del libro/informe/acta de congresoContribución a la conferenciarevisión exhaustiva

2 Citas (Scopus)

Resumen

This article presents the problem of diacritic restoration (or diacritization) in the context of spell-checking, with the focus on an orthographically rich language such as Spanish. We argue that despite the large volume of work published on the topic of diacritization, currently available spell-checking tools have still not found a proper solution to the problem in those cases where both forms of a word are listed in the checker's dictionary. This is the case, for instance, when a word form exists with and without diacritics, such as continuo 'continuous' and continuo 'he/she/it continued', or when different diacritics make other word distinctions, as in continuo 'I continue'. We propose a very simple solution based on a word bigram model derived from correctly typed Spanish texts and evaluate the ability of this model to restore diacritics in artificial as well as real errors. The case of diacritics is only meant to be an example of the possible applications for this idea, yet we believe that the same method could be applied to other kinds of orthographic or even grammatical errors. Moreover, given that no explicit linguistic knowledge is required, the proposed model can be used with other languages provided that a large normative corpus is available.

Idioma originalInglés
Título de la publicación alojadaProceedings of the 8th International Conference on Language Resources and Evaluation, LREC 2012
EditoresMehmet Ugur Dogan, Joseph Mariani, Asuncion Moreno, Sara Goggi, Khalid Choukri, Nicoletta Calzolari, Jan Odijk, Thierry Declerck, Bente Maegaard, Stelios Piperidis, Helene Mazo, Olivier Hamon
EditorialEuropean Language Resources Association (ELRA)
Páginas737-742
Número de páginas6
ISBN (versión digital)9782951740877
EstadoPublicada - 1 ene 2012
Evento8th International Conference on Language Resources and Evaluation, LREC 2012 - Istanbul, Turquía
Duración: 21 may 201227 may 2012

Serie de la publicación

NombreProceedings of the 8th International Conference on Language Resources and Evaluation, LREC 2012

Conferencia

Conferencia8th International Conference on Language Resources and Evaluation, LREC 2012
País/TerritorioTurquía
CiudadIstanbul
Período21/05/1227/05/12

Huella

Profundice en los temas de investigación de 'Spell-checking in Spanish: The case of diacritic accents'. En conjunto forman una huella única.

Citar esto