Irish NLP Grammar

Finite state and Constraint Grammar based analysers, proofing tools and other resources

View the project on GitHub giellalt/lang-gle

Page Content

TTS tokenisation for smj

Requires a recent version of HFST (3.10.0 / git revision>=3aecdbc) Then just:

make
echo "ja, ja" \
| hfst-tokenise --giella-cg tokeniser-disamb-gt-desc.pmhfst

More usage examples:

echo "Juos gorreválggain lea (dárbbašlaš) deavdit gáibádusa \
boasttu olmmoš, man mielde lahtuid." \
| hfst-tokenise --giella-cg tokeniser-disamb-gt-desc.pmhfst
echo "(gáfe) 'ja' ja 3. ja? ц jaja ukjend \"ukjend\"" \
| hfst-tokenise --giella-cg tokeniser-disamb-gt-desc.pmhfst
echo "márffibiillagáffe" \
| hfst-tokenise --giella-cg tokeniser-disamb-gt-desc.pmhfst

Pmatch documentation: https://kitwiki.csc.fi/twiki/bin/view/KitWiki/HfstPmatch

Characters which have analyses in the lexicon, but can appear without spaces before/after, that is, with no context conditions, and adjacent to words:

Whitespace contains ASCII white space and the List contains some unicode white space characters

Apart from what’s in our morphology, there are 1) unknown word-like forms, and 2) unmatched strings We want to give 1) a match, but let 2) be treated specially by hfst-tokenise -a

TODO: Could use something like this, but built-in’s don’t include šžđčŋ:

Simply give an empty reading when something is unknown: hfst-tokenise –giella-cg will treat such empty analyses as unknowns, and remove empty analyses from other readings. Empty readings are also legal in CG, they get a default baseform equal to the wordform, but no tag to check, so it’s safer to let hfst-tokenise handle them.

Needs hfst-tokenise to output things differently depending on the tag they get


This (part of) documentation was generated from tools/tokenisers/tokeniser-tts-cggt-desc.pmscript

Sitemap

Debugging site.pages:

URL: /assets/css/style.css - Title:

URL: /Links.html - Title:

URL: /gle.html - Title: Irish language model documentation

URL: /gramcheck/ - Title: Grammar checker project

URL: /index-header.html - Title: Irish documentation

URL: / - Title: Irish documentation

URL: /src-cg3-functions.cg3.html - Title:

URL: /src-fst-morphology-affixes-nouns.lexc.html - Title:

URL: /src-fst-morphology-affixes-prefixes.lexc.html - Title:

URL: /src-fst-morphology-affixes-propernouns.lexc.html - Title:

URL: /src-fst-morphology-affixes-symbols.lexc.html - Title: Symbol affixes

URL: /src-fst-morphology-affixes-verbs.lexc.html - Title:

URL: /src-fst-morphology-phonology.nounadj.xfscript.html - Title:

URL: /src-fst-morphology-phonology.twolc.html - Title:

URL: /src-fst-morphology-phonology.verb.xfscript.html - Title:

URL: /src-fst-morphology-root-adj.lexc.html - Title:

URL: /src-fst-morphology-root-noun-all.lexc.html - Title:

URL: /src-fst-morphology-root-others.lexc.html - Title:

URL: /src-fst-morphology-root-verb-all.lexc.html - Title:

URL: /src-fst-morphology-root.lexc.html - Title: Irish morphological analyser !

URL: /src-fst-morphology-stems-abbreviations.lexc.html - Title:

URL: /src-fst-morphology-stems-adjectives.lexc.html - Title:

URL: /src-fst-morphology-stems-adpositions.lexc.html - Title:

URL: /src-fst-morphology-stems-adverbs.lexc.html - Title:

URL: /src-fst-morphology-stems-articles.lexc.html - Title:

URL: /src-fst-morphology-stems-conjunctions.lexc.html - Title:

URL: /src-fst-morphology-stems-determiners.lexc.html - Title:

URL: /src-fst-morphology-stems-interjections.lexc.html - Title:

URL: /src-fst-morphology-stems-numerals.lexc.html - Title:

URL: /src-fst-morphology-stems-particles.lexc.html - Title:

URL: /src-fst-morphology-stems-pronouns.lexc.html - Title:

URL: /src-fst-morphology-stems-propernouns.lexc.html - Title:

URL: /src-fst-morphology-stems-punctuations.lexc.html - Title:

URL: /src-fst-morphology-stems-tags.lexc.html - Title:

URL: /src-fst-morphology-stems-tobar.lexc.html - Title:

URL: /src-fst-morphology-stems-verbalnouns.lexc.html - Title:

URL: /src-fst-morphology-stems-verbs.lexc.html - Title:

URL: /src-fst-orthography-urucaps.xfscript.html - Title:

URL: /src-fst-phonetics-txt2ipa.xfscript.html - Title:

URL: /src-fst-transcriptions-transcriptor-abbrevs2text.lexc.html - Title:

URL: /src-fst-transcriptions-transcriptor-numbers-digit2text.lexc.html - Title:

URL: /tools-grammarcheckers-grammarchecker.cg3.html - Title:

URL: /tools-tokenisers-tokeniser-disamb-gt-desc.pmscript.html - Title: Tokeniser for gle

URL: /tools-tokenisers-tokeniser-gramcheck-gt-desc.pmscript.html - Title: Grammar checker tokenisation for gle

URL: /tools-tokenisers-tokeniser-tts-cggt-desc.pmscript.html - Title: TTS tokenisation for smj

Root items:

URL: /Links.html - Title: Links

URL: /gle.html - Title: Irish language model documentation

URL: /gramcheck/ - Title: Grammar checker project

URL: /index-header.html - Title: Irish documentation

URL: / - Title: Irish documentation

URL: /src-cg3-functions.cg3.html - Title: Src-cg3-functions.cg3

URL: /src-fst-morphology-affixes-nouns.lexc.html - Title: Src-fst-morphology-affixes-nouns.lexc

URL: /src-fst-morphology-affixes-prefixes.lexc.html - Title: Src-fst-morphology-affixes-prefixes.lexc

URL: /src-fst-morphology-affixes-propernouns.lexc.html - Title: Src-fst-morphology-affixes-propernouns.lexc

URL: /src-fst-morphology-affixes-symbols.lexc.html - Title: Symbol affixes

URL: /src-fst-morphology-affixes-verbs.lexc.html - Title: Src-fst-morphology-affixes-verbs.lexc

URL: /src-fst-morphology-phonology.nounadj.xfscript.html - Title: Src-fst-morphology-phonology.nounadj.xfscript

URL: /src-fst-morphology-phonology.twolc.html - Title: Src-fst-morphology-phonology.twolc

URL: /src-fst-morphology-phonology.verb.xfscript.html - Title: Src-fst-morphology-phonology.verb.xfscript

URL: /src-fst-morphology-root-adj.lexc.html - Title: Src-fst-morphology-root-adj.lexc

URL: /src-fst-morphology-root-noun-all.lexc.html - Title: Src-fst-morphology-root-noun-all.lexc

URL: /src-fst-morphology-root-others.lexc.html - Title: Src-fst-morphology-root-others.lexc

URL: /src-fst-morphology-root-verb-all.lexc.html - Title: Src-fst-morphology-root-verb-all.lexc

URL: /src-fst-morphology-root.lexc.html - Title: Irish morphological analyser !

URL: /src-fst-morphology-stems-abbreviations.lexc.html - Title: Src-fst-morphology-stems-abbreviations.lexc

URL: /src-fst-morphology-stems-adjectives.lexc.html - Title: Src-fst-morphology-stems-adjectives.lexc

URL: /src-fst-morphology-stems-adpositions.lexc.html - Title: Src-fst-morphology-stems-adpositions.lexc

URL: /src-fst-morphology-stems-adverbs.lexc.html - Title: Src-fst-morphology-stems-adverbs.lexc

URL: /src-fst-morphology-stems-articles.lexc.html - Title: Src-fst-morphology-stems-articles.lexc

URL: /src-fst-morphology-stems-conjunctions.lexc.html - Title: Src-fst-morphology-stems-conjunctions.lexc

URL: /src-fst-morphology-stems-determiners.lexc.html - Title: Src-fst-morphology-stems-determiners.lexc

URL: /src-fst-morphology-stems-interjections.lexc.html - Title: Src-fst-morphology-stems-interjections.lexc

URL: /src-fst-morphology-stems-numerals.lexc.html - Title: Src-fst-morphology-stems-numerals.lexc

URL: /src-fst-morphology-stems-particles.lexc.html - Title: Src-fst-morphology-stems-particles.lexc

URL: /src-fst-morphology-stems-pronouns.lexc.html - Title: Src-fst-morphology-stems-pronouns.lexc

URL: /src-fst-morphology-stems-propernouns.lexc.html - Title: Src-fst-morphology-stems-propernouns.lexc

URL: /src-fst-morphology-stems-punctuations.lexc.html - Title: Src-fst-morphology-stems-punctuations.lexc

URL: /src-fst-morphology-stems-tags.lexc.html - Title: Src-fst-morphology-stems-tags.lexc

URL: /src-fst-morphology-stems-tobar.lexc.html - Title: Src-fst-morphology-stems-tobar.lexc

URL: /src-fst-morphology-stems-verbalnouns.lexc.html - Title: Src-fst-morphology-stems-verbalnouns.lexc

URL: /src-fst-morphology-stems-verbs.lexc.html - Title: Src-fst-morphology-stems-verbs.lexc

URL: /src-fst-orthography-urucaps.xfscript.html - Title: Src-fst-orthography-urucaps.xfscript

URL: /src-fst-phonetics-txt2ipa.xfscript.html - Title: Src-fst-phonetics-txt2ipa.xfscript

URL: /src-fst-transcriptions-transcriptor-abbrevs2text.lexc.html - Title: Src-fst-transcriptions-transcriptor-abbrevs2text.lexc

URL: /src-fst-transcriptions-transcriptor-numbers-digit2text.lexc.html - Title: Src-fst-transcriptions-transcriptor-numbers-digit2text.lexc

URL: /tools-grammarcheckers-grammarchecker.cg3.html - Title: Tools-grammarcheckers-grammarchecker.cg3

URL: /tools-tokenisers-tokeniser-disamb-gt-desc.pmscript.html - Title: Tokeniser for gle

URL: /tools-tokenisers-tokeniser-gramcheck-gt-desc.pmscript.html - Title: Grammar checker tokenisation for gle

URL: /tools-tokenisers-tokeniser-tts-cggt-desc.pmscript.html - Title: TTS tokenisation for smj

Directory items: