Depending on your language there are also other parts that are still needed to make the experience smooth. Like inflection and compounding. In comparison finding a small corpus should not be too hard unless it is a very small language or a language not supported right now.
Sure, but for a starter, this would be nice. Afterwards, with an update, the corpus can be added. But if the software works and only the corpus is missing: give it a try. I am from a country of 7 million people. Getting a corpus for it - i donāt know where and how. So, filling up the words while using the phone sounds like a good alternative for me. 
@dexic: for reference, we have Estonian corpus processed for country of 1.4 m people. 7m should be fine - just look for data. Contact some language institute or lab and ask from them what can they propose. That way Iāve got Estonian corpus.
I sent an e-mail to a friend in Belgrade to take a shot. Wish him luck!
About corpus, quite good corpus could be prepared from wikipedia articles. I have used it for OkBoard together with some university corpuses. But I think wikipedia dump alone is also ok.
This topic was touched briefly in the meeting today.
if Presage could be integrated out-of-process, then licensing issues could be addressed
but if Presage has some UTF limitations, there has been another effort by ljo and his team
so hope for better news in the future for all SFOS languages
@sledges (or anyone at Jolla) could you elaborate a bit more in case something can be done so we can have more languages supported in the phone without the need to install extra stuff (or at least be able to install from the official store whatever you need).
Also -unrelated to the above- iād like to add a way of getting a corpus from Wikipedia in case you have trouble finding one.
Download this: GitHub - attardi/wikiextractor: A tool for extracting plain text from Wikipedia dumps
Download a _locale_wiki-latest-pages-articles.xml file from:
https://dumps.wikimedia.org/_locale_wiki/latest/
and run: python3 WikiExtractor.py --infn _locale_wiki-latest-pages-articles.xml
you will get a large .txt file to use as corpus.
On the above substitute locale with your preferred text one. Ie in the case of Czech use cs and so on. (Index of /cswiki/latest/)
If jolla-keyboard could communicate with the Presage engine without linking it directly (e.g. via D-Bus), then Presageās GPLv2 licence shouldnāt be a problem.
How exactly Presage is not Unicode-aware I do not have the details. I had kept in touch with @ljo, but do not have the latest on their effort.
There is DBus service for it - presage/apps/dbus at master Ā· sailfish-keyboard/presage Ā· GitHub. Although, we maybe missing few extra API calls that we used for predictive keyboard. But those should be easy to add (from the project README, looks like just forget will be missing).
Re unicode - it will be needed for some languages, but many could work without in this context. See current supported languages to judge on applicability assuming that similar languages can be supported as well.
If jolla-keyboard could communicate with the Presage engine without linking it directly (e.g. via D-Bus), then Presageās GPLv2 licence shouldnāt be a problem.
But guys! (Just before investing any expensive time in this area).
Did anyone contacted with Matteo Vescovi (orignal presage author) about a licensing change proposal?
If not than I am more than happy to ask him kindly. What license requirement does Jolla have on integration?
I am not into the licensing business, so if anyone could summarise the reason of the change than it would be helpful.
Resurrecting this thread since we got back XT9 with no updates in languages.
Is there a way to add Greek language somehow in predictive text with proper suggestions and corrections?
I dont have the time to ask that in the next community-meeting. But if you can join, it would be a very good question to ask, since there have been some changes through the release of JP and the announcement of Commodore Callback.
Nothing has changed for having it available in one way, the community way. Since for Greek there is nothing stopping you, or someone you trust with access to data, to create a Greek database and a keyboard for the Presage text predictor now immediately.
Which then can be installed via chum repo by other users as well if you make it available. ![]()
But if the question actually is I want something preinstalled when I have selected my User Interface language upon setup of my device, that should be used whenever I use the default Greek keyboard. Then the advances with Commodore could be a step forward. But most probably for the not immediate future. Then compare this uncertainty of future availabilty of otherās efforts to now with a small effort of your own the community way.
Got it, thank you!
I was curious if there was any progress anywhere started already, but it seems we have to group and start something.
Regarding Greek at some point in the past i had found a corpus -greek version of wikipedia- and tried with martonmiklos to make a pressage predictor. There was some issue making the database and the effort stopped there.
Also greek -and many more languages- are missing from the HW keyboard list on sfos. This should be fixed also.
Yes, there can ofcourse be some bumps especially finding the sweet spot of corpus size and text type mix. But I helped several people including MƔrton to get working databases. It is quite straightforward to create the keyboard if you look at a couple of the most similar ones. So if you can recreate the errors/issues you found on the first attempt it would be good.
It was a long time ago but i seem to recall it was a memory issue or something. The corpus was too big to handle or something along these lines.
Edit: now i noticed that the method i used is described in this message above Predictive text for more languages - #28 by ApB
But if jolla has no interest in adopting pressage there is no point doing any languages. First we need to know what the preferred solution is going to be and then do whatever to have all languages in.
Its one of those things that jolla and the community should do together. It will be idiotic wasting resources and doing two times the work.
Yes, I can imagine. Please donāt hesitate to make it yet smaller than belivable
, since it is better to have something small rather than nothing. And it will probably show in the next step, the live testing that you need to make it even smaller ⦠It will learn quickly from what you write anyway.
[Note to self to fint the tweets where I showed this] ![]()
Comment to edit.
Thanks for the link, looks like the generic instruction.
No, this is not double work. All current and also future new resources will be good for testing. Finding the right corpus mix will be required anyway. And Jolla already said they would accept a rewrite with a lgpl3 license The rewrite will also be extended to work for further language families which is the bigger threshold right now.
What i meant by double work was the community working on a solution (ie pressage) and jolla wanting another solution (ie hunspell since it was mentioned).
So first we need to know what the solution will be. And hopefull fix localization once and for all cause it sucks having an open phone and not being able to use it in your native language.
BTW i also seem to recall that i couldnāt find a way to make the wiki corpus smaller. Its been quite a few years and donāt remember all the details around it.