Version 2.0 · hyw ⇆ eng · NLLB-200 distilled, fine-tuned on 1,532,303 sentence pairs
This is the model from my thesis, running on a shared processor with no graphics card behind it. A short sentence usually comes back in two or three seconds and a paragraph takes longer, though the first request after a quiet spell has to load the model from cold and that one can run to a minute. Nothing is wrong when it does that.
It runs on the free processor Hugging Face hands out, which is why it sleeps between visitors and wakes up slowly, and getting it onto one that stays awake is what the tin is for: PayPal or Buy Me a Coffee.
This is the first machine translation system between Western Armenian and English. It is Meta's No Language Left Behind model, the 600M distilled one, fine-tuned on 1,532,303 sentence pairs and trained in both directions at once. The work started as my master's thesis at KIT under Prof. Jan Niehues and has kept going since.
The corpus behind it runs to about 147,000 real translations, gathered across Germany, Turkey and Armenia from newspapers, school material, printed books, a conversational guide and a handful of constitutions. This model trained on 95,268 of them. The copyright-free part of the collection is published. The remaining 1,437,035 pairs came from backtranslation, which means taking Western Armenian that nobody ever translated, running it backwards through a model to manufacture an English side, and training on the result. The Armenian in those pairs is writing by actual people and the English is machine output, which is the useful way round, since the thing the model most needs to learn is how to produce Armenian.
Ninety-five thousand real pairs is a small foundation, and commercial systems for European languages train on hundreds of millions. You will see the difference in proper names, in dates, in anything carrying a long subordinate clause, and in registers the sources never covered, which mostly means anything technical or legal. It is better on the kind of prose it was built from: journalism, school material and letters.
Nothing you type is stored unless you rate a translation. When you press one of the three buttons, that rating is saved together with your sentence, the translation and the settings it ran with, because rated pairs are where the next version of the corpus comes from. No account, no cookie, nothing else about you.
There is a Classical Armenian model too, trained the same way. It is paused at the moment while I appeal an automated abuse flag that caught it in a sweep for cryptocurrency miners.