Armenian is two standard languages, and almost everything on this page follows from that one fact. The two came out of two nineteenth century centres. Western Armenian was written down in Istanbul and took the speech of that city as its basis, while Eastern Armenian was standardised by the Armenian institutions of Tbilisi, which was where the schools and the presses and the money were, although the dialect it settled on was the Ararat one spoken around Yerevan rather than Tbilisi's own. They have stayed close enough that speakers can mostly follow one another and drifted far enough apart that a system trained on one produces wrong output for the other.
The verbs are formed differently, a good deal of the vocabulary differs, the consonants came out differently, and the spelling has been separate since Eastern Armenian took the Soviet orthographic reform in the 1920s while Western Armenian kept the classical orthography. If you want that properly, Seyfarth and colleagues set the two sound systems side by side with recordings, and the Western Armenian morphological transducer paper is good on the verb paradigms and on the orthography split.
UNESCO's Atlas of the World's Languages in Danger listed Western Armenian in Turkey as definitely endangered, which on that scale means children no longer learn the language as a mother tongue at home. That atlas is from 2010 and has since been replaced, so it is a snapshot rather than a live rating. Ethnologue puts speakers at around 1.58 million, scattered across the globe, and rates the language threatened, which is a step milder. Reporting on the community in Istanbul suggests the number using it daily is very much smaller. All of those can be true at once, and the distance between them is most of the problem.
The language was given its own ISO code in 2018,
hyw, separate from hye. That was a real step and a larger one than it
sounds, because until then no piece of software had any way even in principle to tell the two
apart, and everything that has become possible since, from language identification to labelled
datasets to a locale entry, rests on that code existing. It is also only the first step. Eight
years on, the code is still barely used by the systems that would have to use it.
There is no Western Armenian product. No operating system offers it as a language you can choose, no phone lists it, no translation service has it, and none of that is surprising for a language with no state behind it and its speakers scattered. What there is instead, and what makes this more interesting than a list of absences, is Western Armenian pieces sitting inside other people's products, put there once by somebody and then left alone, while everything that has gone on being maintained was built for Armenian in general, and Armenian in general means Eastern Armenian. Start with typing, since it is the part anybody can check in a minute, and since everything downstream is made out of the characters it produces.
Windows ships four Armenian keyboard layouts, and it is worth knowing what each of them is, because the choice between them is the whole problem in miniature.
The two oldest are marked Legacy, Eastern and Western, and both date from the Windows XP era. They are phonetic layouts, one for each half of the language. Press B on the eastern one and you get բ. Press it on the western one and you get պ. They differ on eight keys out of twenty-six, four swapped pairs, three of which are the consonant shift and the fourth of which is ւ against վ, which is a distinction only somebody writing Western Armenian orthography would think to make. Whoever built these knew the language.
Nobody has been back to either of them since. Neither can produce the verchaged, the mark that ends every Armenian sentence, and both put six of the ten digits out of reach, because the number row went to letters and nobody ever put the digits anywhere else, which is survivable at a desk with a numeric keypad and useless on a laptop. Their punctuation sits where it interrupts a hand rather than where one reaches for it.
The default is Armenian Phonetic, which arrived with Windows 8 in 2012. It is the same idea rebuilt, with the digits left where digits belong, and it is eastern. Nayiri's own instructions describe it as the predominant layout in the Republic of Armenia, roughly corresponding to Eastern Armenian. So the keyboard you get by not choosing is an Eastern Armenian one.
The fourth, also from 2012, is Armenian Typewriter, and it is not QWERTY-shaped at all. It is the traditional arrangement: the top row spells ՃՓԲՍՄՈ and the home row runs through ՄՈՒԿԸ and ԱՆԻ, which is where Armenians get their names for it. It shares not one letter position with any of the other three. It is also the only one of the four that can type every Armenian punctuation mark, and it is the one nobody uses, because it means learning thirty-eight positions that bear no relationship to anything printed on the hardware in front of you, and laptops ship QWERTY or QWERTZ and you cannot go out and buy anything else.
Which leaves a Western Armenian speaker choosing between two of the four, and the choice is worse than it looks, because Western Armenian's sound values are not the eastern ones. Classical Armenian had three series of stops and Western Armenian reduced them to two, so the letters an eastern or a classical reader sounds as b, d and g come out of a western mouth aspirated, as p, t and k, while the letters read as p, t and k in the east are voiced in the west.
So on the default layout, a Western Armenian speaker who wants to write the sound b has to press P. Nayiri puts it in exactly those terms and then explains how to go and change the layout, which is the entire situation contained in one sentence of a help page. Razmig, who had a whole keyboard printed with Armenian legends rather than live with any of it, set out the three ways through: memorise where the letters are, use the roughly phonetic layout, or stick labels on every key. Labels are the cheapest, you can buy them ready made in western and eastern versions, and they solve the part of the problem that goes away by itself in a month anyway. I took the middle option and learned the phonetic layout, as most people do, and I have never made peace with it, because I still cannot look at a key with a B printed on it and think պ. You learn it anyway and after a while you stop noticing, which is what people do and is one of the reasons none of this ever gets reported to anybody.
There is a complication that stops this from being a straightforward complaint, and it is the more interesting half of it. Classical Armenian is on the phonetic layout's side. բ really was b, the eastern values are the older ones, and so by the letter of the earlier language the keyboard is correct and it is the western speakers who moved. That argument is sound and it settles nothing, because you do not get to tell a living language how it ought to sound, and a century and a half of people saying it the other way is not something anybody overrules with a keyboard driver.
So the two options come to this. Use the default and type against your own pronunciation, or switch to the layout with your name on it and lose the ability to end a sentence.
Which is the part that turns this from neglect into something sadder. The western phonetic layout is not missing. It is Armenian Western (Legacy), it has shipped since XP, and somebody who knew the language built it carefully enough to get ւ and վ the right way round. And when Microsoft came back to Armenian in 2012 and rebuilt the phonetic layout properly, with the digits where digits belong, they rebuilt the eastern one and left the western one exactly as it was, missing verchaged and all. So the ask is not a new keyboard. It is this one brought up to the standard of the thing that replaced its eastern twin, which is a few swapped pairs away from something that has already shipped for over a decade.
Whichever layout somebody settles on, the text that comes out carries the choice with it, and that is where a keyboard stops being a personal inconvenience and becomes a data problem.
Armenian ends its sentences with ։, the verchaged,
U+0589. That is not a colon. A colon is :,
U+003A, and at reading size on a screen almost nobody can tell them apart. The
phonetic layout does have the verchaged, on Shift + ;, which is where an English
keyboard puts a colon. That is a reasonable place for it if you already type in English and
think of the mark as a colon-shaped thing, and Armenian has neither a colon nor a semicolon, so
from inside the language there is nothing whatsoever to lead you there. The legacy layouts do
not have it at all and leave the ASCII colon sitting in its place. Between those two facts, real
Western Armenian text is full of U+003A where U+0589 belongs, and
every sentence splitter that meets it does the wrong thing.
The Armenian apostrophe ՚, U+055A, is on neither of the two
layouts people use, and Western Armenian needs it in order to conjugate, because the tense
particle elides before a vowel: կ՚ըսեմ,
կ՚երթամ, կ՚ուտեմ. Since it is not available
people reach for whatever is, so in real text you find a straight quote
' at U+0027, a curly one ’ at
U+2019, and the emphasis mark ՛ at U+055B,
which is a different mark doing a different job, all standing in for a character that already
exists. Four spellings of the same word, a model that has to learn all four, and a search that
finds one of them.
Then there is the dot, which is the case I find most instructive, because nobody did anything
wrong and the result is still a mess. Armenian has a second stop, the michaged, which looks like
a full stop and works roughly like an English colon. Unicode never gave it its own codepoint.
Instead the annotation on U+2024, ONE DOT LEADER, a character otherwise meant for
the dotted line in a table of contents, reads
also used as an Armenian semicolon
(mijaket), spelling the name the eastern way, which is a small illustration of the whole
thing. Microsoft's phonetic layout follows the annotation to the letter and puts
U+2024 on the full stop key. So the layout is doing what the standard says.
Everybody else in the world types U+002E. The two look identical and no tokeniser
or sentence splitter written for anything else is looking for the first one, so I strip it out
of every corpus I build.
Somebody did notice. In 2017 a thread on the Unicode mailing list argued that the unification was wrong and that the michaged deserved its own codepoint, and one of the editors agreed the original justification was thin and invited a formal proposal. Nobody filed one. The codepoint that was free at the time was given to a letter instead, and the question has sat open ever since. This is what a language without an address looks like from the inside: a real orthographic question, raised by the right people in the right forum, and then nine years of nothing, because filing the proposal was nobody's job.
On the phone there is less to examine, because Apple publishes nothing. iOS offers one Armenian keyboard, called Armenian, with no western and no phonetic option, and it is the desktop arrangement compressed into four rows. macOS does ship an input source called Armenian Western QWERTY, so somebody at Apple made the distinction once, though since Apple publishes no layout data I cannot tell you what it produces. The phone has nothing by that or any other name. Armenian does not appear anywhere on Apple's own list of which languages get dictation, Live Text, translation or the newer assistant features. When people complain they complain into a support thread that collects agreement for a few years and is then closed without an answer. What you do instead is install somebody else's keyboard. I use Armenian Keyboard Extension. Nayiriboard is the one built specifically for Western Armenian, and it is still on the version released in September 2020.
Translation is a stranger case. Google, DeepL, Microsoft and Yandex all support Armenian, and
every one of them supports exactly one Armenian, hy, built on the eastern standard.
Not one has ever shipped hyw, eight years after the code was created, and Western
Armenian was not among the
hundred and ten languages Google
added in 2024. In my experience they make a reasonable attempt at Western Armenian going into
English, because most of the vocabulary is shared and they were trained on Armenian without
anybody worrying about which one. Coming back the other way you get Eastern Armenian, because
nothing in them knows there is a difference.
The large language models manage it a good deal better, well enough that I use them for it,
which is a pleasant surprise and an opaque one, because I have no idea how any of these
companies classify Western Armenian internally, whether it is a language to them or a dialect or
an accident of the training data, and there is nobody I could ask.
And underneath all of it there is a layer nobody sees until they go looking for it. Software does not work out for itself how to write a date, or whether the thousands separator should be a comma or a space, or how to sort a list of names. It looks all of that up in CLDR, a shared database that Apple, Google, Microsoft, Java, Python and every browser ship a copy of. Western Armenian is not in CLDR. There are three lines of metadata and no data at all, and because nobody has ever declared what it should fall back to, a system asked to format something in Western Armenian falls back to American English instead of to Armenian, so as things stand, labelling your text correctly makes it behave worse than labelling it wrongly.
Getting into CLDR would not put Western Armenian on anybody's phone, and I do not want to sell it as more than it is. It is plumbing, and it fixes the plumbing only. But it is also the thing every other conversation starts from, because a language that is not in there is a language the rest of the stack has no way of being correct about, and it is the first thing anybody checks.
There is a consequence of all this that matters more than any single defect. A great deal of everyday Western Armenian is not typed in Armenian letters at all. People write it in Latin script in messages, because the keyboard is a fight and the person at the other end can read it anyway. That text is not searchable as Armenian, does not register to a language identifier as Armenian, and cannot be collected as Armenian by anybody at all. So the register in which the language is most alive, which is people talking to each other every day, is the one register invisible to every system that counts. And that is before anybody starts on how to romanise it, where the same man is Mkrtich, Mgrdich, Meguerditch and Mıgırdiç depending on which country wrote his name down.
Every one of those is small on its own, and the reason to add them up is that somebody has to live inside the total. Think about what it takes for a seventeen year old in Marseille to write one paragraph in Western Armenian and put it somewhere. He installs somebody else's keyboard, because the one his phone shipped with has no western option in it. If the one he finds is western he is lucky, and if it is not he learns a layout whose letters do not match the sounds he makes. He writes a sentence and the mark that ends it is the wrong character, so nobody searching for that sentence later will find it. There is a spellchecker for Armenian and it does not work with current Office. There is predictive typing in one of the keyboards and it has not been rebuilt since 2020. So he proofreads his own orthography as he goes, in a language he was taught on Saturday mornings. And when he goes looking for something else in this language, the results are thin, thinner than in French, which was already thinner than in English. In French every one of those steps is free and invisible. In Western Armenian each of them costs him something, and the audience at the end is smaller. Given that arithmetic it is not surprising when people give up, and it is not a failure of love or commitment when they do. The frictions are what to work on, because the frictions are the part that can actually be removed.
None of this is hostility, and it matters to say so, because the version of this page where somebody is to blame would be easier to write and would be false. Nobody at Apple has considered Western Armenian and turned it down. Somebody ticked Armenian, shipped Eastern Armenian and closed the ticket, which was a reasonable thing to do with the information in front of them. Platform support moves on market size, on community advocacy and on institutional pressure, and Western Armenian has no state to supply any of the three and nothing standing in for one.
It is worth being precise about what a state actually buys, because Armenia has one and the difference it makes is smaller and stranger than it sounds. When Microsoft last went back to Armenian, in Windows 8, it added two new keyboard layouts and made one of them the default. That is real maintenance, and it is more attention than Western Armenian has had since Windows XP. Both of the new layouts were built around the eastern standard, and the file with Western on its name was left exactly as it was. Somebody was in the room for Armenian. Nobody was in the room for Western Armenian. That is the whole of what having a state bought and what not having one cost: not recognition, not money, just somebody present when a decision was being made.
The missing body is the part I keep coming back to, because it is the part that could be built. A company that wanted to do the right thing here would need somebody to ask. What is the standard orthography. Which locale data is authoritative. Who signs off on a translation of the interface, and who do we send the next question to. Those questions currently have no addressee, and a language that turns up with a standards body and an institution behind it gets taken seriously while a language that turns up as a handful of individuals with good intentions does not. I do not think that is unfair of them so much as unavoidable.
It also explains why the tools that do exist stay half a step behind. They are not missing. Nayiri has had a spellchecker for Windows for years, Nayiriboard is a real keyboard people use, and Calfa's OCR handles Armenian script better than anything commercial does. What none of them has is a maintenance budget. An operating system ships a major version every year and changes its input and text handling whenever it suits it, and work done in the hours people have left over cannot run at that pace, so anything built for this language ends up a version or two behind the system it has to live inside. That is structural rather than anybody's fault, and it is why voluntary effort on its own will never close this.
It is worth putting that next to what this community did once before, which was harder. In the eighteenth and nineteenth centuries a large part of the Ottoman Armenian population spoke Turkish as its first language, in some towns having no Armenian at all, and the two thousand or so books printed in Turkish in Armenian letters were written for those readers. Then from the middle of the nineteenth century the community built schools at a scale that is hard to credit now, around two thousand of them with a hundred and seventy thousand pupils by 1912, paid for by guilds that each adopted a named school, by merchants who endowed shops whose income went to the teachers, by a community tax, and by village associations abroad that each sent money home to one school.
The obvious reading of that is not the one the historians support. The schools did not simply turn Turkish-speaking Armenians back into Armenian speakers, and Turkish language publishing in Armenian letters was still growing in the 1890s. What changed first was a decision, that being Armenian meant knowing Armenian, and then institutions built to act on it, with reality following unevenly behind. That is the part worth taking. The threat now has a completely different shape and none of the methods transfer, but the order might: somebody decides the language is going to be maintained, then builds the thing that maintains it, and the results arrive later and messier than anyone wanted.
The version raised most often is giving Western Armenian official standing in the Republic of Armenia alongside the eastern standard, on the grounds that it would produce a standards body, a budget line and an addressee in a single move. I do not think that survives much examination. Outside the repatriates, almost nobody in Armenia speaks it, and the Syrian Armenians who arrived after 2012 are the largest western-speaking population the country has had in a century while their children are being educated in Eastern Armenian, which is what happens and what anyone would expect. Official standing without speakers is a letterhead. It is also worth noticing what the Armenian state does when it turns to this field at all: a national AI institute launched in 2025 with Amazon and Mistral, and a national AI strategy discussed through 2026, neither of which mentions language technology for Armenian in any variant.
The more useful thing to notice is that an addressee does not have to be a state. CLDR takes locale data from registered organisations, and the ones doing that for small languages are a tribal government, a university department and Wikimedia. What these processes actually require is an institution willing to be responsible and still there in five years. And there is already one proof that Western Armenian can produce that when it has to.
The hyw code did not appear on its own, and it was not won for its own sake. What
people wanted was a Western Armenian Wikipedia, and Wikipedia would not open one for a language
that did not have a code of its own, so the code became the thing standing in the way. An
application was made in 2011 and rejected. A second one was assembled in 2017 by
Wikimedia
Armenia, the Calouste Gulbenkian Foundation, INALCO and Evertype together, and that one did
the unglamorous work: documenting the schools, the newspapers and the university programmes, and
arguing that Western stands to Eastern Armenian as Bokmål stands to Nynorsk rather than as
American English stands to British. It was accepted on the twenty-third of January 2018. Four
organisations, two attempts and seven years, to obtain three letters.
That is the counterexample to this whole section, and it is also the clearest statement of the problem. There was no standing body whose job it was, so four parties improvised one, won the thing, and went back to their own work. Eight years on, not one of the four major translation engines has adopted the code they won.
The field is small but not empty. Ethnologue now rates Western Armenian's digital language support as ascending, which is the middle of a five point scale and the first rung that means anything is happening at all. Every bit of that was earned by the people and projects below, working outside anybody's mandate and mostly outside anybody's budget. This is the inventory I keep, almost all of it other people's work, and if something is missing from it, tell me and I will add it.
hyw. The headline scores on its model card are measured on Eastern Armenian test
sets; the paper's Western figure is 32% word error. Licensed non-commercially.
hyw label, which is what
lets you separate the two Armenians in a pipeline at all.
It is worth looking at languages that got this done, because several have, and the shape of how they did it is consistent enough to be useful.
At Modern coverage in CLDR, which is the top tier. A spellchecker dating to 1992, a public machine translation service, speech recognition at 2.2% word error rate, and an open large language model trained on 4.2 billion tokens. What produced that was not a plan. It was a university research group, Ixa, that opened in 1988 and never closed, now merged into a centre of about sixty people. The named government programme behind it is €1.68 million over three years, which is not much money. The continuity is what did it.
Every one of them a daily user of a state language, which is very likely more than Western Armenian has by that measure rather than fewer. Its government ran a Language Technology Programme from 2019 to 2023 at 2,255 million krónur, delivered by a foundation and a consortium of nine universities, companies and the national broadcaster, with every output open source. It was renewed. Iceland was a launch partner for GPT-4.
A government action plan from 2018 to 2024, delivered largely by one unit at Bangor University, producing a spellchecker that checks a million words a week, the first Welsh transcriber, synthetic voices, and a Welsh Wikipedia that went from a hundred thousand articles to nearly two hundred and eighty thousand. Bangor is a registered organisation inside CLDR, which is how the locale work gets done.
The closest to our situation and the least comfortable. The Cherokee Nation employs language technologists, joined the Unicode Consortium in 2011, and got the script properly encoded, onto Windows in 2012 and Android in 2015. That took about fifteen years and three named people with a department behind them.
Here is the part I cannot argue around. All four had a government. Basque had a regional one, Iceland and Wales national ones, and the Cherokee Nation is a sovereign government with a legislature that passed an act and attached six and a half million dollars to it. The delivery bodies often look private, and a foundation with banks among its founders or a university unit that describes itself as self-funded can be mistaken for a civil society success, but the money is public in all four. There is no example here of a stateless diaspora doing this out of philanthropy. That is a gap in the evidence rather than proof it cannot be done, and this page would be dishonest if it skipped past it.
What is worth taking from them is narrower than statehood, though, and cheaper. In every case the same three things were present: one body that was accountable, a budget line that renewed rather than a grant that ended, and salaried people who could turn up to the same standards process year after year. Iceland bought that for the price of a mid-sized public building. Cherokee has three employees. Basque got there on a research group that simply never shut. None of that requires a country. It requires somebody deciding to be responsible for it and then still being there in five years, which is the thing Western Armenian has never had. It divides into two kinds of work, though in every case above they lived inside one institution rather than two. In money it is a chair and two or three salaried people, a few hundred thousand euros a year, against Iceland's fifteen million over five and the six and a half million dollars the Cherokee Nation attached to an act. The ask is not large. It has never been made.
Somebody to champion it. A body whose job is to be the address: to hold the
standard, to say which orthography and which locale data are authoritative, to campaign for
hyw wherever a language code can go, and to keep asking the platform companies the
same questions until the answers change. The coalition that won the ISO code proved this is
doable, and also showed what a coalition cannot do, which is still be there afterwards. This is
the job that would have filed the michaged proposal in 2018.
What that work looks like in practice, taking the locale as the example, because it is the piece I know best and because it shows how small each step is and how completely the sequence stalls without somebody owning it.
Get hyw into CLDR at all. The cheapest useful piece is one line
saying that Western Armenian inherits from Armenian. That alone stops the fallback to American
English, and every browser, phone and programming language picks it up on its next release,
because they all ship the same library.
Submit the actual locale data. Month names, day names, sort order, plural rules, number and date formats. It is a few days of work for somebody who knows the language. CLDR takes it from registered organisations, which is why the small languages that have managed it did so through bodies like Cherokee Nation, Bangor University and Wikimedia rather than through individuals.
Translate the interface. This is the gate nobody mentions. No vendor will add a language you can select and which then shows you English, so somebody has to translate thousands of strings of operating system text. That is the largest single piece of work in the sequence and the one that most needs organising rather than funding.
Ask each vendor separately. The list of languages you can choose in Settings is hand maintained by each company and does not come from CLDR. Android's is a file in an open source tree, so there is at least a door. Apple does not publish theirs. Then Samsung, and every other manufacturer shipping its own build.
The first two are small. The third and fourth are not, since translating an operating system is thousands of strings and asking every vendor separately never really finishes. But none of the four is difficult. Notice also that three of the four need something handed to them before they can start. Somebody has to have produced the locale data, the translated strings, the evidence that there are speakers and schools and newspapers worth the vendor's time. Campaigning is not a thing you can do with an empty bag, which is the second job.
Somebody to do the seeding. A group, and a university chair would be the right shape for it, doing the production work that nobody gets thanked for. Ixa is the proof that this is the part which compounds: a research group that opened in 1988 and simply kept going, and thirty-five years later Basque has an open large language model. The work divides about like this.
Benchmarks, first, because everything else follows them. The standard test sets that researchers use to compare systems, FLORES for translation and XNLI and TyDi QA and MMLU for reading and reasoning, contain no Western Armenian at all. A language that cannot be scored is a language nobody publishes on, and a language nobody publishes on stays invisible to the field. Translating one of those properly and building a few thousand question and answer pairs would cost a fraction of what it is worth.
Speech, released rather than collected. Tens of hours from tens of speakers, read and conversational, time-aligned, with honest metadata about which variant is being spoken. Some of this exists already and sits behind a login or an unfinished intention to publish, which is the same as not existing for anyone who was not there. Dictation, captions and voice interfaces all need it and nothing substitutes for it.
Corpora and lexicons, with the normalisation decided once and written down. Which apostrophe. Which dot. Whether classical or reformed orthography is the target and what happens to text in the other one. Everybody working on this currently makes those decisions privately and differently, which is why nobody's data combines with anybody else's.
Baselines that stay current. A fine-tune of whatever this year's models are, so there is always something working to measure against and to build on. These expire, which is exactly why it has to be somebody's standing job rather than somebody's thesis. Spellchecking and hyphenation belong here too, for the same reason: they are not hard, they are just never finished.
And everything the campaign has to carry with it. The locale data itself, the interface strings, a written orthographic standard somebody can be pointed at, terminology for the words the language has not needed until now, a style guide, a count of the schools and the newspapers and the speakers. None of it is research and all of it is what the first job needs in order to have anything to say when it walks into a room. The two halves are the same work seen from opposite ends.
Meanwhile there is one thing that costs almost nothing and would help immediately, which is somewhere that records who is working on what. This is a field of a few dozen people scattered across a dozen countries, who meet occasionally at a conference and otherwise have very little idea what the others are doing, and the failure mode is not laziness. It is two of us building the same thing in parallel and neither finding out until it is finished. A directory would do it: one entry per project, with an owner, a contact, a licence and a status. The same logic runs through how the work gets released. Open licence, in the open, documented well enough that the next person extends it instead of starting again, because there are not enough of us to build the same block twice.
All of that is groundwork, and groundwork on its own settles nothing. Latin is in Google Translate and Latin is dead. It has dictionaries, a Wikipedia and more digitised text than Western Armenian will see this century, and nobody grows up speaking it. Resources are not use. They are what makes use cheap, which is a smaller claim and the one I would actually defend. A language that exists as corpora and benchmarks and papers is a language being studied, and research nobody outside the field ever benefits from is a closed loop that only feels like progress. Whatever gets built at the bottom is worth exactly what comes out at the top, and what comes out at the top has to be things an ordinary person can open and use without knowing any of this exists. All of it, the structure and the tools alike, is there for one purpose, which is to take the friction out of using this language in ordinary modern life. Nothing else on this page matters except as a way of getting to that.
So I build applications alongside the corpus work and mean to keep doing both. The reason to put the structure first is not that it matters more. It is that without it, every tool has to be built from nothing by whoever happens to want it, and then dies when that person stops, which is the history of very nearly everything in this language. Get the layer underneath right and the next keyboard, the next spellchecker, the next captioning tool costs a fraction of what the last one cost, and somebody can make it in a weekend rather than a year.
That layer is also the only realistic way into the models people already use. Large models are trained on whatever is on the internet, and the companies building them are not hostile to Western Armenian so much as unable to see it, since what exists is scattered, unlabelled and largely indistinguishable from Eastern Armenian to anything not looking closely. Put clean, correctly labelled, openly licensed Western Armenian where the crawlers will find it and the next generation is better at it without anybody at those companies deciding anything. That costs a great deal less than asking them to care.
Four things follow, and I hold them fairly firmly.
Open licence or it does not count. A resource nobody is permitted to reuse or extend cannot be built on, so whatever it might have become stops at however much time its author had. Non-commercial clauses look careful and are not, since they quietly exclude most of the systems you would actually want this language to end up inside.
Never collapse Western Armenian into Eastern. Not in the labels, not in the evaluation splits. That is precisely how the language went missing from the data in the first place, and every time it happens the next person inherits the mistake.
Digitising something no longer means scanning it. Fifteen years ago producing a set of page images and putting them somewhere findable was a reasonable definition, and a great deal of Armenian material was rescued that way by people who were right to do it. Digital now means machine readable, searchable, correctly labelled, and in a state where it can go into a training pipeline without a week of cleaning first. A project that ships images has produced an archive rather than a resource. That is not a criticism of anybody who scanned. The requirement moved under them, and the work now is to finish what they started.
A score is not evidence until a speaker has read the output. A great deal of low resource work comes down to a score reported by somebody who does not speak the language, on a test set no speaker has ever read. My own thesis reports BLEU on a test set I built myself, so I would not put much weight on those either. The measure is whether a speaker can tell the output is right, and that is slower and more annoying and there is no substitute for it.
Teachers, weekend schools, publishers, Wikipedia editors, lexicographers and archivists have kept this language going for decades and their problem is the whole language, while mine is one slice of it, which is whether it works on a computer and whether artificial intelligence can work with it. What that slice is for, in the end, is that somebody's child in Marseille or Beirut or Los Angeles should be able to use this language on a phone without needing an extraordinary love of it, which at the moment is roughly what it takes.
There is also something the institutions that already teach this language could do without changing what they are, which is to release what they hold. Schools, publishers and cultural associations sit on textbooks, recordings, corrected prose and decades of graded material, all of it produced by native speakers and most of it in a drawer. Openly licensed, that is a corpus. Nobody has to build anything new to make it one.
So if you work on this, or on another language in the same position, or you are simply interested in this kind of thing, write to me and we can see how to help each other. And if you speak the language and want to tell me what is wrong with the corpus or with a translation, that is the most useful mail I get.