I am from IstanbulՊոլիս and I live in Karlsruhe. I studied computer science at KIT, specialising in artificial intelligence and natural language processing.
My main interest is Western ArmenianԱրեւմտահայերէն, my mother language. Most people do not know that Armenian is two languages, and the one the world means is Eastern Armenian, because Eastern Armenian has a country. Western Armenian is endangered, and everything holding it up was built by the communities themselves: schools, newspapers, dictionaries. None of that gets it inside the platforms. There is no Western Armenian Windows, no Western Armenian iOS, and no organisation whose job it is to ask for one. A language you cannot use on the thing you hold all day becomes a language for occasions.
My master's thesis was the first neural machine translation system between Western Armenian and English. I wrote it to get the language into NLP research, where it had never been. I have kept building data and models for it since, in whatever time I could take from everything else.
Why this matters, and what exists →I have never worked as a translator. I keep doing the job anyway, usually in the figurative sense. At ValueWorks in Karlsruhe I am a data engineer, which mostly means moving data between systems that were never built to talk to each other. Since 2025 I have also advised the Calouste Gulbenkian Foundation's Armenian Communities Department on technology and AI for Western Armenian. Same work again, between people who build things and people who back them. The one thing I have actually built is a translator, in the literal one.
Where I have worked →The first parallel corpus for the pair, and the first translation model. About 147,000 sentence pairs, collected from websites and printed books across Germany, Turkey and Armenia. 52,879 of them are copyright-free and released. The rest I am not allowed to share.
The thesis is the long version: held-out-domain test sets, a study of whether matching domain or matching language matters more, a negative result on custom segmentation, and a chapter on what the work does not show.
It has kept moving since. The version running on this site saw 1,532,303 pairs, most of them backtranslated from Western Armenian that nobody had ever translated into anything, and you can put a sentence through it here without installing something first.
Next is the Classical Armenian model, which is paused while I argue with a robot about whether it is a Monero mine.