Central Asia and the new linguistic infrastructures of AI

In Central Asia, a significant part of the recent expansion of artificial intelligence is not arriving first through large consumer chatbots, but through a more basic and, at the same time, deeper infrastructure: digital dictionaries, speech-recognition systems and linguistic resources capable of turning historically under-represented languages into machine-processable data.

Three recent studies allow us to observe this process from different angles. In Tajikistan, one project proposes using language models to build a digital explanatory dictionary of Tajik. In Kazakhstan, Kyrgyzstan and Uzbekistan, a new automatic speech-recognition model improves machines’ ability to transcribe languages from the region. And in Uzbekistan, a game-based learning proposal turns linguistic practice itself into a mechanism for enriching lexical resources used by AI systems.

Rather than three isolated technical developments, they form a shared sequence: define, listen and learn. They also show that bringing a language into AI means deciding how it is represented, who produces the data that describe it and what new capabilities emerge once a machine begins to process it.

Tajikistan — When an LLM begins to write the dictionary

The first version of Large Language Models for Low-Resource Languages: A Conceptual Framework for an Electronic Explanatory Dictionary of the Tajik Language was published on 4 August. A second version appeared on 6 August. The paper is signed by Mullosharaf K. Arabov, from Kazan Federal University in Russia, and S. S. Pirov and B. Sultonov, from the Tajik National University in Dushanbe.

The proposal sets out a conceptual architecture for developing an electronic explanatory dictionary of Tajik by combining classical lexicographic methods with morphological analysis, lemmatisation, semantic clustering and the generation of entries using large language models.

The starting point is a specific gap. According to the paper, Tajik still lacks a comprehensive digital lexicographic resource comparable in functionality to those available for languages with a stronger computational presence. The proposal aims for the future dictionary to function not only as a tool for people who speak or study the language, but also as infrastructure for machine translation, summarisation, sentiment analysis and other natural-language-processing applications.

A dictionary is never merely a neutral inventory of words. It selects meanings, fixes relationships between terms, establishes variants and participates in the production of a linguistic norm.

When an LLM intervenes in the generation of lexicographic entries, the issue is not only whether the definition it produces is correct. It also matters which corpus was used to construct it, which linguistic uses are represented, which are treated as marginal and how authority is distributed between lexicographers, institutions, speaking communities and generative systems.

In the case of Tajik, there is also a historical layer that is difficult to separate from digital infrastructure. The language has moved through different writing systems during the past century and today officially uses Cyrillic, while maintaining close historical and linguistic ties with Persian and Dari. Digitising a dictionary therefore does not simply mean transferring words into a database. It also means deciding which scripts, variants and historical layers become part of the resource.

It is also worth looking at the institutional genealogy of the work itself. The lead author is based in Kazan, while the two co-authors are affiliated with the Tajik National University. This does not in itself allow conclusions about how authority is actually distributed within the project, but it does remind us that the construction of linguistic infrastructures in Central Asia continues to be shaped by academic networks that extend beyond national borders and, in this case, directly connect Tajikistan with Russia.

The question of meaning therefore appears at two levels at once. It is embedded in the technical architecture, which must decide how definitions are organised and generated, and it is embedded in the institutional network from which that architecture is produced.

What happens to authority over meaning when part of the work of defining a language passes to a model and to transnational academic infrastructures?

Kazakhstan, Kyrgyzstan and Uzbekistan — When machines begin to listen to Central Asia

On 11 July, GigaAM Multilingual: Foundation Model for Underrepresented Languages was published. The paper presents an automatic speech-recognition model designed, among other languages, for Kazakh, Kyrgyz and Uzbek.

The system was pre-trained on approximately two million hours of audio in more than seventy languages and subsequently fine-tuned for speech recognition. The team introduces data-balancing mechanisms intended to reduce one of the recurring problems of multilingual systems: languages with more available data ending up dominating what the model learns.

In the evaluations presented by the team, GigaAM Multilingual improves on the results of open models such as Whisper Large v3 and Omnilingual-1B for the three Central Asian languages studied, with particularly significant gains in spontaneous speech.

Recognising spontaneous speech is anthropologically different from recognising a set of carefully prepared sentences. Everyday conversation contains accents, shifts in register, hesitations, borrowings, forms of respect, language alternation and social markers that rarely appear in the same way in a written corpus.

A machine that can listen more effectively to Kazakh, Kyrgyz or Uzbek expands access to voice interfaces, subtitling, oral archives and accessibility technologies. But that same capacity also reduces the technical cost of turning conversations into structured, analysable text.

Linguistic inclusion and automated listening capacity thus appear as two sides of the same infrastructure. The same technology that can help preserve testimony and widen access to services can also expand the possibilities for recording, classifying and automatically analysing large quantities of speech.

It is equally important to situate who is building this auditory infrastructure. All eight authors of the paper belong to SaluteDevices, linked to the technological ecosystem of the Russian group Sber. GigaAM Multilingual therefore does not emerge from a Central Asian academic or technological institution.

That fact does not invalidate the technical advance, nor does it determine its future uses, but it changes how so-called linguistic inclusion can be read. The capacity to recognise Kazakh, Kyrgyz or Uzbek at scale is being developed through a technological infrastructure located outside those countries. In a region marked by a long history of political, economic and linguistic relations with Russia, the question of who provides the machines’ ears deserves as much attention as their accuracy metrics.

Speech infrastructure allows a language to be handled by systems that previously could barely process it. At the same time, it moves a technical boundary. What was once costly to transcribe can become a stream of text open to classification, search and analysis.

What changes socially and politically when a language moves from being difficult to process automatically to being listenable at scale through infrastructure developed outside the region?

Uzbekistan — Learning a language while building the dataset

UzWordnet and Generative AI for Learning Uzbek by Game Playing, published on arXiv on 6 May 2026, proposes using existing lexical resources and generative AI systems to develop games for learning Uzbek.

The paper is signed by Alessandro Agostini, Saydobid Khusanov and Mirkamol Mirkamilov. Its architecture combines UzWordnet, a large lexical resource for Uzbek, an orthographic dictionary and generative-AI tools. The team designs four educational games with a dual objective. On the one hand, they are intended to support language learning. On the other, they use the dynamics of the games themselves to improve and enrich UzWordnet.

That second element makes the project more interesting than a simple AI-assisted educational application. The interaction of those taking part can contribute to expanding the lexical infrastructure subsequently used by other computational systems.

UzWordnet, moreover, did not begin with this paper. The resource was originally presented in 2021 by Agostini together with T. Usmanov, U. Khamdamov, N. Abdurakhmonova and M. Mamasaidov at the Global WordNet Conference, as the largest Uzbek wordnet built up to that point, with around 28,140 synsets. The 2026 work does not create the infrastructure from scratch. It experiments with a new way of maintaining it, expanding it and putting it into circulation through game mechanics.

The arrangement partly reverses the usual pattern of AI training. People do not appear only as users of a system previously built from large datasets. While playing and learning, they can also produce information that is useful for improving the underlying linguistic resource.

The boundary between learning, participation and data labour therefore becomes less clear.

This model also opens up a different possibility from the large-scale extraction of linguistic information from major platforms. A community of speakers can participate more directly in building the resources that represent its language. But that participation raises questions about ownership, attribution, governance and the eventual destination of the data produced.

The fact that a contribution is voluntary does not make all of those questions disappear. If an answer given within a game ultimately improves a reusable resource for future linguistic systems, that interaction acquires a value different from the simple act of learning.

The proposal is particularly significant because it is applied to an infrastructure with several years of history. This is not about collecting answers for an isolated experiment, but about exploring how the activity of learners can become part of the maintenance and expansion of an existing lexical resource.

When does a cultural or educational activity also become labour for an AI infrastructure, and what return does the community that produces that value receive?

Define, listen, learn

The three cases can be read as layers of the same process.

First, a language must be capable of being defined in a form that allows machines to organise its meanings. It must then be capable of being listened to, even when it appears in spontaneous conversation. Finally, interactions by those who speak or learn it can be used to continue expanding the linguistic infrastructure.

Each layer involves a selection.

The dictionary selects which meanings, variants and scripts are incorporated into a computable norm. Speech recognition selects which accents, registers and forms of speech become processable. The game selects which human interactions become useful information for maintaining and expanding a lexical database.

Taken together, these selections construct a computational representation of Central Asian languages that never automatically coincides with their full social diversity.

A second thread also appears. Two of the three cases contain, in different ways, a Russian connection. The Tajik dictionary project links institutions in Kazan and Dushanbe. GigaAM Multilingual comes directly from SaluteDevices, within the Sber technological ecosystem. The Uzbek project, by contrast, is an international academic collaboration built around a linguistic infrastructure developed over several years.

The point is not to reduce these projects to a single geopolitical explanation. It is to recognise that a language does not enter AI in the abstract. It enters through specific universities, companies, corpora, servers, funding structures and communities.

From this perspective, describing Tajik, Uzbek, Kazakh or Kyrgyz as “low-resource languages” says less about an intrinsic characteristic of the languages themselves than about a historical inequality in the production of digital resources.

The expansion of AI in Central Asia is beginning to alter that situation. The anthropological question, however, does not end when a language finally enters a model. It begins precisely there: who builds its computational representation, who decides what counts as legitimate usage and who can then do new things with the capacity to read, listen to and produce it.


References

Tajikistan

Arabov, M. K., Pirov, S. S. and Sultonov, B. (2026). Large Language Models for Low-Resource Languages: A Conceptual Framework for an Electronic Explanatory Dictionary of the Tajik Language. arXiv:2608.04186. First version, 4 August 2026; v2, 6 August 2026.

Kazakhstan, Kyrgyzstan and Uzbekistan

Kuzmenko, A. et al. (2026). GigaAM Multilingual: Foundation Model for Underrepresented Languages. arXiv:2607.10371. 11 July 2026.

SaluteDevelopers. (2026). GigaAM. Code repository.

Uzbekistan

Agostini, A., Khusanov, S. and Mirkamilov, M. (2026). UzWordnet and Generative AI for Learning Uzbek by Game Playing. arXiv:2607.14104. 6 May 2026.

Alessandro Agostini, Timur Usmanov, Ulugbek Khamdamov, Nilufar Abdurakhmonova and Mukhammadsaid Mamasaidov. (2021). UZWORDNET: A Lexical-Semantic Database for the Uzbek Language. In Proceedings of the 11th Global Wordnet Conference, pp. 8–19. Global Wordnet Association. DOI.


Note on sources

The three main developments have been treated as recent academic publications, rather than as current-affairs news in the journalistic sense. The dates, titles and authorship of the 2026 papers were checked directly on arXiv. In the case of the Tajik dictionary, the first version published on 4 August is distinguished from the revision published on 6 August.

The institutional affiliations of the authors of the Tajik paper are included because they are relevant to understanding the academic networks from which linguistic infrastructure is being built. The origin of GigaAM Multilingual is likewise situated within SaluteDevices and the Sber technological ecosystem. These details describe the institutional production of the projects, but are not used on their own to attribute political intentions or future uses.

UzWordnet is presented as infrastructure that predates the 2026 article. The 2021 reference is used to establish the continuity of the project and to avoid interpreting the new work as the creation of the lexical resource from scratch.

The possibilities of surveillance, value extraction, the fixing of linguistic norms or technological dependency are presented as open anthropological questions arising from the capabilities described, not as uses or consequences demonstrated by the papers themselves.