Anthropology and artificial intelligence
The AI That Still Does Not Know How to Listen

In Accra, the Pan African AI & Innovation Summit shifted part of the debate about African languages. The issue is not only which languages AI understands, but what form of communication it expects to encounter.
The AI That Still Does Not Know How to Listen
In Accra, the Pan African AI & Innovation Summit shifted part of the debate about African languages. The issue is not only which languages an artificial intelligence system understands, but what form of communication it expects to encounter. Much of today’s AI assumes someone who writes, reads and remains connected. In many contexts, linguistic life works differently.
During the second day of the Pan African AI & Innovation Summit 2026 in Accra, one intervention compressed into a few words a problem that often disappears behind discussions of models, computing capacity and digital sovereignty. The day focused on talent and impact, with sessions on skills, African-built solutions and applications in areas such as health, agriculture and climate.
“NLP in Ghana is not a text problem. It is a voice problem.”
The reasoning that followed was straightforward. A language model in a local language may have limited practical value if the people it is meant to reach do not primarily use that language through writing. Some of those who could benefit most from these technologies listen and speak, but do not necessarily read. From that perspective, the starting point is no longer adding another language to a text-based system. It is designing from voice.
The observation sounds modest until the conditions on which much contemporary artificial intelligence is built are made visible. Typing a prompt into a box assumes literacy, a screen, a keyboard or touch interface, familiarity with the logic of querying a system and, usually, enough connectivity to sustain a conversation with remote servers. For some of those designing digital systems, these conditions have become so ordinary that they can disappear from view.
The interface is not neutral. It contains an idea of the person expected to use it.
Before the model come the data
The same presentation showed how linguistic inequality precedes the model itself. According to the speaker, Whisper, one of the best-known speech-recognition systems, was trained on around 680,000 hours of audio. The local project being presented had 16.
“How do you compete?”, the speaker asked.
The answer was not to wait until an equivalent corpus existed. The project had begun producing audio, text and translation pairs and releasing them openly so that other teams could reuse them. Every hour released was an hour the next project would not have to create from scratch.
The distance between 680,000 and 16 hours gives the word scarcity another meaning. A language can be intensely present in conversations, markets, homes, radio stations, celebrations and family networks and still appear data-poor to a machine. Linguistic life and computational availability are not the same thing.
For voice to become training material, it has to be recorded. It may then need transcription, labelling, classification, storage and permission for reuse. What does not pass through that chain remains outside much of the statistical representation of language.
The afternoon added another layer. During a presentation on building a technology company in Ghana, a speaker observed that much of the data feeding AI comes from the internet, described there as predominantly English-speaking and predominantly Western. The same intervention moved directly into everyday multilingualism. A person may use several languages during a single day and change not only language but tone depending on where they are and whom they are speaking to. Translating that movement into English and then back into another language may preserve some information while losing precisely what gave the exchange its social meaning.
A joke is no longer quite the same joke. A way of addressing someone can lose kinship, distance, respect, irony or intimacy. Language is not merely an interchangeable container for meaning.
Data also have context
This difficulty touches a central concern for the anthropology of artificial intelligence. Large models are extraordinarily good at identifying regularities. People, however, do not speak only through grammatical regularities.
They change register. They switch languages. They shorten a sentence because the other person already knows the history behind it. They use an expression whose meaning depends on neighbourhood, generation, kinship or shared experience. They speak differently in an institution, at home or while buying something in the street.
When all of that activity is transformed into normalised text, part of the social structure that accompanied the words can disappear.
Increasing the number of supported languages therefore does not automatically solve the problem. An interface can offer a hundred languages while continuing to imagine the same kind of use. A person writes, the machine processes, and another sequence of text comes back. Multilingualism then becomes a property of the system, but not necessarily a property of the life the system is trying to represent.
Accra suggested another possibility.
Offline before universal
During the Hack-AI-Thon, Haska AI presented an application designed to connect producers of cocoa-industry waste with possible uses and buyers. The presentation emphasised that its models work fully offline and that the application is available in several local languages. The explanation was deliberately practical. Someone in a remote area should not need internet access or be forced to use an English-language mode in order to interact with the tool.
It was not the only project to make offline operation part of its design. Another proposal for medical laboratories also argued for models that could work without internet access, specifically with resource-limited facilities in mind. The repetition matters. Offline operation did not appear as a secondary feature added after the main product. It formed part of the design conditions.
That shift is significant. A technology can present itself as universal because it works in any browser while still depending on conditions that are far from universal.
The judges eventually expressed the criterion in almost anthropological terms. When assessing the proposals, they asked whether a solution could fit into everyday life without forcing people to reorganise their practices simply to accommodate AI. Mobile money and text messaging were offered as examples. A technology is easier to adopt when it enters through behaviour that people have already learnt.
The difference may appear small, but it reverses the usual direction of adaptation. Instead of asking people to learn how artificial intelligence wants to be used, the system begins by learning how the people meant to use it already live.
A machine that listens
The promise of African AI is often framed around local models, data centres, computing capacity, investment and technological sovereignty. All of these were present in Accra and all will matter. Yet the conversations on the second day also exposed a layer much closer to everyday experience.
Who can speak to a machine.
In which language.
In what way.
With what connection.
And how much meaning survives when that voice becomes data.
Building voice-first systems that can operate offline and function in multilingual environments does not mean developing a reduced version of dominant interfaces. It may mean beginning from a different description of reality.
The artificial intelligence we know has become accustomed to waiting for a written prompt. In many places, learning to listen may be a considerably deeper innovation.
Sources
- Pan African AI & Innovation Summit. (2026). Official Programme — Pan African AI & Innovation Summit 2026. Accra, Ghana, 22–23 September 2026.
- Pan African AI & Innovation Summit. (2026, 23 September). Day Two livestream. YouTube. https://www.youtube.com/watch?v=gx6IRu31sUM
- AIthropology Lab transcription of the second day, produced from the official livestream.