Four hundred points are worth one US dollar.

Once a person reaches 2,000 points, they can request the equivalent of five dollars. To earn them, they must complete language tasks, record their voice and wait for other members of the community to review the result.

The recording must be clear. It must match the task and contain no noise that makes it unusable. When it includes slang, code-switching or a regional expression, someone must decide whether the usage is plausible and whether its context has been described correctly.

Only then do the points cease to be pending.

This is the model DataHive Africa is testing in Tanzania. The platform collects voice, text and human judgements to build datasets for speech recognition, language-model fine-tuning and the cultural evaluation of artificial intelligence systems.

The project begins from a simple observation. Language extends beyond what appears in dictionaries, newspapers and institutions. By enabling Kiswahili speakers to produce, review and contextualise data, the pilot opens a route towards democratising the construction of artificial intelligence from Tanzania.

The Kiswahili That Does Not Enter the Corpus

DataHive Africa currently concentrates on Kiswahili as spoken in Tanzania. It aims to record urban expressions, regional accents, emotional language, everyday legal terminology and forms of code-switching between Kiswahili and English.

These are registers that large corpora assembled from encyclopaedias, newspapers and websites represent inadequately. A system may translate a formal word while failing to understand how people bargain in a market, apologise, express anger or alter a phrase inside a messaging group.

DataHive’s own page illustrates the distance between street and formal registers through “kitu kidogo” and “rushwa”, and publishes fragments of everyday speech that allow readers to hear part of what it intends to collect.

The page reports more than 17 verified contributors, 16 dialect categories and several anonymised samples. It also describes a six-stage procedure beginning with recruitment through community networks and ending with structured files delivered for automatic speech-recognition systems.

No consolidated figures have yet been published for completed tasks, hours of audio, rejection rates or contributor retention.

Between those points are prior consent, task allocation according to linguistic profile, human review, automated quality assessment and provenance records.

The distinction matters.

Collecting a language means accumulating recordings and deciding which people represent a variety, which differences deserve a label, which errors must be corrected and which ways of speaking will be recognised as valid data.

AIthropology Lab takeaway — A corpus does not discover a language as it exists. It organises it through tasks, categories, quality thresholds and decisions about who may speak on its behalf.

Speaking Is Also Work

DataHive describes the people who provide material as contributors and recognises recording, reviewing and contextualising as forms of linguistic work.

According to its founder, each approved task generates points that can be converted into money. The platform uses three official tier names inspired by a beehive — Larva, Worker and Queen Bee — and plans to process withdrawals through local mobile-money services. Junior Makwala has clarified that these integrations are still being developed and are not yet fully operational on the production platform.

Its mobile-first design attempts to reduce a common barrier on international platforms. Many people can contribute knowledge, voice or labour from Tanzania without necessarily holding bank accounts compatible with the payment systems used by foreign companies.

Once the planned integrations are operational, mobile money could make participation more accessible and keep part of the economic circulation surrounding data collection within the local environment.

Points assign a price to each task. The company may later combine thousands of tasks, turn them into a dataset and license it to organisations developing commercial products. The difference between the initial payment and the corpus’s later value is part of the economic model.

Beyond payment, it matters who establishes the relationship between a recording and its points, how long each task requires and how contributors can participate in decisions about the corpus.

The points and tiers organise participation and make progress visible within the platform. Their practical effects will become clearer as the pilot develops and more is known about task times and review practices.

The hive metaphor makes visible a collective form of production in which each recording, review and contextual judgement contributes to a shared corpus. The names Larva, Worker and Queen Bee place each contribution within that common effort.

The Community Reviews the Community

Contributions first pass through peer review.

Reviewers assess audio quality, correspondence with the prompt and the absence of noise that would make the recording unusable. Points remain pending until the material is approved.

A second stage addresses linguistic and cultural content. People with experience in Kiswahili selectively audit slang, regional variation, emotional language and code-switching.

This organisation recognises something automated systems cannot resolve on their own.

A phrase may be technically well recorded and culturally implausible. A translation may appear correct while losing the tone of a reproach, a joke or a negotiation. An expression may be used in one city and sound strange in another.

The knowledge required to identify these differences does not come only from formal linguistic qualifications. It is also formed through everyday participation in a speech community.

DataHive turns that knowledge into a function within its production chain. People generate recordings while also verifying, classifying and contextualising the speech of others.

This distinguishes the project from automated extraction from the web. Speakers act as validators and cultural context becomes part of data quality rather than being reduced to noise or exception.

Community review makes visible that linguistic quality also depends on situated knowledge. The challenge will be to turn this initial participation into increasingly shared and transparent criteria as the project grows.

DataHive uses explicit consent and a public agreement rather than scraping web content or hiding collection behind opaque processes. Before accessing tasks, participants must accept DRA-v1.0; the platform records the document version, the time of consent, the IP address and some device information.

The agreement states in direct language that the platform may use recordings, texts, images and metadata to create datasets sold or licensed to companies and research organisations.

The licence is worldwide, sublicensable and royalty-free beyond the initial payment. For contributions already approved and sold, it is perpetual. The document includes uses in speech recognition, language models, commercial and academic research, and aggregated speech-synthesis and voice-cloning systems.

People may stop participating, request a copy of their contributions and seek deletion of certain data. The agreement nevertheless warns that material already sold to third parties cannot be recalled.

The document’s transparency allows contributors to understand the relationship in advance. Its limits are also stated: material already licensed cannot always be recalled, and a voice may retain identifying characteristics even after names and contact details are removed. As new uses emerge, it will be important to keep these terms and contributors’ options understandable.

Building from Tanzania

DataHive Africa launched its operating pilot in May 2026 after a period of design and testing during the first months of the year.

Junior Makwala, the project’s founder, says he is 20 years old and that he designed and coded the infrastructure as a solo founder without initial capital. He also states that the company has completed registration with Tanzania’s Business Registrations and Licensing Agency.

This background illustrates local technical initiative emerging from East Africa’s mobile-first reality. The platform, in turn, depends on the voices, reviews and community relationships that turn the initial design into collective infrastructure.

Makwala belongs to a generation living simultaneously through the expansion of mobile connectivity, the circulation between Kiswahili and English, and the arrival of artificial intelligence systems trained mainly outside East Africa.

He does not observe the language gap from a foreign laboratory. He encounters it in everyday speech and in systems that may recognise standard Kiswahili while failing when faced with an urban phrase, code-switching or a local reference.

Building the platform in Tanzania keeps some technical and economic decisions close to the place of collection and broadens who can design infrastructure for artificial intelligence.

Technological autonomy does not mean isolation either. DataHive aims to offer datasets to AI developers. Its sustainability will therefore depend on contracts, clients, technical standards and markets extending beyond Tanzania.

The project therefore tests a connection between local technological capacity, community participation and a global data economy.

Building Artificial Intelligence from Language

On 6 July 2026, during World Kiswahili Language Day activities in Bujumbura, the East African Legislative Assembly called for greater inclusion of Kiswahili and other African languages in artificial intelligence and digital innovation.

A day later, UNESCO placed Kiswahili at the centre of the creator economy, digital public services and linguistic justice. It stressed that millions of people think, trade and create in Kiswahili, but that the language needs a real place online for that digital future to be fair.

DataHive forms part of this regional movement through a specific task: producing the material models need to interpret forms of speech absent from conventional corpora. Kiswahili appears as a living social practice found in markets, conversations, emotions and code-switching, rather than as a static dictionary resource.

The model opens a path towards a more participatory way of building artificial intelligence. Speakers produce the data, community members review and contextualise it, and compensation creates an explicit economic relationship instead of silent extraction. Public consent and the absence of web scraping are meaningful methodological choices.

The pilot does not yet solve linguistic justice. It does, however, allow speakers to appear as participants in AI production rather than merely as a population represented by models built outside their environment.

The project’s public footprint remains limited. That scale allows us to see decisions that often disappear behind intermediary companies and confidential contracts in larger data supply chains.

Here, it is still possible to see the person who speaks, the person who reviews the recording, the point that remains pending, the licence that is granted and the dataset that begins to take shape.

AIthropology Lab Reading

DataHive Africa shows that a language enters artificial intelligence through a human chain of work.

Someone designs a prompt.

Someone lends their voice.

Someone reviews the result.

Someone contextualises the expression.

Someone assembles the recordings and builds the dataset.

DataHive Africa is testing a way to democratise the construction of artificial intelligence. People who speak Kiswahili do not appear only as the intended users of a technology built elsewhere. They participate in producing the data, reviewing it and providing the cultural interpretation that turns it into useful infrastructure.

The project remains at an early stage, but it makes an important possibility visible. More linguistically diverse AI can be built by recognising the work of speakers, keeping decisions closer to the places where language is lived and replacing silent extraction with an explicit relationship of participation, consent and compensation.

Keeping a language present in artificial intelligence means enabling its communities to participate in how the machine learns to listen.

Sources Consulted