An artificial intelligence system can recognise a ceremonial gate, a city wall, a group of Buddhist sculptures or a landscape carved over centuries. It can connect an image to a place, select an answer from several options and retrieve information that appears relevant.

What it does not resolve as easily is the relationship between that place, its history, the communities that name it and the disputes that sustain its meaning.

That limit sits at the centre of ChinaHeritaQA, a new benchmark for vision-language models focused on World Heritage properties in China. The project brings together 2,279 images drawn from real Weibo posts and 14,133 bilingual questions, in Chinese and English, about 51 UNESCO-listed sites.

The questions cover seven dimensions. They range from visually identifying a site to placing it within a dynasty, interpreting historical context, inferring social function or analysing architectural elements.

The results are clear. Models perform far better when they are asked to recognise an image than when they must relate it to periodisation, function or historical meaning. In the historical-periodisation task, the strongest system reaches 64.42% accuracy.

The finding is not that AI knows nothing about culture. It is more precise: retrieving visual patterns is not the same as understanding situated heritage.

A benchmark also interprets

ChinaHeritaQA is not a neutral test that simply discovers whether a machine understands culture. Like every evaluative instrument, it defines in advance what kind of knowledge counts as understanding.

Its construction starts from a UNESCO-aligned heritage ontology. The descriptions combine Wikipedia texts and UNESCO’s official selection criteria. GPT-4o was used to structure attributes and generate questions, followed by human verification. The team also states that it used Claude to improve the manuscript’s structure and clarity.

None of this invalidates the work. On the contrary, it makes something decisive visible: before a model answers, someone has already decided which categories matter, which sources are legitimate, which answer will count as correct and which forms of knowledge will be left outside the test.

The authors’ institutional composition also matters here. The team is affiliated with LMU Munich, FAU Erlangen-Nuremberg, the Munich Center for Machine Learning, the University of Tübingen and Tübingen AI Center, Sun Yat-sen University, the University of Copenhagen and the University of Maryland.

This is not a matter of discrediting international research. It is a matter of recognising that research on Chinese heritage, built through UNESCO-aligned categories and circulating largely through European and US academic institutions, also participates in defining what should count as a valid cultural interpretation.

The paper itself acknowledges this limit. Its framework, the authors write, may not capture locally defined or contested heritage narratives outside the international UNESCO regime.

When people themselves do not agree

There is another especially revealing result.

The human evaluation involved three college-educated native Chinese speakers, without declared heritage expertise. In historical-periodisation questions, all three gave exactly the same answer in only 16% of cases. Fleiss’ kappa was 0.247, a low level of agreement.

This does not mean that history cannot be known, nor that every interpretation is equally valid. It does show that dating a building or heritage ensemble from an image may remain ambiguous even for informed people.

The lesson matters for any comparison between people and machines. A response deemed correct in a benchmark does not always amount to a stable cultural truth. It may depend on the image selected, the available options, the vocabulary used and the institutional frame through which the question is defined.

The dataset also acknowledges a problem of visibility. All images come from Sina Weibo and the territorial distribution is uneven. Shanxi and Chongqing account for almost 40% of the questions, while Qinghai, Jiangxi and Shandong each contribute fewer than one hundred instances.

The bias is not only technological. It also depends on which places are photographed, who shares them, which landscapes are most visible and which kinds of heritage enter digital circulation.

Heritage is learned with other people

While ChinaHeritaQA attempts to measure the limits of AI in relation to Chinese heritage, UNESCO has presented a regional experience in Latin America and the Caribbean that points in another direction.

The report Education and Intangible Cultural Heritage in Latin America and the Caribbean maps 200 practices across 15 countries. The research and analysis were led by anthropologist Carla Pinochet of the University of Chile.

It does not treat heritage as a collection of isolated objects to be preserved from outside. It presents heritage as languages, techniques, memories, crafts and practices that remain alive because they circulate between people and generations.

Some 49.5% of the experiences take place in non-formal settings. A further 38% are developed in formal educational institutions, while 12.5% combine both settings. Eight out of ten use living heritage to promote appreciation of cultural diversity. Eight out of ten involve communities in transmitting intangible heritage.

Here, context is not an informational layer added after the fact to an image or database. It is a relationship. It lies in who teaches, in which language learning takes place, in which territory a craft is practised and in which community can decide how it wishes to be represented.

This is not about setting communities against technology, nor about imagining local knowledge as intact and untouched by innovation. It is about recognising that culturally responsible technology cannot be limited to extracting images, labelling objects and producing plausible answers.

It must also ask who participates in defining the data, who can challenge it, who retains control over its future uses and who benefits when that knowledge becomes digital infrastructure.

The sky also carries memory

Heritage does not live only in buildings, museums or archives. It can also live in a community’s relationship with the night sky.

A study co-ordinated by the International Astronomical Union Office of Astronomy for Development examines the first global astrotourism community exchange. More than 200 people from 59 countries and territories registered for this gathering, conceived as a space for dialogue among people working in astronomy, tourism, education, community development and cultural management.

The study proposes understanding astrotourism not as niche tourism but as a practice that can link local development, cultural preservation, dark-sky protection and more equitable access to astronomical experience.

Its anthropological interest lies elsewhere. The sky does not appear as an empty backdrop for watching stars, nor as a visual resource available for extraction. It can be orientation, narrative, memory, ecological knowledge, ritual practice and shared territory.

The study stresses the importance of Indigenous and local knowledge, but also identifies a crucial tension: when knowledge becomes a tourist experience, it may create economic opportunities while also becoming exposed to simplification, appropriation or commodification.

The research is exploratory and based on one event with self-selected participants. It does not offer a finished formula. Its value lies in showing that caring for the sky requires more than infrastructure, sensors or anti-light-pollution campaigns.

It requires deciding who tells the story, who controls the uses of that knowledge and who benefits when a night landscape becomes a destination.

What is at stake in Busan

This discussion comes ahead of an important moment for international heritage governance.

From 19 to 29 July, UNESCO’s 48th session of the World Heritage Committee will take place in Busan, Republic of Korea. The Committee will examine new nominations and assess the state of conservation of properties already inscribed.

The List of World Heritage in Danger currently includes 53 properties. That figure is not a news event in itself. It is the context for a broader discussion about which harms become visible, which memories receive international attention and which communities can intervene in decisions affecting their own territories.

The question is not only which monuments will enter or leave a list. It is also which forms of life, narratives, practices of care and conflicts are recognised when heritage is described as universal.

Seeing is not enough

AI can help to document, locate, compare and support conservation work. But its usefulness should not be mistaken for understanding.

A system may recognise a pagoda, a mask, a ruin or a constellation without knowing what relationship a community has with that place, what disputes run through it or what practices of care sustain its continuity.

It is not enough for a machine to see the temple.

Nor is it enough for it to look at the sky.

Si nunca un perro mira al cielo — if a dog never looks at the sky — perhaps the question is not only what it sees, but who decides what counts as looking.

To understand heritage, we need to know who names it, who passes it on, who protects it and who can decide what it means.

Sources