A model can be available online and still not be genuinely within everyone’s reach. It may have downloadable files, a permissive licence and public documentation, while depending on expensive memory, specialised accelerators, high-speed networks, electricity, commercial clouds and state decisions. The opposite can also happen. A system may continue to exist but suddenly become unavailable to part of the world because an authority decides that it should no longer circulate.

In this field, weights are the large numerical files in which a model retains the patterns learned during training. Making them available lets other people run or adapt the model without necessarily relying on a single API. But it does not automatically provide the training data, fine-tuning decisions, infrastructure or permission needed to sustain it over time.

Rather than comparing which model reasons best, it is worth asking who can use it, where, with what infrastructure and under whose permission. Artificial intelligence does not float in a cloud without ground beneath it. It rests on data centres, protocols, cables, energy, rules and people who decide what may cross each border.

An intelligence with a passport

On 9 June 2026, Anthropic launched Fable 5 and Mythos 5. They share the same base model but not the same regime of circulation. Fable 5 was aimed at general use with reinforced safeguards. Mythos 5, with fewer restrictions, was limited to a small number of Project Glasswing organisations for defensive cybersecurity work. Anthropic described that initial allocation.

Three days later, the US government ordered the suspension of access for any foreign national, inside or outside the country, including non-US people working at Anthropic itself. The company said that the letter it received did not specify the particular national-security concern. Since the order took effect immediately and Anthropic had no reliable way to verify nationality in real time, it suspended Fable 5 and Mythos 5 for all customers, including those in the United States. Anthropic’s statement is dated 12 June 2026.

The model did not disappear. What was interrupted was its circulation. An administrative category came to determine who could access a technical capability. The interface stopped being merely a gateway to a service and became a border.

The origin of the episode matters. The directive followed a report by Amazon researchers on a technique for bypassing Fable 5’s safeguards and finding software vulnerabilities. When Anthropic reviewed the case, it argued that less capable models, including Opus 4.8, GPT-5.5 and Kimi K2.7, could identify the same vulnerabilities. According to the company, every tool evaluated could also produce the proof-of-exploit associated with the report. A finding that, according to Anthropic, could be reproduced with already available models led to the withdrawal of two globally available models for eighteen days. Anthropic later set out that assessment.

The company complied with the order but publicly questioned the procedure. It argued that a government may halt unsafe deployments when the process is transparent, fair, clear and grounded in technical evidence. The letter that triggered the shutdown was not published. That opacity is not peripheral. It is part of the problem.

Reopening in layers

The government had approved Mythos 5’s partial return on 26 June. Four days later, on 30 June, the broader controls were lifted. Fable 5 returned on 1 July to Claude Platform, Claude.ai, Claude Code and Claude Cowork for users around the world. But it did not return everywhere in the same way or on the same terms.

Anthropic announced that access through AWS, Google Cloud and Microsoft Foundry would be restored as soon as possible, without setting a date. For Pro, Max, Team and some Enterprise plans, Fable 5 was included only up to 50% of weekly usage limits until 7 July. After that, access depended on additional credits. Mythos 5 returned first for a set of authorised US organisations, while Anthropic continued negotiating its expansion to domestic and international Glasswing partners. The company recorded these conditions.

The border did not reopen all at once. It reopened through products, cloud providers, authorised organisations, usage caps and purchasing power. Graduated access does not appear only when a model has open weights. It already exists inside proprietary models and their mechanisms for reopening.

The security architecture also changed. Anthropic trained a classifier to block the behaviour described in the Amazon report. When it intervenes, the request is redirected to Opus 4.8 and the user receives a notice. The company says the new classifier blocks the specific technique in more than 99% of cases, while acknowledging that this safety margin can increase false positives in legitimate programming and debugging tasks. The technical explanation and its limitations appear in the redeployment statement.

Alongside Amazon, Microsoft, Google and other Glasswing organisations, Anthropic also began to develop a shared framework for assessing the severity of jailbreaks. It proposes measuring capability gain, breadth, ease of weaponisation and ease of discovery. Classifying a risk is not a neutral operation. It determines which findings warrant an urgent response, which models remain available and who takes part in that decision.

Weights are not the whole infrastructure

At the same time, Z.ai presents GLM-5.2 as a scene of openness. Its Hugging Face page distributes the weights under an MIT licence and uses an expressive phrase, “no regional limits, technical access without borders”. The associated GLM-5 repository, by contrast, uses Apache-2.0 for the code. The model card and the repository make that difference visible.

Openness is not a single property. Weights, code, training data, evaluations, fine-tuning tools, deployment capacity and access conditions are different layers. The fact that one is public does not mean that all of them are. Being able to download a model changes the relationship with its producer in a real way. It allows a copy to be retained, other inference environments to be tested and dependence on a single API to be reduced.

GLM-5.2 documents deployment through Transformers, vLLM, SGLang, KTransformers and other frameworks. It is also presented with a one-million-token context window and efficiency improvements for long tasks. Its technical repository explains the architecture and those deployment environments. But downloading is not the same as running a model autonomously. A model at this scale needs memory, accelerators, storage, energy, maintenance and, in many configurations, a network capable of distributing work across several machines.

This is why legal openness must be distinguished from material access. A licence may authorise downloading without guaranteeing that a community can sustain the model in its own territory, with its own resources and under rules that can be debated publicly. Model files are not self-sufficient commodities. They are embedded in physical and administrative infrastructure.

Mistral and sovereignty as infrastructure

Mistral adds an important European counterpoint. The French company does not merely offer models or an API. It combines open models with its own computing infrastructure, designed to train, fine-tune and run AI systems at scale.

Mistral Small 4 was released under Apache-2.0 and can be downloaded, fine-tuned and deployed in different environments. Yet even for deployment, this model called “Small” requires, according to its official documentation, at least four NVIDIA HGX H100 systems, two HGX H200 systems or one DGX B200. The page presents that figure as minimum deployment infrastructure, not as an estimate for training the model from scratch. The distinction matters. A system that is more accessible than GLM-5.2 does not therefore become a household tool or low-cost community infrastructure. Mistral publishes those requirements and the licence.

Mistral’s response is not to deny that dependency, but to turn it into a product. Mistral Compute offers dedicated GPU clusters and tools that distribute tasks across machines, create work queues and set execution priorities. The company aims to provide 200 MW of “sovereign” frontier capacity in the European Union by 2027 and plans to use NVIDIA GB200, GB300 and B300 accelerators. Its platform describes this infrastructure.

Sovereignty therefore appears as an architecture of negotiated dependencies rather than a complete exit from US-centred technology chains. It may widen the room for decision-making over location, compute queues, data control and operational continuity, while retaining dependencies on chips, capital and supply chains it does not fully control.

The framework agreement with France’s Ministry of the Armed Forces makes another tension visible. Mistral said that its solutions would be deployed on French national infrastructure to retain control over sensitive data and technologies. Territorial control can protect public or strategic information, but it also raises questions about which institutions set priorities, which uses receive greater capacity and how public research and civil society participate in those decisions. Reuters reported the agreement in January 2026.

Situated infrastructure

A universal anthropology, attentive to the situated ways in which global relations are lived, should not reduce this landscape to a rivalry between the United States, China and Europe, nor speak of “those left out” as though they were an abstraction. Infrastructure is experienced through particular institutions, languages, energy sources, budgets and historical memories.

LatamGPT offers a different scene. The project, coordinated by Chile’s National Centre for Artificial Intelligence, brings together around two hundred people and more than sixty-five institutions in fifteen countries. Its first version, based on Llama 3.1 with 70 billion parameters, is released as code, data and trained files so that technical teams can adapt it for specific uses. It is not yet available as a mass-market chatbot from ordinary computers or phones. The project’s documentation describes its current state and its published resources.

Its aim is to build regional capacity and improve the representation of Latin American and Caribbean contexts that global models often treat as peripheral. Yet its own documentation also makes the material ambivalence visible. AWS appears among its strategic collaborators, and the project explains that optimising its AWS infrastructure reduced training time from 25 to 9 days. A model can speak from a region while still breathing the infrastructure of a global provider. LatamGPT explains this process in its FAQs.

This does not invalidate the initiative. It is a lesson. Bringing together a region’s data, languages and institutions can open a space for decision-making that global models often ignore. But that opening does not in itself remove dependence on clouds, finance, energy and computing capacity. Sovereignty is built in degrees. It is not obtained through a simple downloadable file.

Agents depend on memory too

The issue becomes even more material when models stop answering a single question and begin to work as agents over dozens or hundreds of turns. In such cases, it is not enough to have the weights and a GPU. The system must recover context, move data between storage and active memory, distribute priorities across machines and retain a record of what it has done.

The DualPath preprint, by Yongtong Wu and twelve co-authors, describes this problem. In multi-turn agentic tasks, inference may be limited less by computation than by reading and moving the KV cache, an operational memory that stores already-processed context and prevents the model from having to reread it from scratch at every turn.

Its proposal opens two loading paths. As well as moving that memory from storage to the machines that reconstruct context before producing a response, it allows the cache to be loaded first on the machines that generate each new fragment of text and then transferred via RDMA. RDMA is a networking method that moves data directly between the memory of different machines, reducing intermediate steps. The result is a more balanced distribution of traffic.

The results must be read at their own scale. The paper reports improvements of up to 1.87 times in offline throughput and an average 1.96 times improvement in online serving capacity, without breaching service-level objectives, in its own system and with the agentic workloads evaluated. It does not show that every agent will double its speed or that the system is ready for every infrastructure.

Even so, it makes visible something that accounts of models often conceal. An agent’s autonomy is not contained in a weights file. It rests on infrastructure that enables it to remember. That infrastructure governs how much context can circulate, which traffic gets priority and who can sustain a long interaction without the system coming to a halt.

Closing

The opposition between closed and open models remains important, but it is insufficient. Between those positions lies an ecology of access.

A proprietary model can be global and suddenly become subject to a letter addressed to Anthropic whose text the company did not publish. An open-weights model can escape a single API while continuing to depend on an extreme concentration of compute. A European company can offer open models while selling GPU capacity and sovereign-deployment agreements. A regional project can broaden cultural representation while needing AWS resources for training. And an inference improvement can make an agent more efficient while revealing that its memory depends on networks, storage and energy.

What matters is not only who owns the model. It is who can make it live, maintain it, debate its rules and repair its harms.

Releasing the weights is not enough. We also need to ask about the memory, network, energy, hardware, permissions and borders that allow — or prevent — a model from remaining available.

Sources

Sources on circulation, infrastructure and memory

Regional and political context