Training a large model outside a data center has stopped being an experiment

the fact
Training a large model has always taken thousands of cards in the same shed, wired to each other. Between 2024 and 2025 several groups finished real models on machines scattered around the world and joined by ordinary internet, and the wall that stood in the way — bandwidth — has been brought down by orders of magnitude.
−2.4%
bitcoin on the day the news was somewhere else
bitcoin’s price around 6 november: −2.4% on the day, a close of $101,346. cointalks archive.

On 6 November 2025 Bitcoin closed at $101,346, −2.4% against the previous day, while the day’s news was about the distributed training of artificial intelligence models.

What actually got finished

The reason any of this interests anybody sits in a single figure: one frontier model costs more than $5 billion today between machines and electricity, with the best cards at $30,000 each. At that price the companies in the world that can afford it are fewer than ten, and everybody else can only ask permission to use what those companies built. The people working on the distributed method say they can bring that cost down by 90 percent using cards that already exist and sit idle half the time.

In November 2024 one group finished a model of 10 billion parameters trained on 112 cards spread over three continents, keeping the machines busy 96 percent of the time under the best conditions. In May 2025 the same group finished one of 32 billion with more than 100 nodes unlike each other, more than 400 hours of training, and the whole process public and repeatable by anybody.

That first model’s scores on the standard tests were no record — thirty-seven and a half out of a hundred on one, seventy-two on the other — and the people who made it never claimed otherwise. The result wasn’t the score: it was that volunteers scattered around the world, with machines coming into the training and dropping out of it while it ran, had carried the thing to the end using 400 times less bandwidth than the normal method.

Meanwhile a project running inside a network of subnetworks has more than 200 cards in production, has finished a model of 1.2 billion parameters and is aiming for 70. These are all figures declared by the people who produced them: checkable in the published work, not redone by us.

The wall was bandwidth, and how they lowered it

The problem has always been the same. Every machine that trains computes corrections to apply to the model, and every machine has to know what the others computed: inside a data center those corrections travel over dedicated, very fast cables; over ordinary internet they take too long. All the research of the last two years has gone into sending fewer of them.

The first blow was to stop synchronizing at every step and to do it every 500: the machines work on their own for a stretch and realign afterwards, and that alone brings the communications down five hundredfold. The second was compressing what gets sent — a transform plus a reduction to a single bit — with reductions declared at between a thousand and 10,000 times against the standard methods.

The third is sharper still: synchronize only 0.1 percent of the parameters at each step, and the model converges all the same. In December 2024 one of these methods trained a model of fifteen billion parameters over eleven thousand steps with machines supplied by four different operators. The wall didn’t fall to a single idea: it fell to three tricks stacked on each other.

every correctiononly every 500 stepscompressed to one bit0.1% of the parameters
how much has to cross the cable: it starts with every correction sent at every step, and each trick stacks on the one before. the reductions are the ones declared in the published work, not measured by us.

The steps that reduce the bandwidth needed: starting from every correction sent at every step, you move to synchronizing every 500 steps, then to compressing the updates down to a single bit, and finally to sending only 0.1 percent of the parameters.

The other road: split the model, not the data

The methods above keep a whole copy of the model on every machine and divide the data. There is a second road, and a harder one: divide the model itself, put different pieces on different machines and make them talk to each other. The trouble is that every piece then waits for the one before it, and over the internet the waits add up.

In September 2024 a group completed a training run of this kind for the first time across different participants and unequal machines: seven and a half billion parameters, three weeks, thirty-six billion words. A year later they showed a model of eight billion trained with its blocks in four different places joined by ordinary internet, with results on a par with a centralized setup — something that before that work was taken to be impossible.

The piece that makes the thing workable is the handling of failures. Another group built a system that looks at which computing routes are free and skips the slow or fallen nodes: it holds 93 percent of its output even when 50 percent of them go down. In a data center a node that vanishes is a fault; here it is a normal operating condition.

Why reinforcement learning falls well here

In early 2025 a Chinese model showed that reinforcement learning — learning by trial and reward — can be used as the main training method and not only for the final polish. That matters a great deal to anyone working distributed, because that method tolerates delays and absences far better than the others.

Out of it came a setup that separates three things that used to sit together: the training, the use of the model to generate answers, and the delivery of the updated weights. Every machine does its own round on its own account and delivers when it has finished, with a light check that verifies it really learned from what it saw rather than copying or inventing.

Another group put a three-beat mechanism on its test network: every node answers on its own, then the nodes criticize each other, then they settle on a common answer, and the reward depends on how close each one was to that agreement. It is a structure that looks more like a newsroom than a computing shed, and it works for the same reason: nobody has to trust anybody.

How you check that somebody is really working

This is where the serious problem sits, and it isn’t technical but a matter of incentives: if I pay machines to train, how do I know they are training rather than sending me random numbers or copying their neighbor’s work? Redoing the computation to check it would cost as much as doing it, so the verification has to be cheap.

There are four different answers going around, and they are worth keeping apart because they fail in different ways. The first is economic: whoever works puts down a bond, whoever cheats loses it, and anybody can report them. The second watches behavior: it checks that the corrections delivered are consistent with the data that machine saw. The third has validators vote on the quality of what arrives, and penalizes the validators who inflate their votes as well.

The fourth, not yet running, tries to make the work itself the proof: every answer generated carries a stamp showing which model produced it. None of the four is settled. One of these networks already came under a real attack over the 2024 holidays, and the defense is the usual chase between whoever attacks and whoever protects.

who checks whom
the waywhat it rests on
a bond put downcheat and you lose it, and anybody can report you
the behaviorthe corrections add up with the data you saw
a vote of the validatorsplus a penalty for whoever inflates the votes
a stamp on the answerthe work itself is the proof · not yet running
the four ways of checking whoever works, and what each of them leans on. none is final.

The four approaches to verification: economic, based on a bond put down that is lost by cheating; behavioral, which checks the consistency between the data seen and the corrections delivered; by vote, where validators judge the quality and whoever inflates the votes is penalized; by stamp, where every answer generated carries cryptographic proof of which model produced it.

The money inside it

The capital raised says how seriously the thing is taken outside the sector: sixty-five million dollars for the best-funded group, with a declared valuation of a billion on a token that doesn’t exist yet; forty-three million for another; twenty and a half for the group behind the two models; seven and six tenths for a fourth. The investors’ names are those of American venture capital funds, not of crypto treasuries.

Of live tokens, for now, there is essentially one: the token of the network of subnetworks, which in December 2024 touched $5 billion of market value and which since February 2025 has had separate tokens for each of its 118 subnetworks. All the others are awaited, and the people waiting for them say so openly: they are capital raises with a token promised later.

nous research$65m · declared valuation $1bn
gensyn$43m · test network, no token
prime intellect$20.5m · no token announced
pluralis$7.6m · shares of the model, not a fixed wage
how much each of them raised, and what it promised in return. from the groups’ own announcements, not from us.

Capital raised by the groups named: nous research, $65m (declared valuation $1bn); gensyn, $43m (test network, no token); prime intellect, $20.5m (no token announced); pluralis, $7.6m (shares of the model, not a fixed wage).

The estimate going around all of this is that the artificial intelligence market will be worth $15 trillion by 2030. It is a partisan estimate, made by people who have put money into these projects, and it is worth what every five-year estimate is worth: it says what the people who wrote it are hoping for, not what will happen.

What doesn’t add up, and what to watch

The limits are four and none of them is small. Bandwidth stays a wall above a hundred billion parameters, even with the compressions above. Security is an open chase: data poisoning, manipulation of the corrections, false identities among the validators. The rules don’t exist yet, and the question of who answers if a model trained by a thousand strangers causes harm has no answer today.

The fourth is the most awkward and concerns real money: if the finished model’s weights are public, who pays for having trained it? One of the groups tries to solve that by giving the participants a share of the model instead of a fee, but that share is worth something only if the model generates revenue, and there is no guarantee at all that it will.

From outside there are four things to watch, and they are all accounting: how many cards actually take part — two hundred today in the project furthest along, and the useful threshold is fifty times that; how many training runs reach the end without being aborted; how much a unit of computing costs here against the traditional suppliers; and how far these models are from the best centralized ones. Bitcoin, in the middle of all this, closed at $101,346, -2.4 percent: the days when a piece of an industry moves are not the days when the price moves.

put bluntly
in a data center a node that vanishes is a fault; here it is the normal condition, and they built on top of it

Put bluntly: in a data center a node that vanishes is a fault to repair; in a network of volunteers it is the normal condition, and all the research went into building on top of that rather than preventing it.

6 November 2025published with the day’s close, recomputed on our archive
the words in this piece · 2
fee
what you pay to use a protocol. it can go to whoever supplies the service, to whoever holds the token, or to both.
token
the unit a protocol issues. it can serve to vote, to pay, to receive revenue, or to do nothing at all.
news · 6 November 2025all the news