From a contest to one shared model
This page reads Incentivised Orchestrated Training Architecture (IOTA): A Technical Primer for Release (Macrocosmos, July 2025) for people who know AI models exist but have never looked at how one is built. It is the design document for IOTA, the system behind the Orion training runs, and it answers the two problems the 2024 subnet 9 paper left open.
Large models are normally trained in one datacenter on matched hardware joined by very fast links.stated Subnet 9 had shown that strangers paid in tokens would pretrain models, but each of them had to hold a whole model, and only the winner was paid. IOTA splits a single model across many miners, keeps it training when machines fail or cheat, and pays each miner for the work they actually did. If pretraining itself is new to you, the ResBM page starts from how a model learns.
Every existing way to train at scale breaks on open, unreliable hardware
What subnet 9 could not do, and why the standard methods could not fix it.
Where subnet 9 left off
By August 2024 subnet 9 had miners pretraining models from 700 million to 14 billion parameters that beat established baselines.stated The primer names the two problems that remained: "every miner had to fit an entire model locally", and "winner-takes-all rewards encouraged model hoarding".stated
The first caps model size at what one participant can hold. The second pays people to keep improvements to themselves. The fix for both is the same: train one model together, with each miner holding a piece and paid for that piece.
Why the standard methods don't transfer
The bandwidth gap
Section 4 of the paperThe internet link barely registers on this scale. Figures from the paper.stated Bars are linear against NVLink, and the internet link is drawn wider than scale so it stays visible: InfiniBand is about 2.8 percent of NVLink, a 200 Mbps link about 0.003 percent.inferred The paper cites a consensus that activations must shrink roughly 100x to 300x to match datacenter transfer times.stated
The question the primer takes on: can one model be trained across many unreliable, untrusted machines on ordinary internet, with everyone paid fairly for their part?
IOTA splits one model across many miners and pays each for its share of the work
The architecture in one idea, and the numbers the paper reports.
IOTA assigns each miner one layer of a single model. Several miners share each layer, so training data streams through the pipeline in parallel. A central orchestrator decides who works on what and when to merge, validators re-run a sample of each miner's work to check it, and every miner is paid for the backward passes it completes. Activations are compressed between layers, and the copies of each layer are merged with a scheme that exposes cheaters and survives dropouts.
Incentivised Orchestrated Training Architecture (IOTA): A Technical Primer for Release, by Felix Quinque, Alan Aboudib, Szymon Fonau, Rodrigo Lopez Portillo Alcocer, Brian McCrindle and Steffen Cruz of Macrocosmos. arXiv 2507.17766, 16 July 2025; the same text was posted on the IOTA site on 30 May 2025. The paper calls its results "preliminary", to be validated in production.
- Activation compression
- 128x1.5B model, up to 400M tokensstated
- Weights merged, 10% of miners down
- > 99%butterfly all-reducestated
- Failure rate tolerated
- 35%about 90% of weights still mergedstated
- Data per miner to merge
- 4W + 2W/Nflat as miners are addedstated
- Roles
- 3orchestrator, miners, validatorsstated
- Mainnet
- 2 Jun 2025planned graduation datestated
The architecture
After the paper's Figure 1Each layer is held by several miners, so the model can be as large as the number of participants allows, not as large as one machine allows. And everything runs through the orchestrator and shared storage: the paper chose hub-and-spoke over peer-to-peer so every interaction can be seen and audited.stated
The next section opens the five mechanisms that make it work, and marks where the evidence is still thin.
A training cycle, a pay rule, a compression block, a merge and an audit
Each mechanism in turn, then where the claims stop.
One epoch, as the orchestrator runs it
After the paper's Figure 2A miner's work counts once it has trained at least B-min batches, and merging starts when enough miners qualify, so slow machines never hold up the rest. Between merges, miners take local steps in the style of DiLoCo.stated
How miners are paid
A validator follows one randomly chosen miner through an epoch and re-computes part of its work, checking forward and backward passes against the miner's by cosine similarity. Miners do not know when they are being watched.stated
The score is the number of backward passes a miner completed and passed. Each score counts for a fixed period γ and then drops to zero, so pay tracks recent work. The paper calls this "fixed compensation per processed activation", which removes the reason to inflate throughput when no one is looking.stated
The trade-off is explicit: because validation depends on reproducing the miner's work exactly, "the design of the incentive landscape does not give power to the miner to innovate algorithmically at this time".stated Miners run the given code; they compete on hardware and uptime.inferred
The bottleneck block
Earlier attempts put a compression layer between transformer blocks and saw convergence fall off. The authors trace the damage to the break in the residual connection, the shortcut that lets gradients pass through a deep network, more than to the smaller tensor.stated
Their bottleneck block, a modified Llama 3 block, keeps partial residual paths flowing through the compressed point. On a 1.5B model with three bottlenecks, trained on up to 400 million FineWeb tokens, compression of 128x for activations and gradients showed "no significant loss in convergence"; moving from 32x to 128x cost only a slight degradation.stated
This is the direct ancestor of ResBM, which formalizes the identity path and tests it at 2B over 26 billion tokens.inferred
Butterfly all-reduce: merging copies of a layer without trusting anyone
Three miners, after the paper's Figure 6With N miners there are N(N−1)/2 pairs and as many shards, so every miner's work is duplicated by every other. In a 50-miner test the 10 deceptive miners stood out in the agreement matrix; with 10 percent of miners failed, over 99 percent of weights still merged, and the scheme tolerates failure rates up to 35 percent.stated A central merger would have to move N·W; each miner here moves roughly four copies of its layer.stated
CLASP: finding the miner who hurts the loss
Toy model, after the paper's Figure 8The red paths share one miner. The orchestrator logs each sample's path and loss; averaging loss per miner makes a bad actor stand out, as in an ablation study. The paper's test is a toy: 5 layers of 5 miners, with a malicious miner assumed to raise loss by 10 percent.stated
What CLASP is, and isn't yet
Contribution Loss Assessment via Sampling of Pathways borrows from Shapley values, a game-theory measure of each player's marginal contribution. Its purpose is to pay miners by their effect on the model and to catch miners who send junk activations.stated
It relies on two assumptions the paper states: that the per-sample losses are reported honestly, and that miners cannot coordinate across layers, since each sees only black-box activations from its neighbors.stated
It is not in the initial release. The paper calls it "an active area of research", with a fuller design and reproducible code promised in a later report.stated
Where the claims stop
The design holds together on paper and in small tests. The last question is what it made possible, and which pieces the team kept building on.
It turned a contest into one training run whose size grows with its participants
What the primer changed, and what it handed on.
Subnet 9 in 2024 could only produce a model as large as its best miner could train, and paid one winner. The primer's design removes both limits: model size scales with the number of participants rather than one machine's memory, and pay follows each miner's verified work.stated It gave consumer GPUs a place in pretraining a model none of them could hold alone.inferred
Three of its pieces carried forward. The bottleneck block became ResBM, tested at larger scale nine months later. The orchestrator, merge and pay rule became the production subnet, which the paper scheduled for mainnet on 2 June 2025.stated And the runs that followed, including Orion-100B and two Orion-16B runs, are recorded with their numbers in the Orion Register.
One trade-off carried forward too: miners who must be reproducible cannot innovate. The team's next paper, Generative Adversarial Mining, is an incentive design built for the opposite case, where miners are paid to find better answers no one can score directly.
Where this paper sits
All four papers- Aug 2024SN9 pretraining whitepaperMiners each train a whole model; the best one takes the reward.Training line
- Jul 2025IOTA technical primerOne model split across miners, each paid for their share of the work.You are here
- Jan 2026Generative Adversarial MiningSide track: an incentive design for Apex, subnet 1, where quality has no scoring rule.Incentive track
- Apr 2026ResBMThe handoff between machines made 128 times smaller.Training line
Glossary · 9 terms
Sources · 4
- Quinque, Aboudib, Fonau, Lopez Portillo Alcocer, McCrindle, Cruz. Incentivised Orchestrated Training Architecture (IOTA): A Technical Primer for Release. arXiv 2507.17766v1, 16 July 2025; also at iota.macrocosmos.ai, dated 30 May 2025, same text. Every stated label on this page is to this paper, except the GAM quote.
- Macrocosmos, Taoverse, Const, Datura. LLM Pretraining: The Use-Case Blockchain Has Been Waiting For? August 2024. Explained here.
- Quinque, Brady, Cruz. Generative Adversarial Mining on Decentralized Networks, January 2026. Source of the "code attestation" description. Explained here.
- Ryabinin, Dettmers, Diskin, Borzunov. SWARM Parallelism, ICML 2023. The pipeline method IOTA builds on.