Pretraining as a competition
This page reads the first research paper from the Macrocosmos team, LLM Pretraining: The Use-Case Blockchain Has Been Waiting For? (August 2024), for people who know AI models exist but have never looked at how one is built. It describes subnet 9 on Bittensor as it ran before IOTA: independent miners each training a complete language model, competing for a reward that goes almost entirely to the best one.
Large models are normally pretrained by a handful of well-funded companies, in their own datacenters, at a cost the paper puts above $100 million for a frontier model.stated The paper's bet is that a blockchain's rewards can pay strangers to do that work instead, and that the result can hold up against models from established labs. If pretraining itself is new to you, the ResBM page starts from how a model learns.
Pretraining is the most expensive step in AI, and few can pay for it
The problem the paper sets out to answer, in its own numbers.
The cost of a frontier model
Pretraining is the stage where a model learns language from raw text by predicting the next token, over and over, across billions of examples. Everything a chatbot does later is built on it, and it is the most compute-hungry stage of development.stated
The paper's figures: Meta's Llama 3.1 8B took 1.46 million H100 GPU hours, the 405B model 30.84 million. OpenAI's CEO has said GPT-4 cost "more than a hundred million dollars" to train, and the compute needed for a state-of-the-art model doubles about every ten months.stated
The consequence the paper draws: only the largest companies can afford to pretrain at the frontier, and they decide privately what data and design go into the result.stated
Bittensor's third way
The paper sorts AI into closed models, built privately and sold through an API, and open models such as Llama and Mistral, released for anyone to build on. Bittensor proposes a third route: a network that pays strangers to improve open models.stated
It works through subnets, each a competition with its own rules. Miners do the work, validators score it, and the chain pays both in TAO, its token, according to those scores. The comparison the paper leans on is Bitcoin, which assembles more computing power than any company and points all of it at one task.stated
Subnet 9 points that machinery at pretraining.
The question the paper takes on: can a reward on a blockchain get independent miners to pretrain models that stand up next to models from established labs?
Subnet 9 ran a contest: every miner trains a whole model, and the best one takes almost everything
The design in one idea, and the numbers it produced.
Each miner pretrains a complete language model on their own hardware and publishes it on Hugging Face. Validators test every model on text drawn at random from a large public dataset, and the model that predicts it best receives nearly all of the subnet's miner rewards. Anyone can then download the winner and try to beat it.
LLM Pretraining: The Use-Case Blockchain Has Been Waiting For? (alternate title "Incentives Are All You Need"), credited to Macrocosmos, Taoverse, Const of Bittensor and Datura, with no individual authors named. Its acknowledgments thank the subnet 9 team: Const, Fish, Sid, Rustic, Alan, Rodrigo, Will and Steffen. Alan, Rodrigo and Steffen are authors on later papers here.
- Largest competition
- 7Bparameters; 700M before itstated
- Share to the top model
- 96%of validator weight, T = 0.01stated
- Improvement to take the lead
- 0.5%the epsilon thresholdstated
- 7B vs falcon-7b
- 8.59vs 9.94 perplexity, FineWeb Edustated
- Paid to miners
- ≈ $5Mestimated lifetime earningsstated
- Running since
- Nov 2023about nine months at writinginferred
Competition, not collaboration
The distinction the paper draws in its section 1.3On the left every miner needs enough hardware for a whole model, so model size is capped by the best single miner. The paper calls the right-hand version, where layers are "spread" across the network, an open question at the time of writing and puts it on its roadmap.stated The reward shares on the left follow the paper's softmax setting; the three small ones are illustrative.inferred
The next section shows how a validator picks a winner, how the rules stop copying, and how the winners compared with known models.
Random tests, a head start for the incumbent, and results against GPT-2 and Falcon
The scoring pipeline, the anti-copying rule, the results, and where the claims stop.
How a validator picks a winner
Section 2.3 of the paperMiners can read the validator's code, so a fixed test set would simply be memorized. Drawing fresh batches from a very large dataset at evaluation time (step 02) is what makes the score hard to game.stated Perplexity measures how surprised a model is by real text; lower is better.
The epsilon rule stops copying
Illustrative lossesEvery model is public, so without a margin anyone could download the leader, nudge its weights and win. A newcomer has to beat the older model by at least ε, set at 0.5 percent.stated A second 7B contest ran at 0.1 percent to test the setting.stated
Rules and exploits
The 700M contest beat GPT-2 Large
Perplexity · lower is better · Table 1| Model | Size | Wikitext103 | Falcon RW | FineWeb Edu |
|---|---|---|---|---|
| gpt2 | 124M | 30.13 | 35.25 | 29.63 |
| gpt2-large | 774M | 19.50 | 23.89 | 19.50 |
| phi-2 | 2.8B | 9.79 | 15.19 | 12.09 |
| net9 miner3 (subnet 9) | 769M | 15.89 | 15.57 | 15.21 |
The subnet's model beats gpt2-large on all three.stated The paper stresses how close it came to phi-2 on Falcon RefinedWeb, then the validation set; on Wikitext103 phi-2 is well ahead.inferred
The 7B contest beat Falcon on one of three
Perplexity · lower is better · Table 2| Model | Size | Wikitext103 | Falcon RW | FineWeb Edu |
|---|---|---|---|---|
| falcon-7b | 6.9B | 6.56 | 11.06 | 9.94 |
| Mistral-7B-v0.1 | 7.2B | 4.93 | 9.11 | 7.18 |
| jw2 (subnet 9) | 6.9B | 7.08 | 13.65 | 8.59 |
The subnet's model beats falcon-7b on FineWeb Edu and trails it on the other two; Mistral leads on all three.stated The loss "continues to improve", per the paper.stated
Where the claims stop
Subnet 9 showed the incentive could buy pretraining. The last question is what that proved, and what the design could never do.
It proved a blockchain could pay for pretraining, and hit the ceiling of one miner's hardware
Back to the opening question, and the step it forced next.
The paper's answer to its own question is a qualified yes. With roughly $5 million in token emissions over about nine months, a small group of miners produced a 700M model that beat GPT-2 Large on every test used and a 7B model that beat Falcon-7B on FineWeb Edu, the subnet's evaluation set.stated For a network with no lab, no hiring and no datacenter of its own, that was the proof of concept it claimed to be.inferred
The design also had a ceiling built in. Every miner had to fit and train a whole model, so the subnet could never produce a model larger than its best single participant could handle, and the winner-takes-all reward paid people to hold models back. The paper's own roadmap ends on the fix: "a decentralized training model where miners are collaborating on model development, rather than each developing their own separate model".stated
That sentence is the brief for the next paper. The IOTA primer splits one model across many miners and pays each for their share.
Where this paper sits
All four papers- Aug 2024SN9 pretraining whitepaperMiners each train a whole model; the best one takes the reward.You are here
- Jul 2025IOTA technical primerOne model split across miners, each paid for their share of the work.Training line
- Jan 2026Generative Adversarial MiningSide track: an incentive design for Apex, subnet 1, where quality has no scoring rule.Incentive track
- Apr 2026ResBMThe handoff between machines made 128 times smaller.Training line
Glossary · 8 terms
Sources · 3
- Macrocosmos, Taoverse, Const, Datura. LLM Pretraining: The Use-Case Blockchain Has Been Waiting For? August 2024, 22 pages, linked from the Research menu on macrocosmos.ai as the "Pretraining Whitepaper". Every stated label on this page is to this paper.
- Quinque et al. IOTA: A Technical Primer for Release, arXiv 2507.17766, July 2025. Cites this paper as its reference [1] and names its two core issues. Explained here.
- Macrocosmos. Subnet 9: Scaling up parameters, Substack, 13 May 2024. The 7B cap as it was announced.