mikyo.one / IOTA / The Macrocosmos papers / SN9 pretraining
SN9 pretraining, explainedWhitepaper · Macrocosmos, Taoverse, Const, Datura · August 2024
Updated 29 Sep 2026
A reader's guide to one paper · 1 of 4

Pretraining as a competition

This page reads the first research paper from the Macrocosmos team, LLM Pretraining: The Use-Case Blockchain Has Been Waiting For? (August 2024), for people who know AI models exist but have never looked at how one is built. It describes subnet 9 on Bittensor as it ran before IOTA: independent miners each training a complete language model, competing for a reward that goes almost entirely to the best one.

Large models are normally pretrained by a handful of well-funded companies, in their own datacenters, at a cost the paper puts above $100 million for a frontier model.stated The paper's bet is that a blockchain's rewards can pay strangers to do that work instead, and that the result can hold up against models from established labs. If pretraining itself is new to you, the ResBM page starts from how a model learns.

statedthe paper says itinferredmy arithmetic or reading of the paperPart of the Macrocosmos papers on mikyo.one
01 · Why

Pretraining is the most expensive step in AI, and few can pay for it

The problem the paper sets out to answer, in its own numbers.

The cost of a frontier model

Pretraining is the stage where a model learns language from raw text by predicting the next token, over and over, across billions of examples. Everything a chatbot does later is built on it, and it is the most compute-hungry stage of development.stated

The paper's figures: Meta's Llama 3.1 8B took 1.46 million H100 GPU hours, the 405B model 30.84 million. OpenAI's CEO has said GPT-4 cost "more than a hundred million dollars" to train, and the compute needed for a state-of-the-art model doubles about every ten months.stated

The consequence the paper draws: only the largest companies can afford to pretrain at the frontier, and they decide privately what data and design go into the result.stated

Bittensor's third way

The paper sorts AI into closed models, built privately and sold through an API, and open models such as Llama and Mistral, released for anyone to build on. Bittensor proposes a third route: a network that pays strangers to improve open models.stated

It works through subnets, each a competition with its own rules. Miners do the work, validators score it, and the chain pays both in TAO, its token, according to those scores. The comparison the paper leans on is Bitcoin, which assembles more computing power than any company and points all of it at one task.stated

Subnet 9 points that machinery at pretraining.

The question the paper takes on: can a reward on a blockchain get independent miners to pretrain models that stand up next to models from established labs?

02 · What

Subnet 9 ran a contest: every miner trains a whole model, and the best one takes almost everything

The design in one idea, and the numbers it produced.

Each miner pretrains a complete language model on their own hardware and publishes it on Hugging Face. Validators test every model on text drawn at random from a large public dataset, and the model that predicts it best receives nearly all of the subnet's miner rewards. Anyone can then download the winner and try to beat it.

LLM Pretraining: The Use-Case Blockchain Has Been Waiting For? (alternate title "Incentives Are All You Need"), credited to Macrocosmos, Taoverse, Const of Bittensor and Datura, with no individual authors named. Its acknowledgments thank the subnet 9 team: Const, Fish, Sid, Rustic, Alan, Rodrigo, Will and Steffen. Alan, Rodrigo and Steffen are authors on later papers here.

Largest competition
7Bparameters; 700M before itstated
Share to the top model
96%of validator weight, T = 0.01stated
Improvement to take the lead
0.5%the epsilon thresholdstated
7B vs falcon-7b
8.59vs 9.94 perplexity, FineWeb Edustated
Paid to miners
≈ $5Mestimated lifetime earningsstated
Running since
Nov 2023about nine months at writinginferred

Competition, not collaboration

The distinction the paper draws in its section 1.3
SUBNET 9 IN 2024 · COMPETITION DECENTRALIZED TRAINING · THE ROADMAP fullmodel fullmodel fullmodel fullmodel miner Aminer Bminer Cminer D ≈ 96%≈ 1%≈ 1%≈ 1% four separate models · each must fit on one miner · best takes nearly all layers 1, 2layers 3, 4layers 5, 6layers 7, 8 miner Aminer Bminer Cminer D each paid for its share one model · can outgrow any single miner · the route IOTA took

On the left every miner needs enough hardware for a whole model, so model size is capped by the best single miner. The paper calls the right-hand version, where layers are "spread" across the network, an open question at the time of writing and puts it on its roadmap.stated The reward shares on the left follow the paper's softmax setting; the three small ones are illustrative.inferred

The next section shows how a validator picks a winner, how the rules stop copying, and how the winners compared with known models.

03 · How

Random tests, a head start for the incumbent, and results against GPT-2 and Falcon

The scoring pipeline, the anti-copying rule, the results, and where the claims stop.

How a validator picks a winner

Section 2.3 of the paper
01Publishmodel on Hugging Facehash committed on chain 02Samplerandom batches ofFineWeb Edu, fresh 03Scoreperplexity per batch;repetitive output fails 04Compareevery pair, every batch;older model gets ε 05Weightsoftmax of win rates,T = 0.01 → 96% to #1 weights go on chain, where Yuma Consensus turns them into TAO emissions

Miners can read the validator's code, so a fixed test set would simply be memorized. Drawing fresh batches from a very large dataset at evaluation time (step 02) is what makes the score hard to game.stated Perplexity measures how surprised a model is by real text; lower is better.

The epsilon rule stops copying

Illustrative losses
better (lower loss) ←→ worse threshold 9.95 leader 10.00 9.97 · loses 9.90 · wins

Every model is public, so without a margin anyone could download the leader, nudge its weights and win. A newcomer has to beat the older model by at least ε, set at 0.5 percent.stated A second 7B contest ran at 0.1 percent to test the setting.stated

Rules and exploits

Mechanism
What the paper says
Rebasing
The winner is public, so every miner can start from it instead of from scratch
Model size caps
186M, then 700M, then about 7B; concurrent 700M, 3B and 7B contests from 12 Aug 2024, 14B planned
Dataset switch
Falcon RefinedWeb was "throttling miner performance"; validation moved to FineWeb Edu
Model hoarding
The leader can sit on better models and release them one at a time; the paper calls this a design flaw and plans a decaying ε
Weight obfuscation
Rescaling weights to cause vanishing or exploding gradients keeps a model usable but poisons anyone who trains on it; watched by monitoring and red-teaming
Cabals
Winner-takes-all removes the payoff from colluding: only one model can win

The 700M contest beat GPT-2 Large

Perplexity · lower is better · Table 1
ModelSizeWikitext103Falcon RWFineWeb Edu
gpt2124M30.1335.2529.63
gpt2-large774M19.5023.8919.50
phi-22.8B9.7915.1912.09
net9 miner3 (subnet 9)769M15.8915.5715.21

The subnet's model beats gpt2-large on all three.stated The paper stresses how close it came to phi-2 on Falcon RefinedWeb, then the validation set; on Wikitext103 phi-2 is well ahead.inferred

The 7B contest beat Falcon on one of three

Perplexity · lower is better · Table 2
ModelSizeWikitext103Falcon RWFineWeb Edu
falcon-7b6.9B6.5611.069.94
Mistral-7B-v0.17.2B4.939.117.18
jw2 (subnet 9)6.9B7.0813.658.59

The subnet's model beats falcon-7b on FineWeb Edu and trails it on the other two; Mistral leads on all three.stated The loss "continues to improve", per the paper.stated

Where the claims stop

The conclusion is broader than Table 2The conclusion says the 7B model matched falcon-7b "across every benchmark test that was used". Table 2 shows it ahead on FineWeb Edu only, and behind on Wikitext103 (7.08 against 6.56) and Falcon RefinedWeb (13.65 against 11.06).stated
Miners train toward the evaluation setMiners know they are tested on FineWeb Edu and, the paper says, will most likely train on it, and FineWeb Edu is where the 7B model leads.inferred The paper argues the dataset is too large to overfit, and answers that its models also do well on benchmarks miners do not train on.stated
Perplexity across tokenizersPerplexity is measured per token, so models that split text differently are not strictly comparable. The paper does not say whether the models in its tables share a tokenizer.inferred
Each miner must hold a whole modelThe design caps model size at what one miner can train, and winner-takes-all rewards hoarding. The paper names hoarding itself; the IOTA primer, eleven months later, names both as the core issues it set out to fix.stated
Credited to organizationsThe paper credits four organizations and no individuals, and is self-published on macrocosmos.ai rather than arXiv.stated

Subnet 9 showed the incentive could buy pretraining. The last question is what that proved, and what the design could never do.

04 · So what

It proved a blockchain could pay for pretraining, and hit the ceiling of one miner's hardware

Back to the opening question, and the step it forced next.

The paper's answer to its own question is a qualified yes. With roughly $5 million in token emissions over about nine months, a small group of miners produced a 700M model that beat GPT-2 Large on every test used and a 7B model that beat Falcon-7B on FineWeb Edu, the subnet's evaluation set.stated For a network with no lab, no hiring and no datacenter of its own, that was the proof of concept it claimed to be.inferred

The design also had a ceiling built in. Every miner had to fit and train a whole model, so the subnet could never produce a model larger than its best single participant could handle, and the winner-takes-all reward paid people to hold models back. The paper's own roadmap ends on the fix: "a decentralized training model where miners are collaborating on model development, rather than each developing their own separate model".stated

That sentence is the brief for the next paper. The IOTA primer splits one model across many miners and pays each for their share.

Where this paper sits

All four papers
  1. Aug 2024SN9 pretraining whitepaperMiners each train a whole model; the best one takes the reward.You are here
  2. Jul 2025IOTA technical primerOne model split across miners, each paid for their share of the work.Training line
  3. Jan 2026Generative Adversarial MiningSide track: an incentive design for Apex, subnet 1, where quality has no scoring rule.Incentive track
  4. Apr 2026ResBMThe handoff between machines made 128 times smaller.Training line
Glossary · 8 terms
Pretraining
Teaching a model language from raw text by having it predict the next token, billions of times.
Parameters
The model's learned numbers. 7B means about seven billion of them.
Perplexity
How surprised a model is by real text, the exponential of its average per-token loss. Lower is better.
Subnet
One competition on Bittensor, with its own task, scoring rules and share of rewards.
Miner, validator
Miners do the work; validators score it. Both are paid in TAO according to those scores.
Yuma Consensus
Bittensor's rule for combining many validators' scores into one payout.
Epsilon (ε)
The margin a newer model must win by to displace an older one, 0.5 percent here.
Softmax, temperature
A way to turn scores into shares. A low temperature, 0.01 here, gives nearly everything to the top score.
Sources · 3
  1. Macrocosmos, Taoverse, Const, Datura. LLM Pretraining: The Use-Case Blockchain Has Been Waiting For? August 2024, 22 pages, linked from the Research menu on macrocosmos.ai as the "Pretraining Whitepaper". Every stated label on this page is to this paper.
  2. Quinque et al. IOTA: A Technical Primer for Release, arXiv 2507.17766, July 2025. Cites this paper as its reference [1] and names its two core issues. Explained here.
  3. Macrocosmos. Subnet 9: Scaling up parameters, Substack, 13 May 2024. The 7B cap as it was announced.