Paying for quality no one can score
This page reads Generative Adversarial Mining on Decentralized Networks (Macrocosmos, January 2026) for people who know AI models exist but have not looked at how networks like Bittensor decide who gets paid. It is a side track in these papers: not about training a model, but about the rule that pays miners on Apex, Macrocosmos's subnet 1.
A company that builds an AI product judges quality with its own staff and tests. An open network has no staff, so every subnet needs a rule a machine can apply to decide which miner did better work. For some tasks that is easy: re-run the job, or wait and see if a prediction came true. For open-ended answers, such as a written explanation or an agent's plan, there is no rule to apply. This paper's answer is to turn judging into a guessing game borrowed from how GANs train.
Every way a subnet pays its miners assumes the work can be checked
The four families of incentive the paper names, and what each one cannot reach.
Four ways to pay a miner
Section 1.3 of the paperNone of the four can pay for a better written answer, a better plan or a more realistic image, where quality is a judgment rather than a check.stated
The idea it borrows
A generative adversarial network, or GAN (Goodfellow et al., 2014), trains two models against each other. A generator makes fake samples; a discriminator tries to tell fakes from real ones. Neither needs a hand-written definition of "realistic": the discriminator's judgment is the definition, and both improve by competing.stated
The paper asks whether the same trick works as a pay rule for people running models on a network, instead of as a training method inside one model.
The question the paper takes on: can an open network pay for quality it cannot score, by letting competition between miners stand in for the score?
Miners are paid for fooling each other, and for not being fooled
One round of the game, and the numbers the paper reports.
A validator poses a challenge. A coin flip decides who answers: a random miner, or the validator's own reference system, which is given more time and tools. A panel of other miners then guesses who wrote the answer. Correct guessers split points; the answering miner earns whatever the panel got wrong. Over time miners learn to produce answers indistinguishable from the validator's expensive ones, at a fraction of the cost.
Generative Adversarial Mining on Decentralized Networks, by Felix Quinque, Kalei Brady and Steffen Cruz of Macrocosmos, dated 9 January 2026. This page reads the Wayback Machine copy captured on 9 January 2026.
- Points per round
- 1strictly zero-sumstated
- Who answers
- 50 / 50a miner or the validatorstated
- Skill correlation
- > 0.9after one answer per miner, 256 miners, panel of 100stated
- Deployed on
- 2 subnetsSN1 Apex, SN34 media detectionstated
- Stable cabals
- 0the only equilibrium, in the modelstated
- Deployment results
- noneno measurements publishedinferred
One round
After the paper's Figure 1The panel's ability to tell a cheap answer from an expensive one is the quality measure. A miner earns as a generator by making answers the panel mistakes for the validator's, and as a discriminator by spotting the difference.stated
The next section shows how the single point is split, why a cabal loses, and how quickly scores settle.
One point per round, burned when the validator wins, and a proof that collusion backfires
The scoring rule, the cabal argument, the convergence result, the deployments, and where the claims stop.
How the point is split
Panel of 4, 3 correctPoints the validator's reference earns are removed from the miners' pool, so miners as a group lose whenever the panel is fooled by the validator.stated The wrong guesser gets nothing.
Why a cabal loses
Suppose some miners collude and tell each other whenever one of them wrote the answer. They can now rule out part of the field, so their guesses get more accurate, and they are right more often against the validator.stated
The paper's analysis shows the catch. The cabal's better guesses "rescue" points that would otherwise be burned, and the cabal ends up behind whenever more than half of those rescued points flow to outsiders, which the paper's plots show at every setting. In every combination of cabal size and classifier accuracy it plots, non-cabal miners score higher than cabal miners.stated
So each colluder does better by leaving, and "the only state in Nash equilibrium is a completely cabal-less state". The burn is essential: if validator points were re-shared among miners instead, colluding to depress others' scores would become the winning strategy.stated
How fast scores find real skill
The paper models each generator and discriminator pairing as a match whose odds depend on the two players' skills, in the Bradley-Terry style used for ranking players from pairwise games. It then asks how closely a miner's win rate tracks its true skill as rounds accumulate.stated
Because each answer is judged by a whole panel, every round produces many matches. With 256 miners, typical subnet capacity, and a panel of 100, the correlation between score and skill passes 0.9 after each miner has generated just once.stated Ranking miners by pay tracks their skill after one turn each.inferred
Resource asymmetry sets the ceiling
A plain GAN saturates once the generator's output can't be told apart from the real thing. Here the "real thing" is the validator's reference, which is deliberately given more: extra processing time, private resources, more context.stated
Miners, held to tighter limits, must find faster and cheaper ways to match it. The paper frames the result as distillation: the network produces lightweight versions of an expensive pipeline, and the size of the resource gap becomes a dial for how hard the game is.stated
Where the claims stop
The paper's support is a game-theory analysis. The last question is why it matters next to the training line.
It rewards better methods, the case IOTA's reproducibility check leaves open
What the paper adds to the picture, and what it leaves open.
The training papers here pay for work that can be reproduced. IOTA checks a miner by re-running its computation, which is why miners there compete on cheap, fast hardware rather than better methods.stated Generative Adversarial Mining is the team's design for the other case: tasks where the valuable thing is a better answer, and there is no way to re-run your way to a score.inferred
If it works as argued, it is a general-purpose pay rule for open-ended AI work on a decentralized network: agent workflows, media, and the paper suggests reinforcement learning, scientific computing and multimodal reasoning next.stated The evidence so far is a game-theory argument and two deployments reported without data.
The paper ends on a next step: capturing the miners' generator and discriminator models themselves, not only their outputs, so the innovations can be sold directly. That is described as under development, for "an upcoming publication".stated The training line resumes with ResBM, three months later.
Where this paper sits
All four papers- Aug 2024SN9 pretraining whitepaperMiners each train a whole model; the best one takes the reward.Training line
- Jul 2025IOTA technical primerOne model split across miners, each paid for their share of the work.Training line
- Jan 2026Generative Adversarial MiningSide track: an incentive design for Apex, subnet 1, where quality has no scoring rule.You are here
- Apr 2026ResBMThe handoff between machines made 128 times smaller.Training line
Glossary · 8 terms
Sources · 3
- Quinque, Brady, Cruz. Generative Adversarial Mining on Decentralized Networks, 9 January 2026, 15 pages. Linked from macrocosmos.ai as the "APEX GAN Whitepaper" at apex.macrocosmos.ai/research/apex_gan.pdf; read from the Wayback Machine copy. Every stated label on this page is to this paper.
- Goodfellow et al. Generative Adversarial Nets, NeurIPS 2014. The idea the mechanism borrows.
- Quinque et al. IOTA: A Technical Primer for Release, arXiv 2507.17766, 2025. The paper's example of code attestation. Explained here.