The paper behind liquid training
This page reads one research paper, ResBM (Macrocosmos, April 2026), for people who know AI models exist but have never looked at how one is built. Macrocosmos calls its approach liquid training: one model trained across scattered, mismatched machines over the public internet.inferred
Large models are normally trained in one place: thousands of matched GPUs in a single datacenter, joined by specialized links (NVLink inside each server, InfiniBand between them) fast enough that splitting one model across many cards costs almost nothing.stated Liquid training drops the building. Most of the cards it draws on are too small to hold a large model on their own, so the model still has to be split across them, and every split is now a handoff over a link a hundred or more times slower.inferred ResBM makes those handoffs 128 times smaller.stated That is what lets a model bigger than any one card train outside the datacenter.inferred
The page does not assume you know what pretraining is. It starts there, shows the exact point where training across the internet breaks down, and then shows what ResBM changes and how well that holds up. Everything is drawn from the paper and labeled by how sure it is.
A model learns by guessing, and big models guess across many machines
Before ResBM makes sense, three things need to: what training is, why a large model has to be split, and what that split costs.
Training is one loop, run trillions of times
Every guess produces a correction, and that correction has to reach every layer of the model before the next guess. Pretraining is this loop repeated over billions to trillions of tokens (word fragments). The probabilities shown are illustrative.
One model, split across machines
This is pipeline parallelism: each GPU holds a slice of the layers, sends its output (the activation) forward, and later receives the correction (the gradient) back. In the paper's test model there are seven boundaries, each carrying a 1024 × 4096 tensor, 8 MiB in bf16, each way.stated
Why the split is necessary
A frontier model has more weights than any single GPU can hold, so training spreads it across hundreds or thousands of accelerators.stated There are two standard ways to do that, and large runs use both.
Data parallelism gives each machine a full copy and a different slice of data, then syncs the copies. Low-communication methods such as DiLoCo and DeMo already make that work over slow links.stated
Pipeline parallelism gives each machine a slice of the model, as in the diagram. It lets a model larger than one card train on cards that each hold a slice, but it needs a fast link at every boundary on every step. The paper calls this the primary bottleneck left for decentralized training.stated
The handoff is where the time goes
One 8 MiB handoff, two linksMoving 8 MiB takes about 6.7 ms at 10 Gbps and about 0.84 s at 80 Mbps, roughly 125 times longer.inferred In the paper's measurements the same 2B model trained at 7,530 tokens a second on a 10 Gbps link and 609 on 80 Mbps, about one twelfth the throughput.stated
What had been tried
Quantization. Sending 8-bit numbers instead of 16-bit saves only 2x.stated
Lossy compression between blocks. Autoencoders, top-k sparsification and low-rank SVD all sit on the residual stream, so error compounds with depth and training fails to converge at aggressive ratios.stated
Subspace Models (Ramasinghe et al., 2025). Constrains each layer to a shared low-rank subspace, which needs a modified AdamW and periodic updates on the Grassmann manifold outside the normal training loop. It works, but training is no longer plain end to end.stated
So the question the paper takes on: can the handoff be made small enough for an internet link without breaking the loop that makes the model learn?
ResBM makes the handoff 128 times smaller and leaves the learning path intact
The paper's answer, in one idea and six numbers.
ResBM puts a small learned encoder at the end of each stage and a decoder at the start of the next. Only the encoder's output crosses the network: 32 numbers per token in place of 4,096. A narrowed copy of the residual path runs alongside, untouched, so the corrections in the backward pass still reach every layer cleanly. The pipeline then runs over consumer-grade links at close to datacenter throughput.
ResBM: Residual Bottleneck Models for Low-Bandwidth Pipeline Parallelism, by Alan Aboudib, Rodrigo Lopez Portillo A., Kalei Brady and Steffen Cruz of Macrocosmos. It is the compression block inside IOTA, the training network on Bittensor subnet 9, and the reason the Orion runs could split one model across ordinary internet connections.
- Activation compression
- 128xd 4,096 to d' 32stated
- Speed on 80 Mbps
- 11.2x to 12.8xMuon 6,795, AdamW 7,837, vs 609 tokens/s; Appendix A.2stated
- Against a 10 Gbps datacenter
- 90% to 104%Muon 6,795, AdamW 7,837, vs 7,530 tokens/sinferred
- Traffic per step
- 896 KiBdown from 112 MiBstated
- Extra parameters
- 3.3%63.4M on a 2B modelstated
- Largest model tested
- 2B8 blocks, 26B tokensstated
Inside one boundary
After the paper's Figure C.1Compare this to the pipeline diagram above. Each arrow that crossed a machine boundary now carries 32 numbers per token instead of 4,096. The green line is why training still works: it is the unobstructed route the backward pass depends on.
The next section opens it up: how the pieces fit, what the tests measured, and where the evidence runs out.
Five changes to a standard model, and what the tests show
The mechanism in detail, the measurements, and the limits of what was tested.
Built from five changes
- Partition. A 2B-parameter Llama-3-style model with qk-norm, 8 blocks, hidden size 4,096, one block per GPU.stated
- Insert the bottleneck. An encoder-decoder pair at each of the seven stage boundaries, 3.3 percent more parameters.stated
- Project the skip connection. The residual is truncated to 32 dimensions on the way out and zero-padded on the way in, so the identity path is narrowed, never routed through a learned layer.stated
- Train end to end. Encoder and decoder weights update on the same step as everything else. No manifold projection, no optimizer modification.stated
- Use Muon to keep the bottleneck higher rank. AdamW tends to collapse a layer's weights onto a few dominant directions, which wastes a 32-wide channel. Muon orthogonalizes each update and spreads energy across dimensions. At the first boundary the bottleneck's effective rank was 19.95 of 32 under Muon against 10.26 under AdamW; the seven-layer means were close, 13.96 against 13.22.stated The authors call this spectral analysis preliminary.
Why it trains
Residual connections let each block add to its input instead of replacing it. That unobstructed path keeps gradients healthy in deep networks, and compressing it is what broke earlier schemes. ResBM folds the encoder into the residual branch of stage l and the decoder into the residual branch of stage l+1, and projects the skip connection with a rectangular identity, so the identity property holds on the dimensions that carry the most signal.stated
The premise comes from prior work: AdamW-trained transformers show rank collapse, so activations already live in a low-dimensional subspace, and the residual stream should be compressible too.stated
x' = P(x) + F(x), with P = I(d',d) a rectangular identity that truncates or zero-padsIt runs at datacenter speed over 80 Mbps
2B model · 8 A10G · GPipe over NCCL · 12 hours on C4Tokens a second, Table 3.stated The AdamW variant beats the uncompressed 10 Gbps baseline; profiling shows both compute and communication get faster, likely because smaller tensors are cheaper to move through memory as well as the network. Muon costs more compute per step, not more traffic. From 800 Mbps to 10 Gbps the gain is flat at about 1x: once compute dominates, compression neither helps nor hurts.stated
And it learns as well as the baseline
26B tokens on C4 · lower is better| Model | Optimizer | Bottleneck d' | Compression | Perplexity |
|---|---|---|---|---|
| Uncompressed baseline | AdamW | 4,096 | 1x | 21.75 |
| ResBM | Muon | 40 | 100x (4,096 to 40) | 21.60 |
| ResBM | Muon | 32 | 128x | 21.77 |
Final perplexity, Table 2.stated 26B tokens is about 65 percent of the compute-optimal budget for this size. The baseline converges faster early; both compressed models match it by the end. Against the authors' implementation of Subspace Models, ResBM wins at both ratios under both optimizers.stated
Where the claims stop
The mechanism holds on a 2B model with caveats. The last question is what it would change if it holds at the scale that matters.
If it scales, a home connection becomes enough to train on
Back to the loop at the top of the page, and what changes about who can run it.
The training loop from the first diagram does not change. What changes is the cost of the arrows that cross machine boundaries. Replica syncing was already solvable over slow links; ResBM addresses the pipeline handoff, the part that decides whether a model too large for one card can be trained on scattered, cheap ones, and it does so inside the architecture, so the training loop stays standard.stated
Scale works in its favor. Per-stage compute grows with the cube of the hidden size while boundary traffic grows with its square (the "square-cube law" of Ryabinin et al.), so the compression needed to hide communication shrinks as models grow.stated If ResBM holds there, the useful GPU is any decent card with a home connection. That is the premise under IOTA and under the "Liquid Compute" pitch Macrocosmos made from the Exploit Summit stage.inferred
Still open: results at 100B and beyond in a paper with a baseline, a head-to-head against a Muon-trained uncompressed model, and non-uniform bottlenecks. The authors suggest giving early layers more dimensions.stated Where Macrocosmos has run it for real, the numbers are in the Orion Register.
Where this paper sits
All four papers- Aug 2024SN9 pretraining whitepaperMiners each train a whole model; the best one takes the reward.Training line
- Jul 2025IOTA technical primerOne model split across miners, each paid for their share of the work.Training line
- Jan 2026Generative Adversarial MiningSide track: an incentive design for Apex, subnet 1, where quality has no scoring rule.Incentive track
- Apr 2026ResBMThe handoff between machines made 128 times smaller.You are here
Glossary · 10 terms
Sources · 4
- Aboudib, Lopez Portillo A., Brady, Cruz. ResBM: Residual Bottleneck Models for Low-Bandwidth Pipeline Parallelism. arXiv 2604.11947v1, 13 April 2026. Every stated label is to this paper, except the Orion-100B figure.
- Cruz. The Orion-100B write-up, 1 June 2026. Source of the 64x figure.
- Ramasinghe et al. Protocol Models, arXiv 2506.01260, 2025. The Subspace Models approach ResBM compares against.
- Quinque et al. IOTA: A Technical Primer for Release, arXiv 2507.17766, 2025. The earlier bottleneck design in IOTA.