30 September 2026
We ran Qwen3.5-397B-A17B, one of the largest open models, split across 20 rented GPUs with 16 GB each. Every connection between computers was forced through a relay, as it would be for machines behind home routers. No single machine held more than 3 of the model's 60 layers.
Result: it works, and gives the same answer every time. Our first run decoded at 0.6 tokens per second. After six rounds of changes it runs at 5.0 to 5.5 tokens per second, with the first token in 1.3 to 6.2 seconds and token-for-token identical output.
| Model | Qwen3.5-397B-A17B, 4-bit (Q4_K_M), published as per-layer slices with checksums |
|---|---|
| Split | 20 stages of 3 layers each; each worker downloads only its own layers |
| Workers | 20 rented GPUs in California (12 × RTX 4090, 8 × RTX 5090), each capped at 16 GB |
| Engine | llama.cpp on CUDA |
| Network | Invitation-only, encrypted; every hop forced through one relay (no direct paths) |
client -> relay -> worker 1 -> relay -> worker 2 -> ... -> relay -> worker 20
|
client <--------------------- next token ------------------------------+
Decode speed in tokens per second, then time to first token:
| Test | First run | Round 2 | Round 4 | Round 5 | Round 6 (now) |
|---|---|---|---|---|---|
| Short answer | 0.58 / 7.5 s | 1.98 / 4.8 s | 2.34 / 2.3 s | 4.14 / 1.6 s | 5.46 / 1.3 s |
| 128-token answer | 0.57 / 7.9 s | 1.81 / 3.6 s | 2.00 / 4.5 s | 4.35 / 2.3 s | 5.42 / 1.9 s |
| Code | 0.67 / 5.8 s | 2.28 / 3.9 s | 2.09 / 2.9 s | 4.24 / 1.9 s | 5.49 / 1.4 s |
| Factual answer | 0.62 / 5.8 s | 2.45 / 3.6 s | 2.25 / 2.8 s | 3.99 / 1.9 s | 4.94 / 1.7 s |
| 99-token prompt | 0.67 / 11.1 s | 2.35 / 6.8 s | 2.05 / 11.5 s | 3.85 / 2.8 s | 5.07 / 2.1 s |
| 237-token prompt | 0.63 / 14.1 s | 2.36 / 11.8 s | 1.94 / 13.8 s | 3.72 / 4.2 s | 4.98 / 4.0 s |
| 459-token prompt | timed out | 2.64 / 22.6 s | 2.16 / 8.5 s | 4.77 / 4.3 s | 5.29 / 6.2 s |
Single runs, greedy decoding. Every configuration produced exactly the same tokens as the first full-precision run.
Correction, 30 September: that held for these prompts and settings, but not in general. A later test found that how tokens are batched can change this model's output at close word choices.
The GPUs spend 21 milliseconds on each token, all 20 together. The model activates only 17 billion of its 397 billion parameters per token, so the computing is cheap. Almost everything else is the network: a token must pass through 20 computers in order, and each hop crosses the relay.
So a 397B model fits on 20 machines with 16 GB each, and those machines sit mostly idle. The limit is how long each message takes to cross the network, not bandwidth or computing power.
Speculative decoding guesses several tokens ahead and has the model check all the guesses in one pass. When the guesses are right, one trip through all 20 workers yields several tokens instead of one. Since our time goes almost entirely to that trip, this was the obvious thing to try.
How it works here:
Correctness: every speculative run produced exactly the same tokens as normal decoding.
Speed: the model accepted few of these guesses, so saving, restoring and replaying on all 20 workers cost more than the trips it saved.
| Test | Guesses accepted | Without speculation | With speculation |
|---|---|---|---|
| Code | 7 of 38 (18%) | 2.1-2.7 tok/s | 1.3-1.4 tok/s |
| Summary | 6 of 24 (25%) | 1.6-2.8 tok/s | 0.6-1.2 tok/s |
| Essay | 2 of 23 (9%) | - | - |
Warm back-to-back runs, measured before round 6, while the network was slower than it is now.
Status: off by default. To pay off it needs better guesses: Qwen3.5 ships its own multi-token-prediction layer, reported to be right 70-80% of the time, which would mean roughly 2-2.4× faster decoding. That layer is not in our published slices yet. Snapshots also need to get cheaper, by saving only the recurrent state and trimming the rest.
Full report, per-machine metrics and raw data: docs/test-results.
Sangama is open source under the MIT License. Contact: hello@sangama.co · Discord