Sangama → Results
Qwen3.5-397B on 20 GPUs, through a relay
30 September 2026
We ran Qwen3.5-397B-A17B, one of the largest open models, split across 20 rented GPUs with 16 GB each. Every connection between computers was forced through a relay, as it would be for machines behind home routers. No single machine held more than 3 of the model's 60 layers.
Result: it works, and gives the same answer every time. Our first run decoded at 0.6 tokens per second. After five rounds of changes it runs at 3.7 to 4.8 tokens per second, with the first token in 1.6 to 4.3 seconds and token-for-token identical output.
Setup
| Model | Qwen3.5-397B-A17B, 4-bit (Q4_K_M), published as per-layer slices with checksums |
| Split | 20 stages of 3 layers each; each worker downloads only its own layers |
| Workers | 20 rented GPUs in California (12 × RTX 4090, 8 × RTX 5090), each capped at 16 GB |
| Engine | llama.cpp on CUDA |
| Network | Invitation-only, encrypted; every hop forced through one relay (no direct paths) |
client -> relay -> worker 1 -> relay -> worker 2 -> ... -> relay -> worker 20
|
client <--------------------- next token ------------------------------+
Speed
Decode speed in tokens per second, then time to first token:
| Test | First run | Round 2 | Round 4 | Round 5 (now) |
| Short answer | 0.58 / 7.5 s | 1.98 / 4.8 s | 2.34 / 2.3 s | 4.14 / 1.6 s |
| 128-token answer | 0.57 / 7.9 s | 1.81 / 3.6 s | 2.00 / 4.5 s | 4.35 / 2.3 s |
| Code | 0.67 / 5.8 s | 2.28 / 3.9 s | 2.09 / 2.9 s | 4.24 / 1.9 s |
| Factual answer | 0.62 / 5.8 s | 2.45 / 3.6 s | 2.25 / 2.8 s | 3.99 / 1.9 s |
| 99-token prompt | 0.67 / 11.1 s | 2.35 / 6.8 s | 2.05 / 11.5 s | 3.85 / 2.8 s |
| 237-token prompt | 0.63 / 14.1 s | 2.36 / 11.8 s | 1.94 / 13.8 s | 3.72 / 4.2 s |
| 459-token prompt | timed out | 2.64 / 22.6 s | 2.16 / 8.5 s | 4.77 / 4.3 s |
Single runs, greedy decoding. Every configuration produced exactly the same tokens as the first full-precision run.
Where the time goes
The GPUs spend 21 milliseconds on each token, all 20 together. The model activates only 17 billion of its 397 billion parameters per token, so the computing is cheap. Almost everything else is the network: a token must pass through 20 computers in order, and each hop crosses the relay.
- First run: about 1,500 ms per token, 98.6% of it network.
- Now: 210-270 ms per token.
- GPUs mostly idle: during generation each GPU used about 12.5 GB of memory, drew about 63 W, and showed near 0% utilisation.
- Tiny traffic: each worker sent and received about 0.1 Mbit/s.
So a 397B model fits on 20 machines with 16 GB each, and those machines sit mostly idle. The limit is how long each message takes to cross the network, not bandwidth or computing power.
What made it faster
- Push-forward (round 2). Each worker used to wait for the rest of the chain to finish before replying, so every token crossed every hop twice. Now each worker passes its result on and replies at once, and the client collects the token from the last worker: 2.2×.
- Half-size activations (round 2). The numbers passed between workers are sent in 16-bit (BF16) instead of 32-bit format; the model was trained in BF16. Another 1.4-2×, and long prompts stopped timing out.
- Reliable reconnects (round 3). After a worker restarted, the client tried an old cached address and never recovered. It now always dials through the relay, and recovered by itself after all 20 workers were restarted.
- Pipelined prompts (round 4). Long prompts are sent in 64-token pieces, so different workers process different pieces at the same time. The first token of a 459-token prompt came 2.7× sooner.
- Nearby workers (round 5). We measured each worker's distance to the relay and replaced the three farthest (36-49 ms away) with machines in the relay's datacenter. Decode speed doubled again.
What the testing caught
- Relay rate limits. Default relay rate limits silently disconnected most workers within minutes when many computers shared one network address. Fixed with explicit limits.
- Discovery traffic. Every worker was asking every other worker what it held, through the relay: about 170 connections where the route needs 39. Each request took about a second. Workers behind a relay now ask the relay only, and a request takes 2-40 ms.
- A dead worker blocked the others. After one worker vanished, the client could not reach the healthy ones until it was restarted. Fixed in round 3.
- Rented machines come and go. One relay box died, and one host reclaimed two GPUs mid-test. Replacements downloaded their layers and joined in 5-15 minutes. The output stayed identical.
Next
- Speculative decoding. Guess a few tokens ahead, so each trip through the chain can confirm several tokens at once.
- Automatic placement. Have the scheduler pick workers close to the relay, or to each other, as we did by hand in round 5.
- Direct connections on home networks. Our rented machines could not accept incoming connections, so hole punching could not be tested here.
Full report, per-machine metrics and raw data: docs/test-results.
Sangama is open source under the MIT License. Contact: hello@sangama.co · Discord