Sangama → Results

Qwen3.5-397B on 20 GPUs, through a relay

30 September 2026

We ran Qwen3.5-397B-A17B, one of the largest open models, split across 20 rented GPUs with 16 GB each. Every connection between computers was forced through a relay, as it would be for machines behind home routers. No single machine held more than 3 of the model's 60 layers.

Result: it works, and gives the same answer every time. Our first run decoded at 0.6 tokens per second. After five rounds of changes it runs at 3.7 to 4.8 tokens per second, with the first token in 1.6 to 4.3 seconds and token-for-token identical output.


Setup

ModelQwen3.5-397B-A17B, 4-bit (Q4_K_M), published as per-layer slices with checksums
Split20 stages of 3 layers each; each worker downloads only its own layers
Workers20 rented GPUs in California (12 × RTX 4090, 8 × RTX 5090), each capped at 16 GB
Enginellama.cpp on CUDA
NetworkInvitation-only, encrypted; every hop forced through one relay (no direct paths)
 client -> relay -> worker 1 -> relay -> worker 2 -> ... -> relay -> worker 20
                                                                        |
 client <--------------------- next token ------------------------------+

Speed

Decode speed in tokens per second, then time to first token:

TestFirst runRound 2Round 4Round 5 (now)
Short answer0.58 / 7.5 s1.98 / 4.8 s2.34 / 2.3 s4.14 / 1.6 s
128-token answer0.57 / 7.9 s1.81 / 3.6 s2.00 / 4.5 s4.35 / 2.3 s
Code0.67 / 5.8 s2.28 / 3.9 s2.09 / 2.9 s4.24 / 1.9 s
Factual answer0.62 / 5.8 s2.45 / 3.6 s2.25 / 2.8 s3.99 / 1.9 s
99-token prompt0.67 / 11.1 s2.35 / 6.8 s2.05 / 11.5 s3.85 / 2.8 s
237-token prompt0.63 / 14.1 s2.36 / 11.8 s1.94 / 13.8 s3.72 / 4.2 s
459-token prompttimed out2.64 / 22.6 s2.16 / 8.5 s4.77 / 4.3 s

Single runs, greedy decoding. Every configuration produced exactly the same tokens as the first full-precision run.

Where the time goes

The GPUs spend 21 milliseconds on each token, all 20 together. The model activates only 17 billion of its 397 billion parameters per token, so the computing is cheap. Almost everything else is the network: a token must pass through 20 computers in order, and each hop crosses the relay.

So a 397B model fits on 20 machines with 16 GB each, and those machines sit mostly idle. The limit is how long each message takes to cross the network, not bandwidth or computing power.

What made it faster

  1. Push-forward (round 2). Each worker used to wait for the rest of the chain to finish before replying, so every token crossed every hop twice. Now each worker passes its result on and replies at once, and the client collects the token from the last worker: 2.2×.
  2. Half-size activations (round 2). The numbers passed between workers are sent in 16-bit (BF16) instead of 32-bit format; the model was trained in BF16. Another 1.4-2×, and long prompts stopped timing out.
  3. Reliable reconnects (round 3). After a worker restarted, the client tried an old cached address and never recovered. It now always dials through the relay, and recovered by itself after all 20 workers were restarted.
  4. Pipelined prompts (round 4). Long prompts are sent in 64-token pieces, so different workers process different pieces at the same time. The first token of a 459-token prompt came 2.7× sooner.
  5. Nearby workers (round 5). We measured each worker's distance to the relay and replaced the three farthest (36-49 ms away) with machines in the relay's datacenter. Decode speed doubled again.

What the testing caught

Next

Full report, per-machine metrics and raw data: docs/test-results.


Sangama is open source under the MIT License. Contact: hello@sangama.co · Discord