30 September 2026
Our earlier tests split Qwen3.5-397B across twenty GPUs with 16 GB each. A token spent almost all its time on the network, crossing 21 hops. This test asks the opposite question: how fast is the same software when the hops are few and close?
We ran the model as three stages of 20 layers on three 96 GB GPUs in one datacenter.
Result:
This is not the home-computer case. These are datacenter cards that rent for about $1.90 an hour each. It shows what the design can do when distance is removed, and where the next limits are.
| Model | Qwen3.5-397B-A17B, 4-bit (Q4_K_M), 3 stages of 20 layers, about 82 GB each |
|---|---|
| Workers | 3 × RTX PRO 6000 Blackwell with 96 GB, in one datacenter |
| Relay and client | One more machine in the same datacenter |
| Network | First through the relay, then with direct connections between all machines |
| Cost | About $6.50 an hour for everything, half of the twenty-machine fleet |
Tokens per second, 128 tokens per run. "Guesses" are tokens the model's own draft layer proposes; the next trip through the machines checks them all at once.
| Guesses per step | Code | Explanation | Factual answer |
|---|---|---|---|
| none | 48-50 | 48-50 | 48-50 |
| 2 | 86 | 74 | 78 |
| 4 | 97-105 | 76-79 | 75-77 |
| 6 | 105-107 | 64-66 | 71-73 |
| Per trip through the three machines | Plain | With 6 guesses |
|---|---|---|
| GPU computing | 14.3 ms | 31.0 ms |
| Waiting inside the workers, including undoing wrong guesses | 0.3 ms | 0.4 ms |
| Network and handling, four hops | 6.2 ms | 5.6 ms |
| Total | 20.7 ms | 36.1 ms |
On the twenty-machine fleets, computing was 1 to 3% of each token's time. Here it is most of it. Checking guesses costs more only because the model does more work per trip.
In our previous test, every machine saved a full copy of its state before each trip that checked guesses, about 36 ms per machine, whether or not any guess was wrong. That is why 5.3 tokens per trip gave only 2.2× the speed.
The model runtime can instead keep the last few states on the GPU and step back to one of them. We now use that. Waiting inside the workers, with the undo included, is 0.2 to 0.4 ms per trip.
| Requests at once | Total tokens per second | Each request | Same tokens as alone |
|---|---|---|---|
| 1 | 51.5 | 51.5 | 1 of 1 |
| 4 | 118.5 | 29.6 | 4 of 4 |
| 16 | 118.8 | 7.4 | 16 of 16 |
| 48 | 118.9 | 2.5 | 48 of 48 |
With 4 guesses per step the total rose to 156 (4 requests), 177 (8) and 183 (16).
The limit is one GPU. We added timing to every stage, and it shows the queue directly: with 48 requests, a frame waited 344 ms at the last machine and almost nothing at the other two. That machine takes 8.1 ms per frame and handles one at a time, which is about 120 per second. The fix is to let it work on several requests in one step.
This is a different limit from the one on the twenty-machine fleet, where the total stopped near 90 with every machine idle. That one is still unexplained.
Machines with open ports can now connect to each other directly and keep the relay as a fallback. All machines here did, and the network share of each trip fell from 6.2 to 5.1 ms. The relay was in the same room, so this test cannot show what direct connections save across a distance. That needs a spread-out fleet.
Full report and raw data: docs/test-results.
Sangama is open source under the MIT License. Contact: hello@sangama.co · Discord