Sangama → Results

Qwen3.5-397B in three large stages

30 September 2026

Our earlier tests split Qwen3.5-397B across twenty GPUs with 16 GB each. A token spent almost all its time on the network, crossing 21 hops. This test asks the opposite question: how fast is the same software when the hops are few and close?

We ran the model as three stages of 20 layers on three 96 GB GPUs in one datacenter.

Result:

This is not the home-computer case. These are datacenter cards that rent for about $1.90 an hour each. It shows what the design can do when distance is removed, and where the next limits are.


Setup

ModelQwen3.5-397B-A17B, 4-bit (Q4_K_M), 3 stages of 20 layers, about 82 GB each
Workers3 × RTX PRO 6000 Blackwell with 96 GB, in one datacenter
Relay and clientOne more machine in the same datacenter
NetworkFirst through the relay, then with direct connections between all machines
CostAbout $6.50 an hour for everything, half of the twenty-machine fleet

One request

Tokens per second, 128 tokens per run. "Guesses" are tokens the model's own draft layer proposes; the next trip through the machines checks them all at once.

Guesses per stepCodeExplanationFactual answer
none48-5048-5048-50
2867478
497-10576-7975-77
6105-10764-6671-73

Where the time goes now

Per trip through the three machinesPlainWith 6 guesses
GPU computing14.3 ms31.0 ms
Waiting inside the workers, including undoing wrong guesses0.3 ms0.4 ms
Network and handling, four hops6.2 ms5.6 ms
Total20.7 ms36.1 ms

On the twenty-machine fleets, computing was 1 to 3% of each token's time. Here it is most of it. Checking guesses costs more only because the model does more work per trip.

Undoing wrong guesses is now free

In our previous test, every machine saved a full copy of its state before each trip that checked guesses, about 36 ms per machine, whether or not any guess was wrong. That is why 5.3 tokens per trip gave only 2.2× the speed.

The model runtime can instead keep the last few states on the GPU and step back to one of them. We now use that. Waiting inside the workers, with the undo included, is 0.2 to 0.4 ms per trip.

Many requests at once

Requests at onceTotal tokens per secondEach requestSame tokens as alone
151.551.51 of 1
4118.529.64 of 4
16118.87.416 of 16
48118.92.548 of 48

With 4 guesses per step the total rose to 156 (4 requests), 177 (8) and 183 (16).

The limit is one GPU. We added timing to every stage, and it shows the queue directly: with 48 requests, a frame waited 344 ms at the last machine and almost nothing at the other two. That machine takes 8.1 ms per frame and handles one at a time, which is about 120 per second. The fix is to let it work on several requests in one step.

This is a different limit from the one on the twenty-machine fleet, where the total stopped near 90 with every machine idle. That one is still unexplained.

Direct connections

Machines with open ports can now connect to each other directly and keep the relay as a fallback. All machines here did, and the network share of each trip fell from 6.2 to 5.1 ms. The relay was in the same room, so this test cannot show what direct connections save across a distance. That needs a spread-out fleet.

What the testing caught

Limits of this test

Next

Full report and raw data: docs/test-results.


Sangama is open source under the MIT License. Contact: hello@sangama.co · Discord