30 September 2026
The previous test ran Qwen3.5-397B across 20 rented GPUs with 16 GB each, every connection through a relay. It served one request at a time, and the GPUs sat almost idle. This test asks three questions about the same pipeline:
Result:
This fleet is slower per request than the previous one (1.25 against 5.4 tokens per second), because its machines are spread over three US states. The previous fleet's provider ended its listings that morning. The comparisons below are within this fleet.
| Model | Qwen3.5-397B-A17B, 4-bit (Q4_K_M), 20 stages of 3 layers each |
|---|---|
| Workers | 20 rented GPUs (RTX 4090, 5090, 3090 Ti, 5080, L4), each capped at 16 GB; the last stage gets 30 GB when it holds the draft layer |
| Network | Every hop through one relay in Nevada; workers in Nevada, California and Utah, 1 to 56 ms from it |
| One request | About 800 ms per token, 23 ms of it GPU computing |
A token spends about 1 ms on each GPU and the rest of its 800 ms on the network, so each machine can work on other requests in the meantime. Each worker now keeps up to 96 separate sessions in memory.
| Requests at once | Total tokens per second | Each request, tokens per second | Same tokens as alone |
|---|---|---|---|
| 1 | 1.3 | 1.25 | 1 of 1 |
| 4 | 6.2 | 1.55 | 4 of 4 |
| 16 | 30 | 1.88 | 16 of 16 |
| 32 | 60 | 1.87 | 32 of 32 |
| 51 | 90.9 | 1.79 | 51 of 51 |
| 64 | 56.4 | 0.88 | 64 of 64 |
| 96 | 67.6 | 0.70 | 96 of 96 |
Total is the sum of each request's decode speed once all are running. Counting prompts and start-up as well, 32 requests gave 35.6 and 96 gave 56.6 tokens per second. The 51-request row is a run of 96 in which 45 clients failed to start; the other 51 ran together.
Qwen3.5 has a small extra layer trained to predict the token after next. The last machine in the pipeline uses it to guess a few tokens ahead and sends the guesses back with each token. The next trip through the 20 machines checks all the guesses at once, keeps the ones the model agrees with, and discards the rest.
| Prompt | Guesses per step | Tokens per trip | Guesses accepted | Tokens per second | Speed-up |
|---|---|---|---|---|---|
| Code | none | 1.0 | - | 1.25 | - |
| 2 | 2.8 | 41 of 42 | 1.82 | 1.5× | |
| 4 | 4.3 | 49 of 55 | 2.29 | 1.8× | |
| 6 | 5.3 | 52 of 66 | 2.69 | 2.2× | |
| Explanation | none | 1.0 | - | 1.24 | - |
| 2 | 2.2 | 35 of 55 | 1.43 | 1.2× | |
| 6 | 3.2 | 44 of 105 | 1.60 | 1.3× | |
| Factual | none | 1.0 | - | 1.22 | - |
| 2 | 2.6 | 39 of 47 | 1.65 | 1.4× | |
| 6 | 2.9 | 42 of 121 | 1.44 | 1.2× |
Many requests at once, each with 4 guesses per step:
| Requests at once | Total tokens per second, no guesses | With guesses |
|---|---|---|
| 16 | 21.7 | 38.9 |
| 32 | 57.1 | 49.9 |
| 32, repeated | 56.5 | 59.6 |
| 48 | 51.3 | 83.1 |
Every request completed, and 56% of guesses were accepted, the same as for one request. But the gain is not consistent: 1.8× at 16 requests, 1.6× at 48, and none at 32. Each figure is a single run, and plain runs on these shared rented machines varied by as much between runs. We need more repeats before quoting a speed-up under load.
Runs with guesses were fluent, but often not token-for-token identical to plain runs. They differ at close word choices, for example "iteratively." against "using an iterative approach.".
Drafting does not cause this. Plain decoding gave different tokens at the same place when we only changed how many prompt tokens were sent at a time (8 instead of 64). Repeating a run with the same settings always gave the same tokens. Processing several positions together changes the arithmetic very slightly, and in this model that can flip a near-tie.
Our earlier page said every configuration produced exactly the same tokens. That was true for the prompts and settings we tested then, but it is not true in general for this model.
Full report and raw data: docs/test-results.
Sangama is open source under the MIT License. Contact: hello@sangama.co · Discord