29 September 2026
Measured on 29 September 2026 with the Qwen2.5-0.5B-Instruct test model, split across two workers. The full list of test scenarios and raw results is in the repository.
| Route (layers 0-11 → 12-23) | Where | Result | Speed |
|---|---|---|---|
| Candle Metal → Candle Metal | One Mac | Reference output | - |
| Candle CUDA → Candle CUDA | Two GPUs, one machine | Identical | 55 tokens/s |
| llama.cpp CUDA → llama.cpp CUDA | Two GPUs, one machine | Identical | 103 tokens/s |
| llama.cpp Vulkan → llama.cpp Vulkan | Two GPUs, one machine | Identical | 106 tokens/s |
| llama.cpp Vulkan → Candle CUDA | Two GPUs, one machine | Identical | 86 tokens/s |
| Candle Metal → Candle CUDA | Home Mac → US datacenter | Identical | 3.1 tokens/s |
| Candle CUDA → Candle Metal | US datacenter → home Mac | Identical | 1.2 tokens/s |
| Candle Metal → Candle CUDA, all through a relay | Home Mac, US GPU, relay | Identical, 7 of 7 | 0.6 tokens/s |
| 4-bit and full-precision workers mixed | Any | Refused, by design | - |
Speeds are single runs of a small test model and include network delay; they are not benchmarks. 4-bit output is only reproducible with the same weight file on the same kind of GPU, so a route must use one published file.
Sangama is open source under the MIT License. Contact: hello@sangama.co · Discord