Sangama → Results

A small model across CPUs, GPUs and the Internet

29 September 2026

Measured on 29 September 2026 with the Qwen2.5-0.5B-Instruct test model, split across two workers. The full list of test scenarios and raw results is in the repository.

Route (layers 0-11 → 12-23)WhereResultSpeed
Candle Metal → Candle MetalOne MacReference output-
Candle CUDA → Candle CUDATwo GPUs, one machineIdentical55 tokens/s
llama.cpp CUDA → llama.cpp CUDATwo GPUs, one machineIdentical103 tokens/s
llama.cpp Vulkan → llama.cpp VulkanTwo GPUs, one machineIdentical106 tokens/s
llama.cpp Vulkan → Candle CUDATwo GPUs, one machineIdentical86 tokens/s
Candle Metal → Candle CUDAHome Mac → US datacenterIdentical3.1 tokens/s
Candle CUDA → Candle MetalUS datacenter → home MacIdentical1.2 tokens/s
Candle Metal → Candle CUDA, all through a relayHome Mac, US GPU, relayIdentical, 7 of 70.6 tokens/s
4-bit and full-precision workers mixedAnyRefused, by design-

Speeds are single runs of a small test model and include network delay; they are not benchmarks. 4-bit output is only reproducible with the same weight file on the same kind of GPU, so a route must use one published file.

What the testing caught


Sangama is open source under the MIT License. Contact: hello@sangama.co · Discord