Sangama (संगम) means confluence: computers coming together to run one AI model.
How it works | Architecture | Results | Hardware | Roadmap | Contact | Source code
Sangama is an open-source program, written in Rust, that splits a language model's layers across several computers: Macs, GPUs and ordinary CPUs. A model too large for any one device can then run across several, with encrypted, invitation-only connections between them.
Status: experimental. It runs on an invite-only network of trusted computers, and is tested with a small model (Qwen2.5-0.5B-Instruct). It is released under the MIT License.
A language model predicts each next piece of text (a token) by passing numbers through a stack of layers. Sangama gives each computer a contiguous slice of those layers.
+----------+ tokens +-------------------+ activations +-------------------+
| Client | ----------> | Worker A | ------------> | Worker B |
| | | layers 0 - 11 | | layers 12 - 23 |
| | | (e.g. a Mac) | | (e.g. a GPU) |
+----------+ +-------------------+ +-------------------+
^ |
| next token, then repeat |
+-----------------------------------------------------------------+
The model's weights stay loaded on the workers; they are not sent over the network for each request.
Sangama keeps three questions apart: who may join, how computers reach each other, and where the model runs.
+---------------------------------------------+
| Portal |
| invitations, signed membership, revocation |
+---------------------------------------------+
| | |
signed membership list, checked every few seconds
v v v
+----------+ +----------+ +----------+
| Client | <----> | Relay | <---> | Workers |
+----------+ +----------+ +----------+
| (only when routers ^
| block a direct path) |
+-------------------------------------+
direct encrypted connection
| Part | What it does |
|---|---|
| Worker | Contributes memory and compute. Loads its share of the model and runs those layers. |
| Client | Finds workers, checks they cover the whole model, and runs requests through them. |
| Portal | Issues single-use invitations, keeps the membership list, and revokes members. |
| Relay | Passes encrypted traffic between computers when their routers block a direct connection. It never runs model layers and cannot read the traffic. |
| Discovery | Workers advertise which part of the model they hold, in signed records. |
Every computer has its own Ed25519 identity, and all traffic is encrypted end to end (Noise over libp2p). Workers check every model file against a pinned SHA-256 list, so a route cannot silently use a different model.
A limit to know about: a worker can see the data it processes, so today's network is meant for trusted groups. Signatures prove who is a member, not that a computation was done correctly.
Measured on 29 September 2026 with the Qwen2.5-0.5B-Instruct test model, split across two workers. The full list of test scenarios and raw results is in the repository.
| Route (layers 0-11 → 12-23) | Where | Result | Speed |
|---|---|---|---|
| Candle Metal → Candle Metal | One Mac | Reference output | - |
| Candle CUDA → Candle CUDA | Two GPUs, one machine | Identical | 55 tokens/s |
| llama.cpp CUDA → llama.cpp CUDA | Two GPUs, one machine | Identical | 103 tokens/s |
| llama.cpp Vulkan → llama.cpp Vulkan | Two GPUs, one machine | Identical | 106 tokens/s |
| llama.cpp Vulkan → Candle CUDA | Two GPUs, one machine | Identical | 86 tokens/s |
| Candle Metal → Candle CUDA | Home Mac → US datacenter | Identical | 3.1 tokens/s |
| Candle CUDA → Candle Metal | US datacenter → home Mac | Identical | 1.2 tokens/s |
| Candle Metal → Candle CUDA, all through a relay | Home Mac, US GPU, relay | Identical, 7 of 7 | 0.6 tokens/s |
| 4-bit and full-precision workers mixed | Any | Refused, by design | - |
Speeds are single runs of a small test model and include network delay; they are not benchmarks. 4-bit output is only reproducible with the same weight file on the same kind of GPU, so a route must use one published file.
| Hardware | How | Status |
|---|---|---|
| Apple Silicon Macs | Metal, using unified memory | Tested |
| NVIDIA GPUs | CUDA or Vulkan on Linux, one GPU per worker | Tested |
| Any CPU | macOS and Linux | Tested |
| AMD, Intel, Qualcomm GPUs | Vulkan or ROCm, through llama.cpp | Built, not yet tested |
| Windows | CUDA or Vulkan | Not yet tested |
Want to contribute a GPU, join a test network, or help build Sangama? Write to hello@sangama.co.
Source code: github.com/devdil/sangama · Contributor guide
Sangama is open source under the MIT License.