Sangama (संगम) means confluence: computers coming together to run one AI model.
How it works | Architecture | Results | Hardware | Roadmap | Contact | Source code | Discord
Sangama is an open-source program, written in Rust, that splits a language model's layers across several computers: Macs, GPUs and ordinary CPUs. A model too large for any one device can then run across several, with encrypted, invitation-only connections between them.
Status: experimental. It runs on an invite-only network of trusted computers. It has run Qwen3.5-397B across 20 GPUs with 16 GB each (results), and is tested in depth with a small model (Qwen2.5-0.5B-Instruct). It is released under the MIT License.
A language model predicts each next piece of text (a token) by passing numbers through a stack of layers. Sangama gives each computer a contiguous slice of those layers.
+----------+ tokens +-------------------+ activations +-------------------+
| Client | ----------> | Worker A | ------------> | Worker B |
| | | layers 0 - 11 | | layers 12 - 23 |
| | | (e.g. a Mac) | | (e.g. a GPU) |
+----------+ +-------------------+ +-------------------+
^ |
| next token, then repeat |
+-----------------------------------------------------------------+
The model's weights stay loaded on the workers; they are not sent over the network for each request.
Sangama keeps three questions apart: who may join, how computers reach each other, and where the model runs.
+---------------------------------------------+
| Portal |
| invitations, signed membership, revocation |
+---------------------------------------------+
| | |
signed membership list, checked every few seconds
v v v
+----------+ +----------+ +----------+
| Client | <----> | Relay | <---> | Workers |
+----------+ +----------+ +----------+
| (only when routers ^
| block a direct path) |
+-------------------------------------+
direct encrypted connection
| Part | What it does |
|---|---|
| Worker | Contributes memory and compute. Loads its share of the model and runs those layers. |
| Client | Finds workers, checks they cover the whole model, and runs requests through them. |
| Portal | Issues single-use invitations, keeps the membership list, and revokes members. |
| Relay | Passes encrypted traffic between computers when their routers block a direct connection. It never runs model layers and cannot read the traffic. |
| Discovery | Workers advertise which part of the model they hold, in signed records. |
Every computer has its own Ed25519 identity, and all traffic is encrypted end to end (Noise over libp2p). Workers check every model file against a pinned SHA-256 list, so a route cannot silently use a different model.
A limit to know about: a worker can see the data it processes, so today's network is meant for trusted groups. Signatures prove who is a member, not that a computation was done correctly.
Test write-ups, newest first. All results.
| Hardware | How | Status |
|---|---|---|
| Apple Silicon Macs | Metal, using unified memory | Tested |
| NVIDIA GPUs | CUDA or Vulkan on Linux, one GPU per worker | Tested |
| Any CPU | macOS and Linux | Tested |
| AMD, Intel, Qualcomm GPUs | Vulkan or ROCm, through llama.cpp | Built, not yet tested |
| Windows | CUDA or Vulkan | Not yet tested |
Want to contribute a GPU, join a test network, or help build Sangama? Join the Discord or write to hello@sangama.co.
Source code: github.com/devdil/sangama · Contributor guide
Sangama is open source under the MIT License.