Sangama

Sangama (संगम) means confluence: computers coming together to run one AI model.

How it works | Architecture | Results | Hardware | Roadmap | Contact | Source code


Sangama is an open-source program, written in Rust, that splits a language model's layers across several computers: Macs, GPUs and ordinary CPUs. A model too large for any one device can then run across several, with encrypted, invitation-only connections between them.

Status: experimental. It runs on an invite-only network of trusted computers, and is tested with a small model (Qwen2.5-0.5B-Instruct). It is released under the MIT License.


How it works

A language model predicts each next piece of text (a token) by passing numbers through a stack of layers. Sangama gives each computer a contiguous slice of those layers.

  +----------+   tokens    +-------------------+  activations  +-------------------+
  |  Client  | ----------> |     Worker A      | ------------> |     Worker B      |
  |          |             |  layers 0 - 11    |               |  layers 12 - 23   |
  |          |             |  (e.g. a Mac)     |               |  (e.g. a GPU)     |
  +----------+             +-------------------+               +-------------------+
       ^                                                                 |
       |                     next token, then repeat                     |
       +-----------------------------------------------------------------+
  1. Contribute. Each worker loads only the slice of the model it serves, sized to the memory it offers.
  2. Route. A client checks that its workers cover every layer in order, reserves them, and sends the prompt through the chain.
  3. Generate. The last worker picks the next token and sends it back. Each worker keeps its own conversation state, so only a few kilobytes cross the network for each token.

The model's weights stay loaded on the workers; they are not sent over the network for each request.


Architecture

Sangama keeps three questions apart: who may join, how computers reach each other, and where the model runs.

                  +---------------------------------------------+
                  |                   Portal                    |
                  | invitations, signed membership, revocation  |
                  +---------------------------------------------+
                        |                   |                  |
                  signed membership list, checked every few seconds
                        v                   v                  v
                  +----------+        +----------+       +----------+
                  |  Client  | <----> |  Relay   | <---> | Workers  |
                  +----------+        +----------+       +----------+
                        |          (only when routers         ^
                        |           block a direct path)      |
                        +-------------------------------------+
                            direct encrypted connection
PartWhat it does
WorkerContributes memory and compute. Loads its share of the model and runs those layers.
ClientFinds workers, checks they cover the whole model, and runs requests through them.
PortalIssues single-use invitations, keeps the membership list, and revokes members.
RelayPasses encrypted traffic between computers when their routers block a direct connection. It never runs model layers and cannot read the traffic.
DiscoveryWorkers advertise which part of the model they hold, in signed records.

Every computer has its own Ed25519 identity, and all traffic is encrypted end to end (Noise over libp2p). Workers check every model file against a pinned SHA-256 list, so a route cannot silently use a different model.

A limit to know about: a worker can see the data it processes, so today's network is meant for trusted groups. Signatures prove who is a member, not that a computation was done correctly.


Results

Measured on 29 September 2026 with the Qwen2.5-0.5B-Instruct test model, split across two workers. The full list of test scenarios and raw results is in the repository.

Route (layers 0-11 → 12-23)WhereResultSpeed
Candle Metal → Candle MetalOne MacReference output-
Candle CUDA → Candle CUDATwo GPUs, one machineIdentical55 tokens/s
llama.cpp CUDA → llama.cpp CUDATwo GPUs, one machineIdentical103 tokens/s
llama.cpp Vulkan → llama.cpp VulkanTwo GPUs, one machineIdentical106 tokens/s
llama.cpp Vulkan → Candle CUDATwo GPUs, one machineIdentical86 tokens/s
Candle Metal → Candle CUDAHome Mac → US datacenterIdentical3.1 tokens/s
Candle CUDA → Candle MetalUS datacenter → home MacIdentical1.2 tokens/s
Candle Metal → Candle CUDA, all through a relayHome Mac, US GPU, relayIdentical, 7 of 70.6 tokens/s
4-bit and full-precision workers mixedAnyRefused, by design-

Speeds are single runs of a small test model and include network delay; they are not benchmarks. 4-bit output is only reproducible with the same weight file on the same kind of GPU, so a route must use one published file.

What the testing caught


Hardware

HardwareHowStatus
Apple Silicon MacsMetal, using unified memoryTested
NVIDIA GPUsCUDA or Vulkan on Linux, one GPU per workerTested
Any CPUmacOS and LinuxTested
AMD, Intel, Qualcomm GPUsVulkan or ROCm, through llama.cppBuilt, not yet tested
WindowsCUDA or VulkanNot yet tested

Roadmap


Contact

Want to contribute a GPU, join a test network, or help build Sangama? Write to hello@sangama.co.

Source code: github.com/devdil/sangama · Contributor guide


Sangama is open source under the MIT License.