Build on the crowd.
Brain speaks the OpenAI Chat Completions format. Change the base URL, set model: "brain/auto", and the router does the rest.
from openai import OpenAI
client = OpenAI(
base_url="https://YOUR_BRAIN_HOST/v1",
api_key="YOUR_BRAIN_API_KEY",
)
res = client.chat.completions.create(
model="brain/auto",
messages=[{"role": "user", "content": "Summarize this audit report."}],
)
print(res.choices[0].message.content)
print(res.model_extra["brain"]["target"]) # BROWSER_NETWORK | CLOUD_FALLBACK | EXTERNAL_MODEL_PROVIDERRequest path
Target status below is read from this server's configuration at request time. Keys stay in server environment variables and are never sent to the browser.
No target can serve LLM requests on this server right now; the gateway returns 503 with the routing trace instead of a fabricated answer.
Routing
Ineligible targets are filtered out, the rest ranked by a weighted score. Weights are configurable per request class. Every response carries the full decision in brain.routing.
Response
Standard OpenAI shape, plus a brain extension that tells you where the request ran and why. SDKs ignore unknown fields.
{
"id": "chatcmpl-3f9c…",
"object": "chat.completion",
"model": "brain/auto",
"choices": [{ "index": 0, "message": { "role": "assistant", "content": "…" }, "finish_reason": "stop" }],
"usage": { "prompt_tokens": 412, "completion_tokens": 233, "total_tokens": 645 },
"brain": {
"target": "CLOUD_FALLBACK",
"provider": "cloud-fallback",
"latencyMs": 1184,
"routing": { "ranked": [ … ], "selected": { … } },
"attempts": [{ "providerId": "cloud-fallback", "ok": true }],
"plan": { "provenance": "simulated", "shards": [ … ] }
}
}Node protocol
What a contributing browser does. Session tokens are random, stored only as hashes, and bound to one node.
Contributors are adversarial.
The server never trusts a client's claimed GPU, score, uptime or results. Rewards follow verified work only.
Benchmarks and jobs are generated server-side from secret seeds. Clients cannot pick their own work.
Compute score = verified ops ÷ server-measured wall time. The client's timing is recorded but never trusted.
The server recomputes 6 randomly chosen rows / blocks it never reveals. One mismatch fails the job.
20% of jobs are small canaries with a fully known answer.
Results returned faster than physically possible for the workload are rejected.
EWMA over outcomes; failures weigh double. Below 0.35 the node is banned.
Per hashed IP, per minute: 600 node calls, 60 API calls, 20 playground runs.
Same unit to N nodes, majority result wins. Policy + comparison implemented; dispatcher wiring lands with LLM shards.