packages/swarm-mcp/src/token.nu — cluster identity + authenticity from one shared --token.
Every node in a cluster is launched with the same --token. The token does two things, both derived deterministically so no node has to coordinate:
token_group_id hashes the token into the 32-byte relaymulticast group id. Clusters with different tokens join different relay groups, so their HELLO/job traffic is mutually invisible even when they share one relay. Without the token you cannot derive the group id, so a stranger cannot even see (let alone join) the cluster's gossip.
token_key derives a 32-byte HMAC key. Every computepayload (a kernel chunk) and every result is tagged with HMAC-SHA256 (token_tag) and verified on receipt (token_untag, constant-time). A worker only acts on token-authentic jobs; a coordinator only accepts token-authentic results. Combined with group isolation this gives mutual authentication of all compute traffic over the dumb (opaque) relay.
The relay itself never needs the token — it forwards opaque bytes — so the secret never leaves the nodes that own it.
@ token_tag_len → iHMAC tag length prepended to authenticated messages (truncated SHA-256 MAC).
@ token_group_id s token → ( Vec u )32-byte relay multicast group id for this cluster (token isolation).
@ token_key s token → ( Vec u )32-byte HMAC key authenticating compute payloads + results.
@ token_tag ( Vec u ) key ( Vec u ) payload → ( Vec u )Prepend a 16-byte HMAC-SHA256 tag to payload. Returns a NEW vec; the caller owns it (and still owns payload).
@ token_untag ( Vec u ) key ( Vec u ) tagged → ?( Vec u )Verify the tag and strip it. Some(payload copy) on a valid tag, None on a missing/short/forged tag. The compare is constant-time. Caller owns the returned payload on the Some arm.
packages/swarm/src/census.nu — lightweight membership gossip for the swarm.
The compute layer (dist/job over dist/ring) needs every node to agree on the live worker set: a worker must know it owns a key, and the coordinator must know which pubkeys to route to. The full SWIM table (net/membership.nu) is the heavy, churn-hardened answer; this is the small one a compute cluster actually needs — a HELLO announcement that feeds the consistent-hash ring.
want, replies with their own HELLO so the newcomer learns them too;
Roles: a WORKER owns keys and executes handlers, so it joins everyone's ring; a CLIENT (the coordinator) only submits, so it is never added to the ring — otherwise it would own keys it has no handler for and silently drop results.
The codec and the membership set are pure; only the pump touches transport.
@ census_hello_t → i@ role_client → i@ role_worker → i@ cap_gpu → iCapability bits a worker advertises in HELLO. cap_gpu: the worker runs wasm chunks with GPU host imports enabled (wt --allow-gpu on real hardware).
@ hello_build i id i role i want ( Vec u ) pubkey i caps → ( Vec u )HELLO wire: [3][id:8][role:1][want:1][pklen:2][pubkey…][caps:1] The caps byte TRAILS the pubkey so a pre-caps decoder (fixed prefix + pklen-delimited pubkey) reads the same fields and ignores the tail; a caps decoder treats a missing tail as caps=0. Mixed-version clusters stay sound.
: Hello { i id i role i want ( Vec u ) pubkey i caps }@ hello_free Hello h → v@ hello_decode ( Vec u ) buf → Hello: Member { ( Vec u ) pubkey i id i caps }: Roster { ( Vec s ) members } // *Member@ roster_new → *Roster@ roster_free * Roster r → v@ roster_has * Roster r ( Vec u ) pubkey → b@ roster_count * Roster r → i@ roster_count_caps * Roster r i mask → iHow many members advertise every capability bit in mask.
@ roster_add * Roster r * Ring ring ( Vec u ) pubkey i id i vnodes i caps → bFold a worker into the roster + ring, once. Returns T if newly added.
packages/swarm-mcp/src/work.nu — the distributed map-reduce the cluster runs.
A task is: evaluate an expression kernel (expr.nu) over an integer range [lo, hi) and fold the results with a reduce op. The coordinator shards the range; dist/ring routes each chunk to its owning worker; the worker parses the kernel once and folds it over its sub-range; the coordinator combines the partial folds. Every reduce op is associative, so sharding is exact.
reduce op : 0 sum · 1 product · 2 min · 3 max · 4 count (of truthy map) dtype : 0 int (i64) · 1 float (f64) chunk : [op:1][dtype:1][lo:8 BE][hi:8 BE][expr bytes…] result : [value:8 BE] — i64, or an f64 bit pattern when dtype=float
Float tasks fold in f64 over the same integer range (x is the index cast to double) and ship the partial as its f64 bit pattern through the same 8-byte result codec; the coordinator reinterprets it (see main.nu tids_combine).
@ red_sum → i@ red_product → i@ red_min → i@ red_max → i@ red_count → i@ red_id i op → iThe identity (empty-range) value for a reduce op.
@ red_fold i op i acc i v → iFold one mapped value into the accumulator (per element).
@ red_combine i op i a i b → iCombine two partial folds (across chunks). count combines by summing the per-chunk counts; every other op combines like its element fold.
@ red_id_f i op → f@ red_fold_f i op f acc f v → f@ red_combine_f i op f a f b → f@ chunk_payload i op i dtype i lo i hi ( Vec u ) expr → ( Vec u )@ result_encode i value → ( Vec u )@ result_decode ( Vec u ) p → i@ kernel_handler ( Vec u ) key → ( @ ( Vec u ) ( Vec u ) ): Chunk { i lo i hi }@ shard i lo i hi i n → ( Vec s )@ shard_free ( Vec s ) chunks → v@ chunk_key i idx → ( Vec u )A ring key for chunk index i: distinct chunks hash to distinct ring points.
@ reduce_op_of s name i fallback → iTrue for a recognised reduce-op name; sets *op. Keeps the MCP layer thin.
@ reduce_op_name i op → spackages/swarm-mcp/src/expr.nu — the phase-1 "kernel": a small, regular integer-expression language the cluster evaluates per element.
A workload is a map-reduce: the coordinator ships an expression in one variable x plus a range and a reduce op; each worker parses the expression once, evaluates it for every x in its sub-range, and folds the results. The language is deliberately small and regular so a language model can write it without docs — and it is the natural precursor to phase 2, where the kernel is arbitrary NURL compiled to wasm instead of interpreted here.
expr := ternary ternary := logic ( '?' expr ':' ternary )? logic := compare ( ('&'|'|') compare ) compare := addsub ( ('<'|'<='|'>'|'>='|'=='|'!=') addsub )? addsub := muldiv ( ('+'|'-') muldiv ) muldiv := unary ( (''|'/'|'%') unary ) unary := '-' unary | primary primary := NUM | 'x' | '(' expr ')' | ('min'|'max') '(' expr ',' expr ')' | 'abs' '(' expr ')' NUM := DIGIT+ ( '.' DIGIT+ )?
The same grammar evaluates in one of two numeric domains, chosen per task by the caller (see work.nu dtype): • int — all arithmetic is i64 (truncated division; div/mod by zero → 0). • float — all arithmetic is f64, x is the integer index cast to double, div/mod by zero → 0.0. A literal with a '.' (e.g. 0.5) is a float literal in either mode (truncated toward zero in int mode). Comparisons and '&' '|' yield 1/0 (1.0/0.0 in float); any non-zero is "true".
Tokens are kept in two parallel int vectors; the parser builds a flat stride-4 node arena (tag,a,b,c) and returns the root index — the resp.nu arena pattern, so nesting needs no per-node allocation. Float literals carry the f64 bit pattern (via floatbits) in the value slot, so the all-int arena holds them losslessly; the float evaluator reinterprets them back.
@ expr_tokenize ( Vec u ) src ( Vec i ) tk ( Vec i ) tv → b: EParser: EParser {
( Vec i ) tk
( Vec i ) tv
i pos
( Vec i ) arena // stride-4: tag,a,b,c
b ok
}
@ expr_parse ( Vec u ) src * EParser p → iParse src → (arena, root). On any error, ok=0; the caller checks it. The arena is returned via the EParser; the root index via the return value.
@ eparser_free * EParser p → vFrees the parser's arena/token vectors AND the struct itself.
@ expr_eval * EParser p i node i x → i@ expr_eval_f * EParser p i node f x → fpackages/swarm-mcp/src/buildwasm.nu — compile NURL source to a wasm module so the MCP server itself can accept a kernel as source (compute_submit_kernel / compute_submit_cuda) instead of requiring the caller to pre-compile.
LOCAL-FIRST: the wasmbuilder package (deps/wasmbuilder) compiles the kernel in-process — nurlc → IR rewrite → the toolchain's bundled zig cc — no network, no build service. Only when the local toolchain can't do it (no nurlc/zig on this box) does it fall back to POSTing {source, filename} to <NURL_BUILD_API>/build_wasm. A direct nurlapi answers with raw wasm bytes; the public playground proxy answers with JSON carrying wasm_base64 (and nurlc_errors on failure). Both are handled.
$NURL_BUILD_API fallback build service base URL (default https://play.nurl-lang.org)
@ build_api_url → String@ wrap_kernel s source i op i dtype i kkind → StringWrap a bare kernel into a complete wasm-ready program for reduce op op. dtype 0 (int): kernel is @ kernel i x → i; the module folds in i64 and prints the partial as a decimal integer. dtype 1 (float): kernel is @ kernel i x → f (x is the integer index, returns a double); the module folds in f64 and prints the partial's f64 BIT PATTERN as a decimal integer — so it rides the same stdout→int wire, and the coordinator reinterprets it (work.nu tids_combine float path).
kkind 0 (element): the generated main folds kernel(x) over [lo, hi). kkind 1 (chunk): the kernel is @ kernel i lo i hi → i (or → f) and the main calls it ONCE with the whole sub-range — the kernel owns the loop. This is the right granularity for kernels with per-invocation setup cost (open a CUDA context, JIT a device kernel, allocate buffers): one setup per CHUNK instead of one per element. The reduce op still combines the chunk partials.
@ compile_to_wasm s source → !( Vec u ) StringCompile NURL source to a wasm module. Ok = module bytes; Err = a human-readable transport/compile error (suitable to hand back to the model).
Local wasmbuilder first. A LOCAL "nurlc failed" is a genuine source error and is returned as-is (the build service would only repeat it); any other local error means the environment can't build wasm (no toolchain, no zig, download forbidden) → fall back to the build API.
packages/swarm-mcp/src/cudakernel.nu — generate a GPU chunk kernel from a bare CUDA-C map function.
The model hands over ONLY the math; cuda_wrap wraps it into a complete, self-contained NURL program in one of three OUTPUT MODES:
scalar (gpu_mode_scalar) — __device__ double f(long long x): grid-stride map over [lo, hi), per-block shared-memory reduction with the task's reduce op, host fold of the block partials, partial's f64 bit pattern printed on stdout (the classic reduce wire). sample (gpu_mode_sample) — same f, but every value comes BACK: out[x-lo] = f(x); the module writes the hi-lo doubles raw (LE) to an output file named on argv — curves, fields, images. hist (gpu_mode_hist) — __device__ long long bin(long long x) plus an optional __device__ double val(long long x) (default 1.0): atomicAdd(&out[bin(x)], val(x)) over K bins (K on argv — the same module serves any K), K doubles written to the output file. The double atomicAdd is a portable CAS loop, so the PTX stays valid for NVRTC's default (pre-sm_60) target.
RUNTIME PARAMS make the kernels dynamic: with has_params the device functions take (long long x, const double* p) and the module parses any argv tail of f64-bit-pattern decimals into a device buffer. Parameter VALUES never touch the generated source, so the module's content hash — and with it the worker-side cache and the coordinator's build cache — survives across a whole parameter scan: no rebuild, no re-upload.
The program's cuda/nvrtc FFI declarations mirror packages/gpu/src/cuda.nu's ABI exactly. Compiled with --ffi-host-imports (the wasm build API always does), every cu/nvrtc symbol becomes an env import that the pure-NURL wasmtime's GPU bridge resolves against the worker's real libcuda/libnvrtc under --allow-gpu. The same source also compiles native (nurl.sh auto-links cuda/nvrtc), so a kernel can be smoke-tested off-cluster.
Argv contracts (must match wasmkernel.nu __wasm_run_gpu): scalar: main lo hi [p…] sample: main lo hi outpath [p…] hist: main lo hi outpath K [p…]
Any CUDA failure exits non-zero WITHOUT producing output — the worker reports the chunk as failed (ok=0) instead of folding silent zeros.
@ cuda_src_ok s src → b── validation ──────────────────────────────────────────────────── The user source is spliced into a generated NURL backtick literal; a backtick would terminate it (NURL strings cannot contain one, escaped or not), so reject it outright. Everything else is escaped below.
@ cuda_wrap s user i op i mode i has_params i has_data → String── the wrapper ─────────────────────────────────────────────────── CUDA map function + reduce op + output mode (+ params flag) → complete NURL chunk-kernel program. The caller validates cuda_src_ok first. has_data: the chunk carries a dataset slice — argv gains an in-file path right after hi, and the device functions receive v = data[x] as their second argument.
packages/swarm-mcp/src/main.nu — swarm-mcp: an MCP-controlled distributed compute engine. A language model sets a workload over MCP — an expression kernel in x plus a range and a reduce op — and the cluster evaluates it distributed; the model reads running tasks and finished results back.
swarm-mcp relay 0.0.0.0 47700 [--v] # the meeting point (one per cluster) swarm-mcp worker <host> <port> # join as a compute node swarm-mcp mcp <host> <port> # MCP server (stdio) → drives the cluster swarm-mcp submit <host> <port> <reduce> <lo> <hi> <expr> # manual CLI submit
The cluster layer (membership → ring → job dispatch) is the swarm package's; here every worker registers ONE generic kernel handler (work.nu) that interprets the submitted expression (expr.nu), so any normal-operation workload runs without recompiling a worker. Phase 2 swaps the interpreter for a NURL→wasm kernel; the protocol and MCP surface stay the same.
@ swarm_vnodes → i@ kind_kernel → i@ pk_from_id i id → ( Vec u )A 32-byte opaque routing pubkey derived deterministically from a node id.
: Swarm: Swarm {
s transport // *Transport
s ring // *Ring
s gpu_ring // *Ring — the GPU capability domain (cap_gpu workers only)
s roster // *Roster
s job // *JobNode
( Vec u ) self_pk
i self_id
i role
i self_caps // capability bits this node advertises (cap_gpu for --gpu workers)
( Vec u ) group // token-derived 32-byte relay multicast group (cluster isolation)
( Vec u ) key // token-derived HMAC key (compute authenticity)
i epoch // bumps on every ring-membership change (block-seed invalidation)
}
@ swarm_new RelayClient rc i id i role i caps s token → *Swarm@ swarm_free * Swarm sw → v@ swarm_join_group * Swarm sw → v@ swarm_announce * Swarm sw i want → v@ swarm_announce_ok * Swarm sw i want → bAnnounce presence; returns whether the broadcast reached the relay. A failed send is the reconnect loop's signal that the relay is gone.
@ swarm_on_hello * Swarm sw Hello h → v@ swarm_pump * Swarm sw i max → v@ swarm_discover * Swarm sw i rounds → v@ parse_relay_list String cs s dhost i dport → ( Vec String )Dial the relay, retrying briefly so a co-located relay that is still binding (or a relay that restarts) doesn't lose its local roles to a startup race. ── relay endpoint list (failover) ─────────────────────────────── --connect accepts a comma-separated list "h1:p1,h2:p2,…"; any one being reachable bootstraps the node, and if the connected relay dies the node rotates to the next. A single endpoint is just a one-element list. The list is kept as raw "host:port" String segments (default filled in when a segment omits either); parse_hostport resolves one at dial time.
@ relay_list_free ( Vec String ) lst → v@ relay_dial_retry s host i port i tries → !RelayClient NetErr@ node_relay s host i port i vflag → vRELAY role: the dumb rendezvous point. Owns the fiber reactor on this thread.
@ node_worker ( Vec String ) relays s dhost i dport s token i vflag i gpu → vWORKER role: join the cluster, register the kernel + wasm handlers (every worker can run wasm modules), and drain cluster traffic forever. Each worker thread takes a fresh random identity, so --workers N spins up N independent ring members in one process. gpu≠0: advertise cap_gpu and register the kind_wasm_gpu handler — wasm chunks of that kind run with --allow-gpu. A worker keeps its identity across reconnects (same ring member) and, when the relay it is on dies, rotates to the next relay in the list and re-forms — no single relay is a point of failure. Death is detected by a periodic heartbeat announce whose SEND fails when the relay is gone; the on-disk block cache survives the switch, so re-seeding is idempotent.
@ cluster_submit * Swarm sw i op i dtype i lo i hi ( Vec u ) expr i nchunks → ( Vec i )@ cluster_submit_wasm * Swarm sw i lo i hi ( Vec u ) wasm i nchunks i kind → ( Vec i )Shard a wasm-kernel task: ship the compiled module + each sub-range to its ring owner under kind (kind_wasm, or kind_wasm_gpu — the GPU capability domain). The module bytes ride every chunk (workers cache by content hash, so it is written once per worker).
@ cluster_submit_wasm_gpu * Swarm sw i mode i lo i hi i kbins ( Vec i ) params ( Vec u ) data ( Vec u ) wasm i nchunks → ( Vec i )Shard a GPU task (payload v2/v3): mode, K, runtime params and the module ride every chunk under kind_wasm_gpu — routed on the GPU capability ring. data is the WHOLE dataset's raw LE f64 bytes (empty when the task has no dataset); each chunk ships exactly its own slice data[clo·8, chi·8) — the split travels with its task, so a worker needs no separate fetch.
@ tids_ready * Swarm sw ( Vec i ) tids → bAll chunk results present? (cluster job results are recorded by task-id.)
: Combined { i value i nfail }Combine all chunk results with the reduce op (assumes ready). For a float task each partial is an f64 bit pattern: decode→f64, combine in f64, and return the combined f64 re-encoded as its bit pattern (the Task stores it that way; task_to_json reinterprets it for the model).
Two result frames ride the wire: the expr kernel's legacy [partial:8], and the wasm kinds' [ok:1][partial:8]. A chunk whose ok=0 (module trap, missing runtime, GPU failure) is COUNTED, not folded — nfail>0 marks the task as failed instead of silently reducing zeros into the answer.
@ tids_combine * Swarm sw i dtype i op ( Vec i ) tids → Combined: CombinedV { ( Vec u ) bytes i nfail }Combine VECTOR results (sample / hist). Frame: [ok:1][count:4 BE][f64 LE × count]. Sample chunks CONCATENATE in shard order (tids order = chunk order); hist chunks add ELEMENTWISE into K bins. A failed or malformed chunk counts toward nfail and contributes nothing.
@ tids_combine_vec * Swarm sw i mode i kbins ( Vec i ) tids → CombinedV@ nchunks_for i nworkers → i@ run_submit s host i port i op i lo i hi s expr s token → i@ nchunks_wasm i nworkers → iChunk count for a wasm task: fewer chunks than the expression path (each chunk ships the module and spawns a wasmtime process), but enough to spread across the ring.
@ run_runwasm s host i port i op i lo i hi s wasmpath s token → i: Task: Task {
i id
( Vec u ) expr // kernel bytes
i dtype // 0 int · 1 float (result holds an f64 bit pattern)
i lo
i hi
i op
i nchunks
( Vec i ) tids
i done
i result
i failed // chunks that reported ok=0 (task status "error" when > 0)
i mode // gpu_mode_scalar/sample/hist (scalar for every non-GPU task)
i kbins // histogram bin count (hist mode)
( Vec u ) vres // vector result bytes (raw LE f64s; sample/hist modes)
String out_file // when set, the finished vector result is written here
i dsid // dataset id the task maps over (0 = none)
i seeded // dataset blocks seeded by THIS submit (0 = all were cached)
}
: McpState: McpState {
s swarm // *Swarm coordinator
( Vec s ) tasks // *Task
i next_id
( Vec s ) wcache // *WasmCached — compiled kernel modules by source hash
( Vec s ) datasets // *Dataset — uploaded data the CUDA tools map over
i next_ds
( Vec String ) seeded // "blockhex|chunk|epoch" — blocks CONFIRMED cached at their owner
}
@ task_new i id ( Vec u ) expr i dtype i lo i hi i op i nchunks ( Vec i ) tids → *TaskConstructor. The params share the field names (lo, hi, …); = . t lo lo is a field store — the field-store/local-shadow miscompile this used to work around was fixed in the compiler (NURL v0.10.4).
: ~ i g_mcp 0: Dataset { i id String name ( Vec u ) bytes ( Vec ( Vec u ) ) blocks }@ ds_count_of * Dataset d → i@ ds_find i id → s@ cluster_submit_wasm_gpu_ds * Swarm sw i mode i rlo i rhi i kbins ( Vec i ) params * Dataset d ( Vec u ) wasm i nchunks_want ( Vec String ) seeded * u nseed_cell → ( Vec i )one-pass min/max/mean of raw LE f64 bytes, attached to a JSON object Dataset-backed GPU submit over CONTENT-ADDRESSED blocks (payload v4). Chunking is block-aligned, so every 8 MiB block belongs to exactly one chunk; blocks not yet confirmed at their owner are seeded first as kind_blob tasks carrying the SAME ring key as the compute chunk (same ring + same key → same worker), and per-connection ordering guarantees the worker handles a seed before the compute that references it. A seed is recorded in seeded (hash|chunk|epoch) only after an OK result, so a lost seed re-seeds on the next submit instead of wedging. Dataset-backed GPU submit over CONTENT-ADDRESSED blocks (payload v4). Chunking is block-aligned, so every block belongs to exactly one chunk. Blocks not yet confirmed at their owner are seeded FIRST, strictly one at a time — submit a block, pump until its OK result, then the next. A single large frame is ever in flight to a given worker, which keeps the relay's forwarding stream from interleaving multi-frame bursts. Only after every referenced block is confirmed cached are the (small) compute chunks submitted; each references its blocks by hash and the worker assembles the slice from its cache, failing visibly on any missing block.
: WasmCached { String hash ( Vec u ) wasm }@ mcp_swarm → *Swarm@ mcp_pump i rounds → v@ task_refresh * Task t → vRefresh one task's status; combine if every chunk has landed. A finished vector task with an out_file writes the raw LE f64s there (once).
@ task_find i id → s@ task_to_json * Task t → JsonBuild the JSON-as-text result body for one task (LLM-readable + parseable).
@ tool_result_json Json o → Json@ tool_submit Json args → Json@ tool_run_wasm Json args → Json@ tool_submit_kernel Json args → Jsoncompute_submit_kernel: the server compiles NURL source → wasm itself (via the build API), then runs it — the model hands over a kernel as plain NURL.
@ tool_iterate Json args → Json@ tool_shuffle Json args → Json@ tool_submit_cuda Json args → Json@ tool_sample_cuda Json args → Jsoncompute_sample_cuda: every f(x) comes back — the cluster gathers the range into one array (curves, fields, images).
@ tool_hist_cuda Json args → Jsoncompute_histogram_cuda: bin(x) picks a bucket, val(x) the weight (default 1.0); K bin sums come back — a whole distribution in one pass.
@ tool_upload_data Json args → Json@ tool_list_data Json args → Json@ tool_list Json args → Json@ tool_result Json args → Json@ ms_prop Json props s name s ty s desc → v@ ms_schema_submit → Json@ ms_schema_result → Json@ ms_schema_empty → Json@ ms_schema_run_wasm → Json@ ms_schema_submit_kernel → Json@ ms_schema_submit_cuda → Json@ ms_schema_shuffle → Json@ ms_schema_iterate → Json@ ms_schema_sample_cuda → Json@ ms_schema_hist_cuda → Json@ ms_schema_upload_data → Json@ build_tools_list → ( Vec Json )@ dispatch_tool s name Json args → Json@ handle_initialize Json id → Json@ handle_ping Json id → Json@ handle_tools_list Json id → Json@ handle_tools_call Json id Json params → Json@ handle_unknown Json id s method → Json@ dispatch Json req → ?Json: ~ i g_mcp_relays 0 // *( Vec String ) as an addressThe coordinator's relay endpoints and its current index — the submit path rebuilds the swarm against the next relay when the current one dies (mcp_ensure_relay below). Held as globals because the MCP handlers are top-level functions over module state, like the swarm itself.
: ~ i g_mcp_from 0 // next relay index to try: ~ s g_mcp_token ``: ~ i g_mcp_reconnects 0@ mcp_reconnect → bBuild a fresh coordinator swarm on the next reachable relay and swap it into the McpState, preserving tasks / datasets / caches. Returns F when no relay in the list is reachable. Called from the submit path when the current relay looks dead, so a relay failure does not take the API down.
@ mcp_ensure_relay → bBefore a submit: if a heartbeat to the current relay fails, the relay is gone — reconnect to the next. Returns T once a live relay is in place.
@ node_mcp ( Vec String ) relays s rhost i rport s mcp_host i mcp_port s cert s key s token → v@ arg_int i idx → i@ arg_eq i idx s lit → b@ has_flag s name → bTrue if name appears anywhere in argv (a boolean role/verbose flag).
@ flag_val s name s deflt → StringThe token following --name in argv, or a copy of deflt. Owned String.
@ flag_int s name i deflt → i: HostPort { String host i port }@ parse_hostport String hp s dhost i dport → HostPortParse "host:port", "host", ":port", or "" with host/port defaults.
@ relay_dial_list ( Vec String ) lst s dhost i dport i from * u idx_cell → !RelayClient NetErrDial the relays in lst starting at from (wrapping once around the whole list), a few quick retries each. Writes the connected index to idx_cell. Returns the live client or F when every endpoint is down.
@ ensure_cert s cert_path s key_path → bEnsure a TLS cert+key exist at the given paths, auto-minting a self-signed EC P-256 pair in pure NURL if absent (std/x509_gen.nu — no openssl, no subprocess), so --mcp works out of the box; pass --tls-cert/--tls-key for a real cert. CN + SAN are localhost (matching the previous openssl invocation minus the IP SAN — MCP clients pin or use --insecure against a self-signed cert either way). Returns T once both files exist.
@ cli_submit → i@ cli_runwasm → i@ usage → v@ run_node → i@ main → ipackages/swarm-mcp/src/blob.nu — content-addressed dataset blocks.
The v3 payload shipped each chunk's dataset slice INSIDE every task, so N submits (or N iteration rounds) over the same data paid the transfer N times. This layer makes data transfer content-addressed and one-shot:
(BLOB_BLOCK_VALS f64 values = 8 MiB; the last block may be short); each block is keyed by its BLAKE3-256 (stdlib/std/hash_blake3.nu),
against the hash on arrival — a re-submit or the next iteration round references the same hashes and moves NO data,
task with the SAME ring key as the compute chunk that needs the block lands on the same worker (consistent hash: same ring + same key → same owner). No new transport protocol, no pull-request re-entrancy — and the cluster HMAC token authenticates blocks exactly like every other payload.
A compute chunk then references its blocks by hash (payload v4, see wasmkernel.nu) and the worker assembles the slice from its cache; a missing or corrupt block fails the chunk VISIBLY (ok=0 → failed_chunks), never silently.
NOTE: siblings are imported by the consumer in dependency order (main.nu: token.nu → blob.nu → …); this file uses token_tag/token_untag from token.nu without importing it, like the other kernel files.
@ blob_block_vals → iThe absolute block grid: block b covers dataset values [b·BLOB_BLOCK_VALS, (b+1)·BLOB_BLOCK_VALS).
@ kind_blob → iJob kind for a block-seed task (kernel=1, wasm=2, wasm_gpu=3).
@ blob_hash ( Vec u ) bytes → ( Vec u )@ blob_hex ( Vec u ) h → String@ blob_cached ( Vec u ) hash → b@ blob_store ( Vec u ) hash ( Vec u ) bytes → bVerify + store a block (idempotent: an already-cached hash is a no-op). F when the bytes do not hash to hash — a corrupt or forged block is never written under a name it doesn't own.
@ blob_append ( Vec u ) hash ( Vec u ) out → bAppend the cached block hash to out. F on miss or a cache file that no longer matches its hash (deleted / truncated / tampered).
@ blob_seed_payload ( Vec u ) hash ( Vec u ) bytes → ( Vec u )@ blob_handler ( Vec u ) key → ( @ ( Vec u ) ( Vec u ) )Worker handler for kind_blob: authenticate, verify content address, cache. Result wire matches the scalar frame ([ok:1][0:8]) so the coordinator's readiness/failure accounting needs no special case.
@ blob_manifest ( Vec u ) data → ( Vec ( Vec u ) )@ blob_manifest_free ( Vec ( Vec u ) ) m → vpackages/swarm-mcp/src/wasmkernel.nu — phase 2: run a NURL-compiled wasm kernel over a chunk.
The language model writes a kernel as ordinary NURL — anything, not just an expression — and compiles it (via the NURL build service) to a wasm32-wasi module whose main reads lo hi from argv, folds the kernel over [lo, hi), and prints the partial. The coordinator ships that module plus a sub-range to each worker; the worker runs it under wasmtime and returns the partial, which the coordinator combines with the reduce op exactly as in phase 1.
The wasmtime here is just a CLI contract — wasmtime run <module> <lo> <hi>, partial on stdout. The toolchain's own pure-NURL runtime (packages/wasmtime) satisfies it exactly, so a worker needs NO external runtime: put that binary on PATH as wasmtime, or point $WASMTIME at it. The Bytecode-Alliance wasmtime works too — whichever $WASMTIME resolves to is used.
CPU chunk (kind_wasm) : [lo:8 BE][hi:8 BE][wasm…] GPU chunk (kind_wasm_gpu) : payload v2 — see wasm_gpu_chunk_payload below (adds an output mode, K, and runtime params)
The reduce op is baked into the module by the wrapper, so the partial is final per chunk. A scalar module prints its partial as a decimal integer on stdout; for a float task that integer is the partial's f64 BIT PATTERN (see buildwasm.nu), so the same int wire carries it unchanged — the coordinator reinterprets it. Vector modes (sample/hist) write raw little-endian f64s to a worker-named output file instead (fast: one fwrite, no per-value stdout).
The module is cached on each worker by a content hash, so re-running the same kernel (every chunk of a task, and across tasks — runtime params don't touch the module) writes the .wasm only once.
@ kind_wasm → i@ kind_wasm_gpu → iGPU wasm chunks are a SEPARATE job kind, routed on the GPU capability ring (dist/job job_set_ring): only --gpu workers own these keys, and only they register this handler — a non-GPU worker can never receive (or garble) one.
@ gpu_mode_scalar → iGPU task output modes (baked into the generated program AND carried in the chunk so the worker knows how to invoke the module and collect its output).
@ gpu_mode_sample → i@ gpu_mode_hist → i@ gpu_mode_vecreduce → i@ gpu_mode_shuffle_map → i@ gpu_mode_shuffle_reduce → i@ wasm_chunk_payload i lo i hi ( Vec u ) wasm → ( Vec u )@ wasm_gpu_chunk_payload i mode i lo i hi i kbins ( Vec i ) params ( Vec u ) data ( Vec u ) wasm → ( Vec u )── GPU chunk payload (v2 / v3) ────────────────────────────────── v2: [2][mode:1][lo:8][hi:8][K:4][nparams:2][f64 bits:8×n][wasm…] v3: [3][mode:1][lo:8][hi:8][K:4][nparams:2][f64 bits:8×n] [dlen:4][data slice: raw LE f64s][wasm…] mode selects the module's output contract; K is the histogram bin count (0 otherwise); params are runtime kernel arguments passed on the module's argv as f64-bit-pattern decimals — the SAME module (same content hash, no rebuild, warm worker cache) serves any parameter values. v3 additionally carries the chunk's DATASET SLICE — data[lo..hi) as raw doubles, following the map-reduce rule that a split travels with its task. dlen is exact (8·(hi−lo)) so a truncated frame is rejected, and dlen=0 means "no data" (the encoder then emits plain v2).
@ wasm_gpu_chunk_payload_blobs i mode i lo i hi i kbins ( Vec i ) params i b0 ( Vec ( Vec u ) ) hashes ( Vec u ) wasm → ( Vec u )v4: the chunk references its dataset blocks by CONTENT HASH instead of carrying the slice — [4][mode:1][lo:8][hi:8][K:4][nparams:2][f64 bits:8×n] [b0:4][nblob:2][hash32 × nblob][wasm…] b0 is the absolute index of the first referenced block on the dataset's block grid (blob.nu); the worker assembles data[lo..hi) from its block cache (seeded via kind_blob) and FAILS the chunk on any missing block.
: GpuChunk: GpuChunk {
i ok
i mode
i lo
i hi
i kbins
( Vec i ) params
( Vec u ) data // raw LE f64 slice (empty without a dataset)
i b0 // v4: first referenced block index on the dataset grid
( Vec ( Vec u ) ) blobs // v4: 32-byte block hashes (empty otherwise)
( Vec u ) wasm
}
@ gpu_chunk_free GpuChunk c → v@ wasm_gpu_chunk_decode ( Vec u ) body → GpuChunk@ gpu_chunk_assemble inout GpuChunk c → bv4: assemble data[lo..hi) into c.data from the content-addressed block cache. A v2/v3 chunk (no blob references) is a no-op T. F on ANY missing or corrupt block, or a reference window that does not cover the range — the caller fails the chunk VISIBLY instead of computing on garbage.
@ _wasm_hash ( Vec u ) v → StringFNV-1a/64 content hash → 16 lowercase hex chars, for the cache filename.
: WasmRun { i ok i value }Run a cached module over [lo, hi). ok=1 iff the runtime ran the module to a zero exit — a failed chunk (missing wasmtime, module trap, GPU error path exiting non-zero) is REPORTED, not folded into the reduce as a silent zero. allow_gpu≠0 passes --allow-gpu so the module's env imports (CUDA/NVRTC) resolve to real hardware — this needs the pure-NURL packages/wasmtime.
@ wasm_handler ( Vec u ) key → ( @ ( Vec u ) ( Vec u ) )The worker handler for a (CPU) wasm chunk: verify the cluster HMAC tag, cache the module by content hash, run it under wasmtime, return the partial tagged for the coordinator. An untrusted/forged payload yields a tagged ok=0 — a stranger can't make a worker fetch-and-run an arbitrary module. key is the cluster HMAC key (token_key), captured by the handler closure.
: GpuOut { i ok i scalar ( Vec u ) bytes }@ wasm_gpu_handler ( Vec u ) key → ( @ ( Vec u ) ( Vec u ) )GPU worker handler: decode payload v2, cache the module, run per mode, and reply the mode's result frame — scalar [ok:1][bits:8], vector [ok:1][count:4 BE][f64 LE × count]. Forged/undecodable payloads → ok=0.