packages/swarm-mcp/src/token.nu — cluster identity + authenticity from one shared --token.
Every node in a cluster is launched with the same --token. The token does two things, both derived deterministically so no node has to coordinate:
token_group_id hashes the token into the 32-byte relaymulticast group id. Clusters with different tokens join different relay groups, so their HELLO/job traffic is mutually invisible even when they share one relay. Without the token you cannot derive the group id, so a stranger cannot even see (let alone join) the cluster's gossip.
token_key derives a 32-byte HMAC key. Every computepayload (a kernel chunk) and every result is tagged with HMAC-SHA256 (token_tag) and verified on receipt (token_untag, constant-time). A worker only acts on token-authentic jobs; a coordinator only accepts token-authentic results. Combined with group isolation this gives mutual authentication of all compute traffic over the dumb (opaque) relay.
The relay itself never needs the token — it forwards opaque bytes — so the secret never leaves the nodes that own it.
@ token_tag_len → iHMAC tag length prepended to authenticated messages (truncated SHA-256 MAC).
@ token_group_id s token → ( Vec u )32-byte relay multicast group id for this cluster (token isolation).
@ token_key s token → ( Vec u )32-byte HMAC key authenticating compute payloads + results.
@ token_tag ( Vec u ) key ( Vec u ) payload → ( Vec u )Prepend a 16-byte HMAC-SHA256 tag to payload. Returns a NEW vec; the caller owns it (and still owns payload).
@ token_untag ( Vec u ) key ( Vec u ) tagged → ?( Vec u )Verify the tag and strip it. Some(payload copy) on a valid tag, None on a missing/short/forged tag. The compare is constant-time. Caller owns the returned payload on the Some arm.
packages/swarm/src/census.nu — lightweight membership gossip for the swarm.
The compute layer (dist/job over dist/ring) needs every node to agree on the live worker set: a worker must know it owns a key, and the coordinator must know which pubkeys to route to. The full SWIM table (net/membership.nu) is the heavy, churn-hardened answer; this is the small one a compute cluster actually needs — a HELLO announcement that feeds the consistent-hash ring.
want, replies with their own HELLO so the newcomer learns them too;
Roles: a WORKER owns keys and executes handlers, so it joins everyone's ring; a CLIENT (the coordinator) only submits, so it is never added to the ring — otherwise it would own keys it has no handler for and silently drop results.
The codec and the membership set are pure; only the pump touches transport.
@ census_hello_t → i@ role_client → i@ role_worker → i@ cap_gpu → iCapability bits a worker advertises in HELLO. cap_gpu: the worker runs wasm chunks with GPU host imports enabled (wt --allow-gpu on real hardware).
@ hello_build i id i role i want ( Vec u ) pubkey i caps → ( Vec u )HELLO wire: [3][id:8][role:1][want:1][pklen:2][pubkey…][caps:1] The caps byte TRAILS the pubkey so a pre-caps decoder (fixed prefix + pklen-delimited pubkey) reads the same fields and ignores the tail; a caps decoder treats a missing tail as caps=0. Mixed-version clusters stay sound.
: Hello { i id i role i want ( Vec u ) pubkey i caps }@ hello_free Hello h → v@ hello_decode ( Vec u ) buf → Hello: Member { ( Vec u ) pubkey i id i caps }: Roster { ( Vec s ) members } // *Member@ roster_new → *Roster@ roster_free * Roster r → v@ roster_has * Roster r ( Vec u ) pubkey → b@ roster_count * Roster r → i@ roster_count_caps * Roster r i mask → iHow many members advertise every capability bit in mask.
@ roster_add * Roster r * Ring ring ( Vec u ) pubkey i id i vnodes i caps → bFold a worker into the roster + ring, once. Returns T if newly added.
packages/swarm-mcp/src/work.nu — the distributed map-reduce the cluster runs.
A task is: evaluate an expression kernel (expr.nu) over an integer range [lo, hi) and fold the results with a reduce op. The coordinator shards the range; dist/ring routes each chunk to its owning worker; the worker parses the kernel once and folds it over its sub-range; the coordinator combines the partial folds. Every reduce op is associative, so sharding is exact.
reduce op : 0 sum · 1 product · 2 min · 3 max · 4 count (of truthy map) dtype : 0 int (i64) · 1 float (f64) chunk : [op:1][dtype:1][lo:8 BE][hi:8 BE][expr bytes…] result : [value:8 BE] — i64, or an f64 bit pattern when dtype=float
Float tasks fold in f64 over the same integer range (x is the index cast to double) and ship the partial as its f64 bit pattern through the same 8-byte result codec; the coordinator reinterprets it (see main.nu tids_combine).
@ red_sum → i@ red_product → i@ red_min → i@ red_max → i@ red_count → i@ red_id i op → iThe identity (empty-range) value for a reduce op.
@ red_fold i op i acc i v → iFold one mapped value into the accumulator (per element).
@ red_combine i op i a i b → iCombine two partial folds (across chunks). count combines by summing the per-chunk counts; every other op combines like its element fold.
@ red_id_f i op → f@ red_fold_f i op f acc f v → f@ red_combine_f i op f a f b → f@ chunk_payload i op i dtype i lo i hi ( Vec u ) expr → ( Vec u )@ result_encode i value → ( Vec u )@ result_decode ( Vec u ) p → i@ kernel_handler ( Vec u ) key → ( @ ( Vec u ) ( Vec u ) ): Chunk { i lo i hi }@ shard i lo i hi i n → ( Vec s )@ shard_free ( Vec s ) chunks → v@ chunk_key i idx → ( Vec u )A ring key for chunk index i: distinct chunks hash to distinct ring points.
@ reduce_op_of s name i fallback → iTrue for a recognised reduce-op name; sets *op. Keeps the MCP layer thin.
@ reduce_op_name i op → spackages/swarm-mcp/src/expr.nu — the phase-1 "kernel": a small, regular integer-expression language the cluster evaluates per element.
A workload is a map-reduce: the coordinator ships an expression in one variable x plus a range and a reduce op; each worker parses the expression once, evaluates it for every x in its sub-range, and folds the results. The language is deliberately small and regular so a language model can write it without docs — and it is the natural precursor to phase 2, where the kernel is arbitrary NURL compiled to wasm instead of interpreted here.
expr := ternary ternary := logic ( '?' expr ':' ternary )? logic := compare ( ('&'|'|') compare ) compare := addsub ( ('<'|'<='|'>'|'>='|'=='|'!=') addsub )? addsub := muldiv ( ('+'|'-') muldiv ) muldiv := unary ( (''|'/'|'%') unary ) unary := '-' unary | primary primary := NUM | 'x' | '(' expr ')' | ('min'|'max') '(' expr ',' expr ')' | 'abs' '(' expr ')' NUM := DIGIT+ ( '.' DIGIT+ )?
The same grammar evaluates in one of two numeric domains, chosen per task by the caller (see work.nu dtype): • int — all arithmetic is i64 (truncated division; div/mod by zero → 0). • float — all arithmetic is f64, x is the integer index cast to double, div/mod by zero → 0.0. A literal with a '.' (e.g. 0.5) is a float literal in either mode (truncated toward zero in int mode). Comparisons and '&' '|' yield 1/0 (1.0/0.0 in float); any non-zero is "true".
Tokens are kept in two parallel int vectors; the parser builds a flat stride-4 node arena (tag,a,b,c) and returns the root index — the resp.nu arena pattern, so nesting needs no per-node allocation. Float literals carry the f64 bit pattern (via floatbits) in the value slot, so the all-int arena holds them losslessly; the float evaluator reinterprets them back.
@ expr_tokenize ( Vec u ) src ( Vec i ) tk ( Vec i ) tv → b: EParser: EParser {
( Vec i ) tk
( Vec i ) tv
i pos
( Vec i ) arena // stride-4: tag,a,b,c
b ok
}
@ expr_parse ( Vec u ) src * EParser p → iParse src → (arena, root). On any error, ok=0; the caller checks it. The arena is returned via the EParser; the root index via the return value.
@ eparser_free * EParser p → vFrees the parser's arena/token vectors AND the struct itself.
@ expr_eval * EParser p i node i x → i@ expr_eval_f * EParser p i node f x → fpackages/swarm-mcp/src/buildwasm.nu — compile NURL source to a wasm module so the MCP server itself can accept a kernel as source (compute_submit_kernel / compute_submit_cuda) instead of requiring the caller to pre-compile.
LOCAL-FIRST: the wasmbuilder package (deps/wasmbuilder) compiles the kernel in-process — nurlc → IR rewrite → the toolchain's bundled zig cc — no network, no build service. Only when the local toolchain can't do it (no nurlc/zig on this box) does it fall back to POSTing {source, filename} to <NURL_BUILD_API>/build_wasm. A direct nurlapi answers with raw wasm bytes; the public playground proxy answers with JSON carrying wasm_base64 (and nurlc_errors on failure). Both are handled.
$NURL_BUILD_API fallback build service base URL (default https://play.nurl-lang.org)
@ build_api_url → String@ wrap_kernel s source i op i dtype i kkind → StringWrap a bare kernel into a complete wasm-ready program for reduce op op. dtype 0 (int): kernel is @ kernel i x → i; the module folds in i64 and prints the partial as a decimal integer. dtype 1 (float): kernel is @ kernel i x → f (x is the integer index, returns a double); the module folds in f64 and prints the partial's f64 BIT PATTERN as a decimal integer — so it rides the same stdout→int wire, and the coordinator reinterprets it (work.nu tids_combine float path).
kkind 0 (element): the generated main folds kernel(x) over [lo, hi). kkind 1 (chunk): the kernel is @ kernel i lo i hi → i (or → f) and the main calls it ONCE with the whole sub-range — the kernel owns the loop. This is the right granularity for kernels with per-invocation setup cost (open a CUDA context, JIT a device kernel, allocate buffers): one setup per CHUNK instead of one per element. The reduce op still combines the chunk partials.
@ compile_to_wasm s source → !( Vec u ) StringCompile NURL source to a wasm module. Ok = module bytes; Err = a human-readable transport/compile error (suitable to hand back to the model).
Local wasmbuilder first. A LOCAL "nurlc failed" is a genuine source error and is returned as-is (the build service would only repeat it); any other local error means the environment can't build wasm (no toolchain, no zig, download forbidden) → fall back to the build API.
packages/swarm-mcp/src/cudakernel.nu — generate a GPU chunk kernel from a bare CUDA-C map function.
The model hands over ONLY the math; cuda_wrap wraps it into a complete, self-contained NURL program in one of three OUTPUT MODES:
scalar (gpu_mode_scalar) — __device__ double f(long long x): grid-stride map over [lo, hi), per-block shared-memory reduction with the task's reduce op, host fold of the block partials, partial's f64 bit pattern printed on stdout (the classic reduce wire). sample (gpu_mode_sample) — same f, but every value comes BACK: out[x-lo] = f(x); the module writes the hi-lo doubles raw (LE) to an output file named on argv — curves, fields, images. hist (gpu_mode_hist) — __device__ long long bin(long long x) plus an optional __device__ double val(long long x) (default 1.0): atomicAdd(&out[bin(x)], val(x)) over K bins (K on argv — the same module serves any K), K doubles written to the output file. The double atomicAdd is a portable CAS loop, so the PTX stays valid for NVRTC's default (pre-sm_60) target.
RUNTIME PARAMS make the kernels dynamic: with has_params the device functions take (long long x, const double* p) and the module parses any argv tail of f64-bit-pattern decimals into a device buffer. Parameter VALUES never touch the generated source, so the module's content hash — and with it the worker-side cache and the coordinator's build cache — survives across a whole parameter scan: no rebuild, no re-upload.
The program's cuda/nvrtc FFI declarations mirror packages/gpu/src/cuda.nu's ABI exactly. Compiled with --ffi-host-imports (the wasm build API always does), every cu/nvrtc symbol becomes an env import that the pure-NURL wasmtime's GPU bridge resolves against the worker's real libcuda/libnvrtc under --allow-gpu. The same source also compiles native (nurl.sh auto-links cuda/nvrtc), so a kernel can be smoke-tested off-cluster.
Argv contracts (must match wasmkernel.nu __wasm_run_gpu): scalar: main lo hi [p…] sample: main lo hi outpath [p…] hist: main lo hi outpath K [p…]
Any CUDA failure exits non-zero WITHOUT producing output — the worker reports the chunk as failed (ok=0) instead of folding silent zeros.
@ cuda_src_ok s src → b── validation ──────────────────────────────────────────────────── The user source is spliced into a generated NURL backtick literal; a backtick would terminate it (NURL strings cannot contain one, escaped or not), so reject it outright. Everything else is escaped below.
@ cuda_wrap s user i op i mode i has_params i has_data → String── the wrapper ─────────────────────────────────────────────────── CUDA map function + reduce op + output mode (+ params flag) → complete NURL chunk-kernel program. The caller validates cuda_src_ok first. has_data: the chunk carries a dataset slice — argv gains an in-file path right after hi, and the device functions receive v = data[x] as their second argument.
packages/swarm-mcp/src/main.nu — swarm-mcp: an MCP-controlled distributed compute engine. A language model sets a workload over MCP — an expression kernel in x plus a range and a reduce op — and the cluster evaluates it distributed; the model reads running tasks and finished results back.
swarm-mcp relay 0.0.0.0 47700 [--v] # the meeting point (one per cluster) swarm-mcp worker <host> <port> # join as a compute node swarm-mcp mcp <host> <port> # MCP server (stdio) → drives the cluster swarm-mcp submit <host> <port> <reduce> <lo> <hi> <expr> # manual CLI submit
The cluster layer (membership → ring → job dispatch) is the swarm package's; here every worker registers ONE generic kernel handler (work.nu) that interprets the submitted expression (expr.nu), so any normal-operation workload runs without recompiling a worker. Phase 2 swaps the interpreter for a NURL→wasm kernel; the protocol and MCP surface stay the same.
@ swarm_vnodes → i@ kind_kernel → i@ pk_from_id i id → ( Vec u )A 32-byte opaque routing pubkey derived deterministically from a node id.
: Swarm: Swarm {
s transport // *Transport
s ring // *Ring
s gpu_ring // *Ring — the GPU capability domain (cap_gpu workers only)
s roster // *Roster
s job // *JobNode
( Vec u ) self_pk
i self_id
i role
i self_caps // capability bits this node advertises (cap_gpu for --gpu workers)
( Vec u ) group // token-derived 32-byte relay multicast group (cluster isolation)
( Vec u ) key // token-derived HMAC key (compute authenticity)
}
@ swarm_new RelayClient rc i id i role i caps s token → *Swarm@ swarm_free * Swarm sw → v@ swarm_join_group * Swarm sw → v@ swarm_announce * Swarm sw i want → v@ swarm_on_hello * Swarm sw Hello h → v@ swarm_pump * Swarm sw i max → v@ swarm_discover * Swarm sw i rounds → v@ relay_dial_retry s host i port i tries → !RelayClient NetErrDial the relay, retrying briefly so a co-located relay that is still binding (or a relay that restarts) doesn't lose its local roles to a startup race.
@ node_relay s host i port i vflag → vRELAY role: the dumb rendezvous point. Owns the fiber reactor on this thread.
@ node_worker s host i port s token i vflag i gpu → vWORKER role: join the cluster, register the kernel + wasm handlers (every worker can run wasm modules), and drain cluster traffic forever. Each worker thread takes a fresh random identity, so --workers N spins up N independent ring members in one process. gpu≠0: advertise cap_gpu and register the kind_wasm_gpu handler — wasm chunks of that kind run with --allow-gpu.
@ cluster_submit * Swarm sw i op i dtype i lo i hi ( Vec u ) expr i nchunks → ( Vec i )@ cluster_submit_wasm * Swarm sw i lo i hi ( Vec u ) wasm i nchunks i kind → ( Vec i )Shard a wasm-kernel task: ship the compiled module + each sub-range to its ring owner under kind (kind_wasm, or kind_wasm_gpu — the GPU capability domain). The module bytes ride every chunk (workers cache by content hash, so it is written once per worker).
@ cluster_submit_wasm_gpu * Swarm sw i mode i lo i hi i kbins ( Vec i ) params ( Vec u ) data ( Vec u ) wasm i nchunks → ( Vec i )Shard a GPU task (payload v2/v3): mode, K, runtime params and the module ride every chunk under kind_wasm_gpu — routed on the GPU capability ring. data is the WHOLE dataset's raw LE f64 bytes (empty when the task has no dataset); each chunk ships exactly its own slice data[clo·8, chi·8) — the split travels with its task, so a worker needs no separate fetch.
@ tids_ready * Swarm sw ( Vec i ) tids → bAll chunk results present? (cluster job results are recorded by task-id.)
: Combined { i value i nfail }Combine all chunk results with the reduce op (assumes ready). For a float task each partial is an f64 bit pattern: decode→f64, combine in f64, and return the combined f64 re-encoded as its bit pattern (the Task stores it that way; task_to_json reinterprets it for the model).
Two result frames ride the wire: the expr kernel's legacy [partial:8], and the wasm kinds' [ok:1][partial:8]. A chunk whose ok=0 (module trap, missing runtime, GPU failure) is COUNTED, not folded — nfail>0 marks the task as failed instead of silently reducing zeros into the answer.
@ tids_combine * Swarm sw i dtype i op ( Vec i ) tids → Combined: CombinedV { ( Vec u ) bytes i nfail }Combine VECTOR results (sample / hist). Frame: [ok:1][count:4 BE][f64 LE × count]. Sample chunks CONCATENATE in shard order (tids order = chunk order); hist chunks add ELEMENTWISE into K bins. A failed or malformed chunk counts toward nfail and contributes nothing.
@ tids_combine_vec * Swarm sw i mode i kbins ( Vec i ) tids → CombinedV@ nchunks_for i nworkers → i@ run_submit s host i port i op i lo i hi s expr s token → i@ nchunks_wasm i nworkers → iChunk count for a wasm task: fewer chunks than the expression path (each chunk ships the module and spawns a wasmtime process), but enough to spread across the ring.
@ run_runwasm s host i port i op i lo i hi s wasmpath s token → i: Task: Task {
i id
( Vec u ) expr // kernel bytes
i dtype // 0 int · 1 float (result holds an f64 bit pattern)
i lo
i hi
i op
i nchunks
( Vec i ) tids
i done
i result
i failed // chunks that reported ok=0 (task status "error" when > 0)
i mode // gpu_mode_scalar/sample/hist (scalar for every non-GPU task)
i kbins // histogram bin count (hist mode)
( Vec u ) vres // vector result bytes (raw LE f64s; sample/hist modes)
String out_file // when set, the finished vector result is written here
i dsid // dataset id the task maps over (0 = none)
}
: McpState: McpState {
s swarm // *Swarm coordinator
( Vec s ) tasks // *Task
i next_id
( Vec s ) wcache // *WasmCached — compiled kernel modules by source hash
( Vec s ) datasets // *Dataset — uploaded data the CUDA tools map over
i next_ds
}
@ task_new i id ( Vec u ) expr i dtype i lo i hi i op i nchunks ( Vec i ) tids → *TaskConstructor. The params share the field names (lo, hi, …); = . t lo lo is a field store — the field-store/local-shadow miscompile this used to work around was fixed in the compiler (NURL v0.10.4).
: ~ i g_mcp 0: Dataset { i id String name ( Vec u ) bytes }@ ds_count_of * Dataset d → i@ ds_find i id → s: WasmCached { String hash ( Vec u ) wasm }@ mcp_swarm → *Swarm@ mcp_pump i rounds → v@ task_refresh * Task t → vRefresh one task's status; combine if every chunk has landed. A finished vector task with an out_file writes the raw LE f64s there (once).
@ task_find i id → s@ task_to_json * Task t → JsonBuild the JSON-as-text result body for one task (LLM-readable + parseable).
@ tool_result_json Json o → Json@ tool_submit Json args → Json@ tool_run_wasm Json args → Json@ tool_submit_kernel Json args → Jsoncompute_submit_kernel: the server compiles NURL source → wasm itself (via the build API), then runs it — the model hands over a kernel as plain NURL.
@ tool_submit_cuda Json args → Json@ tool_sample_cuda Json args → Jsoncompute_sample_cuda: every f(x) comes back — the cluster gathers the range into one array (curves, fields, images).
@ tool_hist_cuda Json args → Jsoncompute_histogram_cuda: bin(x) picks a bucket, val(x) the weight (default 1.0); K bin sums come back — a whole distribution in one pass.
@ tool_upload_data Json args → Json@ tool_list_data Json args → Json@ tool_list Json args → Json@ tool_result Json args → Json@ ms_prop Json props s name s ty s desc → v@ ms_schema_submit → Json@ ms_schema_result → Json@ ms_schema_empty → Json@ ms_schema_run_wasm → Json@ ms_schema_submit_kernel → Json@ ms_schema_submit_cuda → Json@ ms_schema_sample_cuda → Json@ ms_schema_hist_cuda → Json@ ms_schema_upload_data → Json@ build_tools_list → ( Vec Json )@ dispatch_tool s name Json args → Json@ handle_initialize Json id → Json@ handle_ping Json id → Json@ handle_tools_list Json id → Json@ handle_tools_call Json id Json params → Json@ handle_unknown Json id s method → Json@ dispatch Json req → ?Json@ node_mcp s rhost i rport s mcp_host i mcp_port s cert s key s token → v@ arg_int i idx → i@ arg_eq i idx s lit → b@ has_flag s name → bTrue if name appears anywhere in argv (a boolean role/verbose flag).
@ flag_val s name s deflt → StringThe token following --name in argv, or a copy of deflt. Owned String.
@ flag_int s name i deflt → i: HostPort { String host i port }@ parse_hostport String hp s dhost i dport → HostPortParse "host:port", "host", ":port", or "" with host/port defaults.
@ ensure_cert s cert_path s key_path → bEnsure a TLS cert+key exist at the given paths, auto-minting a self-signed EC P-256 pair in pure NURL if absent (std/x509_gen.nu — no openssl, no subprocess), so --mcp works out of the box; pass --tls-cert/--tls-key for a real cert. CN + SAN are localhost (matching the previous openssl invocation minus the IP SAN — MCP clients pin or use --insecure against a self-signed cert either way). Returns T once both files exist.
@ cli_submit → i@ cli_runwasm → i@ usage → v@ run_node → i@ main → ipackages/swarm-mcp/src/wasmkernel.nu — phase 2: run a NURL-compiled wasm kernel over a chunk.
The language model writes a kernel as ordinary NURL — anything, not just an expression — and compiles it (via the NURL build service) to a wasm32-wasi module whose main reads lo hi from argv, folds the kernel over [lo, hi), and prints the partial. The coordinator ships that module plus a sub-range to each worker; the worker runs it under wasmtime and returns the partial, which the coordinator combines with the reduce op exactly as in phase 1.
The wasmtime here is just a CLI contract — wasmtime run <module> <lo> <hi>, partial on stdout. The toolchain's own pure-NURL runtime (packages/wasmtime) satisfies it exactly, so a worker needs NO external runtime: put that binary on PATH as wasmtime, or point $WASMTIME at it. The Bytecode-Alliance wasmtime works too — whichever $WASMTIME resolves to is used.
CPU chunk (kind_wasm) : [lo:8 BE][hi:8 BE][wasm…] GPU chunk (kind_wasm_gpu) : payload v2 — see wasm_gpu_chunk_payload below (adds an output mode, K, and runtime params)
The reduce op is baked into the module by the wrapper, so the partial is final per chunk. A scalar module prints its partial as a decimal integer on stdout; for a float task that integer is the partial's f64 BIT PATTERN (see buildwasm.nu), so the same int wire carries it unchanged — the coordinator reinterprets it. Vector modes (sample/hist) write raw little-endian f64s to a worker-named output file instead (fast: one fwrite, no per-value stdout).
The module is cached on each worker by a content hash, so re-running the same kernel (every chunk of a task, and across tasks — runtime params don't touch the module) writes the .wasm only once.
@ kind_wasm → i@ kind_wasm_gpu → iGPU wasm chunks are a SEPARATE job kind, routed on the GPU capability ring (dist/job job_set_ring): only --gpu workers own these keys, and only they register this handler — a non-GPU worker can never receive (or garble) one.
@ gpu_mode_scalar → iGPU task output modes (baked into the generated program AND carried in the chunk so the worker knows how to invoke the module and collect its output).
@ gpu_mode_sample → i@ gpu_mode_hist → i@ wasm_chunk_payload i lo i hi ( Vec u ) wasm → ( Vec u )@ wasm_gpu_chunk_payload i mode i lo i hi i kbins ( Vec i ) params ( Vec u ) data ( Vec u ) wasm → ( Vec u )── GPU chunk payload (v2 / v3) ────────────────────────────────── v2: [2][mode:1][lo:8][hi:8][K:4][nparams:2][f64 bits:8×n][wasm…] v3: [3][mode:1][lo:8][hi:8][K:4][nparams:2][f64 bits:8×n] [dlen:4][data slice: raw LE f64s][wasm…] mode selects the module's output contract; K is the histogram bin count (0 otherwise); params are runtime kernel arguments passed on the module's argv as f64-bit-pattern decimals — the SAME module (same content hash, no rebuild, warm worker cache) serves any parameter values. v3 additionally carries the chunk's DATASET SLICE — data[lo..hi) as raw doubles, following the map-reduce rule that a split travels with its task. dlen is exact (8·(hi−lo)) so a truncated frame is rejected, and dlen=0 means "no data" (the encoder then emits plain v2).
: GpuChunk: GpuChunk {
i ok
i mode
i lo
i hi
i kbins
( Vec i ) params
( Vec u ) data // raw LE f64 slice (empty without a dataset)
( Vec u ) wasm
}
@ gpu_chunk_free GpuChunk c → v@ wasm_gpu_chunk_decode ( Vec u ) body → GpuChunk@ _wasm_hash ( Vec u ) v → StringFNV-1a/64 content hash → 16 lowercase hex chars, for the cache filename.
: WasmRun { i ok i value }Run a cached module over [lo, hi). ok=1 iff the runtime ran the module to a zero exit — a failed chunk (missing wasmtime, module trap, GPU error path exiting non-zero) is REPORTED, not folded into the reduce as a silent zero. allow_gpu≠0 passes --allow-gpu so the module's env imports (CUDA/NVRTC) resolve to real hardware — this needs the pure-NURL packages/wasmtime.
@ wasm_handler ( Vec u ) key → ( @ ( Vec u ) ( Vec u ) )The worker handler for a (CPU) wasm chunk: verify the cluster HMAC tag, cache the module by content hash, run it under wasmtime, return the partial tagged for the coordinator. An untrusted/forged payload yields a tagged ok=0 — a stranger can't make a worker fetch-and-run an arbitrary module. key is the cluster HMAC key (token_key), captured by the handler closure.
: GpuOut { i ok i scalar ( Vec u ) bytes }@ wasm_gpu_handler ( Vec u ) key → ( @ ( Vec u ) ( Vec u ) )GPU worker handler: decode payload v2, cache the module, run per mode, and reply the mode's result frame — scalar [ok:1][bits:8], vector [ok:1][count:4 BE][f64 LE × count]. Forged/undecodable payloads → ok=0.