A distributed backend framework where crashes self-heal, code hot-swaps, faults speak a protocol, and continuations drive conversations.
Pure Chez Scheme · Erlang-style actors · libuv event loop · MIT
$ npm i igropyrHTTP requests run in a supervised worker pool. Handlers don't defend—they crash, and the system recovers. (WebSocket sessions own their own processes, so a crash kills only that connection, leaving the pool untouched.)
Write the happy path. The supervisor owns the sad one.
(app-get app "/crash" (lambda (req res) ;; the worker dies; the supervisor retries on a ;; fresh worker -- the pool refills itself (raise 'handler-crashed))) (app-get app "/stuck" (lambda (req res) ;; a hot loop cannot freeze the system: preemption ;; keeps serving, the ticker kills this worker (let loop ((n 0)) (loop (+ n 1)))))
Swap the entire handler—or just patch a single route—on a live server. The TCP listener, open connections, and the worker pool remain completely untouched. In-flight requests gracefully drain on the old code, while new traffic instantly hits the new logic.
Routes live in a mutable registry behind the worker pool. Re-registering a path patches it atomically for the very next request; http-swap! replaces the entire top-level handler with the exact same guarantee.
Combined with graceful shutdown (http-shutdown! safely drains in-flight work) and SO_REUSEPORT multi-process listening, zero-downtime operation isn't an engineering project—it's the default.
(app-get app "/version" (lambda (req res) (send-text! res "v1"))) ;; hit /upgrade on the LIVE server: (app-get app "/upgrade" (lambda (req res) (app-get app "/version" ; re-register = (lambda (req res) ; hot replace (send-text! res "v2 (hot swapped)"))) (send-text! res "upgraded")))
When retries are exhausted or a stuck worker is killed, Igropyr doesn't just throw a black-box 500 error—it tells the client exactly what happened, on a connection that stays open.
The on-failure hook returns a structured JSON fault only after the stuck worker is physically dead. When the client hears stuck, it comes with an absolute guarantee: there is no execution left in flight. The state is definite.
Keep-alive survives the fault. The client resubmits on the very same connection and gets a fresh retry round. Dial down the stuck-ms limit, and a user who once stared at a spinner for 30 seconds now transparently cycles through several informed retries in the exact same time. Failures become invisible at the UI.
(app-listen app 8080 `((stuck-ms . 3000) ; fail fast (check-ms . 1000) (on-failure . ,(make-fault-handler)))) ;; the client receives, connection kept alive: ;; {"fault":"crash","attempts":4,"retryable":true} ;; {"fault":"stuck","elapsed-ms":3012,...} ;; unset? the plain 500 remains. zero breakage.
A multi-request workflow—a checkout wizard, a booking, a fund transfer—runs as one single green process. Its local lexical bindings are the conversation state. This allows it to hold resources that a disconnected session store could never serialize: like a live, open database transaction spanning across multiple network roundtrips.
"The user is at the confirm step" literally means the process is parked at that exact line of code. A sequence of events that the code cannot express simply cannot happen—there is no external state machine to get wrong, and no replay attack to defend against.
The gone guarantee
The transaction declares its point of no return through the commit! primitive, giving the framework a razor-sharp boundary to judge any death.
commit!: A dead process means a dropped connection, which means the database itself automatically rolled back. The framework answers gone—absolute physical proof that nothing committed. This is the one status you may safely retry on.commit! (or an unknown kill/missing record): The framework answers unknown. It refuses to guess. Reconcile; never resubmit.Combined with the fault protocols above, the client always knows the definite server state. Just as importantly, it is told plainly—and honestly—when the state cannot be known. The complete remote transaction ring is finally closed.
(conversation-start! (lambda (req suspend! commit!) (let ((tx (begin-tx!))) ; live, across requests (guard (e (#t (rollback! tx) (raise e))) (let ((req2 (suspend! confirm-page))) (commit! (lambda () (commit-tx! tx))) done)))) req) (conversation-resume! id token req) ;; => (values reply next-token) | 'stale | 'gone ;; the token names the reply you are answering; ;; repeating one replays its answer, once. ;; commit through commit!: a failure after it is ;; 'unknown, never the retryable 'gone.
When the client is Scheme too, requests and replies are pure s-expressions. There is no codec to design, agree upon, or debug. (igropyr sexpr) acts as the safe boundary parser, while app-rpc dispatches exactly one datum per message—seamlessly across HTTP, WebSocket, or SSE.
Exact ratios and bignums cross the network intact. There is no lossy JSON floating-point approximation anywhere in the stack. A call like (rpc "/rpc" '(add 1 2 1/2)) comes back as (ok 7/2)—the mathematical ratio perfectly preserved.
The WebAssembly peer
The natural partner to this backend is Goeteia, a Scheme compiler running natively in WebAssembly. Its browser-side (web rpc), (web ws), and (web sse) modules speak this exact same wire format, turning the browser into a first-class Scheme runtime.
This very site is written in pure Scheme and compiled down to bare HTML and CSS. That honeycomb fire effect above? It is compiled and rendered in real time, directly in your browser, by Goeteia.
;; Igropyr: one s-expression per message (app-rpc app "/rpc" `((add . ,(lambda (args) (apply + args))) (get-user . ,(lambda (args) (find-user (car args)))))) ;; a Scheme browser -- Goeteia -- calls it. no JSON, no codec: (rpc "/rpc" '(add 1 2 1/2)) ;; the Igropyr server returns (ok 7/2) ;; -- exact ratio intact
Nodes discover each other and wire up a full, true mesh—no central coordinator, and no fragile registry to babysit. Links self-heal, and work fluidly spreads across every live member.
Point cluster-start at a discovery strategy, and it keeps the topology honest: it actively dials any member it isn't linked to yet, and mercilessly drops anyone that leaves. The static strategy uses a fixed list; with redis, nodes heartbeat themselves into a transient set. If a node stops beating, it simply falls out of the mesh. There is no central bookkeeping to drift out of sync.
Secure, name-based routing
Underneath lies a pure node-to-node distribution layer. A mutual HMAC-SHA256 handshake strictly gates who may join. Once inside, rsend and rcall reach registered processes on remote machines purely by name. monitor-node watches members come and go, while the links themselves stubbornly reconnect through network blips.
Distributed execution pools
(igropyr dpool) rides on top of this mesh. Submit a task, and it lands on an available live node. If that node suffers a physical death mid-execution, the mesh notices, and the work instantly reappears elsewhere. You get a guaranteed at-least-once execution primitive, seamlessly stretched across the entire cluster.
(node-start! 'web-1 secret 8888 "0.0.0.0") ;; discover peers via redis; nodes heartbeat ;; themselves in and expire on their own (cluster-start `((name . "render-farm") (discover . (redis ,conn "10.0.0.1" 8888)))) ;; fan work across every live member; a node ;; dying mid-task -> the task reruns elsewhere (define pool (dpool-start '(web-1 web-2 web-3) 'render)) (dpool-await pool (dpool-submit pool #(resize "x.png" 800)))
Every line is Scheme — R6RS libraries in .sc, no C shim. libuv, zlib and the crypto for MySQL auth are reached through Chez's FFI directly. Whole-program compilation folds the framework and your app into one optimized binary.
Green processes with spawn / send / receive, link and monitor, a process registry, gen-server and PubSub. One OS thread, preemptive scheduling, pure message passing — no shared state, no locks.
One event loop feeds thousands of parked processes. DNS, file reads and database round-trips park the calling process, never the thread. Non-blocking HTTP/WebSocket clients and Redis, MySQL and PostgreSQL drivers included.
(http-listen port (lambda (req res) ...)); the bundled (igropyr express) layer (create-app, app-get, send-json!, ...) is optional, and alternative frameworks can be built on the same corespawn / send / receive / link / monitor; no shared state between processeson-failure handler answers a structured JSON fault instead of the plain 500, on the same keep-alive connection — the client resubmits (changed parameters, carried state) and gets a fresh retry round; unset, the plain 500 remainssuspend! answers and parks, conversation-resume! continues — carrying a token that names the reply it answers, so a double click or a retried request replays that answer instead of taking a step nobody asked for; the transaction commits through commit!, so a death that left the flow before it is the rollback guarantee — a later resume gets gone and may be retried — while one after it is unknown, which may notres-begin!/res-write!/res-end!; Server-Sent Events helpers on topgen-server (call/cast/info), a process registry (register/whereis), and topic PubSub with automatic cleanup of dead subscribersread; full escape and surrogate handling) and writer(igropyr sexpr) is a safe whitelisted parser (no read, depth-limited), and app-rpc / send-sexpr! / ws-send-sexpr! / sse-send-sexpr! carry one datum per message — exact ratios and bignums cross intact. The browser end is Goeteia's (web rpc/ws/sse)req-form parses urlencoded and multipart bodies (file uploads included); req-cookie / set-cookie!Transfer-Encoding: chunked request bodies are decoded transparentlyhttp-get / http-post and ws-connect, both with async DNS (libuv thread pool) and the same park-the-caller modelAccept-Encoding; static files cache their compressed form/metrics endpointhttp-stats (live connection/request/pool counters), http-shutdown! (drain in-flight requests, refuse new connections)SO_REUSEPORT bind option for kernel-balanced multi-process listening on Linux (pair with pm2 or systemd)ab -n 50000 -c 500, zero failed requests), on an Apple M4 ProIgropyr is built on Chez Scheme — the fastest Scheme compiler, with a first-class FFI that reaches libuv directly. With deep gratitude for Kent Dybvig's life work, and to Cisco for open-sourcing it.
The primary inspirations: Node.js is the event-loop server on libuv, and the lean core / optional-framework split that Node and Express made the norm. The actor model, the supervisor, and Let It Crash come from Erlang/OTP; Swish — a Chez Scheme system built on those ideas — was the concrete blueprint for the scheduler, the receive macro, and the supervisor. The conversation model is the actor-native take on web programming with continuations — a great idea from the Scheme and functional-programming community.