Skip to content

Specification §8

Resilience: supervisors, links and monitors

language-design.md §8 · 223 lines · 12 min read

@supervisor(strategy: one_for_one)
spawn {
    spawn ingest(src) catch |e| {
        log("ingest died: {{e}}")
    }
    spawn flush(out)
    spawn stats(tick)
}
The way your spawns nest is the tree. A process that crashes dies alone, and catch decides what happens next.

Your language inherits the resilience model from BEAM (Erlang/Elixir): isolated processes, with no shared memory, that fail early and quietly (let it crash) instead of spreading corrupted state, plus a supervision tree that restarts what fell. Whoever comes from BEAM knows three central concepts (links, monitors and supervisors), and this section shows exactly how each one is expressed here. The summary, to anchor it: supervisors and the tree emerge from the structure of the code (not from separately declared behaviours), links are not exposed raw (only through the supervision tree), and monitors are not a new construct, they are the catch of the supervised spawn. The rest of the section details each one.

Before links and monitors, you have to be exact about what happens when a process dies, because links and monitors are just reactions to that event. The sequence is fixed and ordered:

  1. The process’s internal defers run (LIFO, section 14): the deterministic cleanup the process registered (close a file, release a lock, @unpin of an object given to C). It is the process’s last chance to tidy up the house, and it runs inside the dying process, while the stack still exists.
  2. The process’s @mm is released: its memory goes back to its MemoryManager (section 5). Since each process has its own memory context (isolation), death cleans everything up at once. There is no leak between processes, and the dead one’s arena/GC simply ceases to exist.
  3. The error that describes the death is produced, and this is the point that connects everything: a process’s death is an error value (error-as-value, section 7). Did it crash? the error carries the reason (which variant, which message). Did it leave cleanly (finished the work)? the “death” is a normal exit. That error is what will be delivered to whoever observes the death.

The three in that order (cleanup, memory, notification) guarantee that, when someone is notified of the death (step 3), the process has already fully unwound (steps 1 and 2). You never observe a “half-dead” process.

main is the root of the tree, and the program ends when main ends – but the processes still alive at that moment do not simply vanish: each one goes through the same death sequence. When main’s body and its own defers have finished, the runtime stops every process that is still alive (as p.kill() would): each runs its defers, releases its @mm, and its death reaches its supervisor as error.Killed – so main’s own handlers for its children run too – and only then does the program exit. A child’s defer flush() therefore always runs, whether the child finished before main or was still working.

If you want a child to finish rather than be stopped, wait for it before main returns (p.wait(), or runtime.wait_children() for all of them); if you want it stopped earlier, p.kill() / runtime.kill_children(). The shutdown only takes care of what is left.

A process that cannot stop – one blocked inside a C call, which has no safepoint (section 17) – is given a grace period (5 seconds by default, @shutdown(grace: 10s) on fn main to change it). When it runs out, the program exits anyway and reports on stderr which process did not stop; that process’s defers did not run.

In BEAM, a link is a symmetric tie that propagates death: if A and B are linked and one dies, the other receives an exit signal and, by default, dies along with it. It is bidirectional (the tie is mutual, not “who links whom”) and lethal by default. It is the raw mechanism that makes groups of processes live and die as a unit.

Your language does not expose raw links, there is no link(pid). The reason is the one that guided the rest: a death cascade is hard to reason about (A kills B which kills C linked to D… who is left?), and a loose link is a non-local effect that does not appear in the structure. Instead, the only way to couple fates is the supervision tree: processes in the same spawn-block (or under the same supervisor) share a fate because the structure of the code says so, not because someone called link in a distant place.

And the supervisor is a linked process, only special: in BEAM, a supervisor traps the exit signal (trap_exit) instead of dying with it, and decides what to do. Here it is exactly what @supervisor does: it is coupled to the children (receives their death), but it traps that death, which becomes the error of the catch instead of bringing the supervisor down. Links exist, underneath; you only touch them through the discipline of the tree.

Monitors: the catch of the supervised spawn

Section titled “Monitors: the catch of the supervised spawn”

In BEAM, a monitor is the asymmetric opposite of a link: A monitors B, and if B dies, A receives a message notifying it, but A does not die. It is unidirectional (B does not even know) and non-lethal (B’s death becomes data for A, not a sentence). It is what you use to know that something fell without falling along.

Here is the central unification of the section: a monitor is not a new construct, it is the catch |e| of the supervised spawn.

@supervisor(configs)
spawn process(values) catch |e| {
// 'process' died, and I (the supervisor) did not die along with it.
// 'e' is the reason for the death. I react, without propagating.
}

Compare with the definition of a monitor (“being notified of another process’s death, without dying along, receiving the reason”) and it is exactly this catch: the child dies, the supervisor does not die and receives the error in the |e|. “Reporting the death without propagating it” is, word for word, “handling the error without re-raising it”. The catch you already use for ordinary errors is the monitoring channel: death arrives by the same mechanism as any error, because a process’s death is an error. (Reconciles with section 7: here the catch is in a statement position, not in a value-binding. There is no v to fill, so the handler reacts and continues, without the obligation to diverge that the trailer carries when it binds a value.)

This is what makes links/monitors “already solved” here: BEAM needs two separate concepts (link couples fate, monitor observes without coupling) because there both are runtime primitives. Here, coupling fate is the tree and observing without coupling is the catch, and both already exist for other reasons.

(Honesty about the scope: this ties observation to the spawn relation, that is, you observe the death of the children you spawned. It is a deliberate narrowing relative to BEAM’s monitor, which observes any pid. Observing a process that another spawned is done via the normal message path, where a process sends to you on a channel, not via a monitor over an arbitrary pid.)

The |e| is the reason for the death, the same error from step 3. And since it is an ordinary error, you discriminate it with the match |e| that already exists (section 7), distinguishing how the process died:

@supervisor(configs)
spawn process(values) match |e| {
Normal => log("process finished its work") // clean exit
Crashed(msg) => restart() // crash with reason
Timeout => escalate() // stuck
OutOfMemory => alert_ops() // resource exhausted
Killed => log("stopped on purpose") // p.kill(), or the shutdown at the end of main
}

Killed is the death of a process that was stopped: by p.kill() (section 3) or by the orderly shutdown at the end of main (below). It is neither a clean exit nor a crash – someone decided it – so an automatic @supervisor policy (one without a handler) does not restart it, exactly as it does not restart a Normal exit; a handler may still restart() it if it wants.

It is match |e| directly on the spawn, not catch |e| with a match inside: the match |e| already captures and discriminates at once (section 7), and the two together would be the redundant double-binding. Clean death (finished the work) and death by crash arrive through the same |e|, discriminated by the variants, and you restart, escalate, log or ignore according to the reason. There is no one channel for “finished normally” and another for “crashed”: a single error, with variants, matched by the usual match. (The variants Normal, Crashed and the like are illustrative; the point is that the reason is a structured value, not an opaque code.)

When a process fails:

  1. Recovery attempt: the process tries to restore its last valid state
  2. Fresh start: if recovery fails (corruption, etc.), the process restarts clean from the entry point

The supervision tree emerges from the structure of the code, not from separately declared modules or behaviours. Erlang/OTP’s three strategies map naturally onto the syntax:

one_for_one: independent processes with individual catch

Section titled “one_for_one: independent processes with individual catch”

Each process has its own handler. If B dies, A does not know and does not care.

spawn worker_a(data) catch |e| {
// A failed: restart? log? drop? the caller's decision
}
spawn worker_b(data) catch |e| { ... }

With no annotation at all, the default behavior is one_for_one with no restart limit.

one_for_all: coordinated group with a spawn-block

Section titled “one_for_all: coordinated group with a spawn-block”

The block communicates that the processes live and die together. If one member fails, the coordinator proactively kills the rest before restarting the whole group.

spawn {
worker_a(data_a) catch |e| {
// Layer 1: A failed specifically: local notification
// B and C do not pass through here when they are terminated by the coordinator
log("worker_a down: {{e}}")
}
worker_b(data_b) catch |e| { ... }
worker_c(data_c) catch |e| { ... }
} catch |e| {
// Layer 2: the whole group is exhausted (max_restarts exceeded)
// Here you decide: escalate? alert? drop?
notify_ops("critical group died: {{e}}")
}

rest_for_one: pipeline with order dependency

Section titled “rest_for_one: pipeline with order dependency”

The only case that requires an explicit annotation, because the “restart the ones that came after” semantics is not inferable from the structure:

@supervisor(strategy: rest_for_one)
spawn {
db_connection() // dies → restarts all three
db_writer() // dies → restarts writer + reader
db_reader() // dies → restarts only reader
} catch |e| { ... }

@supervisor is optional and only appears when you want to change something from the default:

// With an attempt limit
@supervisor(max_restarts: 3, window: 10s)
spawn worker(data) catch |e| { ... }
// Combined with rest_for_one
@supervisor(strategy: rest_for_one, max_restarts: 5, window: 30s)
spawn {
stage_a(data)
stage_b(data)
} catch |e| { ... }

When max_restarts is exceeded, the block’s outer catch (or that of the individual spawn) fires with the accumulated error. From there, the caller decides; there is no implicit automatic escalation.

Restart comes in two forms, and which one you get depends on whether you wrote a handler:

  • @supervisor(...) on a spawn with no catch — the runtime restarts, on its own, following the strategy and bounded by max_restarts. This is the automatic form: you declared the policy and the runtime applies it.
  • @supervisor(...) on a spawn with a catch — you took control. The runtime does what the handler says: restart() forces a restart now, escalate() pushes the failure up, and doing neither leaves the child dead.

That is why the same catch |e| is both the monitor and the supervisor: the difference is only whether the handler acts. And it is why a clean exit is not restarted by accident — a handler that matches Normal and merely logs it has decided, and the runtime does not second-guess it.

runtime.supervisor.{restart, escalate}
@supervisor(max_restarts: 3, window: 10s)
spawn process(values) catch |e| {
match |e| {
Normal => log("the process finished its work") // decided: stay dead
Crashed(msg) => restart() // force a restart
Timeout => escalate() // hand it upward
OutOfMemory => alert_ops() // decided: stay dead, but shout
}
}

restart() takes no argument: the runtime kept the spawn’s function and arguments, and re-invokes it from its initial state. The error is already bound by the |e| of the handler you are standing in, so passing it would be repeating yourself. The restarted process is a new one internally, but a registered name survives (section 3).

Because the restart re-invokes with the same arguments, a channel handed to the process as a spawn argument reconnects: it comes back live, the endpoint moves onto the restarted process so its counts never dip, and a peer holding the other end — the parent that spawned it, which never died — keeps talking to it across the restart without ever seeing a spurious ProcessDown. What does not reconnect is a raw handle that some other process stored to this one: that handle named the dead internal identity, so it goes stale and a send on it yields ProcessDown. Reaching a restartable process from the outside is therefore by its registered name (section 3), never by a stored handle. This is BEAM’s model exactly — same child-spec arguments, a new internal id, a surviving name.

escalate() takes no argument either. It kills the current process with the observed error, delivering it to whoever supervises it — the supervisor saying “this is not mine to handle”. The ordinary death sequence runs (defers, @mm, error). With nobody above, it is an unsupervised failure: if that process is main, the program exits with the error. Escalation is always explicit, which is the other side of “there is no implicit automatic escalation” above.

Both live in runtime.supervisor, not as bare globals: nothing is too fundamental to import (section 1). They sit apart from the read-only introspection of runtime (section 19) because they only mean anything inside the handler of a supervised spawn — restart() on its own has nothing to restart.

The tree emerges naturally from the nesting of the processes. There is no separate concept of “supervisor process”: any process that spawns children is implicitly their supervisor:

// Root: implicit one_for_one (the main program)
spawn http_server(cfg) catch |e| { restart() }
@supervisor(strategy: rest_for_one)
spawn {
// Sub-group: rest_for_one
db_connection()
db_writer()
db_reader()
} catch |e| { notify_ops(e) }

Retry is explicit and defined by the caller, not by the callee. The function executed by the process is clean, with no built-in retry policy:

@supervisor(max_restarts: 3, window: 10s)
spawn worker(data) catch |e| {
log("worker failed after 3 attempts: {{e}}")
}

The reason: if the callee defines its own retries, it becomes hard to reuse and the caller loses control in critical scenarios.

From Erlang/OTP to here: translation table

Section titled “From Erlang/OTP to here: translation table”

For whoever comes from BEAM, the direct mapping:

Erlang/OTP In your language
spawn_link (couple fate) processes in the same spawn-block or under the same @supervisor; the tree couples, not a call
monitor/2 (observe without coupling) the catch |e| of the supervised spawn
{'DOWN', ref, process, pid, reason} the |e| of the catch, that is, the error that describes the death
trap_exit (supervisor traps the signal) the @supervisor trapping the child’s death
exit(reason) / exit reason the error produced on death (step 3)
Supervisor behaviour + child spec @supervisor(configs) over the spawn-block
one_for_one / one_for_all / rest_for_one individual spawn / spawn-block / @supervisor(strategy: rest_for_one)
restart intensity (max_restarts, period) @supervisor(max_restarts: N, window: T)

The philosophical difference in summary: BEAM gives runtime primitives (link, monitor, trap_exit) that you compose to build supervision; your language makes supervision emerge from the structure and treats death as error-as-value, so BEAM’s concepts become consequences of things that already exist (the tree, the catch) instead of separate primitives.