Skip to content

Specification §19

Observability

language-design.md §19 · 109 lines · 8 min read

The language has isolated and supervised processes (section 8), and a system like that needs to be observable in production without being taken down for it. Observability here covers two things: introspection of the runtime (which processes exist, how much they consume, how the tree looks, the state of the channels) and live tracing (the flow of events, messages, deaths, GC and scheduling, over time). What it does not cover is application telemetry (logs, metrics and spans that you emit): that is a library and stdlib, solved as in any language, and stays out of here.

The principle that governs everything is the opt-in test (section 1) applied to cost: observability is pull and default-off, and you do not pay for what you do not turn on. Reading the state (introspection) is cheap and on-demand; tracing events (trace) costs, so it stays off and you turn on only what you want, at the frequency you want. Did not use it, did not pay; wanted everything all the time, paid for everything; need three things, pay only for them. It is control and transparency, not always-on magic.

What the architecture already exposes for free

Section titled “What the architecture already exposes for free”

A good part of observability you already produced without calling it observability. The architecture generates the data as a byproduct, and the introspection just gathers it:

  • Memory per process, with strategy: each process has its @mm (section 5), be it arena, gc, none or c. The runtime already knows exactly how much memory each process uses and of which strategy, because it needs that to operate. It is a metric that falls into your lap, and that most platforms do not have so cleanly (a shared heap does not tell you “this process uses X”). You have memory isolation, so you have per-process memory accounting.
  • The supervision tree: it is the structure of your code (section 8). Who supervises whom, who spawned whom, restart rate, all already exist structurally, they do not need to be reconstructed by instrumentation.
  • The reason for the death: the |e| of section 8 is already a structured observable event, because each death carries the error that caused it. “Why did process X die” does not require tracing, it is data that the supervision model already delivers.

Introspection does not invent this data; it exposes it with a uniform interface. Tracing adds the temporal dimension (the flow, not just the snapshot).

Observability splits into two APIs because they have opposite costs:

  • Introspection (snapshot): reads state that already exists. Cheap, pull at any time, with no continuous overhead, a snapshot of the now, taken when you ask. Listing processes, seeing memory, the depth of a channel: none of this needs instrumentation turned on.
  • Trace (stream): captures a flow of events over time. It costs (the runtime intercepts and records continuously), so it is opt-in and default-off, and you pay only for what you turn on.

A snapshot is reading what is there (free); a trace is instrumenting to capture what passes (pays while on).

Everything lives in the runtime module:

procs := runtime.processes() // snapshot: list of process handles
p := runtime.process("db_writer") // a process by registered NAME (restart-stable, section 3)
grp := runtime.group("handlers") // a registered group → iterable collection, [i] by index
chans := runtime.channels() // channel handles
stats := runtime.stats() // global aggregates of the runtime

A snapshot of a running process is taken at a safepoint: the runtime synchronizes to read consistent state, without a race with the process itself. The data already exists (the architecture produces it, section 19 above), and the consistency is just the read point.

Each process handle (Process) carries the maximum, modeled on Erlang’s process_info (the reference in introspection) plus your @mm differential:

id (internal runtime identity, for display, not a lookup key and does not survive a restart), name (the registered name, if any, the stable key; unregistered ones appear without a name), state (running, runnable, waiting in a recv, suspended, exiting or collecting), current_fn (plus stacktrace), spawn_point (with which function it was created), memory (plus heap/stack), mm_strategy (gc/arena/none/c) plus allocated and tracked objects, work (work done, the reductions), supervisor plus children (the tree), restarts, uptime, priority, gc_stats.

Each channel handle: id, depth (queued now), capacity, blocked_senders / blocked_receivers (who waits, valuable for deadlock), sent / received, closed, type and direction.

The global one (stats): process_count, channel_count, workers (schedulers), memory_total plus memory_by_strategy, run_queue, gc_stats, uptime.

Lookup is by name (runtime.process("name") returns a Process) or by group (runtime.group("name") returns a collection, [i] by index), never by raw id: the id is a point in time and ages on the restart, the name crosses (section 3). From inside a process, runtime.self() returns its own reference, and binding a spawn (p := spawn worker()) returns the new process’s. Beyond inspection, a handle can be waited for and stopped (p.wait(), p.kill(), section 3). Since everything is pull, this maximum surface costs zero until you call it, so “more data is always better” comes for free in introspection.

Trace: cheap capture, configurable consumption

Section titled “Trace: cheap capture, configurable consumption”

A trace turns on a flow of events:

runtime.trace(target, events, sink, options)
  • target: a process name, a group name, or .all (and runtime.self() for self-trace).
  • events: .send/.receive (messages), .spawn/.exit (lifecycle), .schedule_in/.schedule_out (scheduling), .gc, .restart, .channel_block (a process got stuck on a channel), and the expensive and optional .call/.return and .alloc/.free. (It is the taxonomy of Erlang’s trace flags.)
  • options: sample: N (1 in N), duration: T (time-boxed), inherit (the trace is inherited by the target’s children, the set_on_spawn), timestamp.

The capture is always a ring buffer in the runtime: each event writes a fixed record, without scheduling anyone. It is the floor of overhead (every nanosecond counts in the capture; it is what Erlang’s native erl_tracer does so as not to pay a message per event). What changes is how you consume the buffer, and here come the sinks, now as a choice of consumption only:

  • .buffer (default): events accumulate in the circular buffer and you read them on demand (a snapshot of the recent window). It is the cheapest, because the capture writes and you read whenever you want. Ideal for “leave it running, and if something goes wrong I inspect the history”.
  • .channel: a consumer process drains the buffer and streams it on a channel, consumed with the usual loop/match/<-. The transmission (even with a queue) is not critical in nanoseconds like the capture, so it can pay the channel.
  • .callback(fn): each event calls a function of yours; the app reacts to its own trace.
  • .endpoint: events drain to the native port, for external consumption.
ev := runtime.trace("worker", [.send, .receive, .channel_block], sink: .channel, sample: 100)
loop e in ev {
match e {
Send(m) => ...
Receive(m) => ...
ChannelBlock(ch) => alert("worker stuck at {{ch.id}}")
}
}

The stream’s element type is Stream[Event{events...}]: it follows the comptime list of events, so the match covers exactly what you asked for (a set of variants, section 7), exhaustive without _. A list chosen at runtime does not narrow: there e is the full Event.

The cheap capture (buffer) decouples from the consumption (sink), and that is the key to “pay only for what you turn on”: turning on the capture is cheap and uniform; the cost of transmitting and formatting lives in the sink, which you choose. The three active sinks (.channel/.callback/.endpoint) oppose self-observation to external debug; only the consumer changes, not the mechanism.

Each event (and each stack frame) carries its provenance: user (your code), runtime (scheduler, GC, channel machinery) or tracer (the cost of the tracing itself). It is a label the runtime assigns at the capture, and it unlocks filters in the consumer:

  • hide_runtime: hides the internal frames of the runtime; you see only your program, without the noise of the machinery.
  • hide_tracer: hides the cost of the tracing itself; you do not confuse the overhead of measuring with the program’s behavior (the anti-heisenbug in reading, where the profile shows your app, not the observer).

Both are toggles of whoever consumes (on the server or in the app that reads), not of whoever captures: the capture takes everything (it is cheap), the consumer filters. About formats: the capability of stack-sampling (sampling the stack plus the scheduling events) is the runtime’s; the design (flamegraph, call graph) is the consumer’s and the tooling’s. The runtime delivers the raw labeled data; the tool renders and filters.

An honesty the design assumes: observing a concurrent system alters it, and this does not eliminate, it only reduces. Measuring has a cost, the cost changes the timing, and the timing changes concurrency bugs. No one solves this (not BEAM, not Go, not kernel profilers). The design minimizes the perturbation by four routes: sampling (sample: N perturbs less than measuring everything), ring buffer (writing to the buffer perturbs less than formatting and transmitting per event, so the expensive thing stays out of the hot path), predictable cost (uniform overhead is less treacherous than variable), and default-off (in normal production there is no perturbation; you pay the heisenbug only during the investigation, consciously). It is not an anti-heisenbug feature; it is a set of choices that reduce the perturbation, and the design says so instead of pretending it solves it.

The endpoint serves introspection and trace in the language’s native serialization, without coupling to Prometheus, OpenTelemetry or any third-party format. The reason is the usual one (control, not dependency): the language does not die if a tooling format goes out of fashion, because the raw API survives any one of them.

And the gain is that, with the API in code, the programmer builds the adapter for what they use. Want Prometheus? A handler reads runtime.processes()/runtime.stats() and spits it out in the Prometheus format. Want OpenTelemetry? Same thing, another format. Adapters are user library over the raw API, not the runtime’s responsibility. The language does not marry any format, but does not prevent any: Prometheus becomes a lib that someone writes (the same principle as cdeps in the FFI, which does not couple to any C lib and lets you use them all).

Respecting that tooling comes later, what enters now into the runtime and the language (what this section defines) is the introspection API (functions and fields), the trace primitive (runtime.trace, buffer capture, sinks, provenance), and the endpoint serving the data in native serialization. What is left for the tooling phase is the observer (the UI/CLI that connects to the endpoint and shows everything, Erlang’s dbg/:observer, the flamegraph and graph rendering) and the exact wire-format. You get the trace (the primitive) now; the dbg (the tool) in the tooling phase.