Skip to content

Specification §18

Inline assembly

language-design.md §18 · 120 lines · 7 min read

Inline assembly is the low-level extension of the FFI (section 17): just as @cimport brings C inside, assembly brings raw machine instructions, for whoever needs to descend to the metal (a syscall, an rdtsc, a SIMD instruction the compiler does not emit, a memory barrier). It is the lowest boundary the language offers, used by a minority, and like every boundary out of the guarantees, it is unsafe: the judgment “this assembly is correct” is inexpressible to the compiler, it is the human case-B of assume (section 5).

There are two forms, which share the same decorators: the assembly-function (the whole function is asm, with params and return bound to registers) and the inline block (instructions in the middle of normal code). The rule for choosing is simple. If you want a value of asm, named and reusable, use the function; if you want to run asm in the middle of normal code, touching variables that already exist, use the block.

The function’s body is assembly, and the params and the return bind to registers via @asm(in, …)/@asm(out, …) placed on the param itself (the binding stays next to what it binds):

@asm(x86, intel)
unsafe fn add10(value: u32 @asm(in, rax)) -> u32 @asm(out, rbx) {
mov rbx, rax
add rbx, 10
}
// use: value-form, like any function:
result := unsafe add10(x) assume "x fits in u32 and rbx is not used by the ABI here"

@asm(x86, intel) decorates the function with architecture and dialect (and options, below). value: u32 @asm(in, rax) binds the param to rax; -> u32 @asm(out, rbx) says the return is rbx. The same @asm carries two roles without ambiguity: on the function it takes (architecture, dialect); on a param or local it takes (in|out|clobber, register), and the position plus the first argument distinguish. The clobbers are implicit here: the function is a call boundary, so the ABI already defines which registers it dirties, and you list nothing. The function returns a value (result := add10(x)), is named, reusable, testable, and with inlining has zero cost. It is the form for “asm that produces a value”.

To run asm in the middle of a normal function, use the block. The decorators are stacked above the unsafe {} (a decorator always above what it decorates), and the block contains only instructions:

var lo: u32
var hi: u32
@asm(x86, intel)
@asm(out, rax) lo
@asm(out, rdx) hi
@asm(clobber, "memory")
unsafe {
rdtsc // puts the timestamp in EDX:EAX
}
// 'lo' and 'hi' have the value now

The header binds locals in scope to registers (@asm(out, rax) lo), and the unsafe {} (which is already an expression-scope, section 5) carries the instructions. An operand is not restricted to a bare local: it may be a place-expression — a struct field (@asm(out, rax) result.lo) or an array/slice element (@asm(in, rax) a[i]) — and the compiler generates the copy automatically (an in place is loaded into its register before the block; an out place is written back from its register after). This lets the asm read from and write straight into an aggregate, for instance filling a result struct’s fields (@asm(out, rax) r.lo and @asm(out, rdx) r.hi) with no hand-written temporaries. There are three differences relative to the function, each for a reason:

  1. Binding by header, not by param: the block has no signature where to put the binding, so the @asm(in, …)/@asm(out, …) go above, binding locals.
  2. Explicit clobbers (@asm(clobber, …)), and here it is correctness, not verbosity: the block is in the middle of code that uses registers, so it has to declare what it dirties (registers, "memory"), otherwise the compiler assumes they are intact and generates wrong code. The function does not need it because the ABI covers it; the block needs it because it protects the surrounding code.
  3. It is a statement, not an expression: the block writes to the output locals (@asm(out, rax) lo writes to lo), it does not evaluate to a value. Whoever wants asm as a value uses the assembly-function; the block covers “asm in the middle of code”. (There is no expression-form: it would be an anonymous assembly-function, duplicating the function without being able to reuse nor name it, and the result := at the top, separated from the unsafe {} down below by three decorators, reads worse.)

Each binding names a specific register (rax) or the class xreg (the compiler chooses a free register, and optimizes):

@asm(x86, intel)
unsafe fn double(value: u32 @asm(in, xreg)) -> u32 @asm(out, xreg) {
add value, value // refers by NAME: you do not know which register it is
}

The difference in the body: with an explicit register you write the raw register (rax) or the operand’s name, and both resolve, because you know which it is. With xreg you have to use the symbolic name (value), because the register is chosen by the compiler and you have no way to name it. Explicit is for when the instruction requires a specific register (many syscalls); xreg is for when you just need “a register” and want to let the allocator optimize.

The dialect lives in the second argument of the @asm, and it is per architecture (each machine has its own syntax):

  • x86: intel or att (the same instruction, different spellings, both parsed by the backend).
  • ARM and RISC-V: the native dialects of each.
  • WASM: the two syntaxes of WebAssembly text (folded/S-expression and linear/stack), accepted in the same block without a toggle, because the folded form is just sugar for the linear one (the WAT parser expands one into the other).
@asm(x86, att) // same addition, AT&T dialect
unsafe fn add10(value: u32 @asm(in, rax)) -> u32 @asm(out, rbx) {
movl %eax, %ebx
addl $10, %ebx
}

WASM is a stack machine, it has no registers. So the whole model of binding by register (@asm(in, rax)) does not apply: there is no rax. But that is less work, not more, because WASM is already higher-level, with natively typed params and locals. A WASM assembly-function binds by its own params (no @asm(in/out)); a WASM block binds by locals:

@asm(wasm, wasm) // no @asm(in/out): WASM already has typed params/result
unsafe fn add(a: u32, b: u32) -> u32 {
local.get a
local.get b
i32.add
}

The asymmetry is part of the design: binding by register on x86, ARM and RISC-V, and native binding by param or local on WASM. The @asm has operand semantics dependent on the architecture: on register machines you bind registers, on the stack machine the params are already the operands.

Since asm is architecture-specific, you write one version per architecture you support and the compiler chooses the target’s, via the arch package of the stdlib (comptime constants: .x86_64, .arm, .riscv, .wasm and the like) with the comptime match/comptime if that already exists:

use arch
fn timestamp() -> u64 {
comptime if arch.current == .x86_64 {
return unsafe rdtsc_x86() assume "rdtsc available on this target"
}
return portable_clock() // other target: normal path, no asm
}

The comptime prefix is what makes the selection resolve at compile time: the whole if is evaluated by the compiler, one branch is chosen, and ONLY that branch is emitted (dead-branch elimination). This is the same prefix comptime match and comptime loop use, and it is what a plain runtime if could not do – a runtime if would still emit both branches, so an untaken branch’s foreign-architecture @asm would reach codegen. The selection runs at comptime (the arch.current is a compile-time constant), so the target’s binary contains only the path of its architecture. And an architecture mismatch is a compile error: compiling an @asm(x86, ...) block for an ARM or WASM target does not compile. You have to provide the block of the target architecture, with no magic fallback (it is the nature of asm; a mov rax does not exist on ARM). It is the same rigidity as Rust and Zig: without the arch’s asm, it does not compile.

Beyond the inline, assembly can also live in its own files (.s / .asm / .extasm), for large pieces (a whole crypto routine, an extensive hot path) that would pollute the code if inline. The files follow the same dialects by architecture, and are linked by the build like the FFI’s .c.

The assembly is emitted by LLVM for now, a pragmatic choice that enables the multiple targets (including WASM) and the cross-compile without rewriting an assembler per architecture. An own backend is in the plans (the path Zig walked: LLVM first, native backends later), but it is future work.

And what does not enter is an own normalized assembly syntax (what Go did with Plan9, a single cross-arch grammar). With support for the N real dialects (Intel, AT&T, ARM, RISC-V, WASM), inventing one more unified syntax is work with no return: you write in the native dialect of each architecture, which is what each ISA’s tools and documentation already use.