Skip to content

Specification §17

FFI with C

language-design.md §17 · 166 lines · 15 min read

@cimport("header.h") brings C inside. It runs at comptime, invokes a C compiler embedded in the toolchain, translates the whole header (functions, structs, types, macros) and returns a module. You re-declare no prototype: you point to the .h and use it. It is Zig’s approach (@cImport), for the same reason. It opens the whole C ecosystem “for free”: SDL2, Vulkan, SQLite and libcurl just start working, because the compiler reads the official header instead of you transcribing signatures (and getting it wrong, turning it into UB). The cost is a C compiler inside the toolchain, real weight, paid by the ergonomics.

Two caveats. The first: @cimport is not user comptime (section 10), it is a compiler phase, which reads the .h and runs the embedded C. The “pure comptime, no IO” holds for your code, not for this phase. The second: since the C compiler and libc/musl are shipped and located with the toolchain (our compiler is also a C compiler), the build is deterministic given the pinned toolchain, without depending on the C-compiler nor the system headers. A @cimport of a system header, that one does tie the build to the environment, and it is recorded as a dependency (as the mk.sum locks deps, section 15).

C has bitfields (unsigned x : 3;), but the rule “sub-byte width at the C boundary is an error” (section 14) holds here too: a C struct with a bitfield, imported, becomes an opaque type. You hold the pointer and pass it back to C, but you do not read nor write the fields through the language (it is what Zig’s translate-c does, demote-to-opaque). The “for free” promise has that caveat: structs with bitfields come as an opaque handle, not field by field. (bindgen-style generated accessors, with getters and setters that mask the bits, stay as a future extension if there is demand; for now, opaque.)

@cimport has full macro support — object-like #defines become constants, and function-like macros are translated too. This covers the case of errno, which in modern libc is not a plain global but a macro (#define errno (*__errno_location())): @cimport recognizes it and exposes c.<mod>.errno as a readable lvalue (reading it evaluates *__errno_location()), so the ubiquitous C error-checking idiom (if c.errno.errno == c.errno.EINTR { … }) works without you reaching for __errno_location() by hand.

Each @cimport creates a submodule under the root-namespace c, the root of all FFI, which arises with the first @cimport. The submodule’s name is the header’s basename (no path, no extension), automatic, without you declaring it:

@cimport("stdio.h") // → c.stdio
@cimport("SDL2/SDL.h") // → c.SDL
@cimport("SDL2/SDL.h") as gfx // → c.gfx (renames the submodule; 'as' for collision or clarity)
unsafe c.stdio.printf(fmt, x) // all qualified: c.stdio.printf, c.SDL.SDL_Init, c.SDL.SDL_INIT_VIDEO...

Everything from C stays under c.: functions, types, constants, structs, with no exception. No FFI in your global scope (the pollution of #include). The c. on each use marks “this is C”, the boundary stays visible at every point of contact, and your namespace stays clean (FFI does not pollute whoever does not do FFI, the opt-in test of section 1). The as is the same rule as use ... as (section 14), for a basename collision (two foo.h in different paths) or clarity. And since @cimport is resolved at compile time, it falls into the comptime and modules that already exist (section 15): it is the conditional use, applied to C.

Types at the boundary: all qualified, including the fundamental ones

Section titled “Types at the boundary: all qualified, including the fundamental ones”

C’s types are not your language’s types, they are members of the C module, just as c.stdio.printf is a function of it. There is no loose c_int/c_char; everything comes qualified from a @cimport. This holds even for the fundamental types (int, char, size_t), for a reason that is the premise of FFI with C: those types depend on the platform. There is no “C’s int”, there is “the int under a given ABI”. And the axis that actually decides the width is the psABI (platform + architecture), not the compiler: on one ABI gcc and clang agree (int==32, long==64 on System V AMD64), and it is the ABI that differs — long is 64 on System V but 32 on Windows x64 (LLP64), plain char is signed on x86 but unsigned on ARM/RISC-V, and long double is x87 80-bit on System V yet IEEE quad-128 on AArch64. So the canonical spelling names the ABI:

@cimport("systemv") // → c.systemv (also: win64, aarch64_aapcs, riscv64, wasm32)
size: c.systemv.size_t = n // the size_t under the System V AMD64 psABI: correct size/ABI
count: c.systemv.int // 32-bit; c.win64.long would be 32, c.systemv.long is 64

The compiler names remain as aliases: c.gcc.int / c.clang.int resolve to the ABI of the target that compiler is configured for, so c.gcc.long is 64 when compiling for Linux and 32 when compiling for Windows x64 — the compiler is the agent that carries the target’s ABI, and naming it is shorthand for “the int as gcc builds it for this target”. (As a side effect, cross-compilation becomes swapping the target, and the sizes follow, something a fixed c.int could not express.) It is the single rule, without exception: every C symbol, including the base types, comes from a @cimport and lives under c.<source>., where <source> is an ABI name (target-independent) or a compiler/runtime/header name (bound to the compile target).

Pointers and null. C’s T* becomes *c.<...>.T; and since C has NULL and you have Optional (section 6), a C pointer that can be null arrives as Optional[*c.gcc.char]: C’s null is the Optional’s None, matched normally. You do not deref an Optional without handling it, so C’s null possibility becomes explicit handling on your side.

Structs and layout. Your structs are optimized by default: the compiler reorders fields to minimize padding. This is the opposite of the C layout (fields in the order written), so passing a struct to C requires @repr(c), which turns off the reordering and gives a layout faithful to C:

@repr(c)
decl Point { x: c.gcc.int; y: c.gcc.int } // faithful order, C padding: safe to cross the boundary

The symmetry holds on the other side: structs coming from C (via @cimport) already arrive with C layout by origin, they are @repr(c) automatically, you mark nothing. Only your structs, going to C, need the marker. (Since @repr(c) turns off the optimization, the default without a marker remains the fast common case; the marker appears only in the minority that crosses the boundary.)

Enums at the boundary. A C enum is an int, so a C enum imported through @cimport is int-compatible at the boundary: its constants take part in bitwise |/& and pass where a C int is expected, with no extra marker — this is what makes the ubiquitous flag idiom (open(path, c.fcntl.O_RDONLY | c.fcntl.O_WRONLY)) just work. Going the other way, exporting one of your enums to C (@extern) depends on its shape: a payloadless enum (decl Color { Red; Green; Blue }) crosses as its integer discriminant (0/1/2); a payloaded one (decl Shape { Circle(float); … }) has no plain-int form and must be given an explicit @repr(c) tagged layout (a struct of tag + union) or it is rejected. A native Makoto enum stays a strict ADT — it is exactly one variant, and bitmasking would break that invariant — so A | B is not valid on a bare enum. When you genuinely want a native flag set, mark it opt-in with @flags (@flags decl Perms { Read; Write; Exec }): that makes it int-backed, so Read | Write is a valid value of the type (the bitflags/Zig model). Only C enums and @flags enums are bitmaskable; a plain enum never is.

Unions at the boundary. A raw union (section 5) is already a pure-overlap layout identical to C’s, so it crosses freely; the explicit @repr(c) union { … } spelling is accepted as the marker that pins that C-ABI overlap (a no-op today, since native and C union layout coincide, but the honest future-proof marker should the native union ever grow a hidden tag or alignment — exactly parallel to @repr(c) on a struct).

The other marker is @repr(packed), the Zig-style bit-packed layout (section 14): the struct becomes a single host integer and each field takes its exact bit width, so a u3 and a u5 share one byte. It is for bit flags and wire/hardware records, not the C boundary (a bit-packed field has no C ABI, and no *T, since it does not sit on a byte). The layout of any type is observable at compile time through mem.size_of[T](), mem.align_of[T](), and mem.offset_of[T]("field") (the offset follows the reordering). These are ordinary generic functions in the mem package written in Makoto over compiler.reflect (not compiler builtins): size_of/align_of fold to a constant yet stay callable at runtime, while offset_of is a comptime fn (bind its result to a const), and offset_of of a sub-byte or packed field is a compile error, because such a field has no byte offset.

C does not know your @mm. There are three situations, all reusing section 5.

Memory that C gives you arrives marked @mm(c), a MemoryManager (section 5) whose implementation points to C’s free (the deallocator), not yours. This carries the information that @mm(none) would lose: C memory is not freed with your free, it is freed with C’s, and the @mm(c) knows it. To bring the memory into your world (stop depending on C), use @promote, which already means “copy to another MM”:

data: *Record @mm(c) = c.sqlite3.get_record(...) // C memory, under @mm(c)
mine := data @promote(gc) // repatriate: copies to my GC, decouples me from C

@promote from @mm(c) to your MM is literally “take it from C, bring it to me”. The @mm(c) is just another MemoryManager, so @transfer and @promote handle it with nothing new.

Your memory that goes to C: if the object is non-movable (@mm(arena)/none/c), it crosses directly (it is already stable). But a @mm(gc) object is the dangerous case, because the collector can move or free the object while C holds the pointer, and C reads garbage. The rule is hard and single: passing a @mm(gc) object to C requires @pin, and without @pin it is a compile error. The @pin fixes the object (the GC does not move nor collect while the pin lives); @unpin releases:

@pin handle // fixes: the GC does not touch it while pinned
unsafe c.SDL.SDL_AddEventWatch(cb, handle) // C can hold the pointer safely
// ... later, when C is done:
@unpin handle // releases: on YOUR account

Releasing the pin is your responsibility, and forgetting is a debt, like leaking @mm(none) memory. The compiler forces you to put the @pin (otherwise it does not compile), but it does not hold your hand to release it: it is unsafe territory, and you signed up when you called C. A non-movable object does without the @pin.

Every call to C is a black box that can violate any invariant of yours, so calling C is an unsafe operation (section 5). A C function is an unsafe function of ours: you kill it with assert or assume on the operation, or group it in an unsafe block:

n := unsafe c.unistd.read(fd, buf, len) assume "fd valid and buf holds len bytes"
unsafe {
c.SDL.SDL_Init(c.SDL.SDL_INIT_VIDEO)
win := c.SDL.SDL_CreateWindow(...)
}

And the C return is handled at the boundary, but without magic conversions. A pointer that C can return as NULL arrives as an Optional (section 6), matched normally — that mapping is structural (a nullable pointer is an Optional). A const char* (a C string), however, arrives raw, as a *c.<...>.char: it is bytes, with no Display and no string invariant, not a string. Turning it into a string is an explicit call, from_cstring(p), which validates UTF-8 and returns a Result[string, error] (the inverse of .to_cstring(), section 16). The boundary does not silently turn C bytes into a string — you ask for the conversion, and you handle the Result. After that validation, you operate with the normal guarantees.

Variadics (c.stdio.printf(fmt, ...)) are transparent: the @cimport saw the ... in the prototype, so c.stdio.printf accepts N arguments and you just pass them, without annotating variadicity (it came from the header). But C’s variadic does not check types (getting it wrong is UB) and the signature does not say the expected types, so the types of the arguments are your responsibility. It is a classic case-B of unsafe (section 5): the judgment “the arguments match the format string” is human and inexpressible, killed by the assume or by the unsafe block:

unsafe c.stdio.printf("%d items, %s", count, name) // YOU guarantee that the types match the format

Callbacks: @callback, @extern and spawn over the bridge

Section titled “Callbacks: @callback, @extern and spawn over the bridge”

There are two senses of “C calls back”, and two markers:

  • @callback marks a function of yours to be passed as a callback to a C function (c.stdlib.qsort(...)). It gains C ABI so C can call it.
  • @extern exposes a function of yours with C ABI to be a lib that C consumes (the sense “being called by C”).
@callback fn compare(a: *c.gcc.void, b: *c.gcc.void) -> c.gcc.int {
// void* → typed pointer: the dedicated unsafe primitive (both type-args explicit)
ia := unsafe mem.reinterpret[c.gcc.void, i32](a) // *void → *i32
ib := unsafe mem.reinterpret[c.gcc.void, i32](b)
return (*ia - *ib) as c.gcc.int
}
unsafe c.stdlib.qsort(arr, n, sz, compare)
@extern fn my_library_init() -> c.gcc.int { ... } // external C can link and call this

Reinterpreting a pointer at the boundary: mem.reinterpret. A C comparator receives void* and must read the real type behind it — the ubiquitous FFI move (qsort, bsearch, pthread userdata, generic containers). This is a pointer reinterpretation, and it is neither as (which is value conversion only — width/repr, never pointers) nor a union (which is byte type-punning of values, section 5). It has its own dedicated, unsafe-gated primitive: unsafe fn mem.reinterpret[T, V](p: *T) -> *V, taking a *T and returning a *V. It is declared unsafe, so every call sits in an unsafe context, exactly like a C call; and it is named (greppable — every pointer reinterpretation in the program is a reinterpret, never a silent cast). Both type-args are explicit at the call site ([c.gcc.void, i32], source-first: “from void to i32”): the source T could be inferred from the argument, but the language does not allow partial type-argument application (all inferred, or all explicit; section 14) — Makoto always answers “where did this type come from”.

A @callback/@extern function may return an aggregate by value (a @repr(c) struct returned directly, not only through an out-pointer): it rides the same C-ABI return machinery the inbound direction already uses — small aggregates in registers, larger ones via the ABI’s sret hidden pointer — so the two directions are symmetric (a C function imported through @cimport can already return a struct by value, and yours exported to C can too).

The thorny case is inside a callback. C calls your callback from a raw C thread, without your scheduler, your isolation nor your process runtime (section 8). If the callback does spawn, it cannot create a process there. The solution is the runtime as the middle-man: the spawn in the callback does not create the thread locally, it becomes a request to your runtime (“create the process worker(x) in the right world and give me back the channel”), and the runtime, on its side, creates the process in the right place (your scheduler, your isolation) and returns the channel.

The same principle closes the @callback contract: the body does not suspend and does not allocate. Local heap allocation and io.* in there are compile errors, because there is no @mm nor process on the raw C thread to sustain them (the @mm comes from the process, section 5; the io.* suspension needs a process to suspend, section 12). What is allowed is talking to the runtime and the processes: sending on a channel, and the spawn-over-the-bridge above. Even if the message allocates, whoever allocates is the runtime on the other side, not the callback. The line is this: communication with the runtime and processes is an allowed bridge; local heap allocation or IO is forbidden. A @callback is a thin bridge; anything “real” (allocate, do IO, create a process) crosses through the middle-man.

The syntax is identical inside and outside: you write spawn worker(x) the same, and at the boundary the spawn is just routed by the runtime. And since that route can fail (runtime unavailable, resource exhausted), it uses the form the language already has, the catch of the spawn (section 3):

@callback fn on_event(ev: *c.SDL.SDL_Event) -> c.gcc.int {
spawn handle_event(ev) catch |e| { return c.SDL.SDL_FALSE } // spawn over the bridge; can fail
return c.SDL.SDL_TRUE
}

A spawn that crosses the boundary is just a spawn with one more way to fail, and the catch covers it, wherever the failure comes from. The raw C thread never operates your scheduler; it requests, and the runtime executes. (It is BEAM’s model for NIFs and Go’s for cgo: foreign code does not become a process nor a goroutine, it communicates with one, and that is already your model of channels between isolated worlds, section 4.)

The compiler compiles the .c together (Zig style, and it is what makes sqlite.c “just work”: the .c enters the build, it is not a pre-compiled binary that you pray exists). How it discovers what to link follows the reproducibility that governed strings and modules: explicit declaration, not auto-discover.

Auto-discover (the compiler scans the @cimport and guesses the -lSDL2 of the environment) is magic when it works and breaks reproducibility when it does not. It runs on your machine, fails on CI with another version of the lib, and there is nowhere to record which version. It is the problem that mk.sum solved. So C deps are declared in cdeps in the mk.project (section 15), like the Git deps: name, version, checksum:

mk.project
cdeps {
SDL2 2.30.1 sha256:...
sqlite3 3.45.0 sha256:...
}

The @cimport("SDL2/SDL.h") says which header; the cdeps says which SDL2, from where, with what hash, the same separation “code points, manifest governs version and proof” as the rest. The build becomes reproducible. The exception is the omnipresent system libs (libc and company): always present, not versioned by your project, do without declaration. The rest (SDL2, SQLite, Vulkan) is declared.

The FFI lets a Makoto program call C; the build interop lets a Makoto module drop into a C/C++ codebase without that codebase rewriting its build. This is the adoption path: a team with an existing CMake or Ninja tree wants one module in Makoto, incrementally, not a rewrite. It goes both ways, and neither way invents a second build system — each side keeps its own and calls the other at a thin boundary.

A Makoto build calling cmake/ninja is just the task runner (section 22). A task declares the subprocess it runs, so building a C dependency that ships its own CMake, or invoking a project’s ninja, is a [tasks.<name>] with run = cmake --build build and cmake (or ninja) in its subprocess list. The declared-subprocess gate already governs it; there is nothing new to learn, and the hermetic-by-convention rule (section 22) still holds — an undeclared tool is refused.

A cmake/ninja build calling Makoto is mk export. Since mk build --emit=lib already produces a C-callable static library plus a generated header (section 17, the two-way boundary), the export is thin glue that shells out to it at build time:

  • mk export cmake writes a MakotoConfig.cmake — a reusable module a C/C++ project include()s (or find_package(Makoto)s), giving it makoto_add_library(<target> <entry.mko>) and makoto_add_executable(...). The library builds through mk during the cmake build and links like any other IMPORTED target: target_link_libraries(theirapp PRIVATE mymod). It accepts the build’s own knobs (RELEASE, TARGET <arch>, LIBC <name>).
  • mk export ninja writes a ninja file with a makoto_lib rule and a build edge for the project’s entry lib<name>.a + header, run standalone (ninja -f makoto.ninja) or subninja’d into a larger graph.

The C side compiles with mk cc (the drop-in of section 17, so it links against the same hermetic toolchain the library was built with — set CMAKE_C_COMPILER to it), or, for a project that links with the host toolchain, the library is built LIBC system to match. Either way the discipline is the design’s throughout: Makoto owns the Makoto build (one compiler, reproducible), the foreign build owns its own, and the two meet at a declared, versionable boundary — the same mk.project/mk.mod that governs a pure-Makoto build, now reachable from the outside.