Skip to content

Specification §16

Strings

language-design.md §16 · 172 lines · 13 min read

The spelling string/string[u16] deceives on purpose, so it is worth saying exactly what string is in the type system, because it is none of the three things it appears to be:

  • It is not a scalar primitive (like int/u8): scalars fit in a register and are atomic; string is an aggregate, with a length, living in memory under @mm, with internal structure.
  • It is not an alias of []byte: it has an invariant that []byte does not have (being valid UTF-8) and per-unit access (.chars/.codepoints) that []byte does not have. That is why converting between the two is an operation (safe copy or unsafe reinterpretation), not a free cast.
  • It is not a generic like List[T], and here the spelling lies the most: there is no type string[u16]. string is always UTF-8, it is the only string type that exists. The “other encodings” are not instances of a generic; they are just []byte (bytes in that encoding). The [u16] does not parameterize the type, but rather the conversion methods (to/from).

Affirmatively: string is a builtin composite type (of the array and slice family, builtin, not defined by you), with an invariant (valid UTF-8) and unit-aware access. It stores UTF-8 bytes and presents three views: .bytes (for byte), .codepoints (for codepoint) and .chars (for grapheme). It is not []byte (it has the invariant), nor []char (it does not hold loose graphemes, but rather bytes it decodes on demand). It is its own builtin thing.

Everything follows from this. If you have a string, you have valid UTF-8 text, period; arbitrary or possibly-invalid bytes are []byte. The validation lives at the boundary (the conversion, where the error is a Result), and the core of the program never asks “is this UTF-8?”. And since string[u16] is not a type, pure string is the only one that circulates in every signature (fn log(msg: string)); the [E] appears only in the boundary conversion methods, never in an internal signature. The “Reading-1” is not a convention, it is structure (there is no string[E] to circulate because it does not exist). It is what Go, Rust and Swift do: a single string, UTF-8, and the rest is bytes plus conversion.

The three views correspond to three builtin types, the atoms of text:

  • byte (u8): the storage unit. UTF-8 uses 1 to 4 bytes per codepoint.
  • codepoint (u32): the Unicode scalar.
  • grapheme: the visual character (what the human sees). It can be several codepoints (👨‍👩‍👧 is a grapheme of seven), so it is variable-size, since it carries its bytes and does not fit in a register like byte/codepoint.

And char is an alias for grapheme (the familiar name of the common unit, since what people want most of the time is the visual character), not a generic char[E]. Want a codepoint? Write codepoint; there is no char[codepoint]. The same discipline of Reading-1 holds: char/grapheme is what circulates in the core, and raw codepoint/byte appear when you descend to the scalar (parsing, interop). The symmetry closes the system: string is to char as []byte is to byte.

A literal 'w' is a comptime character with no fixed type: just as 5 is a comptime integer that becomes u8/i64/int according to context, 'w' resolves to byte/codepoint/grapheme by context:

k: codepoint = 'w' // context defines → codepoint (U+0077)
match key { 'w' => move(); _ => {} } // 'w' is a codepoint; '_' = default (codepoint is an open domain)
b: byte = 'w' // → byte (0x77; valid, 'w' fits in 1 byte)

Where context does not disambiguate (loose x := 'w'), you qualify the unit ('w'.codepoint, 'w'.byte, 'w'.grapheme), with validity checking ('é'.byte is an error: ‘é’ is 2 bytes; '👨‍👩‍👧'.codepoint is an error: it is several). It is the same discipline as string access (explicit unit when ambiguous), applied to the literal. char is an alias of grapheme, so 'w'.char equals 'w'.grapheme.

There is no pure s[i] nor s[2..6]/s[2:6]: they are compile errors. The reason is what bans s[i] in every language that tries it: s[2..6] is “over WHAT?”, byte, codepoint or grapheme? The ambiguity does not resolve by a default (byte is the natural one for whoever implements and the most treacherous for whoever uses: s[2..6] expecting 4 letters and receiving 4 bytes cuts an emoji in half). The unit is always qualified.

Each access has a method and a shorthand, paired for singular and range, with no orphan case:

Shorthand Method What
s.bytes[2] s.byte_at(2) one, by byte
s.bytes[2..6] s.bytes_view(2, 6) view (range), by byte
s.bytes[2:6] s.bytes_copy(2, 6) copy (range), by byte
s.chars[2] s.char_at(2) one grapheme
s.chars[2..6] s.chars_view(2, 6) view (range), by grapheme
s.chars[2:6] s.chars_copy(2, 6) copy (range), by grapheme
s.codepoints[2] s.codepoint_at(2) one codepoint
s.codepoints[2..6] s.codepoints_view(2, 6) view (range)
s.codepoints[2:6] s.codepoints_copy(2, 6) copy (range)

Singular uses the unit in the singular (byte_at); range uses the plural (bytes_view), mirroring the shorthand. The _view/_copy carries the cost in the name (O(1) vs O(n)), just as ../: carries it in the shorthand. Whoever comes from Go or Python tries s[2..6] and takes a compile error, the same friction-that-teaches as the never-mutated var, and it saves them from the real bug (cutting a grapheme by accident).

Views and copies: one primitive, explicit cost

Section titled “Views and copies: one primitive, explicit cost”

Slicing a sequence has two costs, and the language makes them two explicit syntaxes, instead of hiding one:

nums[2..6] // VIEW: ptr+len into nums; O(1); keeps nums alive (let by default)
nums[2:6] // COPY: new slice, independent; O(n); allocates

.. is view, : is copy, and it is not arbitrary: .. is already the range operator (loop i in 0..8), and a range is a light description of indices, which is what a view is. A familiar and light symbol for the light operation; : (new in that context) for the one that allocates. It is the same primitive for array, slice and string. Array and slice do not need a qualifier (there is only “element”, so arr[2..6] is unambiguous); a string is the only type where the unit is ambiguous and the only one where the qualifier (.bytes/.chars/.codepoints) is mandatory, by the rule “require the unit when it is ambiguous”. The view keeps the original alive while it exists, but that is the same scope rule as @mm (“do not free what still has a live reference”), no new mechanism.

The UTF-8 bytes come out of the .bytes view itself: they are already UTF-8, it is just a view or a copy, with no conversion. Only other encodings go through to[E]/from[E], and always explicitly (transcoding UTF-8UTF-16 allocates and traverses the whole string; hiding that behind an argument passing would contradict the language’s explicitness):

s.bytes // UTF-8: view of the bytes (no conversion)
s.bytes_copy(0, n) // UTF-8: copy of the bytes
s.to[u16]() -> []byte // other encoding: transcodes (allocates)
f(s.to[u32]()) // passing another encoding is explicit, never f(s)

To build a string from bytes, there is the safe form (validates, can fail) and the unsafe form (reinterprets without a copy, killed with assume):

string.from[u8](bytes) -> Result[string, error] // UTF-8 bytes → string (safe: validates)
string.from[u16](bytes) -> Result[string, error] // other encoding → string (transcodes+validates)
bytes.as_string_unchecked() // UTF-8 bytes → string (unsafe, no copy):
// kills with assume "these bytes are valid UTF-8"; the _unchecked shouts the danger

The unsafe form reuses the unsafe/assume of section 5: the bytes may not be UTF-8, that is, case-B, killed with assume "reason". The set of boundary encodings: ASCII, Latin-1, UTF-8, UTF-16, UTF-32, raw bytes.

The UTF-16/UTF-32 bytes are little-endian, no BOM: that is Makoto’s canonical order (the native one of x86/ARM), so to[u16]/from[u16] always speak little-endian and stay spec-clean. Talking to a system that wants the other byte order or a byte-order mark is the unicode package’s job (opt-in, use unicode): unicode.swap16(b)/unicode.swap32(b) flip the byte order (little to big and back, the same call both ways), unicode.with_bom16(b, .little)/unicode.with_bom32(b, .big) prefix a BOM for a byte order, and unicode.strip_bom16(b)/unicode.strip_bom32(b) remove a leading BOM and hand back canonical little-endian bytes, ready to feed straight into from[u16]. The order is Endianness (.little/.big). The core carries one order; the byte-order zoo lives in the opt-in package, so a program that only ever stays inside Makoto never meets it.

Comparison: byte operators, semantics by method

Section titled “Comparison: byte operators, semantics by method”

A string has the comparison operators (==, !=, <, >, <=, >=), and all of them operate byte-by-byte: ==/!= is byte equality (the common case: protocols, paths, identifiers), and </> are lexicographic byte order (what Go, Zig and Rust do, and what a sort expects). This is what separates comparison from indexing: the unit of s[i] is ambiguous (byte? codepoint? grapheme?), but byte equality and order are not: they are well-defined and are what almost all code expects. What does not exist is semantic comparison by operator; Unicode normalization and localized collation are a rare case and become an explicit method:

a == b // byte equality (the common one, fast)
a < b // lexicographic byte order, feeds a sort directly
a.equals(b, .semantic) // equality with Unicode normalization (composed "é" == precomposed "é")
locale.Collator.new(.sv).order(a, b) // localized alphabetical order (collation), 'locale' subpackage

The byte-vs-semantic choice is a method (.equals(_, .semantic), the Collator), not a second operator, which kills === before it is born (JS’s === only exists because JS spent == on coercion; here == is already unambiguous byte equality, with no second meaning to undo). Alphabetical and localized comparison (collation) lives in the locale subpackage, pure opt-in: whoever never compares accented text by alphabetical order never imports locale, and the weight of the Unicode tables stays out of the program that does not need it. The s[i] remains banned (the unit is ambiguous); the comparison operator is allowed because byte-equality and byte-order do not have that ambiguity.

Collation needs a locale (sorting in Swedish is not sorting in German or English), so it is a Collator built for one and then reused: locale.Collator.new(.sv) from the common-locale enum Locale (.pt, .en_us, .sv, …), or locale.Collator.from_tag("pt-BR") for any BCP 47 tag outside the enum (an unparseable tag is an Err). The collator answers order(a, b) -> Ordering and equals(a, b) -> bool. This makes the locale explicit rather than a hidden global: the locale is the collator you name, not ambient state the comparison silently reads.

The .semantic there is a leading-dot enum literal: it names a variant (semantic) whose enum (a CompareMode, with .byte the default) the compiler infers from the expected type, the Zig-style inferred variant. It is general to any enum, not special to the compare mode: where the expected type is a known enum, .Variant is that enum’s variant (paint(.Red), let c: Color = .Blue, return .Idle), so you do not repeat the enum name the signature already fixed. Where no enum is in reach (a bare binding let x := .Red, an open error that declares no members), it does not infer and you name it in full (Color.Red, error.NotFound, CompareMode.semantic). An un-inferrable leading dot is a compile error, never a silent guess.

Mutating a string is the usual mutability rule (section 6): var mutates, :=/let does not. There is no “mutable string” as a type, but rather a var binding:

var s := "abc"
s.append("d") // mutates in-place; follows the units (append by char/byte/codepoint)

Growing (append in a loop) reallocates by doubling: on reaching capacity, the buffer doubles, amortizing the cost. And the compiler optimizes the known-size case: if it proves how much the string will grow, it pre-allocates the final size and skips the reallocations. There is no separate StringBuilder: a single type (var string with doubling) delivers what the builder would deliver, without the second type everyone would have to learn (the opt-in test).

Every literal "abc" is comptime-in-source (known at compile time), but what it becomes depends on the binding, and this follows the rule of all types (section 6 separates “immutable” from “comptime”):

x := "abc" // RUNTIME immutable string (normal value under @mm)
var x := "abc" // runtime mutable string
const x := "abc" // comptime string: embedded in the binary, immutable

It is the const that makes the string comptime (in the binary), not the :=, just as 5 becomes a comptime int in a const and a runtime int in a :=. Strings are no exception: the literal is comptime at the source, and the binding decides whether it stays comptime or is materialized at runtime.

Interpolation is named, with {{ }}:

greeting := "Hello {{name}}, you are {{age}} years old"

Inside {{ }} any expression that produces a value is valid: a name, field access, function call, arithmetic:

"{{user.name.trim()}} placed {{count}} orders, total {{total.fixed(2)}}"

An unbalanced hole is a Syntax Error: a {{ with no matching }} (a mistyped "x {{ a }" with a single closing brace, or a hole left open) is rejected at its position, not silently reinterpreted as literal text. A literal brace is \{{ (above); an interpolation hole must be closed.

A string literal inside a hole needs no escaping: the hole delimiter {{ }} is distinct from the string quote ", so a nested "..." reads cleanly, "cmp: {{ "a".cmp("b") }}". This follows the modern interpolation norm (Swift, Kotlin, Ruby, C#, and Python 3.12 via PEP 701 all allow it). The escaped form still works for anyone who prefers it (\"a\"), and a nested string may even contain }}. Quoting stays simple: " is only for strings, ' is only for char/rune/grapheme, and there are no backticks.

The value becomes text via the Display interface (below): {{x}} is sugar for x.to_string(). Interpolating a type without Display is a compile error, with no leaked <Object@0x7f...>; you implement Display and then the type can be interpolated. With comptime, {{name}} is checked at compile time (does the name exist? does the type format?), becoming a compile error instead of a runtime surprise.

Formatting is by method, not by a mini-language: since {{ }} already accepts any expression, decimal places, padding and hex are calls ({{pi.fixed(2)}}, {{x.pad(8)}}, {{n.hex()}}), instead of a second grammar ({:.2}) that everyone would have to learn. Zero new grammar, discoverable by autocomplete. A literal {{ escapes with a backslash (\{{), reusing the escape mechanism the string already has (\n, \t, \r, \", \\, \0). A Unicode code point is written directly as UTF-8 source ("é"), or, when the source must stay ASCII or name a non-printing point, with \u{HEX} ("caf\u{e9}" is café, '\u{1F600}' is 😀); the value must be a Unicode scalar (0…10FFFF, no surrogates) or it is a compile error. Two sibling forms name a raw byte/scalar in another base with the same brace syntax: \x{HH} reads it in hex and \o{OOO} in octal (so "\x{41}" is A, "\o{101}" is A); like \u{}, an out-of-range value is a compile error. A genuinely unknown escape is a compile error (\q does not exist): the lexer is not lenient — only the meaningful escapes above are special, and anything else is rejected at its position rather than silently dropping the backslash. This keeps typos from passing (\x41 written without braces, a mistyped \z) instead of being swallowed.

Display is the interface for “this type knows how to represent itself as text”, with one method, to_string():

decl Display {
fn to_string() -> string
}

It is the same interface used in any generic function over Display and in interpolation. Primitive types (int, bool, float…) implement it out of the box; for your types, you implement to_string() and gain both interpolation and any generic function over Display.

A char/grapheme/codepoint interpolates as its glyph — the visual character — not as its numeric codepoint: {{ 'A' }} yields A, {{ 'é' }} yields é (never 65/233). The glyph is the ergonomic default (it is what interpolation is for). But because a character value is also, underneath, a number, the compiler emits a warning suggesting you make the intent explicit when it matters — leave it bare (or .char/.grapheme) for the visual character, or descend to .codepoint/.byte and format that when you want the number — so a reader is never left guessing which you meant. It is a warning, not an error: the default already does the useful thing; the numeric reading is one qualifier away ({{ c.codepoint }} displays 233, since a codepoint is a u32).

C uses null-terminated strings (char*); a string is length-prefixed. The bridge is a method, .to_cstring(), not a new type: a CString is just “these bytes plus a \0”, and its problem (who frees it? C or you?) is ownership, which @mm/@transfer already governs; a CString type would reopen that resolved question. s.to_cstring() returns null-terminated bytes under the current @mm, and @transfer passes ownership to C when appropriate. The inverse direction — a C string coming in — is from_cstring(p): it takes a char* (raw bytes from C), validates UTF-8, and returns a Result[string, error]; the boundary performs no silent bytestring conversion (section 17). The detail is deferred with the FFI (section 17); here only the direction.