Strings
What string is
Section titled “What string is”The spelling string/string[u16] deceives on purpose, so it is worth saying exactly what string is in the type system, because it is none of the three things it appears to be:
- It is not a scalar primitive (like
int/u8): scalars fit in a register and are atomic;stringis an aggregate, with a length, living in memory under@mm, with internal structure. - It is not an alias of
[]byte: it has an invariant that[]bytedoes not have (being valid UTF-8) and per-unit access (.chars/.codepoints) that[]bytedoes not have. That is why converting between the two is an operation (safe copy or unsafe reinterpretation), not a free cast. - It is not a generic like
List[T], and here the spelling lies the most: there is no typestring[u16].stringis always UTF-8, it is the only string type that exists. The “other encodings” are not instances of a generic; they are just[]byte(bytes in that encoding). The[u16]does not parameterize the type, but rather the conversion methods (to/from).
Affirmatively: string is a builtin composite type (of the array and slice family, builtin, not defined by you), with an invariant (valid UTF-8) and unit-aware access. It stores UTF-8 bytes and presents three views: .bytes (for byte), .codepoints (for codepoint) and .chars (for grapheme). It is not []byte (it has the invariant), nor []char (it does not hold loose graphemes, but rather bytes it decodes on demand). It is its own builtin thing.
Everything follows from this. If you have a string, you have valid UTF-8 text, period; arbitrary or possibly-invalid bytes are []byte. The validation lives at the boundary (the conversion, where the error is a Result), and the core of the program never asks “is this UTF-8?”. And since string[u16] is not a type, pure string is the only one that circulates in every signature (fn log(msg: string)); the [E] appears only in the boundary conversion methods, never in an internal signature. The “Reading-1” is not a convention, it is structure (there is no string[E] to circulate because it does not exist). It is what Go, Rust and Swift do: a single string, UTF-8, and the rest is bytes plus conversion.
The elementary types and char
Section titled “The elementary types and char”The three views correspond to three builtin types, the atoms of text:
byte(u8): the storage unit. UTF-8 uses 1 to 4 bytes per codepoint.codepoint(u32): the Unicode scalar.grapheme: the visual character (what the human sees). It can be several codepoints (👨👩👧 is a grapheme of seven), so it is variable-size, since it carries its bytes and does not fit in a register likebyte/codepoint.
And char is an alias for grapheme (the familiar name of the common unit, since what people want most of the time is the visual character), not a generic char[E]. Want a codepoint? Write codepoint; there is no char[codepoint]. The same discipline of Reading-1 holds: char/grapheme is what circulates in the core, and raw codepoint/byte appear when you descend to the scalar (parsing, interop). The symmetry closes the system: string is to char as []byte is to byte.
Character literals
Section titled “Character literals”A literal 'w' is a comptime character with no fixed type: just as 5 is a comptime integer that becomes u8/i64/int according to context, 'w' resolves to byte/codepoint/grapheme by context:
k: codepoint = 'w' // context defines → codepoint (U+0077)match key { 'w' => move(); _ => {} } // 'w' is a codepoint; '_' = default (codepoint is an open domain)b: byte = 'w' // → byte (0x77; valid, 'w' fits in 1 byte)Where context does not disambiguate (loose x := 'w'), you qualify the unit ('w'.codepoint, 'w'.byte, 'w'.grapheme), with validity checking ('é'.byte is an error: ‘é’ is 2 bytes; '👨👩👧'.codepoint is an error: it is several). It is the same discipline as string access (explicit unit when ambiguous), applied to the literal. char is an alias of grapheme, so 'w'.char equals 'w'.grapheme.
Access: the unit is always explicit
Section titled “Access: the unit is always explicit”There is no pure s[i] nor s[2..6]/s[2:6]: they are compile errors. The reason is what bans s[i] in every language that tries it: s[2..6] is “over WHAT?”, byte, codepoint or grapheme? The ambiguity does not resolve by a default (byte is the natural one for whoever implements and the most treacherous for whoever uses: s[2..6] expecting 4 letters and receiving 4 bytes cuts an emoji in half). The unit is always qualified.
Each access has a method and a shorthand, paired for singular and range, with no orphan case:
| Shorthand | Method | What |
|---|---|---|
s.bytes[2] |
s.byte_at(2) |
one, by byte |
s.bytes[2..6] |
s.bytes_view(2, 6) |
view (range), by byte |
s.bytes[2:6] |
s.bytes_copy(2, 6) |
copy (range), by byte |
s.chars[2] |
s.char_at(2) |
one grapheme |
s.chars[2..6] |
s.chars_view(2, 6) |
view (range), by grapheme |
s.chars[2:6] |
s.chars_copy(2, 6) |
copy (range), by grapheme |
s.codepoints[2] |
s.codepoint_at(2) |
one codepoint |
s.codepoints[2..6] |
s.codepoints_view(2, 6) |
view (range) |
s.codepoints[2:6] |
s.codepoints_copy(2, 6) |
copy (range) |
Singular uses the unit in the singular (byte_at); range uses the plural (bytes_view), mirroring the shorthand. The _view/_copy carries the cost in the name (O(1) vs O(n)), just as ../: carries it in the shorthand. Whoever comes from Go or Python tries s[2..6] and takes a compile error, the same friction-that-teaches as the never-mutated var, and it saves them from the real bug (cutting a grapheme by accident).
Views and copies: one primitive, explicit cost
Section titled “Views and copies: one primitive, explicit cost”Slicing a sequence has two costs, and the language makes them two explicit syntaxes, instead of hiding one:
nums[2..6] // VIEW: ptr+len into nums; O(1); keeps nums alive (let by default)nums[2:6] // COPY: new slice, independent; O(n); allocates.. is view, : is copy, and it is not arbitrary: .. is already the range operator (loop i in 0..8), and a range is a light description of indices, which is what a view is. A familiar and light symbol for the light operation; : (new in that context) for the one that allocates. It is the same primitive for array, slice and string. Array and slice do not need a qualifier (there is only “element”, so arr[2..6] is unambiguous); a string is the only type where the unit is ambiguous and the only one where the qualifier (.bytes/.chars/.codepoints) is mandatory, by the rule “require the unit when it is ambiguous”. The view keeps the original alive while it exists, but that is the same scope rule as @mm (“do not free what still has a live reference”), no new mechanism.
Boundary conversions
Section titled “Boundary conversions”The UTF-8 bytes come out of the .bytes view itself: they are already UTF-8, it is just a view or a copy, with no conversion. Only other encodings go through to[E]/from[E], and always explicitly (transcoding UTF-8UTF-16 allocates and traverses the whole string; hiding that behind an argument passing would contradict the language’s explicitness):
s.bytes // UTF-8: view of the bytes (no conversion)s.bytes_copy(0, n) // UTF-8: copy of the bytess.to[u16]() -> []byte // other encoding: transcodes (allocates)f(s.to[u32]()) // passing another encoding is explicit, never f(s)To build a string from bytes, there is the safe form (validates, can fail) and the unsafe form (reinterprets without a copy, killed with assume):
string.from[u8](bytes) -> Result[string, error] // UTF-8 bytes → string (safe: validates)string.from[u16](bytes) -> Result[string, error] // other encoding → string (transcodes+validates)bytes.as_string_unchecked() // UTF-8 bytes → string (unsafe, no copy): // kills with assume "these bytes are valid UTF-8"; the _unchecked shouts the dangerThe unsafe form reuses the unsafe/assume of section 5: the bytes may not be UTF-8, that is, case-B, killed with assume "reason". The set of boundary encodings: ASCII, Latin-1, UTF-8, UTF-16, UTF-32, raw bytes.
The UTF-16/UTF-32 bytes are little-endian, no BOM: that is Makoto’s canonical order (the native one of x86/ARM), so to[u16]/from[u16] always speak little-endian and stay spec-clean. Talking to a system that wants the other byte order or a byte-order mark is the unicode package’s job (opt-in, use unicode): unicode.swap16(b)/unicode.swap32(b) flip the byte order (little to big and back, the same call both ways), unicode.with_bom16(b, .little)/unicode.with_bom32(b, .big) prefix a BOM for a byte order, and unicode.strip_bom16(b)/unicode.strip_bom32(b) remove a leading BOM and hand back canonical little-endian bytes, ready to feed straight into from[u16]. The order is Endianness (.little/.big). The core carries one order; the byte-order zoo lives in the opt-in package, so a program that only ever stays inside Makoto never meets it.
Comparison: byte operators, semantics by method
Section titled “Comparison: byte operators, semantics by method”A string has the comparison operators (==, !=, <, >, <=, >=), and all of them operate byte-by-byte: ==/!= is byte equality (the common case: protocols, paths, identifiers), and </> are lexicographic byte order (what Go, Zig and Rust do, and what a sort expects). This is what separates comparison from indexing: the unit of s[i] is ambiguous (byte? codepoint? grapheme?), but byte equality and order are not: they are well-defined and are what almost all code expects. What does not exist is semantic comparison by operator; Unicode normalization and localized collation are a rare case and become an explicit method:
a == b // byte equality (the common one, fast)a < b // lexicographic byte order, feeds a sort directlya.equals(b, .semantic) // equality with Unicode normalization (composed "é" == precomposed "é")locale.Collator.new(.sv).order(a, b) // localized alphabetical order (collation), 'locale' subpackageThe byte-vs-semantic choice is a method (.equals(_, .semantic), the Collator), not a second operator, which kills === before it is born (JS’s === only exists because JS spent == on coercion; here == is already unambiguous byte equality, with no second meaning to undo). Alphabetical and localized comparison (collation) lives in the locale subpackage, pure opt-in: whoever never compares accented text by alphabetical order never imports locale, and the weight of the Unicode tables stays out of the program that does not need it. The s[i] remains banned (the unit is ambiguous); the comparison operator is allowed because byte-equality and byte-order do not have that ambiguity.
Collation needs a locale (sorting in Swedish is not sorting in German or English), so it is a Collator built for one and then reused: locale.Collator.new(.sv) from the common-locale enum Locale (.pt, .en_us, .sv, …), or locale.Collator.from_tag("pt-BR") for any BCP 47 tag outside the enum (an unparseable tag is an Err). The collator answers order(a, b) -> Ordering and equals(a, b) -> bool. This makes the locale explicit rather than a hidden global: the locale is the collator you name, not ambient state the comparison silently reads.
The .semantic there is a leading-dot enum literal: it names a variant (semantic) whose enum (a CompareMode, with .byte the default) the compiler infers from the expected type, the Zig-style inferred variant. It is general to any enum, not special to the compare mode: where the expected type is a known enum, .Variant is that enum’s variant (paint(.Red), let c: Color = .Blue, return .Idle), so you do not repeat the enum name the signature already fixed. Where no enum is in reach (a bare binding let x := .Red, an open error that declares no members), it does not infer and you name it in full (Color.Red, error.NotFound, CompareMode.semantic). An un-inferrable leading dot is a compile error, never a silent guess.
Mutation and growth
Section titled “Mutation and growth”Mutating a string is the usual mutability rule (section 6): var mutates, :=/let does not. There is no “mutable string” as a type, but rather a var binding:
var s := "abc"s.append("d") // mutates in-place; follows the units (append by char/byte/codepoint)Growing (append in a loop) reallocates by doubling: on reaching capacity, the buffer doubles, amortizing the cost. And the compiler optimizes the known-size case: if it proves how much the string will grow, it pre-allocates the final size and skips the reallocations. There is no separate StringBuilder: a single type (var string with doubling) delivers what the builder would deliver, without the second type everyone would have to learn (the opt-in test).
Literals and comptime
Section titled “Literals and comptime”Every literal "abc" is comptime-in-source (known at compile time), but what it becomes depends on the binding, and this follows the rule of all types (section 6 separates “immutable” from “comptime”):
x := "abc" // RUNTIME immutable string (normal value under @mm)var x := "abc" // runtime mutable stringconst x := "abc" // comptime string: embedded in the binary, immutableIt is the const that makes the string comptime (in the binary), not the :=, just as 5 becomes a comptime int in a const and a runtime int in a :=. Strings are no exception: the literal is comptime at the source, and the binding decides whether it stays comptime or is materialized at runtime.
Interpolation
Section titled “Interpolation”Interpolation is named, with {{ }}:
greeting := "Hello {{name}}, you are {{age}} years old"Inside {{ }} any expression that produces a value is valid: a name, field access, function call, arithmetic:
"{{user.name.trim()}} placed {{count}} orders, total {{total.fixed(2)}}"An unbalanced hole is a Syntax Error: a {{ with no matching }} (a mistyped "x {{ a }" with a single closing brace, or a hole left open) is rejected at its position, not silently reinterpreted as literal text. A literal brace is \{{ (above); an interpolation hole must be closed.
A string literal inside a hole needs no escaping: the hole delimiter {{ }} is distinct from the string quote ", so a nested "..." reads cleanly, "cmp: {{ "a".cmp("b") }}". This follows the modern interpolation norm (Swift, Kotlin, Ruby, C#, and Python 3.12 via PEP 701 all allow it). The escaped form still works for anyone who prefers it (\"a\"), and a nested string may even contain }}. Quoting stays simple: " is only for strings, ' is only for char/rune/grapheme, and there are no backticks.
The value becomes text via the Display interface (below): {{x}} is sugar for x.to_string(). Interpolating a type without Display is a compile error, with no leaked <Object@0x7f...>; you implement Display and then the type can be interpolated. With comptime, {{name}} is checked at compile time (does the name exist? does the type format?), becoming a compile error instead of a runtime surprise.
Formatting is by method, not by a mini-language: since {{ }} already accepts any expression, decimal places, padding and hex are calls ({{pi.fixed(2)}}, {{x.pad(8)}}, {{n.hex()}}), instead of a second grammar ({:.2}) that everyone would have to learn. Zero new grammar, discoverable by autocomplete. A literal {{ escapes with a backslash (\{{), reusing the escape mechanism the string already has (\n, \t, \r, \", \\, \0). A Unicode code point is written directly as UTF-8 source ("é"), or, when the source must stay ASCII or name a non-printing point, with \u{HEX} ("caf\u{e9}" is café, '\u{1F600}' is 😀); the value must be a Unicode scalar (0…10FFFF, no surrogates) or it is a compile error. Two sibling forms name a raw byte/scalar in another base with the same brace syntax: \x{HH} reads it in hex and \o{OOO} in octal (so "\x{41}" is A, "\o{101}" is A); like \u{}, an out-of-range value is a compile error. A genuinely unknown escape is a compile error (\q does not exist): the lexer is not lenient — only the meaningful escapes above are special, and anything else is rejected at its position rather than silently dropping the backslash. This keeps typos from passing (\x41 written without braces, a mistyped \z) instead of being swallowed.
The Display interface
Section titled “The Display interface”Display is the interface for “this type knows how to represent itself as text”, with one method, to_string():
decl Display { fn to_string() -> string}It is the same interface used in any generic function over Display and in interpolation. Primitive types (int, bool, float…) implement it out of the box; for your types, you implement to_string() and gain both interpolation and any generic function over Display.
A char/grapheme/codepoint interpolates as its glyph — the visual character — not as its numeric codepoint: {{ 'A' }} yields A, {{ 'é' }} yields é (never 65/233). The glyph is the ergonomic default (it is what interpolation is for). But because a character value is also, underneath, a number, the compiler emits a warning suggesting you make the intent explicit when it matters — leave it bare (or .char/.grapheme) for the visual character, or descend to .codepoint/.byte and format that when you want the number — so a reader is never left guessing which you meant. It is a warning, not an error: the default already does the useful thing; the numeric reading is one qualifier away ({{ c.codepoint }} displays 233, since a codepoint is a u32).
FFI and CString
Section titled “FFI and CString”C uses null-terminated strings (char*); a string is length-prefixed. The bridge is a method, .to_cstring(), not a new type: a CString is just “these bytes plus a \0”, and its problem (who frees it? C or you?) is ownership, which @mm/@transfer already governs; a CString type would reopen that resolved question. s.to_cstring() returns null-terminated bytes under the current @mm, and @transfer passes ownership to C when appropriate. The inverse direction — a C string coming in — is from_cstring(p): it takes a char* (raw bytes from C), validates UTF-8, and returns a Result[string, error]; the boundary performs no silent bytestring conversion (section 17). The detail is deferred with the FFI (section 17); here only the direction.