Skip to content

Rationale · essay 09

Strings

rationale.md · 75 lines · 3 min read

The redefinition: the string is invariant UTF-8 with an explicit unit (s[i] is forbidden) and comparing is a method, because “equal” is ambiguous.

A string seems simple until you remember that UTF-8 has three distinct units (bytes, codepoints and graphemes, the visual characters) and that “are these two strings equal?” depends on normalization and locale. Most languages hide this: s[i] gives you something (byte? char?), and == does some comparison. The comfort is a surface lie.

A string is always valid UTF-8 and exposes three views: .bytes, .codepoints, .chars (grapheme). Raw s[i] is forbidden: you qualify the unit (s.bytes[i], s.chars[i]), because indexing without saying what is ambiguous, and the ambiguity hides the cost (a byte is O(1); a grapheme is O(n)).

Comparison has a deliberate cut: the operators ==/!=/</> exist, but they are byte-by-byte (byte equality, lexicographic order, the common and fast case, which feeds a direct sort). The Unicode semantics (normalization, collation, case-insensitive) is a method (a.equals(b, .semantic)), not an operator. And there is no ===: a second equality operator would only create the question “which is the real one?”.

Interpolation is {{ }} via the Display interface, not a mini format language inside the string.

The criterion is the same one that governs the rest of the language: what is trivial and mappable in your head is an operator; what is an environment-dependent semantic decision is a method. Comparing bytes is trivial: you know exactly what happens. Comparing “São” with “Sao” respecting locale is not trivial: it depends on normalization, on Unicode tables, on a choice. Hiding that choice behind a == would be magic; exposing it as equals(_, mode) is honesty, and it signals, by the very fact of being a method, that complexity lives there (and that the Unicode tables are heavy, so they are opt-in).

Forbidding s[i] is the same principle applied to access: the string forces you to choose the unit, because there is no single, cheap answer. It is the honest name (ch. 12) taken to indexing: s.bytes and s.chars say exactly what you get and how much it costs.

  • s[i] allowed. The programmer would never know whether they pay O(1) (byte) or O(n) (grapheme), and which unit they receive. It violates visible-linearity; the explicit unit solves it.
  • Transparent Unicode comparison in the operator (Swift et al.). It hides the locale/ normalization decision behind a < that looks trivial, and when the platform does collation where you wanted bytes (or vice versa), you are in a minefield with no warning. A method makes the decision visible.
  • === for bit-exact equality. Two equalities force you to ask which is the canonical one; it is “letting the wrong answer be cheap”. == is already unambiguous byte-equality; the rest is a method.
  • A mini format language in the string (%d, {:.2}). A second grammar inside the string. {{ expr }} via Display reuses the interface and lets each type control its own rendering.

The “localization hell”: the == that compares right in your locale and wrong in the user’s; the s[i] that splits an emoji in the middle because it indexed a byte thinking it was a character; the .length that counts UTF-16 code units and lies about how many characters there are. Each one is the ambiguous unit or the hidden comparison charging the price, almost always in production, almost always in a language that is not the developer’s.

Comparing a string is an explicit method because “equal” is ambiguous. And indexing requires saying the unit, because there is no single, cheap answer.

You type more: s.bytes.len instead of s.len, a.equals(b, .semantic) instead of a == b when you want Unicode. The common case (bytes) stays short, but the semantic one costs a call, on purpose, because it is where the real decision lives. The string refuses the surface comfort in exchange for you always knowing what you are comparing and how much you are paying.

Next: 10 · FFI