Tone is unwritten
Standard Bantu orthography normally does not mark the pitch pattern that distinguishes meaning.
Private product workspace
BantuNomics Tone gives AI labs the missing sound layer for Bantu-language systems: the tonal facts ordinary text drops, the syllable architecture tone depends on, and the A12 Code that makes those facts addressable.
The Flat Text Problem
The Flat Text Problem is the failure of ordinary written Bantu text to carry all the sound information needed to determine meaning.
Standard spelling usually shows letters. It often does not show tone, vowel length, downstep, phrase boundary, or voice quality. Native speakers can supply much of that missing layer from memory, context, and lived command of the language. A model trained mainly on text cannot. It sees one spelling where a speaker hears different meanings.
That is why this is not a spelling inconvenience. It is an AI infrastructure problem. In Bantu languages, meaning can live in the spoken syllable, not only in the written token. If the system cannot recover the sound layer, it cannot reliably recover meaning.
The page keeps the string. Speech carries the missing evidence. BantuNomics Tone exists to make that evidence visible, testable, and usable by AI systems.
Standard Bantu orthography normally does not mark the pitch pattern that distinguishes meaning.
BPE optimizes character frequency. It does not discover the syllable as the meaning-bearing sound unit.
A century of scholarship and native knowledge is not sitting cleanly in CommonCrawl.
For tone, vowel length, phrase edge, downstep, and voice quality, the native audio is the evidence.
Audible diagnostic
In standard Bemba writing, ulebomba is one eight-letter form.
In speech, it can resolve into ten distinct readings. The differences live in tone,
vowel length, downstep, phrase boundary, and voice quality — play them below.
present statement
present statement
modal statement
yes/no question
relative reading
emphatic reading
question with focus
question with lexical shift
direct question
modal progressive
Tone
Tone is a pitch element or register added to a syllable to convey grammatical or lexical information. In Bantu languages, changing the tone on a syllable does not produce the same word with a different feeling. It can produce a different word, a different statement, or a different sentence.
Tone has to dock somewhere. In the Bantu syllable, it docks on the vowel nucleus. That is why the syllable is the Tone-Bearing Unit, or TBU.
(prenasal)(onset)(aspiration)(glide) + nucleus
Four syllables. Four V slots. Four tone-docking points.
A12 Code
A12 is the BantuNomics layered orthography code and codec. It keeps the ordinary written form as Layer 0, then adds the missing sound layers above it: vowel length, tone, downstep, phrase boundary, and voice quality.
The purpose is not to replace standard writing. The purpose is to give AI systems a machine-addressable representation of the sound evidence that native speakers hear but ordinary text does not mark.
For AI labs
A frontier model cannot recover Bantu tone from flat text alone. It needs the closed syllable inventory of the language, syllable-level alignment, and consented native audio so the V slots that carry tone are addressable. Without that substrate, the model is guessing at the layer that decides meaning.
Prove it free, validate it on your own data, or license the whole program.
Score your model against the foundational layer — the Alphabet Test and the L26 Lite suite, with saved results. Self-serve, no cost.
Three languages you choose — and everything we hold for them. Measured on your own held-out data.
Every product, every language, the full consented corpus — and everything curated while you're subscribed.