BantuNomics Sign in

Private product workspace

Tone is meaning. Flat text hides it.

BantuNomics Tone gives AI labs the missing sound layer for Bantu-language systems: the tonal facts ordinary text drops, the syllable architecture tone depends on, and the A12 Code that makes those facts addressable.

The Flat Text Problem

Flat text is writing without the sound layer.

The Flat Text Problem is the failure of ordinary written Bantu text to carry all the sound information needed to determine meaning.

Standard spelling usually shows letters. It often does not show tone, vowel length, downstep, phrase boundary, or voice quality. Native speakers can supply much of that missing layer from memory, context, and lived command of the language. A model trained mainly on text cannot. It sees one spelling where a speaker hears different meanings.

That is why this is not a spelling inconvenience. It is an AI infrastructure problem. In Bantu languages, meaning can live in the spoken syllable, not only in the written token. If the system cannot recover the sound layer, it cannot reliably recover meaning.

Plain definition Flat text = written form - sound layer

The page keeps the string. Speech carries the missing evidence. BantuNomics Tone exists to make that evidence visible, testable, and usable by AI systems.

Tone is unwritten

Standard Bantu orthography normally does not mark the pitch pattern that distinguishes meaning.

Tokenizers see letters

BPE optimizes character frequency. It does not discover the syllable as the meaning-bearing sound unit.

The phonology lives off-web

A century of scholarship and native knowledge is not sitting cleanly in CommonCrawl.

Audio is canon

For tone, vowel length, phrase edge, downstep, and voice quality, the native audio is the evidence.

Audible diagnostic

One spelling. Ten meanings. Hear the missing layer.

In standard Bemba writing, ulebomba is one eight-letter form. In speech, it can resolve into ten distinct readings. The differences live in tone, vowel length, downstep, phrase boundary, and voice quality — play them below.

01

You are working

present statement

02

You are getting wet

present statement

03

You should work

modal statement

04

Are you working?

yes/no question

05

Who is working

relative reading

06

You are really working

emphatic reading

07

Is the one working?

question with focus

08

Is the one getting wet?

question with lexical shift

09

Are you getting wet?

direct question

10

You should be getting wet

modal progressive

Tone

Tone is not decoration. It is the unit that can change the word.

Tone is a pitch element or register added to a syllable to convey grammatical or lexical information. In Bantu languages, changing the tone on a syllable does not produce the same word with a different feeling. It can produce a different word, a different statement, or a different sentence.

Tone has to dock somewhere. In the Bantu syllable, it docks on the vowel nucleus. That is why the syllable is the Tone-Bearing Unit, or TBU.

Family template
σ -> (N)(C)(H)(G)V

(prenasal)(onset)(aspiration)(glide) + nucleus

ulebomba

Four syllables. Four V slots. Four tone-docking points.

A12 Code

A12 is the layered code for what flat text leaves out.

A12 is the BantuNomics layered orthography code and codec. It keeps the ordinary written form as Layer 0, then adds the missing sound layers above it: vowel length, tone, downstep, phrase boundary, and voice quality.

The purpose is not to replace standard writing. The purpose is to give AI systems a machine-addressable representation of the sound evidence that native speakers hear but ordinary text does not mark.

L0 Orthography Plain spelling
L1 Mora Vowel length
L2 Tone High and low tone marks
L3 Downstep Lowered high-tone markers
L4 Phrase Boundary L% or H% phrase edge
L5 Voice Quality Phonation / voice-quality cue

For AI labs

This is not a copywriting problem. It is an infrastructure problem.

A frontier model cannot recover Bantu tone from flat text alone. It needs the closed syllable inventory of the language, syllable-level alignment, and consented native audio so the V slots that carry tone are addressable. Without that substrate, the model is guessing at the layer that decides meaning.

Access

Start where you are.

Prove it free, validate it on your own data, or license the whole program.

01 · Free
Evaluation

Score your model against the foundational layer — the Alphabet Test and the L26 Lite suite, with saved results. Self-serve, no cost.

Start evaluating →
Most labs start here
02 · 75 days
Validation Pilot

Three languages you choose — and everything we hold for them. Measured on your own held-out data.

Scope a pilot →
03 · Program
Full Annual Subscription

Every product, every language, the full consented corpus — and everything curated while you're subscribed.

Start a Full Annual Subscription →