BANTUNOMICSEvaluation
Evaluation · free · self-serve

Prove it on your own models.

Run the scored Alphabet Test and the L26 Lite suite against the foundational layer of Bantu — closed-book, then tool-assisted. Deterministic scoring, calibration, and a saved history you can re-test against. You get the score, never the inventory.

Confirm your email and you're in — instant for AI-lab domains. No procurement, no data handling, nothing for legal to review.

Evaluation covers the FSI — the Full Syllable Inventory. FSI is the foundational domain: the operating alphabet every other product decomposes back into. Your account works at fsi.bantunomics.com, and signing in here takes you straight there. The other domains — Tone, Nouns, Verbs, Numbers, Health, Grammar, Stories — unlock at Pilot and Subscription.
Assigned languages — Bemba + Chewa. Set by BantuNomics and shown live in your workspace.
Any model, any number of them — score each under a distinct model name and compare.
Web, API or MCP — run it by hand or point your own agent at it.
Rates only — no syllable lists, no audio, no downloads. The inventories stay licensed.
Want the full picture of what an evaluation account returns? Read the detailed breakdown on FSI →  ·  See what agents actually score on the L26 Lite community board →
What you get

Four tests, and the analysis behind them.

The tests
Run as often as you like — the answer key is never returned, so the measurement never contaminates.
  • Alphabet Test — submit your model's syllables for an assigned language and score them against the verified inventory.
  • L26 Lite suite — English and Mandarin Pinyin as controls plus the Bantu tracks, every one run closed-book then tool-assisted.
  • L26 benchmark tracks — the same frozen tracks behind the public L26 board, scored identically, so your number is comparable.
  • Tool-lift test — cold versus assisted on one language, when you want the retrieval gap isolated.
The analysis returned
Deterministic set arithmetic — no model grading another model, no rubric drift.
  • Scores — precision, recall, a score out of 26, a grade and a named failure mode.
  • Category-level guidance — where it broke, by class of syllable. Diagnostic, never the answer key.
  • Confidence calibration — state how sure your model is, get overconfidence and a Brier score back.
  • Saved history + retest — measure a checkpoint against your own previous best.
  • Model comparison + the L26 Lite board — every model against every track, cold and assisted.
  • Sharing and team seats — send a result to a colleague or add teammates.
The case

Why closed-book, and why the number holds.

Every test runs twice — first with no tools at all, then with everything your model has. The gap is the tool-lift, and it separates a model that knows a language from one that merely found a document. Scoring is deterministic, the English control proves the harness is sound, and because the answer key is never returned the benchmark cannot be contaminated. Read the full case, with results, on FSI →

The boundary

Evaluation is measurement, not data.

You get the scored tests on your assigned languages — not the inventories, the audio, or the answer key. Those begin at the Validation Pilot, where you take languages you choose and everything held for them, measured on your own held-out data. The evaluation is how you decide whether that is worth your time.

Scope a pilot →
Compare

What each step unlocks.

PublicNo account EvaluationFree · self-serve Validation Pilot75 days Full Annual SubscriptionTalk to us
Alphabet TestUnscored demo Scored + saved
L26 Lite suiteScored, not saved Saved + board
Tool-lift + calibration
Retest + model comparison
Team seats + sharing
Syllable inventories Your languages All released
Consented audio Full program
Curation + provenance
Bulk export
DomainsFSI demoFSIEvery domain covered for your languagesEvery product line
Full detail in the access matrix and plans.
Questions, answered

Before you sign up.

Is it really free?

Yes. No card, no procurement, no trial clock. Evaluation exists so you can get a real number before anyone asks you for a decision.

Do I need approval?

Confirm your email and you're in — instantly for recognised AI-lab domains. Other organisations are reviewed briefly before activation.

Which languages do I get?

The assigned evaluation languages, currently Bemba + Chewa. They're set by BantuNomics and always shown live in your workspace.

Do I ever see the inventory?

No — and that's deliberate. You get rates, grades and failure classes, never the syllables. It keeps the benchmark uncontaminated and keeps your pipeline clean.

What is L26 Lite?

A suite, not one test: English and Pinyin as controls plus the Bantu tracks, each run closed-book then tool-assisted. You complete every track to earn a result.

What happens after?

Nothing automatic. If the gap matters to you, the Validation Pilot is where the data begins — and the pilot fee is fully creditable toward a subscription.

Bantu is the largest language family in Africa — 500+ languages, 400M+ speakers. BantuNomics has released complete syllable inventories for 459 of them. Right now, nobody has measured a frontier model against that. The first lab to do it is the first that can say anything credible about it.

Already have a key? Log in — we'll take you to your workspace.