When evaluating top frontier models on basic conversational phrases in Santomean Creole, hallucination rates exceeded 80%. We started ForroLLM to discover how modern fine-tuning techniques can preserve low-resource languages with minimal parameter budgets.
The tokenizer bottleneck
Standard byte-pair encoding (BPE) tokenizers used in mainstream foundation models splinter Creole words into meaningless single-character byte sequences. This causes exorbitant compute consumption and degrades semantic coherence.
Our initial experiments focused on customized phonetic vocabulary augmentation, merging dialectal root tokens and training specialized embedding layers.
“If your tokenizer destroys the morphological rhythm of a language, no amount of reinforcement learning will restore its soul.”
Curating authentic oral corpora
Because written literature in Forro is scarce, we collaborated directly with native speakers and elders, digitizing oral histories and creating synchronized audio-text alignments. Every token in our evaluation set represents verified communal knowledge.
Written by Henriques
Henriques is the founder of LIVLU Studio, exploring product design, Creole language AI, and creative photography.
Respond or start a conversation →