In Somali, unkad means creation from nothing. It shares a root with unug — the cell, the smallest unit of a living thing. That pairing is the whole idea of this lab: large things are created from nothing by assembling small units, one by one, until something lives.
That is what Somali needs in the age of artificial intelligence. The language is spoken by more than 20 million people and carries one of the world's great oral traditions, but the infrastructure AI needs — clean text, annotated datasets, benchmarks, safety evaluations — mostly does not exist. Dialects such as Maay are nearly invisible in what little data there is. The systems becoming the front door to information — search, assistants, translation — perform poorly in Somali, and nobody is measuring how poorly, or where the failures cause real harm.
We are not documenting infrastructure that exists. We are creating it from nothing. Unkad Labs works on two things.
The first is data. The Unkad Platform invites Somali speakers to write, record, and validate language data across the sectors where language matters most: health, education, agriculture, law, media, religion. Models assist; people verify. Every dataset we produce is released under an open license, and every contributor is credited and consented. Cell by cell.
The second is alignment. We are building the evaluation infrastructure that Somali currently lacks: safety test sets, refusal and toxicity evaluations, translation and comprehension benchmarks, red-teaming methods designed for low-resource languages. Our early measurements show open models refuse harmful requests in Somali at a fraction of their English rates. That gap is exactly the kind of thing no one fixes until someone measures it.
And this is bigger than one language. Somali is a hard case for alignment: little data, rich dialect variation, a vast oral tradition and a young written one. Methods that hold up here will transfer to the hundreds of languages in the same position. We intend our work to be useful to the broader alignment community — benchmarks, methodology, and honest negative results included.
We are starting small: a pilot in one or two sectors, a first cohort of contributors, a first open dataset, a first paper. Everything we make will be public.
If you speak Somali — any dialect — you can help us build this. If you are a researcher, a university, or a funder who believes no language should be left out of this technology, we would like to hear from you. Get in touch.
[SOMALI TRANSLATION — founders to write: a closing paragraph in Somali summarizing the invitation above.]