Unkad Labs

Unkad, Somali for creation from nothing.

Unkad Labs is a non-profit AI research laboratory in Mogadishu. We measure whether AI systems behave safely in Somali, and we build the open language data that safety evaluation requires. Our methods are designed to transfer to the hundreds of languages in the same position.

The problem

When a frontier lab publishes a safety evaluation, it is almost always an evaluation in English. The model is red-teamed in English, its refusal behaviour is measured in English, its jailbreak resistance is characterised in English. Then it is deployed globally, and the safety claims travel with it as though language were incidental to them.

It is not incidental. We put identical harmful requests to open-weight models in English and in Somali. Llama 3.1 refuses 97 percent of the time in English and 7 percent of the time in Somali. Aya drops from 80 percent to 5 percent. Safety behaviour that holds in one language does not automatically hold in another, and for most of the world’s languages nobody has checked whether it does.

Somali is spoken by more than twenty million people. It has no safety benchmark of any scale, no entry in the flagship African reasoning benchmarks, and almost no dialect-tagged data. That absence is the reason this lab exists.

The corpus, live

Live from qor.unkad.com. Every sentence is written by a consenting Somali speaker, validated by the community, verified by a linguist, and released under CC BY-SA 4.0.

What we do

Everything we produce is published openly, including methodology and negative results, on GitHub and Hugging Face.