Bridging the gap from global foundational AI to local African authenticity.
LughaGen is a student-led research initiative under JHUB Africa at Jomo Kenyatta University of Agriculture and Technology (JKUAT), funded by the NVIDIA Academic Grant Program.
Our mission is to advance data-efficient, culturally appropriate, and trustworthy foundational Large Language Models for low-resource African languages, with a primary focus on Kenyan and East African languages.
We combine rigorous corpus engineering, compute-efficient training, safety & cultural alignment, and scalable deployment to create models that truly serve African contexts.
| Role | Lead | Focus Area |
|---|---|---|
| Data Leads | Nicolette Nkirote & Baraka Innocent | Corpus curation, cleaning & governance |
| Model Lead | Derek Mayabi | Efficient pretraining & customization |
| Trust & Evaluation | Kevin Musembi | Safety, cultural alignment & explainability |
| Deployment Lead | Godfrey Koros | Scalable inference & real-world deployment |
Principal Investigator: Dr. Lawrence Nderu Host Institution: JHUB Africa Center of Excellence, JKUAT
All datasets and models are released with clear provenance, licensing (primarily CC-BY-SA 4.0), and ethical documentation. We prioritize transparency, community benefit, and respect for the linguistic heritage of Kenyan and African communities.
Powered by NVIDIA A100 GPUs through the NVIDIA Academic Grant Program.
We deeply thank the creators of all open African language resources and the Kenyan language communities whose linguistic data makes this work possible.