From September 2026
Doctoral researcher, Blekinge Institute of Technology
Cybercampus Graduate School, Karlskrona. Project: Distributed Intelligence for Cybersecurity in Industry 5.0 and Critical Infrastructures — From Data Fusion to Intelligence Fusion. Supervised by Prof. Kurt Tutschku and Dr. Jianguo Ding.
From May 2026
Chief Technology Officer, Zoner Health
Kigali. Technical architecture, engineering team and product roadmap for the Zoner Pharmacy Management System, built for the African pharmacy sector.
Two tracks, one habit of mind: reading scattered, incomplete evidence and working out how much you can trust the conclusion.
Data scarcity is the whole problem.
Large language models are trained on text contributed by low-resource language communities, then reachable mostly through commercial APIs those same communities can't easily use. My work sits in the layer underneath. Most of it was done as a research associate at Carnegie Mellon University Africa with Prasenjit Mitra, and with Roald Eiselen at North-West University's Centre for Text Technology.
Mapping what already exists
A systematic catalog of public text and speech resources for both languages — parallel corpora, monolingual text, speech datasets, pre-trained models, benchmarks — recording size, domain, format, licence, and whether you can actually obtain it. Plenty of surveys stop at citations. This one stops at usability.
Making data where there is none
Fine-tuning MMS-300M with CTC loss on a curated Fɔ̀ngbè speech set reached 9.48% WER on the ALFFA benchmark and 3.96% CER, down from a 44.04% prior state of the art — with Fɔ̀ngbè-specific characters and tone diacritics fully preserved. The pipeline transcribed 45.5 hours across 424 videos into roughly 6,770 audio-text segments, and merged the ALFFA and Zenodo corpora into a unified 12.3-hour dataset with zero leakage, published to the HuggingFace Hub.
Testing whether it works
Prompting strategies for pulling usable text out of LLMs produced over 31,000 target-language words; four commercial models were benchmarked at up to 10,000 sentences with native-speaker evaluation. The uncomfortable result is that automatic metrics and human judgement disagree, and rankings flip by language — so a single leaderboard number tells you very little about either.
Preprints
-
From Speech to Text Corpora: Evaluating ASR-Based Data Acquisition for Low-Resource Fongbe and Hausa
A 78% relative WER reduction on Fɔ̀ngbè, with tone diacritics preserved.
-
Model rankings flip by language, and the metrics don't agree with the humans.
-
Augmentation helps by task, not by language — and for NER, not at all.
Research AssociateCarnegie Mellon University Africa, Kigali. ASR fine-tuning, corpus construction and LLM evaluation — the work above.
Graduate Teaching AssistantCMU-Africa. Supported 190 incoming students through orientation, built Java assignments and ran office hours for an online cohort, contributed to workshops at Abomey-Calavi and Gaston Berger, mentored Arduino and C++ IoT projects.
Professional TrainerTrainingcred and Indepth Research Institute, Kigali. A 10-day intensive on HR analytics, plus courses on GIS and remote sensing.QGIS · Google Earth Engine · PostGIS · PostgreSQL
Software engineeringChanels Innovation, MTN Benin, Adaptative Research, BJFarmers, Trellix. A microservice firmware-update system with versioned releases and audit trails; Python, SQL and Bash automation across enterprise data systems at a major West African telecom; React front ends over REST APIs.
MS, Information TechnologyCarnegie Mellon University, Pittsburgh.
BSc, Computer ScienceUniversity of Abomey-Calavi, Cotonou.
Paper Research AssistantA multilingual RAG assistant for semantic paper search and summarisation, containerised and deployed on GCP.LangChain · ChromaDB · Streamlit · Docker
Utterance-to-phoneme ASRA sequence-to-sequence pipeline over MFCC features reaching a validation Levenshtein distance of 5.85, with training automation and reproducible configs.PyTorch · pBLSTM · CTC · beam search · WandB
Medical image infrastructureWith Rwanda Biomedical Center: a backend serving image-classification models, PACS configured for DICOM management, PostgreSQL replication.FastAPI · PostgreSQL · Docker
Lung cancer detectionAn ensemble on the LIDC-IDRI CT dataset reaching 97.28% accuracy and 0.9921 AUC-ROC.TensorFlow · EfficientNet-B7
- Speech
- CTC loss, beam search, pBLSTM, MMS-300M, Whisper, torchaudio, MFCC, HuggingFace
- NLP & LLM
- Transformers, LangChain, RAG, ChromaDB, OpenAI API
- ML & systems
- PyTorch, TensorFlow, scikit-learn, FastAPI, Docker, Kubernetes, GCP
- Code
- Python, C, C++, Java, JavaScript, R, SQL, Bash
- Spoken
- French and Fɔ̀ngbè (native), English (professional), German (basic)
Get in touch.
I'm glad to hear from anyone working on African language resources, speech, or evaluation — and from people thinking about distributed intelligence in critical infrastructure.
- Emailmadjovi@alumni.cmu.edu
- Scholarscholar.google.com
- GitHubgithub.com/Pericles001
- ORCID0009-0006-1284-5563
- LinkedInin/periclesadjovi