Pericles Adjovi

Karlskrona, Sweden · Kigali, Rwanda · Cotonou, Benin

MahounanPericles Adjovi

Fɔ̀ngbè is my first language. It is also one of the least resourced languages in speech and language technology. Most of what I publish is an attempt to answer one question: when there is almost no data, what actually works?

Publicly available NLP resources
Fɔ̀ngbè Hausa ≈2M speakers · Benin · tonal 80–100M speakers · non-tonal

Fɔ̀ngbè tone marks change meaning. Most pipelines strip them. Keeping them intact is a design requirement, not a nice-to-have.

Now

From September 2026

Doctoral researcher, Blekinge Institute of Technology

Cybercampus Graduate School, Karlskrona. Project: Distributed Intelligence for Cybersecurity in Industry 5.0 and Critical Infrastructures — From Data Fusion to Intelligence Fusion. Supervised by Prof. Kurt Tutschku and Dr. Jianguo Ding.

From May 2026

Chief Technology Officer, Zoner Health

Kigali. Technical architecture, engineering team and product roadmap for the Zoner Pharmacy Management System, built for the African pharmacy sector.

Two tracks, one habit of mind: reading scattered, incomplete evidence and working out how much you can trust the conclusion.

ResearchLanguage technology

Data scarcity is the whole problem.

Large language models are trained on text contributed by low-resource language communities, then reachable mostly through commercial APIs those same communities can't easily use. My work sits in the layer underneath. Most of it was done as a research associate at Carnegie Mellon University Africa with Prasenjit Mitra, and with Roald Eiselen at North-West University's Centre for Text Technology.

Mapping what already exists

A systematic catalog of public text and speech resources for both languages — parallel corpora, monolingual text, speech datasets, pre-trained models, benchmarks — recording size, domain, format, licence, and whether you can actually obtain it. Plenty of surveys stop at citations. This one stops at usability.

Making data where there is none

Fine-tuning MMS-300M with CTC loss on a curated Fɔ̀ngbè speech set reached 9.48% WER on the ALFFA benchmark and 3.96% CER, down from a 44.04% prior state of the art — with Fɔ̀ngbè-specific characters and tone diacritics fully preserved. The pipeline transcribed 45.5 hours across 424 videos into roughly 6,770 audio-text segments, and merged the ALFFA and Zenodo corpora into a unified 12.3-hour dataset with zero leakage, published to the HuggingFace Hub.

Testing whether it works

Prompting strategies for pulling usable text out of LLMs produced over 31,000 target-language words; four commercial models were benchmarked at up to 10,000 sentences with native-speaker evaluation. The uncomfortable result is that automatic metrics and human judgement disagree, and rankings flip by language — so a single leaderboard number tells you very little about either.

PublicationsAccepted, 2026
  1. From Translation to Retrieval: Evaluating LLM-Based Information Retrieval for Hausa and Fongbe

    SIGIR 2026 · Low-Resource Language Track Adjovi, Eiselen, Mitra Melbourne
  2. A Survey of Text and Speech Resources for Hausa and Fongbe: Availability, Quality, and Gaps for NLP Development

    IEEE SDS 2026 Adjovi, Olufemi, Eiselen, Mitra Zurich
  3. Mining Large Language Models for Low-Resource Language Data: Comparing Elicitation Strategies for Hausa and Fongbe

    RAIL 2026 · co-located with LREC Adjovi, Eiselen, Mitra Mallorca

Preprints

  1. From Speech to Text Corpora: Evaluating ASR-Based Data Acquisition for Low-Resource Fongbe and Hausa

    Adjovi, Olufemi, Eiselen, Mitra arXiv:2606.22274

    A 78% relative WER reduction on Fɔ̀ngbè, with tone diacritics preserved.

  2. Evaluating Large Language Models for Hausa and Fongbe Machine Translation: Benchmarks, Failures, and Metric Reliability

    Adjovi, Eiselen, Mitra arXiv:2606.22269

    Model rankings flip by language, and the metrics don't agree with the humans.

  3. When Does Data Augmentation Help? Evaluating LLM and Back-Translation Methods for Hausa and Fongbe NLP

    Adjovi, Eiselen, Mitra arXiv:2604.12540

    Augmentation helps by task, not by language — and for NER, not at all.

BackgroundRoles
2025 – 2026

Research AssociateCarnegie Mellon University Africa, Kigali. ASR fine-tuning, corpus construction and LLM evaluation — the work above.

2024 – 2025

Graduate Teaching AssistantCMU-Africa. Supported 190 incoming students through orientation, built Java assignments and ran office hours for an online cohort, contributed to workshops at Abomey-Calavi and Gaston Berger, mentored Arduino and C++ IoT projects.

2025

Professional TrainerTrainingcred and Indepth Research Institute, Kigali. A 10-day intensive on HR analytics, plus courses on GIS and remote sensing.QGIS · Google Earth Engine · PostGIS · PostgreSQL

2021 – 2024

Software engineeringChanels Innovation, MTN Benin, Adaptative Research, BJFarmers, Trellix. A microservice firmware-update system with versioned releases and audit trails; Python, SQL and Bash automation across enterprise data systems at a major West African telecom; React front ends over REST APIs.

Education
2023 – 2025

MS, Information TechnologyCarnegie Mellon University, Pittsburgh.

2019 – 2022

BSc, Computer ScienceUniversity of Abomey-Calavi, Cotonou.

Selected work
2025

Paper Research AssistantA multilingual RAG assistant for semantic paper search and summarisation, containerised and deployed on GCP.LangChain · ChromaDB · Streamlit · Docker

2024

Utterance-to-phoneme ASRA sequence-to-sequence pipeline over MFCC features reaching a validation Levenshtein distance of 5.85, with training automation and reproducible configs.PyTorch · pBLSTM · CTC · beam search · WandB

2024

Medical image infrastructureWith Rwanda Biomedical Center: a backend serving image-classification models, PACS configured for DICOM management, PostgreSQL replication.FastAPI · PostgreSQL · Docker

2024

Lung cancer detectionAn ensemble on the LIDC-IDRI CT dataset reaching 97.28% accuracy and 0.9921 AUC-ROC.TensorFlow · EfficientNet-B7

Toolkit
Speech
CTC loss, beam search, pBLSTM, MMS-300M, Whisper, torchaudio, MFCC, HuggingFace
NLP & LLM
Transformers, LangChain, RAG, ChromaDB, OpenAI API
ML & systems
PyTorch, TensorFlow, scikit-learn, FastAPI, Docker, Kubernetes, GCP
Code
Python, C, C++, Java, JavaScript, R, SQL, Bash
Spoken
French and Fɔ̀ngbè (native), English (professional), German (basic)
Contact

Get in touch.

I'm glad to hear from anyone working on African language resources, speech, or evaluation — and from people thinking about distributed intelligence in critical infrastructure.