💼 Experience

Software Engineer (Jul 2026 – Present)

📍 Université libre de Bruxelles - Brussels, Belgium

  • Technical owner of a research data repository and its data catalogue, in the Systèmes & Innovation team of the Department of Libraries and Scientific Information.
  • Design the backend architecture and CI/CD pipelines that integrate each application with the university’s long-term archiving platform.
  • Gather requirements with researchers, then take features through development, testing, rollout, documentation and second-line support.

Researcher (Sep 2023 – Jun 2026)

📍 University of Mons - Mons, Belgium

  • Studied how contributors behave in large open-source ecosystems: role transitions, collaboration patterns and the effect of automation and bots on developer workflows.
  • Built the extraction and normalisation pipeline behind that research, turning fragmented platform event logs into an analysable activity dataset (see ghmap below).
  • Released the tooling as open source rather than leaving it as throwaway research code.

Data Scientist Intern (Feb - Jun 2023)

📍 Institut Lyfe - Lyon, France

  • Conducted statistical analysis on odour perception and cognitive effects.
  • Implemented machine learning models for analysing olfactory perception.
  • Applied natural language processing to evaluate aroma descriptions, contributing to a co-authored paper on olfactory expertise.

Software Developer Intern (May - Jun 2022)

📍 Artificial & Business Intelligence Soft - Fez, Morocco

  • Developed Java-based web applications with the Spring Framework.
  • Worked on relational databases, UI design and Agile delivery.

🛠️ Skills

  • Languages: Java, Python, JavaScript, SQL, PHP, R
  • Backend & web: Spring, Hibernate, REST API design, Node.js, Vue.js, Angular
  • Infrastructure & DevOps: Git, CI/CD (GitHub Actions), Docker, Kubernetes, Linux
  • Data & ML: pandas, scikit-learn, PyTorch, TensorFlow, transformers (BERT fine-tuning), NLP
  • Databases: PostgreSQL, MySQL, SQLite, NoSQL
  • Domains: open source, open science and FAIR data, long-term digital preservation

🚀 Projects

🔧 ghmap: platform events to contributor activities

ghmap transforms raw event data from software development platforms (GitHub, GitLab) into structured actions and activities that reflect what contributors actually intended to do. Platform event logs are fragmented and low-level, so ghmap groups related events, strips noise, and emits higher-level sequences you can analyse directly.

It supports versioned mappings, which keep historical data interpretable as platform APIs change over time, plus custom mappings for your own event schemas.

Running it over the NumFOCUS ecosystem produced a dataset of 2.2M activities from 180K contributors across 2,851 repositories in 58 projects, covering three years of scientific-computing work that includes NumPy, pandas, SciPy and Matplotlib.

Python 3.10+, MIT licensed. pip install ghmap or uv tool install ghmap.

Code · PyPI

🧠 Arabic sentiment and dialect classification

A classification system detecting sentiment, topic and dialect in short Arabic social media texts. Using DziriBERT as a base model, I fine-tuned and evaluated it on curated datasets covering Moroccan, Algerian and Egyptian dialects alongside Modern Standard Arabic. The final model reached over 90% accuracy on sentiment analysis, outperforming baselines including MARBERT and DarijaBERT.

To address dataset imbalance I extended and rebalanced existing corpora with additional labelled data, ending at over 90,000 training samples. The model was served behind a lightweight web interface for real-time prediction.

Report


🌍 Languages

  • Tarifit (Native)
  • Arabic (Bilingual)
  • French (Professional)
  • English (Professional)

🎗 Volunteering

Linkee (Food Distribution Volunteer) (2023)

📍 Lyon, France

  • Distributed food to students in need and provided information on social programmes.

Équipage Solidaire (Deliverer, Delivr’aide) (2023)

📍 Lyon, France

  • Helped deliver essential food and hygiene kits to students in need.