Knowledge 4 All Foundation Joins London’s AI Ecosystem at The AI Summit London

Knowledge 4 All Foundation participated in The AI Summit London, one of the world’s longest-running and most influential AI conferences, bringing together thousands of business leaders, researchers, policymakers, investors, and technology innovators at Tobacco Dock. The Summit featured more than 300 speakers across multiple stages and tracks, showcasing practical applications of AI, emerging trends, governance challenges, and opportunities for social impact. K4A engaged with leaders from across the AI ecosystem, exploring new opportunities for collaboration and responsible AI adoption.

A key milestone for K4A during the event was strengthening its presence within London’s rapidly growing AI community. As a registered UK charity dedicated to advancing knowledge and education for public benefit, K4A connected with organisations working at the intersection of artificial intelligence, research, education, and social good. These conversations highlighted the increasing importance of civil society organisations in shaping an inclusive and beneficial AI future.

K4A would like to express its sincere appreciation to the organisers of The AI Summit London for welcoming charities and non-profit organisations through their Charity Pass programme. The Foundation was pleased to participate as part of London’s AI ecosystem and looks forward to building new partnerships that help ensure AI serves the public interest and creates opportunities for all.

AI4D blog series: The First Tunisian Arabizi Sentiment Analysis Dataset

Motivation

On social media, Arabic speakers tend to express themselves in their own local dialect. To do so, Tunisians use “Tunisian Arabizi”, which consists in supplementing numerals to the Latin script rather than the Arabic alphabet.

In the African continent, analytical studies based on Deep Learning are data hungry. To the best of our knowledge, no annotated Tunisian Arabizi dataset exists.

Twitter, Facebook and other micro-blogging systems are becoming a rich source of feedback information in several vital sectors, such as politics, economics, sports and other matters of general interest. Our dataset is taken from people expressing themselves in their own Tunisian Dialect using Tunisian Arabizi.

TUNIZI is composed of one instance presented as text comments collected from Social Media, annotated as positive, negative or neutral. This data does not include any confidential information. However, negative comments may include offensive or insulting content.

TUNIZI dataset is used in all iCompass products that are using the Tunisian Dialect. TUNIZI is used in a Sentiment Analysis project dedicated for the e-reputation and also for all Tunisian chatbots that are able to understand the Tunisian Arabizi and reply using it.

Team

 TUNIZI Dataset is collected, preprocessed and annotated by iCompass team, the Tunisian Startup speciallized in NLP/NLU. The team composed of academics and engineers specialized in Information technology, mathematics and linguistics were all dedicated to ensure the success of the project. iCompass can be contacted through emails or through the website: www.icompass.tn

Implementation

  1. Data Collection: TUNIZI is collected from comments on Social Media platforms. All data was directly observable and did not require other data to be inferred from. Our dataset is taken from people expressing themselves in their own Tunisian Dialect using Arabizi. This dataset relates directly to Tunisians from different regions, different ages and different genders. Our dataset is collected anonymously and contains no information about users identity.
  2. Data Preprocessing & Annotation: TUNIZI was preprocessed by removing links, emoji symbols and punctuation. Annotation was then performed by five Tunisian native speakers, three males and two females at a higher education level (Master/PhD).
  3. Distribution and Maintenance: TUNIZI dataset is made public for all upcoming research and development activitieson Github. TUNIZI is maintained by iCompass team that can be contacted through emails or through the Github repository. Updates will be available on the same Github link.
  4. Conclusion: As the interest in Natural Language Processing, particularly for African languages is growing, a natural future step would involve building Arabizi datasets for other underrepresented north African dialects such as Algerian and Moroccan.