Skip to main content
NC State Home

Fihris.org

The AI-Powered Arabic Manuscript Archive

The homepage of Fihris.org, reading "Explore Arab American History" with a search bar.
The homepage for Fihris, showing the search bar used to navigate the database.

The Khayrallah Center is thrilled to announce that Fihris.org — your resource for searching and translating the archive’s Arabic manuscripts — is now available for public use.

Fihris revolutionizes the Arabic archive as we know it. It renders previously inaccessible documents immediately accessible from anywhere in the world by not only hosting manuscript images but also transcribing them into digital text, summarizing them, and translating them from Arabic. “This is the first ever full-text searchable Arabic manuscript archive, made possible through the proprietary AI model that we have developed at the Khayrallah Center over the past two years,” explains Khayrallah Center director Dr. Akram Khater. “Now, you can enter a word or more–in Arabic, English, French, Spanish, or Portuguese–and Fihris will retrieve the dozens, hundreds, or thousands of documents that contain that/those word(s).” From there, you are able to see the original document, read the transcribed Arabic text, and ask for its translation.

“The fact that you can search not only in Arabic but also in English or Spanish means that anyone can find and read these historical manuscripts, regardless of whether they know Arabic or not.”

In the near future, users will be able to submit images of their own Arabic manuscripts, and Fihris will transcribe them into digital text and provide translation if needed. They will also be able to type in a question, and our AI model will provide them with a historically contextualized answer, as well as a list of relevant primary sources to answer their query. This opens the Arabic archive to a far wider audience; with Fihris, students, the general public, and scholars alike will be able to read and utilize Arabic archival material in an unprecedented way.

Fihris comes to fruition as one of many efforts to achieve the Khayrallah Center’s mission: to share the extensive knowledge of the Lebanese diaspora that we have access to. “One of the main reasons we launched this project is to give Arab Americans wider access to the sources of their own history,” says Dr. Khater. “Across the Americas, many Arab Americans have handwritten Arabic letters, diaries, notebooks, and other material left to them by their ancestral families. Yet they cannot access these rich memories of their families and communities because most do not know how to read Arabic, and certainly not handwritten text. So, by developing Fihris, we are able to give them complete access to these hidden-in-plain-sight stories. Now they can read it, write about it, and pass it on to the next generation.”

Timeline

Over the past two years of development, the Khayrallah Center programming team has:

  • Developed ScribeArabic for creating a dataset.
  • Created/compiled Muharaf, the largest Arabic HTR dataset with a large variety of different types of handwriting, for training AI.
  • Generated transcriptions via our own HTR model + manual correction via ScribeArabic.
  • Generated translations, summaries, etc. via OpenAI’s GPT.
  • Developed the Fihris website and app for making the dataset searchable. 

More information on the development process can be found on Archival Technologies Specialist Mehreen Saeed’s GitHub. For a further explanation of the program and how it came to be, view our webinar discussing the project: