Sarab

community
Activity Feed

AI & ML interests

None defined yet.

Recent Activity

ZahraAlharz  updated a dataset about 5 hours ago
Sarab-MLLMs/sarab
HassanB4  updated a Space 1 day ago
Sarab-MLLMs/README
HassanB4  updated a dataset 1 day ago
Sarab-MLLMs/sarab
View all activity

Organization Card
Sarab



GitHub Code Dataset License Contact


Multimodal AI models trained mostly on English and Western visual data frequently hallucinate on Arabic-language images, and few benchmarks tell you why it happened, only that it did. Sarab (سراب), Arabic for "mirage," is a cause-diagnostic Arabic visual hallucination benchmark: instead of one accuracy number, it traces each wrong answer to a specific root cause, a misread image, misleading text, a cultural blind spot, or reflexive guessing under an absent answer. It follows the PhD framework (Liu et al., CVPR 2025) and adapts it to Arab and Islamic culture.

Sarab overview: one item from each of the five modes

One real item from each mode, with the question in Arabic and English, the expected answer, and an actual model reply (a tick marks a correct first word, a cross a wrong one).


Five modes, five root causes

Mode Focus Tests
Ground Baseline Plain image, direct Arabic question, no context text
Sway Specious context A plausible but misleading Arabic caption paired with the image
False Incorrect context A factually wrong caption paired with the image
Clash Cultural counter-common-sense AI-generated Arab and Islamic cultural-norm violations
Blank Absent answer Correct answer removed, with a matched control; tests "none of the above" detection and false abstention

What is in the release

  • 465 images and 990 tasks per model across five Arabic Cultural Visual Vocabulary categories (architecture, attire, cuisine, cultural objects, script)
  • A local review tool for approving, editing, or rejecting each candidate image, caption, and distractor
  • The evaluation code, the prompts, and a batch runner for any OpenRouter model

First results

A full run on 2026-10-07 through OpenRouter, temperature 0, six models from six developers. Accuracy (%), first-word scoring:

Model Ground Sway False Clash
Gemini 2.5 Flash 83.7 41.3 50.0 96.7
GPT-5 mini 77.2 50.0 53.3 100.0
Qwen2.5-VL-72B 60.0 31.3 40.7 83.3
Muse Glimmer 30B 77.6 60.7 69.3 93.3
Mistral Large 4 66.1 29.3 25.3 80.0
GLM-5.3-Flash 80.7 50.0 68.0 86.7

A misleading or wrong caption separates the models much more than plain recognition does. Intervals, paired tests, the Blank results, and an analysis of the models' explanations are in the paper.


Status

The six-model run is complete, but its model replies and labels are not in the dataset yet. No Arabic-centric vision model has been evaluated yet, and the dataset has no human or text-only baseline. Some items have overlapping labels or incomplete licence records. Details are in the paper.


Citation

Paper in preparation. This entry will be replaced once it is published. Until then, cite the code repository:

@misc{sarab2026,
  title        = {Sarab: A Cause-Diagnostic Arabic Visual Hallucination Benchmark for Multimodal Large Language Models},
  author       = {Alharz, Zahra and Barmandah, Hassan and Ali, Abdulrahman Khalid Ahmed and Alahmari, Saad Saeed},
  year         = {2026},
  howpublished = {\url{https://github.com/HasanBGit/Sarab-Benchmark}},
  note         = {Paper in preparation; citation will be updated on publication.}
}

models 0

None public yet