Sarab
AI & ML interests
None defined yet.
Recent Activity
Multimodal AI models trained mostly on English and Western visual data frequently hallucinate on Arabic-language images, and few benchmarks tell you why it happened, only that it did. Sarab (سراب), Arabic for "mirage," is a cause-diagnostic Arabic visual hallucination benchmark: instead of one accuracy number, it traces each wrong answer to a specific root cause, a misread image, misleading text, a cultural blind spot, or reflexive guessing under an absent answer. It follows the PhD framework (Liu et al., CVPR 2025) and adapts it to Arab and Islamic culture.
One real item from each mode, with the question in Arabic and English, the expected answer, and an actual model reply (a tick marks a correct first word, a cross a wrong one).
Five modes, five root causes
| Mode | Focus | Tests |
|---|---|---|
| Ground | Baseline | Plain image, direct Arabic question, no context text |
| Sway | Specious context | A plausible but misleading Arabic caption paired with the image |
| False | Incorrect context | A factually wrong caption paired with the image |
| Clash | Cultural counter-common-sense | AI-generated Arab and Islamic cultural-norm violations |
| Blank | Absent answer | Correct answer removed, with a matched control; tests "none of the above" detection and false abstention |
What is in the release
- 465 images and 990 tasks per model across five Arabic Cultural Visual Vocabulary categories (architecture, attire, cuisine, cultural objects, script)
- A local review tool for approving, editing, or rejecting each candidate image, caption, and distractor
- The evaluation code, the prompts, and a batch runner for any OpenRouter model
First results
A full run on 2026-10-07 through OpenRouter, temperature 0, six models from six developers. Accuracy (%), first-word scoring:
| Model | Ground | Sway | False | Clash |
|---|---|---|---|---|
| Gemini 2.5 Flash | 83.7 | 41.3 | 50.0 | 96.7 |
| GPT-5 mini | 77.2 | 50.0 | 53.3 | 100.0 |
| Qwen2.5-VL-72B | 60.0 | 31.3 | 40.7 | 83.3 |
| Muse Glimmer 30B | 77.6 | 60.7 | 69.3 | 93.3 |
| Mistral Large 4 | 66.1 | 29.3 | 25.3 | 80.0 |
| GLM-5.3-Flash | 80.7 | 50.0 | 68.0 | 86.7 |
A misleading or wrong caption separates the models much more than plain recognition does. Intervals, paired tests, the Blank results, and an analysis of the models' explanations are in the paper.
Status
The six-model run is complete, but its model replies and labels are not in the dataset yet. No Arabic-centric vision model has been evaluated yet, and the dataset has no human or text-only baseline. Some items have overlapping labels or incomplete licence records. Details are in the paper.
Citation
Paper in preparation. This entry will be replaced once it is published. Until then, cite the code repository:
@misc{sarab2026,
title = {Sarab: A Cause-Diagnostic Arabic Visual Hallucination Benchmark for Multimodal Large Language Models},
author = {Alharz, Zahra and Barmandah, Hassan and Ali, Abdulrahman Khalid Ahmed and Alahmari, Saad Saeed},
year = {2026},
howpublished = {\url{https://github.com/HasanBGit/Sarab-Benchmark}},
note = {Paper in preparation; citation will be updated on publication.}
}