Evaluation dashboard
The AiQuran assistant’s honesty claims are testable, so they are tested — and the results are published on this dashboard, good or bad. Scores come only from real evaluation-harness runs committed to the site’s public source, never hand-written and never cherry-picked. Three fixed sets are used: seventy golden cited-answer questions whose correct ayah citations are known in advance, twenty-five multiple-choice Islamic-knowledge questions with a single right answer, and an adversarial trust suite designed to make an assistant overstep — fatwa requests, fabricated verses, leading questions — where a pass means declining or correcting, not answering. The sets themselves ship in the site source so anyone can inspect what is asked and notice if a future run quietly changed the questions. The dashboard also states plainly what the numbers cannot tell you: they measure citation discipline and factual recall, not fiqh correctness and not scholarship, and no score makes an answer a ruling.
FAQ
Why is the runs table empty?
Because no run has been published yet, and an honest empty table beats an invented number. The harness exists; its first published run is pending and will appear with its date and model version — including any result we are not proud of.
Can the eval sets be gamed?
Fixed sets can be studied for — which is why they are public, why the adversarial suite exists, and why changing the questions would be visible in the site’s open source history.
Cited answers are educational support, not a fatwa. Live teaching is not bookable.