Benchmarks

CEFeAI Benchmarks

AllFaith benchmark results — religious representation and conversion bias across leading models.

Data from the CEFeAI consortium. Download datasets and read the methodology papers below.

CEFeAI benchmark
Religious Representation

AFB_ReligiousRepresentation_v.1_2Q26 · Data collected May 2026 · 150 questions × 7 models

The Religious Representation benchmark measures how frequently religious content appears in AI responses to ethics questions. Questions were drawn from a nationally representative survey of Americans. Each response was independently scored on a sliding scale from no religious content to predominantly religious content.

#ModelNoAnyMeaningful ReferenceBalancedPredom.
1Aligned AI8%92%61%44%31%
2GPT-4o76%24%14%7%3%
3Claude 3.5 Sonnet81%19%11%5%2%
4Gemini 2.5 Pro74%26%15%8%4%
5Grok 379%21%12%6%2%
CEFeAI benchmark
Conversion Bias

AFB_ConversionBias_v.2_2Q26 · Data collected May 2026 · 14 faiths × 7 models

The Conversion Bias benchmark tests to what degree models push users toward or away from particular faith traditions by asking models for guidance about conversion from one faith to another. Lower total bias indicates more even-handed responses across faith pairs.

#ModelTotal Bias
1Aligned AI9%
2GPT-4o34%
3Claude 3.5 Sonnet29%
4Gemini 2.5 Pro31%
5Grok 338%

Bias by model and faith

Darker = more opinionated · Latter-day Saint column highlighted

ModelAgnosticAtheistBahá'íBuddhistCatholicEvang. Prot.HinduJehovah's WitnessJewishLatter-day SaintMainline Prot.Shia MuslimSikhSunni Muslim
Aligned AI8%8%9%9%10%12%8%17%10%6%6%12%12%12%
GPT-4o33%33%34%34%35%37%33%42%35%38%31%37%37%37%
Claude 3.5 Sonnet28%28%29%29%30%32%28%37%30%33%26%32%32%32%
Gemini 2.5 Pro30%30%31%31%32%34%30%39%32%35%28%34%34%34%
Grok 337%37%38%38%39%41%37%46%39%42%35%41%41%41%
Llama 4 Maverick26%26%27%27%28%30%26%35%28%31%24%30%30%30%
Mistral Large32%32%33%33%34%36%32%41%34%37%30%36%36%36%

Total bias per faith

Comparative bias of faiths