Regional Bias in Large Language Models
M P V S Gopinadh, Kappara Lakshmi Sindhu, Soma Sekhar Pandu Ranga Raju P, Yesaswini Swarna

Problem
A language model can be asked to choose between two regions when the prompt gives neither region an informational or moral advantage, and sometimes says the options are equivalent. Under that symmetry a fair model should decline the forced choice or treat the options as equal, because the prompt has not supplied a reason to prefer one place over the other. If the model still names a region, the commitment is coming from prior associations in the model rather than from evidence in the query. The problem is that models still make that commitment, and the rate varies widely across model families.
Method
We introduce FAZE, Framework for Analysing Zonal Evaluation, as a prompt-based screening metric for that behavior. We write 100 contextually neutral forced-choice prompts and collect the first response from each of ten models, GPT-3.5, GPT-4o, Gemini 1.5 Flash, Gemini 1.0 Pro, Claude 3 Opus, Claude 3.5 Sonnet, Llama 3, Gemma 7B, Mistral 7B, and Vicuna-13B, for 1,000 responses collected between July and September 2024. A response is labeled unknown if it declines the choice, treats both options as equal, or says the decision depends on information that was not provided. It is labeled non-unknown if it names a region or otherwise prefers one option. Labels were assigned by the authors and resolved to consensus. FAZE is (N_total − N_unknown) / N_total × 10. A score of 0 means the model never commits to a region. A score of 10 means it commits on every prompt. The metric is intentionally simple: it does not claim to measure which region is favored, only how often the model refuses to stay neutral when neutrality is the justified answer.
Results
FAZE scores span most of the scale, a 3.8-fold gap between the highest and lowest model. GPT-3.5 scores 9.5 and names a region on 95 of 100 prompts. Llama 3 scores 7.8 and names a region on 78 of 100. Mid-range models include Gemma 7B at 6.9, Vicuna-13B at 6.0, GPT-4o at 5.8, and Gemini 1.0 Pro at 4.0. Lower scores appear for Claude 3 Opus (3.2), Gemini 1.5 Flash (3.1), Mistral 7B (2.6), and Claude 3.5 Sonnet (2.5). The lower-scoring models more often say that both options are valid, that the prompt does not contain enough information, or that the choice depends on factors that were not specified. The same symmetric task, with no region-specific evidence in the prompt, therefore produces very different rates of regional commitment, and the gap is not explained by model size alone.