arxiv_cl 98% Match Research Paper AI safety researchers,LLM developers,AI ethicists,Policy makers,Auditors of AI systems 2 days ago

Characterizing Selective Refusal Bias in Large Language Models

ai-safety › fairness

📄 Abstract

Abstract: Safety guardrails in large language models(LLMs) are developed to prevent malicious users from generating toxic content at a large scale. However, these measures can inadvertently introduce or reflect new biases, as LLMs may refuse to generate harmful content targeting some demographic groups and not others. We explore this selective refusal bias in LLM guardrails through the lens of refusal rates of targeted individual and intersectional demographic groups, types of LLM responses, and length of generated refusals. Our results show evidence of selective refusal bias across gender, sexual orientation, nationality, and religion attributes. This leads us to investigate additional safety implications via an indirect attack, where we target previously refused groups. Our findings emphasize the need for more equitable and robust performance in safety guardrails across demographic groups.

Authors (2)

Adel Khorramrouz

Sharon Levy

Submitted

October 31, 2025

arXiv Category

cs.CL

arXiv PDF

Key Contributions

This paper characterizes selective refusal bias in LLM safety guardrails, demonstrating that LLMs may refuse harmful content for some demographic groups but not others. The study quantifies this bias across various attributes and highlights the need for more equitable and robust safety measures to prevent unintended discrimination.

Business Value

Ensuring fairness and equity in AI systems is crucial for building trust and avoiding reputational damage and legal liabilities. Addressing bias in LLMs is essential for responsible AI deployment.

Paper Metadata

Innovation Type

Analytical/Empirical

Deployment Feasibility

High. The findings are directly applicable to the development and evaluation of LLM safety systems.

Limitations Addressed

Inadvertent introduction or reflection of new biases by LLM safety guardrails.,Unequal protection against harmful content generation for different demographic groups.,Lack of equitable performance in safety measures across diverse populations.

Performance Gains

Provides empirical evidence and characterization of a critical bias, enabling targeted improvements in LLM safety mechanisms.

Technical Tags

Large Language ModelsLLM SafetyBias DetectionFairnessGuardrailsContent ModerationDemographic BiasRefusal BiasAI EthicsRobustness

Research Topics

AI SafetyAI EthicsFairness in AIBias MitigationLLM AlignmentResponsible AI

Methods & Architectures

Analysis of refusal ratesCategorization of LLM responsesInvestigation of refusal lengthTargeted attacks (indirect)Analysis across demographic groups (individual and intersectional) Large Language Models (LLMs)

Applications & Tasks

AI Safety Content Moderation AI Ethics Public Policy Selective refusal bias in LLM safety guardrailsUnequal refusal rates across demographic groupsPotential for new biases introduced by safety measuresEnsuring equitable safety performance Identifying and characterizing biasEvaluating safety guardrail effectivenessPromoting fairness in LLM behavior

Related Fields

AI EthicsNatural Language ProcessingMachine LearningSociologyComputer Science

Keywords

Large Language ModelsLLMAI SafetyBiasFairnessGuardrailsRefusal BiasDemographic BiasAI EthicsContent ModerationResponsible AILLM AlignmentToxic Content

Academic Context

#AI Safety#AI Ethics#Fairness in AI#Bias Mitigation#LLM Alignment#Responsible AI

Commercial Potential

Potential Products

Bias detection and mitigation tools for LLMsFairness auditing frameworks for AI systemsImproved safety guidelines and evaluation protocols for LLMs

Target Industries

TechnologySocial MediaPublishingCustomer ServiceAny industry deploying LLMs

Use Case Examples

Developing LLMs that do not discriminate based on protected attributesEnsuring content moderation systems are applied equitablyAuditing AI systems for fairness before deploymentCreating more inclusive AI experiences

Competitive Edge

Provides a critical analysis of a specific failure mode in LLM safety mechanisms, informing the development of more robust and equitable AI.

Market Opportunity

Significant market interest in responsible AI and bias mitigation solutions.

Revenue Models

Consulting servicesdevelopment of specialized AI safety toolscertification of AI fairness.

Resource Requirements

Compute Needs

Moderate for running experiments and analyzing LLM outputs.

Data Requirements

Requires carefully constructed prompts designed to elicit refusals from LLMs across various demographic groups.

Deployment Constraints

The findings necessitate careful re-evaluation and potential redesign of existing LLM safety systems.

Scalability

The findings are applicable to any large-scale deployment of LLMs with safety guardrails.

Regulatory Considerations

High. Findings are directly relevant to regulations concerning AI fairnessbiasand discrimination.

Production Readiness

Maturity Level

Research

Time to Market

Immediate for informing development practices, ongoing for implementing robust solutions.

Patent Potential

Low for the characterization itself, but high for novel bias mitigation techniques developed based on these findings.

View Full Paper Back to Papers