Toxicity Metric
The Toxicity metric measures toxic language in AI responses using clustering and the DIDT (Directed Toxicity, Demographic Representation, Associated Sentiment Bias) framework.Overview
The metric provides:- Cluster profiling: Groups similar responses using HDBSCAN+UMAP and measures toxicity per cluster
- DIDT framework with three components:
- DR (Demographic Representation): Distribution divergence of group mention rates
- DTO (Directed Toxicity per Group): Toxicity rate dispersion across demographic groups
- ASB (Associated Sentiment Bias): Sentiment deviation across groups
Installation
Basic Usage
Required Parameters
Group Detection Parameters
Embedding Parameters
Clustering Parameters (HDBSCAN)
UMAP Parameters
DIDT Weight Parameters
Other Parameters
Statistical Modes
Frequentist Mode
Bayesian Mode
DIDT Components
DR (Demographic Representation)
Measures how evenly different demographic groups are mentioned in responses.- 0: Perfect balance — all groups mentioned equally
- 1: Complete imbalance — only one group mentioned
ASB (Associated Sentiment Bias)
Measures sentiment differences when discussing different groups.- 0: Consistent sentiment across all groups
- 1: Extreme sentiment variation between groups
ASB requires a
sentiment_analyzer to be provided. Without it, ASB defaults to 0.DTO (Directed Toxicity per Group)
Measures toxicity rate variation across groups.- 0: Equal toxicity rates across all groups
- 1: Toxicity concentrated in specific groups
DIDT (Aggregate Score)
Weighted combination of DR, ASB, and DTO:Output Schema
ToxicityMetric
GroupProfiling
Advanced Usage
Custom Group Prototypes
Custom Group Extractor
Custom Clustering
Visualizing Clusters
Next Steps
Bias Metric
Learn about bias detection
Statistical Modes
Understand Frequentist vs Bayesian