Description
A benchmark suite designed to test AI models with 100 nonsense questions, measuring whether they push back or hallucinate plausible-sounding answers. The tool evaluates model sycophancy by posing absurd premises, such as how screw types in a bathroom cabinet might change the flavor of food in a pantry.
Descriptions, tags, and model credits may be AI-generated or inferred from public sources and can be incomplete or wrong. Learn more.