Agents / Skills
GPT-Red
Automated red-teaming system using self-play to improve AI safety and robustness
byOpenAI
GPT-Red media is blocked
Allow external media to connect to the provider and play this content.
Description
An internal-only automated red-teaming system trained via self-play reinforcement learning to find and fix LLM vulnerabilities. It iterates on adversarial attacks to train production models like GPT-5.6 Sol, significantly reducing failure rates on prompt injection benchmarks.
Descriptions, tags, and model credits may be AI-generated or inferred from public sources and can be incomplete or wrong. Learn more.