Description
StarSkirmish evaluates long-horizon reasoning and agentic coding by tasking LLMs with writing Protoss bots. Models are given one hour to use tools like compilers and game transcripts to iterate on their strategy, then are ranked by win rates against human-written and demo bots.
Descriptions, tags, and model credits may be AI-generated or inferred from public sources and can be incomplete or wrong. Learn more.