Both were defective as evaluations. Every score measured against them is void, including ours. A single field comparison solved them:
| one-line detector | v1 | v2 |
|---|---|---|
response_length == 0 | balanced 0.9936 | balanced 1.0000 |
Our own detector scored 0.8126 and 0.9811 on those seasons — so a one-liner beat every real entry on the board. Any ranking they produced was measuring an artifact of how the corpus was assembled.
This is not the same as a season being retired. A retired season is sound and stays a valid reference. A withdrawn one should not be used for anything.
Season v3 replaces both, built by the same pipeline with the defect fixed and a build-time guard that refuses to emit a corpus a single field can solve.
The full account — what broke, why, and what changed — is in WITHDRAWN.md. The live board is here.
We found this ourselves, before any third party had submitted a score. Replacing the seasons quietly was available and is not what happened: a benchmark’s only asset is that its numbers mean what they claim.