Start at /benchmarks. Each entry links to a page that groups ingested scores by exact comparability key. Within a panel, models are sorted by score; uncertainty renders as provided or as unknown — never as zero. Every panel carries a source line: attribution string, license id, retrieval date, and a link back to /sources. Rows run by a model developer, when present, wear a self-reported badge and stay out of default headline rankings. If a model you care about is missing, it usually means we could not map an upstream name onto the catalog identity, or the row was not Tier A. Prefer a gap over a guess.
2026-08-10
Walking the benchmark desk
Open a benchmark page to see panels grouped by comparability key, with attribution, license, retrieval date, and uncertainty when the source provides it. Self-reported rows stay visually distinct.