ADHD wins 5 of 6 problems head-to-head. The biggest gap is trap detection — single-shot baselines almost never name the seductive-but-broken ideas, while ADHD’s separate critic pass routinely flags 15–20 of them with mechanistic reasons.
Full per-problem verdicts live in
EVALS.md; full transcripts in bench/results.json. Run date: 2026-05-25.
Independent external benchmarks
- A measured duel vs. single-shot by Shichinomiya — an independent blind-scored benchmark (2 problems, LLM-as-judge, A/B positions swapped). ADHD won both, with the biggest gains in novelty (4.5→9.0) and trap detection (5.0→9.0), at a real cost of ~2.3× time and ~1.9× output.
- Han plugin compatibility analysis — an evidence-based review using Han’s own
/researchskill: 11 sources, 8 validation rounds. Findings tracked openly as issues #16, #17, and #18.
Known limitations
Stated plainly, because reviewers will find them anyway:- Same-model judging. The judge is the same model family as the generator (familiarity bias). Cross-model judging is on the roadmap (issue #6).
- Small set. Six problems, all engineering-shaped.
- Scale gap. Evals run at K=5 branches; the academic diversity literature measures at K=100. Bridging this is tracked in issue #18.
- Human-in-the-loop applicability is unproven. A controlled CHI 2025 study found no significant benefit from LLM problem-reframing with human designers. ADHD’s LLM-to-LLM context differs, but this is honest counter-evidence, tracked in issue #16.
Reproduce the numbers
Methodology, how the judge is prompted, and how to run the suite yourself.