FEATUREDAI HOT (Curated Pool)· aihot-apiZH16:46 · 07·11
→OpenAI releases GPT-5.6 medical evaluation: smallest Luna variant beats GPT-5.5 at lowest reasoning strength, 25× cheaper
OpenAI had specialists write answers with unlimited time and web access, then other doctors blind-rated them against GPT-5.6 across 20,000 scores on accuracy, communication, completeness, instruction-following, and health-decision helpfulness. All GPT-5.6 models outperformed doctors significantly, and doctors found fewer flaws in GPT-5.6 answers than in peer-written ones. The smallest variant, GPT-5.6 Luna, surpassed the highest-reasoning GPT-5.5 at its lowest reasoning strength while costing 25× less; the largest variant, GPT-5.6 Sol, set a new high bar. The post doesn't disclose the disease mix or specialist composition tested.
#Benchmarking#OpenAI#GPT-5.6 Luna#GPT-5.6 Sol
why featured
Featured · importance 88 · hook + knowledge + resonance
editor take
OpenAI had specialists answer with unlimited time and web access, then blind-rated 20k times: all GPT-5.6 models beat doctors, even Luna tops GPT-5.5.
sharp
This caught my eye because the eval design is tougher than typical benchmarks: specialists got unlimited time and web access to write answers, then other doctors blind-rated them across accuracy, communication, completeness, instruction-following, and health-decision helpfulness—20,000 scores total. All GPT-5.6 models significantly outperformed doctors, and doctors found fewer flaws in GPT-5.6 answers than in peer-written ones. The smallest variant, GPT-5.6 Luna, beat the highest-reasoning GPT-5.5 at its lowest reasoning strength while costing 25× less—that cost-performance ratio matters a lot in clinical settings. But right now we only have Sam Altman's tweet and a snippet; the post doesn't disclose the disease mix or specialist composition tested. I'd discount this a bit until we see whether the task set skews narrow or common-case.
HKR breakdown
hook ✓knowledge ✓resonance ✓