OpenAI finds roughly 30 percent of popular AI coding test is broken
OpenAI reviewed SWE-Bench Pro, a widely used test for measuring AI models' programming skills, and found roughly 30 percent of its tasks are broken. The company is pulling its earlier endorsement of the benchmark. The article OpenAI finds roughly 30 percent of popular AI coding test is broken appeared first on The Decoder .
The Decoder
·Maximilian Schreiner
·
// relacionados
Leia também
Blog
Flight attendants freaked out that Google is buying tons of Spirit employee data
Blog
Anthropic says any lab can now let a language model agent run the whole protein design stack
Blog
Valid Per-Field Selective Risk Control for Document Extraction: Three Failure Modes, a Validity Ladder, and When Conditioning Pays
Blog