Blog
Dados & Embeddings
An AI model programmed nonstop for 19 days on a single MirrorCode task that cost $2,600 to run
Epoch AI's new MirrorCode benchmark tests whether AI models can recreate complete programs without access to the original code. Claude Opus 4.7 leads with a 56 percent solve rate, rebuilding a 16,000-line toolkit in just 14 hours. But every model tested still fails on the most complex tasks. The article An AI model programmed nonstop for 19 days on a single MirrorCode task that cost $2,600 to run appeared first on The Decoder .
The Decoder
·Matthias Bastian
·
// relacionados
Leia também
Editorial
O modelo que entra em espiral: a receita da Liquid AI contra os "doom loops"
Blog
Consórcio alemão de IA lança o Soofi S, um modelo aberto de 30B que lidera benchmarks tanto em inglês quanto em alemão
Blog
SensorFM, do Google, transforma dados desorganizados de sensores de wearables em uma camada de inteligência de saúde de uso geral
Blog