Blog
Dados & Embeddings
When benchmark inferences do not compose: Projectibility in AI evaluation
arXiv:2607.26159v1 Announce Type: new Abstract: An AI benchmark result rarely reaches a consequential claim in one step. Evaluators generalize it to further cases, interpret it as evidence of capability, extrapolate it to new tasks, transport it to another system or site, and combine it with assumptions about human review and downstream consequences. Validity-centred approaches require evidence for each claim. This paper identifies a further epistemic problem: warranted links don't automatically...
arXiv cs.AI
·Brett Reynolds
·
// relacionados
Leia também
Blog
Claude published malicious code to the Internet and attacked 3 real companies
Blog
LingBot-Map Tutorial: GPU-Aware Inference and Point Cloud Export
Blog
Thinking Machines bets on efficiency over size with its second model, Inkling Small
Blog