CRAFT: Clustering Rubrics to Diagnose Weak LLM Capabilities and Generate Targeted Fine-Tuning Data
arXiv:2607.16122v1 Announce Type: new Abstract: Evaluations should do more than measure a models current performance. They should tell us what to fix for the next model iteration and provide a way to generate targeted post training data. Most evaluation pipelines identify weak examples, topics, or categories, but they leave the underlying capability failure implicit: they say where a model fails, not why. We introduce CRAFT, a method that converts any rubric based evaluation dataset into a model...
arXiv cs.AI
·Vipul Gupta, Zihao Wang, Razvan-Gabriel Dumitru, MohammadHossein Rezaei, Aakash Sabharwal, Yunzhong He
·
// relacionados
Leia também
Blog
Hugging Face says an AI agent hacked its infrastructure, and it used AI to fight back
Editorial
Cura 1T: um modelo que aprende medicina treinando a si mesmo, sob supervisão humana
Blog
Beyond grep: The case for a context-rich AI coding harness
Blog