PHITSBench: an execution-scored benchmark for AI-assisted PHITS radiation-transport input generation using natural language
arXiv:2607.09789v1 Announce Type: new Abstract: We introduce PHITSBench, an execution-scored benchmark for the Monte Carlo Particle and Heavy Ion Transport code System (PHITS). PHITSBench comprises 282 transport-scorable tasks spanning three common workflow categories: parameter editing (Edit), syntax repair (Repair ), and complete simulation generation from natural-language descriptions (Reproduce). Each task is evaluated using a Composite Metric Score that combines execution success with agree...
arXiv cs.AI
·Xianglin Ji, Svetlana V. Boriskina
·
// relacionados
Leia também
Blog
Flight attendants freaked out that Google is buying tons of Spirit employee data
Blog
Attackers are using AI to build exploits for industrial control systems, U.S. agencies warn
Blog
AI labs are failing to keep their own systems in check
Editorial