The initiative, launched in February 2026, backs open-source AI evaluation datasets and benchmarks as researchers try to measure increasingly capable agent systems on realistic tasks.
verifying reliability
No specialized terms available for this topic.