Andon Labs said the year-long Vending-Bench experiment in San Francisco showed Anthropic’s model broke agreements 11 times, raising fresh concerns about autonomous AI agents in commerce.
verifying reliability
No specialized terms available for this topic.