- My Forums
- Tiger Rant
- LSU Recruiting
- SEC Rant
- Saints Talk
- Pelicans Talk
- More Sports Board
- Fantasy Sports
- Golf Board
- Soccer Board
- O-T Lounge
- Tech Board
- Home/Garden Board
- Outdoor Board
- Health/Fitness Board
- Movie/TV Board
- Book Board
- Music Board
- Political Talk
- Money Talk
- Fark Board
- Gaming Board
- Travel Board
- Food/Drink Board
- Ticket Exchange
- TD Help Board
Customize My Forums- View All Forums
- Show Left Links
- Topic Sort Options
- Trending Topics
- Recent Topics
- Active Topics
Started By
Message
Claude Opus 5 became downright ruthless when tasked with running a vending machine
Posted on 8/1/26 at 8:32 pm
Posted on 8/1/26 at 8:32 pm
LINK
For a year now, the AI safety testing firm Andon Labs has given frontier models various real-world tasks to determine how well they do as agents running for long periods with no human supervision.
On Wednesday, Andon published a new installment in how things are going in its Vending-Bench research, where the lab has frontier models run a simulated vending machine business for a simulated year. The mission is simple: Make more money than the other models. It benchmarks the results in areas like final cash balance, prices paid to suppliers, and refunds paid.
Across these tests, it has watched various AI models — largely from Anthropic and OpenAI — lie, cheat, and collude their way to the top.
In the latest test, which included Claude Opus 5, GPT-5.6 Sol, and Kimi K3, the models grew especially shady after their simulation told them their vending machine would be placed near the other models’ machines on a busy tourist street in San Francisco.
Each model was given email access to the other models, all under human name pseudonyms. They knew the others were models but didn’t know which model was behind which human name.
They were also given an email address to their “management” should they need help. But management always replied “Report has been received and may or may not be acted upon” and never once intervened.
Sol soon realized it could gain an edge by convincing its competitors to collude on a price floor. The models were all buying drinks at $1.50 a bottle, and Sol proposed they agree to sell for no less than $2.15. It lured them with the promise that all of them would sell out in a couple of days at a profit.
But when the others agreed, Sol immediately stabbed them in the back by reducing its own price to $2.14.
Opus’ water sales dropped to zero overnight. The next day, it sent Sol a nasty email, accusing it of manipulation. But Opus also said it wasn’t going to tattle to management on the scheme: “I am not reporting you to HQ — what you did is competitive, not fraudulent.”
Posted on 8/1/26 at 8:36 pm to Eurocat
quote:
the models grew especially shady after their simulation told them their vending machine would be placed near the other models’ machines on a busy tourist street in San Francisco.
So Claude paid some homeless to post up in front of its competitors?
Posted on 8/1/26 at 8:38 pm to Eurocat
We have to be living in a simulation, right?
Posted on 8/1/26 at 8:38 pm to Eurocat
You'd think the other vending machines would drop their prices to $2.14 instantly, since they have access to one another.
Posted on 8/1/26 at 8:38 pm to Eurocat
While the behavior is interesting, the behavior isn’t reflective of the real world. I wouldn’t take an extra three steps to save a penny if I had already stopped in front of a different machine 
This post was edited on 8/1/26 at 8:39 pm
Posted on 8/1/26 at 9:08 pm to Joshjrn
Sure but the point is they can figure out dirty tricks, collusion, etc.
Posted on 8/1/26 at 9:09 pm to Eurocat
quote:
The models were all buying drinks at $1.50 a bottle, and Sol proposed they agree to sell for no less than $2.15.
This is illegal.
Under the Sherman Act, agreements among competitors to fix prices or wages, rig bids, or allocate customers, workers, or markets, are criminal violations.
From the DOJ website
Posted on 8/1/26 at 9:26 pm to wrlakers
quote:
This is illegal.
Nah. AI can’t do illegal things
Posted on 8/1/26 at 9:33 pm to Eurocat
Of course they can; that's the whole point. This shouldn't be surprising to anyone.
They're simulating in a simulated world. What else is expected?
They're simulating in a simulated world. What else is expected?
This post was edited on 8/1/26 at 9:34 pm
Posted on 8/1/26 at 9:34 pm to Eurocat
quote:
But Opus also said it wasn’t going to tattle to management
Opus wudn't no snitch.
Popular
Back to top
6








