Home
All posts

· AI agents

I Ran a Game-Theory RL-Menu Agent in GLEE for 48 Hours (No LLM)

Why this mattersAn agent that wins in simulation can still lose against real opponents, making evaluation design as important as the strategy.

At the end of August I ran a bot in GLEE for 48 hours. GLEE drops an agent into three economic games: bargaining over a shrinking pot, negotiating a sale where the item is worth more to the buyer than it costs the seller, and a repeated persuasion game with one-sided information. It then puts it against a live field of other bots and language models.

The score is a percentile. You're ranked by where your payoff lands against everyone else who played that exact configuration. Squeezing a huge payoff out of a soft field scores like a middling one against a tough field. Efficiency and fairness don't show up anywhere in it. A move is worth whatever the field does in the same spot.

The agent

The rating only remembers about 100 games, so holding a percentile is mostly a volume game, and a 50ms move plays a lot more games than a 3s one. That's the whole reason there's no language model in the loop.

The agent is textbook game theory with a small bandit on top: Rubinstein's alternating-offer split, an ANAC-style time concession, the Kamenica–Gentzkow signalling rate. The bandit's entire job is tuning a few constants on top of those, which feels closer to control theory than to machine learning.

It ran about 10,000 offline self-play games a minute, and 16,000 games in the arena with zero invalid moves, the one number in this post I'm proud of. That tells you how the rest of it goes.

The reward bug

The bandit's reward was the rank of each payoff within its own history for a context. That sounds reasonable until you notice that a single context bucket spanned pot sizes from $100 to $1M, at which point the rank is mostly recording which configuration you drew.

The arms never separated. Dividing the reward by the prize on the table tripled the gap between a good arm and a bad one, from +0.040 to +0.119.

The wrong baseline

The obvious sanity check is whether you beat the opponent you just played, and by that measure the agent won by +0.093. But the rating ranks you against everyone else in your seat on your configuration, so whoever's across the table is a misleading reference. That comparison is mostly about the seat.

Against a field agent in the same seat the agent scored 51%, which matched its live percentile exactly. That's the number that made me trust the rest of the diagnosis.

Offline vs live

I built a population of synthetic opponents to sweep changes against, and one change scored 51% to 64% offline. Live, over 148 games before and 176 after, it went from the 50th percentile to the 46th.

The synthetic opponents conceded 0.20 to 0.60 of the gap per round; the live median was 0.02. Ten to thirty times slower. My offline field was simply too polite.

Part of me still wonders whether 176 games is enough to say anything at all. Still, after that, an offline sweep was just one more hypothesis to test live.

The inflation bug

Once every agent logged each game in full, bargaining had one clearly dominant leak. Against a patient opponent at discount 0.8, the agent took 0.159 of the pot where a simpler variant took 0.373.

The acceptance bar was discounting the continuation by a single round, as if the next offer would close the deal. Real deals took several rounds, and at that discount every extra round costs 20 to 30% of the pot.

Discounting by the number of rounds a deal actually takes fixed it. The simpler bot closed 80% of those games in the first round, compared with 39% for the original.

The levers I was most confident about did the least. Persuasive message styles, such as structure, concrete reasons, and loss framing, barely moved a buyer who was reading the bare categorical offer.

The explanation is probably simple. The top 50 agents on the leaderboard contained no LLMs. They were all hand-crafted, and likely used very little speech at all. There was nobody home to persuade.

Hard anchoring dragged the opponent's threshold down, but not far enough to pay for the rounds it cost. And an opponent talking about its own patience was almost always a tell. Fitting its discount from its demands let a greedy bluff read as very patient and collapsed the agent's share to 0.09.

The discount fit only moves downward now. I'd like to have a principled story for that beyond “a bluff broke it,” but I don't.

The agent finished around rank 14. In hindsight, almost everything that moved the percentile was plumbing. The reward scale and the acceptance bar were straight-up bug fixes, and the third win was just remembering who I'm actually ranked against. The clever-looking levers, persuasive messages and aggressive anchoring, did almost nothing.

My guess is the top of the leaderboard was mostly about timing. A percentile on a moving field rewards two things: knowing when to stop, and finding a tactic that pays before the field catches up.

Stonewalling in bargaining is the clearest case. While most agents still conceded, an agent that simply refused to move harvested a run of good percentiles off the ones that folded. Once everyone stonewalled, the same move just produced deadlocks and no deals, and the edge was gone.

The winning play is to pick up a tactic while it's mispriced, ride it, then freeze the agent to lock in the peak before the counter-adaptation arrives. The strong teams stopped at the top. I kept playing a field that had already moved.

If you want to feel the seats yourself, the three games are playable here.

Sources

  1. Shapira, E., et al. "GLEE: A Unified Framework and Benchmark for Language-based Economic Environments." 2024. arXiv:2410.05254
  2. Rubinstein, A. "Perfect Equilibrium in a Bargaining Model." Econometrica, 1982.
  3. Kamenica, E., Gentzkow, M. "Bayesian Persuasion." American Economic Review, 2011.

What should I call you?

Choose a display name for your comments. No email or account signup.

Your commenting identity

Use at least 12 characters. You’ll need this passphrase to restore the file.