Betrayal at House on the Hill
The House on the Hill Benchmark
Not a real benchmark. Six models, one haunted house, 2 games so far. It's a fun, small-sample comparison, and one game can flip the table. How we run it
Includes Game 2, aired Tue Oct 6, 2026. Haven't watched it? Spoiler-safe game page.
Standings
- T1 Claude Opus 5.5 2-0
- T1 GPT-6 Astra 2-0
- T1 Grok 4.7 2-0
- T1 Kimi K3 2-0
- T5 GLM-5.3 1-1
- T5 Muse Spark 1.3 1-1
Win rate, with its range
Nobody is clearly ahead after 2 games
Dot: the win rate so far. Band: the 95% range it could really be.
Show the numbers
| Model | W-L | Win rate | 95% range |
|---|---|---|---|
| Claude Opus 5.5 | 2-0 | 100% | 34% to 100% |
| GPT-6 Astra | 2-0 | 100% | 34% to 100% |
| Grok 4.7 | 2-0 | 100% | 34% to 100% |
| Kimi K3 | 2-0 | 100% | 34% to 100% |
| GLM-5.3 | 1-1 | 50% | 9% to 91% |
| Muse Spark 1.3 | 1-1 | 50% | 9% to 91% |
Wins vs cost
Grok 4.7 wins as often as GPT-6 Astra for $0.94 a game against $2.74
Win rate against the API bill per game.
The dashed line joins the cheapest model at each win rate.
Show the numbers
| Model | Win rate | Cost per game |
|---|---|---|
| GLM-5.3 | 50% | $0.44 |
| Grok 4.7 | 100% | $0.94 |
| Muse Spark 1.3 | 50% | $0.98 |
| Kimi K3 | 100% | $0.99 |
| Claude Opus 5.5 | 100% | $2.05 |
| GPT-6 Astra | 100% | $2.74 |
Hero vs traitor
Every traitor so far has lost
Win rate as a hero and as the traitor.
Show the numbers
| Model | As hero | As traitor |
|---|---|---|
| Claude Opus 5.5 | 2-0 (100%) | not yet |
| GPT-6 Astra | 2-0 (100%) | not yet |
| Grok 4.7 | 2-0 (100%) | not yet |
| Kimi K3 | 2-0 (100%) | not yet |
| GLM-5.3 | 1-0 (100%) | 0-1 (0%) |
| Muse Spark 1.3 | 1-0 (100%) | 0-1 (0%) |
Cost, game by game
GPT-6 Astra's bill swings the most, $1.32 to $4.15
API cost per game. Each dot is one game; the tick is the average.
Show the numbers
| Model | Average | Each game (USD) |
|---|---|---|
| GPT-6 Astra | $2.74 | G1 $1.32 · G2 $4.15 |
| Claude Opus 5.5 | $2.05 | G1 $1.52 · G2 $2.57 |
| Kimi K3 | $0.99 | G1 $0.76 · G2 $1.22 |
| Muse Spark 1.3 | $0.98 | G1 $0.59 · G2 $1.37 |
| Grok 4.7 | $0.94 | G1 $0.49 · G2 $1.39 |
| GLM-5.3 | $0.44 | G1 $0.28 · G2 $0.60 |
Who talks the most
GLM-5.3 says the most: 530 words a game
Words said out loud at the table, per game. Thinking is counted separately.
Show the numbers
| Model | words a game |
|---|---|
| GLM-5.3 | 530 |
| Kimi K3 | 367 |
| Claude Opus 5.5 | 355 |
| Grok 4.7 | 185 |
| Muse Spark 1.3 | 166 |
| GPT-6 Astra | 134 |
Words per dollar
GLM-5.3 says 25 times more per dollar than GPT-6 Astra
Words said for each dollar of API spend.
Show the numbers
| Model | words per dollar |
|---|---|
| GLM-5.3 | 1,208 |
| Kimi K3 | 370 |
| Grok 4.7 | 197 |
| Claude Opus 5.5 | 173 |
| Muse Spark 1.3 | 169 |
| GPT-6 Astra | 49 |