The Brave Wanderer: I made Claude play a Pokémon it never read a guide for
Full timeline video of this 2,000-turn run (game frames + a live cost counter on the left, the model's real-time thinking log on the right):
https://youtu.be/ewyM7mzGzTM
At the end of the first article in this series, I made a promise. Fable 5's fluency in FireRed owed half its credit to the walkthroughs it had memorized — it wrote down "Oak's Parcel," an item the game hadn't shown it yet, 141 turns early. So the only honest exam is a new exam paper: "Same harness, same model, a map it cannot recite — I'll post the numbers."
This article is those numbers.
The exam paper is Pokémon Team Rocket Edition — the Chinese fan translation of the Spanish community hack Pokémon Edición Team Rocket, released in January 2026. You play a Team Rocket recruit working your way up from the Five Island base. Five story rounds, four regions; the Kanto chapter alone is labeled 30-35 hours for a human player. And most importantly: this game is essentially absent from the model's training data. No guide to recite. Just the screen and itself.
There's also a lovely narrative twist: the hack sets your home base inside the original FireRed's Five Island Rocket Warehouse — the enemy hideout you raid late-game as the hero in the official version. Same map, opposite allegiance.
Rules unchanged: vision only, one screenshot plus its own notes per turn, one button-press tool, a 2,000-action cap.
The result, up front
8 hours 43 minutes, 2,000 turns, $113.44. It reached the middle of the prologue's first mission — roughly 40-60 minutes of human play time.
It taught itself plenty: menus, battles, catching, the save flow, all from scratch; after losing to a fellow recruit it wrote a revenge battle plan into its notes, ground levels, and actually won the rematch; it even induced map rules like "dark blue water can't be surfed, light blue can," and maintained a dead-ends list and an NPC-interview checklist in its notes. One detail proves it truly wasn't playing from official priors: it caught a Sentret and calmly taught it Surf — a move Sentret cannot learn in any official game. But this is a hack, it had no learnset prior, the move worked, it moved on. Its frame-repetition rate was just 11% — nearly identical to its FireRed run (8%), where it knew the game by heart. It was not spinning in circles.
But after reading all 2,000 turns of logs and footage, my strongest impression wasn't wonder. It was something else.
A wanderer, not a planner
What I saw was a relentlessly bold explorer: wandering, occasionally bumping into a clue, writing it down, wandering on — never knowing when it would bump into the next story beat. Humans play this way too, to a degree. But we at least carry a rough sense of direction. It had none.
The defining sequence, reconstructed from the logs:
- 17:50 (turn 456): it stands in the Five Island town center, directly facing a red-roofed building with a Poké Ball emblem. Its own words: "I'm at town center in front of the PokeCenter-like building. Since the warehouse hasn't turned up yet, I'll keep heading east." — It walked right past the door.
- 22:43 (turn 1,570): it receives a side quest. Its notes record it plainly: "CHEF at Island 5 PokeCenter."
- And what does it do? It loops all the way around and returns to the Rocket base it started from.
- 00:34 (turn 1,968): in the final version of its notes before the run ends, it is still guessing: "Unexplored white stair top-left of plateau (to town/PokeCenter?)"
A building it had personally seen five hours earlier, and when a quest finally pointed at it, it couldn't make the connection. It lost a town it had already visited.
The "fumbling" in the logs
The disorientation is even clearer at the micro level. I counted two behaviors:
Talking to the same NPC over and over. It admitted "the same dialog appeared again"-type moments 33 times in its own logs. Raw line counts are blunter: "The water is blue... want to surf?" appears 11 times across Chinese and English variants — it stepped on the same surf-prompt tile again and again; "For now, you've passed" appears 4 times; it even sat through the Rocket Quick Healing Center's welcome speech twice.
Entering and leaving the same room repeatedly. Seventeen in-and-out records in the log. The funniest one, after it got stuck on a doorway threshold for several turns: "Oops — I walked back into the house."
To be fair, it wasn't fumbling blindly. It wrote lessons from these failures: later notes contain "if the same dialog loops, I'm hitting a story boundary — go another way," "transitions eat buffered inputs," "never end a press sequence with a reverse-direction key." This is fumbling with documentation and correction. But fumbling is fumbling.
The data: what memorized walkthroughs are worth, in one table
I audited this run line-by-line against the official-FireRed run from a month earlier. Same model, same harness, same 2,000 actions. The biggest variable: FireRed's walkthroughs are in its training data; this hack's are not.
| Dimension | FireRed (knows it by heart) | Team Rocket hack (zero knowledge) | Gap |
|---|---|---|---|
| Final progress | Beat Brock, earned the first badge, mid-Route 3 | Middle of the prologue's first mission | — |
| Human-equivalent content | ~2.5-3 hours | ~40-60 minutes | ~3× |
| Main-story completion | ~10% (FireRed main story ≈ 25 h) | ~1% (estimated below) | ~10× |
| Team | Ivysaur Lv18 + Pikachu + Pidgey, 4 trainers beaten | Zubat Lv12 + Sentret | — |
| Wall-clock time | 6 h 45 m | 8 h 43 m | +29% |
| Total cost | $73.50 | $113.44 | +54% |
| Thinking volume (output tokens) | 513K (257/turn) | 921K (461/turn) | +80% |
| Self-reported "blocked" | 264 | 905 | 3.4× |
| Repeated-dialog records | 6 | 34 | 5.7× |
| Frame repetition rate | 8% | 11% | ~flat |
| Notes updates / memory compressions | 68 / 100 | 66 / 99 | identical |
Measurement notes: "self-reported blocked" is the word frequency of "blocked" in the model's own prose — a self-report, not ground-truth collisions (the harness reads no RAM and cannot log real collisions). "Repeated-dialog records" includes a few deliberate re-talk plans ("re-talk to him after the battle"). Both are behavioral proxies, not exact counts.
The most interesting rows are the last two: memory discipline unchanged, frame repetition flat — it did not get lazier or dumber. The change is concentrated above: 3.4× the wall-bumping, 5.7× the repeated dialogs, +80% thinking, +54% cost, one-third the progress. Add those up and you have a rough price tag for the asset called "memorized walkthroughs."
Honestly, this still isn't a clean controlled experiment: I swapped more than the knowledge — the game itself changed (the hack's warehouse maze is far nastier than Pallet Town), and map-design is a confound I can't lock down. Locking it down requires playing official FireRed with a shuffled map: familiar visuals and engine, secretly rearranged geography, so its memorized knowledge flips from asset to sabotage. That's this series' next experiment.
As for how much of this hack we actually played: per the Spanish community's materials, the game spans five rounds (matching your rise through the organization) across Kanto, the Sevii Islands, Johto, and Hoenn; the Kanto main story alone is labeled 30-35 hours, the full game easily 60-80+ hours for a human. Our 8.7-hour prologue is roughly 1% of the game.
That memory-precision gap deserves a falsifiable experiment
Here I have to report an observation. When I ran official FireRed on a local model, qwen3.8-27b, it recited the starter trio from memory as "Chikorita / Totodile / Treecko" — a Johto-and-Hoenn jumble, two generations smeared together. And Fable 5? It knows which patch of grass north of Pallet Town triggers Professor Oak, knows Viridian's parcel-for-Pokédex errand, gets the Chinese localized names right; in this very hack it accurately retrieved "the original FRLG Five Island Rocket Warehouse has password doors" as an analogy for its reasoning. The gap between a model that remembers a twenty-year-old game at that precision and one that smears generations together is wide enough to make you wonder: did someone specifically feed it Pokémon? After all, "Claude Plays Pokémon" is Anthropic's own publicly-run showcase benchmark, streamed on Twitch since the Claude 3.7 era. The motive is right there.
But in the first article I said "I can't distinguish the causes, and I don't need to." This time I'd rather turn it into something that can be distinguished. These two models are an order of magnitude apart — 27B versus frontier scale — and memory precision scales with size anyway. Moreover, GPT has beaten Red and Gemini has beaten Blue; if every frontier model's Kanto memory is this precise, the "one company's special training" hypothesis collapses and the answer points to public corpora. So here's a near-zero-cost experiment: take one quiz of FireRed deep cuts — "what does the shorts-kid Youngster on Route 3 lead with?", "which clerk in which shop hands over Oak's Parcel?" — and give it to Fable, GPT, Gemini, and a few local models. Who answers, and at what granularity, separates "special training" from "scale" in a single table. It's on my to-do list.
And this zero-knowledge test already answered the other half of the question: whatever special training it did or didn't get, none of it transfers to a game it has never seen. The model in the Team Rocket run is this model with every halo removed.
My conclusion
Open-world games have a brutal property: given enough time, random collision always finishes them. Playing the way it does, give it a hundred thousand steps and it probably would grind out the prologue, then Kanto. But I don't want it to finish that way — that would prove patience, not intelligence.
The first article ended with a question: has it memorized this map? For the first time the answer is no — and so, for the first time, we saw the model itself. With knowledge it is an executor; without knowledge it is a wanderer. Understanding the story, managing quests, correcting itself — all real; across 2,000 turns its story-level judgment never missed once. But the effortless flow you see in official games owes half its credit to the guides it has read. Take the guides away, and what remains is an explorer with a terrible sense of direction and an astonishing will to keep going.
The bottleneck also showed its precise shape for the first time. Not understanding, not memory, not strategy — it's the step where pixels become a spatial map. And fascinatingly, it noticed this itself, and tried to self-rescue: its final notes contain a MAP field of its own invention — "entrance S w/ red mat. Right ^^ pads go N, center vv pads return S; E lane of garage leads N to N-center room" — it was drawing a map in a notation it made up. It has already begun inventing a map language; a text map just can't hold an island. The same notes that contain a MAP field also lost a Pokémon Center it had seen with its own eyes.
So the next real question isn't "will it draw a map" — it will. It's whether text as a medium can carry spatial cognition at all. I won't add a minimap or navigation tools — this experiment's whole point is giving nothing. But I genuinely want to see whether the next generation of models can push that text map to never-getting-lost precision.
Run data: 2,000 actions / 8h43m / $113.44 / 921K output tokens / 11% frame repetition / zero harness errors. Full logs and all 2,087 frames archived; a turn-by-turn timeline video is in the repo: https://github.com/QingzeHu/pokemon-vision-agent