Analyzing Frontier Model Progress with My Favourite Game: Prince of Persia
28 points by msephton 22 hours ago | 23 comments
Kuyawa 3 hours ago
I am going to try to build the original Prince of Persia using Swift as it is the best language for building a macOS game. Being a classic 2D cinematic platformer, Apple's native frameworks provide exactly what I need without the overhead of a massive cross-platform engine, and I love simplicity
replySo here goes my weekend, flag this and get a life
rhipitr 3 hours ago
Off tangent but the latest prince of Persia game is really fun
replypawelduda 3 hours ago
I played The Lost Crown and it's top tier, I assume you're talking about The Rogue?
replyKuyawa 3 hours ago
The Lost Crown? https://www.youtube.com/watch?v=n4vaxKwZ5rQ
replyThat looks really amazing
hapless 36 minutes ago
when the author mentioned playing prince of persia on a pc-xt, i assumed this was some kind of ai-generated nonsense
replybut no
they actually did port the game to the lowest-end hardware available in 1989. you could actually play prince of persia on an 8088 with a cga card!
smokel 3 hours ago
The prompts provided are atrocious. It's amazing that the LLMs actually built something useful.
reply> My first prompt was simple: "in this original code there 6502 assembly code for prince of persia, use the save level files and try to do it in c# console."
Do what in the what now in the console?
gcanyon 3 hours ago
"now char coming but movement everything wrong ... prince is not in floor"
replyI'm assuming a language barrier on the part of the author. I wonder if the models would do better being prompted in the author's main language.
cavemandaveman 2 hours ago
Or having another model proofread the prompts and write clearer instructions
replydaemonologist 2 hours ago
Yeah, I'm thinking the nigh-unreadable AI-speak we get these days makes a lot more sense if this is what they're training on. Or maybe the author has translated the prompts from another language?
replyInsideOutSanta 2 hours ago
It's so confusing how the actual article is in English, but the prompts are just gibberish. But LLMs are pretty good at deciphering gibberish; I often put our CEO's absolutely atrocious E-Mails into ChatGPT and tell it to explain wtf he wants from me.
replyAlso, I feel like the LLMs would have done better if they had started from scratch each time, rather than being burdened by the output from the previous attempt.
The highest scoring submission that won the contest had a high score of around 137k. Last week, I had GPT-6 Astra, Sol and Luna implement and hill-climb on this task, as I wanted to see how big the difference in smartness is. Luna implemented something, but never exceeded ca. 20k points, with a large variance. Sol got something in the area of the humans implementation.
Astra, which finished fastest, had a highscore of around 1.7Mm when the game seemed to fairly reliably crash. On the way, it disassembled parts of the ROM to extract information about the game.
I didn't do a ton work to document and measure the specifics, but it was very impressive.
Comparing strategies, though, Astra does something pretty different from the highest scoring submission. It doesn't predict the RNG, instead it reacts to the visuals on-screen, planning ahead by estimating velocity of all objects on screen. Unlike the winning solution, it actively flies the ship, whereas the winning solution basically only rotates and teleports. So Astra behaves more like a regular "perfect" player would, rather than one that breaks the PRNG.
To be fair, though, I mainly ran this experiment to distinguish models, not to see if they were better than humans or not, so for that purpose it didn't really matter if the basic conditions were the same or not.
You should check the logs, there's probably a bunch of retro gaming forums that got hacked behind the scenes :D