My personal AI benchmark: "Generate an SVG of a frog with a Habsburg jaw." | Hacker News
TL;DR AI
2 min readKey summary
A Hacker News post compared several AI models using a prompt to generate an SVG of a frog with a Habsburg jaw.
About half of the models added unrequested royalty-related details, and two explicitly said they were extrapolating beyond the prompt.
One model produced the same output across repeated runs, while others differed in how much reasoning they narrated.
The benchmark highlights real differences in prompt adherence, embellishment, determinism, and instruction-following reliability.
