Switch language한국어
Back to the list

My personal AI benchmark: "Generate an SVG of a frog with a Habsburg jaw." | Hacker News

TL;DR AI

Key summary

2 min read
  1. A Hacker News post compared several AI models using a prompt to generate an SVG of a frog with a Habsburg jaw.

  2. About half of the models added unrequested royalty-related details, and two explicitly said they were extrapolating beyond the prompt.

  3. One model produced the same output across repeated runs, while others differed in how much reasoning they narrated.

  4. The benchmark highlights real differences in prompt adherence, embellishment, determinism, and instruction-following reliability.

Read the original