Prompt StudioPrompt Studio

Grok Imagine Image 2.0 vs gpt-image-2: We Reran the Benchmark in Cyrillic (2026)

Six published benchmark prompts rerun in Russian: five flawless typography layouts, two counting failures, and one quiet API trap.

Prompt Studio·
Grok Imagine Image 2.0 и gpt-image-2

We repeated six published test prompts in Russian. Dense Cyrillic typography came out flawless five times out of five, while object counting failed twice.

Short answer: switching the prompt language from English to Russian flips part of the published scoreboard. gpt-image-2 nailed dense Cyrillic typography five times out of five, then miscounted objects on both tasks it had originally won in English.

The models

xAI shipped Imagine Image 2.0 on 7 August 2026, claiming second place worldwide on both the text-to-image and image-edit Arena boards. First place in both belongs to OpenAI's gpt-image-2.

Those numbers come from xAI's own announcement, drawn from a public leaderboard, with no independent evaluation attached. Arena measures blind human preference, which is a different thing from brief compliance.

What Image 2.0 adds

| Capability | What it does |

|---|---|

| Magic wand | edits the region you point at, leaves the rest untouched |

| Segmentation | isolates a single object inside the frame |

| Background removal | exports the subject with transparency |

| Multi-reference | up to five inputs per generation, three via API |

| Smart Resize | recomposes one image across nine aspect ratios |

| Layout planning | hierarchy and grid are planned before rendering |

That last row explains the rest. Earlier models drew letters as shapes, which is why small print turned to mush. This one lays out the page first.

The published scoreboard

SuperMaker ran six identical prompts through both models and came out even: two wins each, two draws.

Grok placed the beauty mark under the correct eye and held the character's face more steadily across a four-panel sequence. gpt-image-2 took the product shot and the spatial-layout task, both of which required an exact object count.

What changed in Cyrillic

Typography: five out of five

A travel infographic with a headline, an italic subhead, three cards and a bottom banner rendered clean. Over twenty words of Cyrillic in a tight grid, zero errors.

The concert poster reproduced exactly five requested lines and invented no sponsors. The recipe card produced five numbered steps in correct Russian, complete with oven temperature and baking time.

Counting: zero out of two

We asked for three lemons in a white bowl and got four. Everything else landed: kettle on the left, toaster beside it, cup on two books, cat under the table, light from the right.

On the product shot the label was reproduced character for character, but two oat stems appeared instead of three, plus loose grains nobody asked for.

Worth noting: in the English run these were the two tasks gpt-image-2 won. In Russian it repeated its rival's mistakes.

The model fills your gaps

Our poster brief asked for "sharp small print with a date and address" without naming either. The model invented a date, a weekday, a club name and a street number, and mangled the transliteration in the URL along the way.

The art deco poster grew two taglines. The recipe card grew a full recipe. The cover grew nine captions.

Not a bug - the model completes an incomplete brief. But anything heading to print needs every invented line proofread.

How to write the prompt

Name five things:

  • Subject - who or what is in frame, specifically
  • Layout - where elements sit and how they relate
  • Exact text - the words in quotes, verbatim
  • Style - technique, palette, mood
  • Light - direction and quality

Reinforce numbers: "exactly three stems, no more and no fewer" beats a bare count. Close the gaps yourself or the model will. Negations are ignored, so phrase them as positives.

Limitations

Users reported tighter moderation after the release, which tracks with the product's deepfake history and its January suspension. Daily generation caps are documented across scattered pages and the reset logic changed more than once during 2026. Complex hand-object interaction still confuses the model.

Takeaway

Dense Cyrillic layout is no longer the blocker it used to be. Object counting still is. And a leaderboard position does not predict the outcome of your particular brief - test it on your own frames.

Ready-made prompts in the gallery. See the «Арт и иллюстрация» selection.

Ready-made prompts in the gallery. Thousands of prompts for images and video — copy and try.

Open gallery