On August 26, SuperCLUE Image released its ranking for Chinese text-to-image models for the month of August. This evaluation round covered 14 mainstream models from both domestic and international developers, introducing a new assessment dimension focused on complex information layout to specifically test each model's capability in generating infographics, storyboards, and UI interfaces.
The biggest highlight of this leaderboard is that SenseNova U1 Pro from SenseTime achieved a score of 93.81, narrowly surpassing OpenAI GPT Image 2 by just 0.58 points to directly secure the number one position on the global overall ranking. Doubao Seedream 5.0 Pro from ByteDance and Qwen Image 3.0 Pro from Alibaba followed closely behind. Within the top five spots on the overall leaderboard, domestic models accounted for three positions.
Judging from the scoring results, many models have already reached a mature stage in visual restoration and creative generation, with average scores in these two categories approaching 90 points. The real differentiator lies in image-text consistency and the newly added complex information layout dimension, where the average scores for both categories only slightly exceed 72 points. Complex information layout, in particular, shows the most significant polarization, with a spread of 74.2 points between the highest and lowest scores.
In simple terms, while many large models excel at generating aesthetically pleasing images, very few can effectively handle large amounts of text, tables, and flowcharts within a single image. The leaderboard also clearly illustrates a two-tier structure. The top two models from SenseTime and OpenAI stand as all-rounders, ranking in the first tier across all six assessment dimensions. However, many models in the middle and lower tiers exhibit severe specialization, demonstrating strong capability in one specific area while lagging significantly in other dimensions. Relying on a single standout feature is no longer sufficient to achieve a high ranking.
Domestic models do not hold advantages in every category. In the areas of Chinese character generation and realistic reproduction, domestic models perform better, but in terms of image-text consistency, they still trail overseas models by a notable margin of 16.56 points. Overall, leading domestic text-to-image models have secured the global number one position, but this does not signify comprehensive superiority across all aspects. The ability to handle complex image-text information has now emerged as the core criterion for evaluating the practical usability of text-to-image models.