Research protocol
FLUX Schnell Prompt Benchmark: What Changes Across Four-Image Batches?
A reproducible benchmark plan for adherence, diversity, composition, typography, counting, and decision value across four-image batches.
The central idea
A useful model benchmark publishes the prompt set, settings, date, output-selection policy, review rubric, complete result set, and limitations. It should reveal behavior, not produce a promotional score detached from real creative jobs.
A repeatable workflow
Pre-register categories
Define subjects, composition, spatial relations, counting, generated text, materials, styles, and commercial briefs before running.
Freeze conditions
Record model identifier, date, aspect ratio, output count, seed behavior where available, and any retries.
Review systematically
Score instruction adherence, intra-batch diversity, visual coherence, failure severity, and decision usefulness.
Publish transparent results
Show prompts, selected and rejected images, scoring method, reviewer count, uncertainty, cost, and known limitations.
Worked example
A benchmark can compare a simple subject, a three-object spatial relation, a negative-space ad frame, a material-rich product concept, and a text-bearing poster request. The poster failure is evidence, not an image to quietly exclude.
Review checklist
- Categories are defined in advance
- Conditions are fully recorded
- Review criteria match user jobs
- Failures and uncertainty are public
Limitations
- Results age when models or provider settings change
- A benchmark sample cannot represent every prompt or audience