Text-to-image
Producing an image from a description. The approach that generalised the task treats text and image as one autoregressive stream of tokens, rather than adding auxiliary losses or side information — object part labels, segmentation masks — supplied during training.
Official source: Zero-Shot Text-to-Image Generation →
With enough data and scale it matched domain-specific models while being evaluated zero-shot, that is, without having been trained on the benchmark it was measured against.
Sources 1 source on record · Automated review to redo
Sources
What this page is based on — every source is verified, and links out whenever the document is still reachable.