About this tool
Build a reusable image-prompt style sheet across eight axes, with a coverage score, de-duplication and a CLIP token estimate.
AI Art Style Reference Sheet builds a reusable image-prompt style definition across eight axes — medium, lighting, lens, composition, colour, era, surface finish and mood — and scores how many of them you have actually named. It de-duplicates phrases, applies (phrase:weight) attention syntax to one axis when you want to push it, and estimates the prompt against CLIP's 77-token context so you know when the tail of your prompt stops being read. For anyone who gets one good image and then cannot reproduce the look.
Open AI Art Style Reference Sheet on AltFTool — it loads instantly in your browser.
Enter your Subject, then tap phrases under each of the eight axes — Medium, Lighting, Lens and camera, Composition, Colour, Era or movement, Surface and finish, and Mood — and add your own comma-separated phrases.
Optionally pick one axis under 'Emphasise one axis' and set the Emphasis weight between 0.5 and 2.0 to apply (phrase:weight) syntax, choose an aspect ratio, and toggle the negative prompt presets.
Watch Style coverage, the style-phrase count, Axes named and Estimated CLIP tokens against the 75 usable tokens, then press 'Copy sheet' to take the definition from 'Your reference sheet'.
Coverage tells you which parts of the look you have left to the model's imagination.
The sheet is a saved definition you paste again, so the same look survives across sessions.
An estimate against CLIP's 75 usable tokens warns you before the end of the prompt gets ignored.
CLIP, the text encoder used by Stable Diffusion 1.x and 2.x, has a context length of 77 tokens, two of which are start and end markers — so 75 tokens of prompt text per chunk. Interfaces like AUTOMATIC1111 split longer prompts into further chunks, but the first 75 tokens carry the most influence.
It is attention weighting in AUTOMATIC1111 and ComfyUI: the number multiplies how strongly that phrase pulls the image. This tool allows 0.5 to 2.0. Most useful results stay between 0.5 and 1.5 — past about 1.6 the phrase starts dominating and distorting everything else, so save the top of the range for a phrase that genuinely needs to win.
Usually because some style axes are unstated, so the model fills them in differently each time. Naming the medium, lighting, lens, composition, colour, era, finish and mood explicitly removes most of that drift; a fixed seed removes the rest.
They help for specific, recurring failures — hands, watermarks, text, jpeg artifacts — because they steer away from a concept the model has clearly learned. Long generic negative lists mostly waste tokens; six to eight targeted phrases is usually enough.