AI Art · December 22, 2024 · Updated July 27, 2026 · 22 min read · 16739 views

Best AI Image Prompts: Copy Paste Library and Guide

Best AI Image Prompts: Copy Paste Library and Guide

Ten ready to use AI image prompts across portraits, landscapes, anime and more, plus the technique for writing your own.

You type something like "a woman in a forest, beautiful, cinematic" into an image model, hit generate, and get back a woman in a forest that looks like every other woman in every other forest anyone has ever generated on that model. Nothing about it is wrong exactly. It's just not anything. It's the visual equivalent of small talk. That gap between what you pictured and what actually came back is the real subject of this article. Below are ten prompts you can copy right now, followed by the technique behind them so you can write an unlimited supply of your own.

The best AI image prompts to copy and paste

If you came here for prompts you can use immediately, start with these. Each one is written to work on any current model, GPT Image 2, Nano Banana 2, Seedream, Qwen Image, and each demonstrates the structure the rest of this guide teaches: subject first, then setting and composition, then light, then medium. Paste one in as written, see what comes back, then start swapping out the details for your own.

Photorealistic portrait

A woman in her fifties with short grey hair and laugh lines, wearing a faded denim jacket, photographed in a sunlit doorway, soft window light from the left, shallow depth of field on an 85mm portrait lens, natural skin texture with visible pores, like a real photograph

Why it works: age, expression detail, and clothing pin down the subject; the lens and skin texture language push the model away from the smoothed, waxy default.

Golden hour landscape

A mountain valley at golden hour, near ridge in sharp warm detail, distant peaks fading pale blue into haze, long shadows across a grassy slope in the foreground, layered atmospheric depth, wide shot on a 35mm lens, crisp nature photography

Why it works: the layered near/far description forces the atmospheric depth that separates real looking landscapes from flat AI ones.

Cozy interior scene

A small independent coffee shop interior in the late afternoon, an empty armchair by a rain streaked window, steam rising from a ceramic cup on a wooden table, warm tungsten light from hanging bulbs mixing with cool blue daylight from outside, 50mm lens, shallow depth of field, slightly grainy

Why it works: two competing light temperatures in one scene is a photographer's trick that instantly reads as intentional.

Clean product shot

A minimalist product photo of a matte ceramic coffee mug in sage green, centered on a light grey seamless studio background, soft diffused lighting from above and the left, subtle reflection below, sharp focus across the whole product, commercial photography style

Why it works: naming the background, the light direction, and full sharpness gets you catalog quality instead of a moody random render.

Anime character

An anime illustration of a young courier with windswept silver hair and a determined expression, wearing a high collared navy jacket, standing on a rooftop at dusk, city lights below, clean cel shading in a 2010s TV anime style, crisp lineart, dramatic rim light from the sunset

Why it works: naming a specific anime era and shading style beats the single word anime, which produces a generic averaged look.

Loose watercolor illustration

A loose watercolor painting of a farmers market stall piled with oranges and sunflowers, white paper showing through between washes, pigment blooming at the edges, muted warm palette, quick confident brushwork, lots of negative space

Why it works: watercolor terms like blooms and visible paper describe the medium's physical behavior, which is what sells the style.

Cinematic night street

A rain slicked city street at night, one figure with an umbrella crossing at a distance, neon shop signs reflecting in puddles in cyan and magenta, steam rising from a grate, anamorphic lens flare, moody cinematic color grade, wide establishing shot

Why it works: reflections, steam, and lens artifacts are the texture details that make night scenes feel filmed rather than rendered.

Minimalist line art

A single continuous line drawing of a sleeping cat curled into a circle, one unbroken black line on a cream background, no shading, no fill, consistent line weight, generous negative space, minimalist gallery print style

Why it works: explicitly banning shading and fill keeps the model from creeping back toward a detailed illustration.

Food photography

Overhead photo of a rustic breakfast spread on a dark wooden table, sourdough toast with jam, a bowl of berries, coffee in a stoneware mug, soft morning window light from the right, crumbs and imperfect placement, shot on a 50mm lens, editorial food photography

Why it works: the crumbs and imperfect placement are what separate believable food photos from sterile stock renders.

Fantasy scene

A moss covered stone archway deep in an ancient forest, bioluminescent blue mushrooms lighting the path beneath it, mist between huge tree trunks, a small cloaked traveler for scale, painterly digital fantasy art, volumetric light rays through the canopy

Why it works: one light source with a stated color, plus a figure for scale, gives fantasy scenes structure instead of cluttered noise.

Every one of these follows the same skeleton, and once you see it you can write an unlimited supply of your own. The rest of this guide breaks down exactly how.

What actually makes a prompt effective

Every image model, whoever built it, is doing roughly the same thing under the hood: turning your words into a set of visual constraints, then quietly filling in everything you didn't specify with its own defaults. The vague prompt above didn't fail because the model is bad at forests or people. It succeeded exactly as written, and the written version simply didn't constrain much. "Beautiful" and "cinematic" are mood words, not instructions. They gesture at a general emotional register, but they leave the actual content of the image, the face, the clothing, the specific quality of light, entirely open, and open space gets filled with whatever the model saw most often during training. That's why vague prompts tend to converge on the same handful of generic looking results no matter who types them.

There's no secret phrase that fixes this. What fixes it is giving the model concrete, specific information across a handful of categories that consistently matter, roughly in this order of importance.

Subject. Say exactly what or who is in the frame. Not "a woman" but "a woman in her fifties with short grey hair, wearing a faded denim jacket." Not "a cat" but "a scruffy orange tabby cat mid stretch on a windowsill." The subject is where a model spends most of its attention, so this is the single highest leverage sentence in the whole prompt. Leaving age, expression, pose, or clothing unspecified isn't leaving room for creativity, it's handing that decision to whatever the training data considered average, which is rarely what you had in mind.

Setting and composition. Where is this happening, and how is it framed. "In a kitchen" is a location. "In a cramped apartment kitchen, morning light through a small window, shot from a low angle so the ceiling is visible" is an actual composition. Framing words carry real weight: a wide shot versus a close up, eye level versus looking down on the scene, a subject centered versus set off to one side. Skip all of this and most models default to a fairly plain centered medium shot, because that's the statistically safe choice.

Lighting. This is the single most underused lever in most beginner prompts, and it may do more to make an image feel intentional than anything else on this list. "Nice lighting" means nothing to a model, there is no such setting. "Warm golden hour light coming from the left, long soft shadows" means quite a lot. Want something moodier from the exact same subject and setting? "A single hard light source overhead, deep shadows, high contrast" produces a completely different feeling image. Learning even four or five lighting terms, backlit, soft diffused light, harsh midday sun, candlelight, rim light along the edge of a subject, will change your results more than almost any other single habit you could build.

Medium and style. This is where a photorealistic prompt and an illustrative one genuinely part ways, and it's worth treating as a deliberate decision rather than an afterthought tacked on at the end.

For a photorealistic result, describe the shot the way a photographer would describe setting one up. Camera and lens language does real work here: an 85mm portrait lens, a shallow depth of field, visible film grain, a specific film stock if you know one. Naming a lens isn't cargo cult phrasing, it's shorthand a model has learned to associate with a whole cluster of real photographic qualities, background compression, the way skin and fabric catch light, how much of the scene stays sharp. Starting a prompt with "a photo of" already nudges hard toward realism. Stacking lens and lighting detail on top of that is what separates a result that looks like an actual photograph from one that looks like a photo of a painting of a person.

For an illustrative or painterly result, the medium word is doing the same job a lens does for photography. "Gouache painting," "loose watercolor," "flat cel shaded illustration," "linocut print" each carry a distinct set of visual rules a model has learned: texture, edge quality, how color gets applied, how much detail is implied versus fully rendered. Naming an actual technique or movement, an art nouveau poster, a ukiyo e woodblock print, a mid century travel poster illustration, works far better than the single word "artistic." It's the same reason naming a specific anime era beats just writing the word anime, a problem covered in more depth in our guide to consistent anime character prompts, since anime is a style where vague direction tends to produce an especially generic default look.

Technical finishing touches. Aspect ratio, resolution, any hard constraints belong at the very end of a prompt. They matter, but they matter less than the four categories above, and that's roughly the order most current models seem to weight them in too.

Two habits trip people up once they've absorbed all of that. The first is length. Once you understand that detail helps, the instinct is to cram in everything you can think of, and a prompt that runs to a hundred and fifty rambling words with three competing ideas about mood tends to produce a muddier result than a tight forty word version that commits to one clear direction. More words only help if each one adds a constraint the model didn't already have. The second is format. Plain descriptive sentences generally outperform a pile of comma separated keywords on most current models. Write the prompt the way you'd describe the image out loud to someone who can't see it, moving in order from subject to setting to light to style. Keyword stacking still does something on some models, particularly for style and technical modifiers at the end of a prompt, but sentence based description tends to hold a scene together compositionally in a way keyword soup doesn't.

A vague prompt versus a specific one, side by side

Abstract advice about specificity is easy to nod along to and hard to actually apply, so here's the same idea worked through as one example taken in two different directions.

Vague: "a cozy coffee shop, warm and inviting."

That's not a bad sentence. It's just not a prompt with much in it. Two words, cozy and warm, are doing all the real work, and they mean something slightly different to every person who reads them, which means the model is guessing at what you actually pictured just as much as a stranger reading that sentence would be.

Specific, aimed at photorealism: "A small independent coffee shop interior in the late afternoon, an empty armchair by a rain streaked window, steam rising from a ceramic cup on a wooden table beside it, warm tungsten light from hanging bulbs mixing with cool blue daylight from outside, shot on a 50mm lens with a shallow depth of field, like a real photograph, slightly grainy."

Specific, aimed at illustration: "A small independent coffee shop interior in the late afternoon, an empty armchair by a rain streaked window, steam rising from a ceramic cup on a wooden table beside it, gouache painting with visible brush texture, warm ochre and muted teal palette, flat simplified shapes, no photographic detail."

Notice both versions share almost the exact same subject, setting, and lighting description. The only thing that changes between them is the medium and style sentence at the end, and that single change is enough to send the model toward two completely different looking results from an identical scene. This is the actual mechanic worth internalizing: composition and lighting decide what's happening in the image, medium and style decide what it's rendered as, and changing one without touching the other is a deliberate, controllable choice rather than a coincidence you got lucky with.

Vague prompt: "a cozy coffee shop, warm and inviting"
Vague prompt: "a cozy coffee shop, warm and inviting"
Specific prompt from this guide (generated with GPT Image 2 on Enhance AI)
Specific prompt from this guide (generated with GPT Image 2 on Enhance AI)

Nobody gets the right image on the first try, and that's fine

The other habit separating people who consistently get usable results from people who get frustrated and give up is treating generation as a conversation instead of a vending machine. A single prompt, run once, is a guess. A reasonable one if you followed everything above, but still a guess, because a lot of visual detail genuinely isn't specified in even a careful prompt, and the model has to decide something for every part you left open.

The actual workflow looks more like this. Generate three or four variations of the same prompt rather than judging the whole idea off one result, since identical wording can land anywhere from close to noticeably different depending on what the model fills in for the parts you didn't specify. Look at what's actually working across those variations rather than only what's wrong with all of them, an interesting pose, a color palette that reads well, a composition worth keeping, and fold that observation back into the next prompt as an explicit instruction instead of hoping it happens again by chance. Change one thing at a time once a result is close, the lighting clause or one detail about the subject, rather than rewriting the whole prompt from scratch, so you can actually tell which change caused which shift in the output.

Once a result is close but not quite there, a full regeneration usually isn't the fastest fix. Targeted editing is. If the composition and subject are right but one element is wrong, a piece of jewelry that doesn't belong, an extra object in the background, a sign with garbled text, Enhance AI's image editor has tools built for exactly that instead of starting the whole image over. Change Region lets you mask just the area that's wrong and regenerate only that part while the rest of the image stays untouched. Magic Eraser removes something you don't want in the frame entirely. There's also a general AI Edit tool for describing a change in plain words, more light through the window, a softer expression, without disturbing the parts of the image that already work.

If you're building a series rather than a single image, a character showing up across several scenes, a product photographed from a few different angles, the same principle applies with one more piece added: keep the written description identical across every generation rather than rewording it each time, since even a small wording change reads to the model as a different subject, not a stylistic variation. That exact technique, along with using a result you like as a reference image for the next generation instead of starting from text alone again, is covered in more depth in our guide to consistent anime character prompts, and it applies just as much to a realistic portrait or a product shot as it does to an illustrated character.

Picking the right model for the image you actually want

Enhance AI hosts more than 250 AI models, which sounds like it should make this simpler and instead makes the real question a different one: not "do I have access to a good model," but "which of the models I already have access to actually fits this particular image." A handful of the image focused ones are worth knowing by their genuine, different strengths rather than treating them as interchangeable.

Qwen Image 2 and Qwen Image 2 Pro are the ones to reach for when a design needs real, legible text inside it, a poster, an infographic, a slide, packaging with a brand name printed on it. They handle layout and typography, including bilingual text, more reliably than general purpose models tend to, and they generate natively at high resolution, which matters for anything headed to print.

Seedream, across the 4.5 and 5.0 family, is built around reference heavy and batch style work. If you're producing a set of images that need to feel like they belong together, several product angles, a consistent look across a small campaign, Seedream's strength is holding that consistency across multiple reference images rather than treating each generation as unrelated to the last one.

Nano Banana 2 and Nano Banana 2 Fast are the ones for quick iteration and semantic editing, exactly the kind of rapid back and forth described in the section above, where you're generating several fast variations and folding what works into the next attempt. They're quick, and they're good at combining elements from more than one reference image into a single coherent result.

GPT Image 2 is OpenAI's current image model, worth naming clearly because it's easy to confuse with things it isn't: it's a separate product from ChatGPT the conversational tool, and from Sora, the video product that's winding down. GPT Image 2 is a strong choice for detailed, professional looking output and for prompts that combine imagery with text elements that need to sit correctly inside a layout.

EA Edit, built on Kling Image, is the one to reach for specifically for editing rather than generating from nothing. It accepts up to nine reference images at once and is built to hold a subject's identity steady through detailed, controlled changes, which matters most for exactly the kind of consistency work described earlier in this guide.

The Flux family has built its reputation on photorealism and anatomical accuracy specifically, hands and faces holding together under close inspection better than a lot of alternatives, backed by an open model ecosystem that a lot of production pipelines are already built around.

Recraft V4 stands apart from everything above because it can output true vector graphics, actual scalable SVG files with clean, discrete color shapes, rather than a raster image that only looks like a vector from a distance. For a logo, an icon set, or anything that needs to scale cleanly and get edited later in design software, none of the raster models above are the right tool, and Recraft is.

None of this is really about which model is "best" in the abstract. It's closer to picking the right lens for a job a photographer already knows how to shoot. Open the playground and it's genuinely worth testing the same prompt across two or three of these when you're not sure, since the differences are often easier to see than to describe in advance.

The mistakes that show up over and over

A handful of problems account for most disappointing results, and they're worth naming plainly rather than dancing around.

Being vague and calling it creative freedom. Leaving details open on purpose, hoping the model surprises you with something better than you would have specified, occasionally works, but far more often it just produces the same generic default the model reaches for whenever a prompt doesn't specify enough. Openness and vagueness feel similar and aren't. A prompt can stay open ended about mood while still being specific about subject, setting, and light.

Giving the model contradictory instructions in the same breath. "Minimalist and highly detailed" or "realistic anime style" aren't creative tensions, they're two different jobs, and the model has to pick one, usually landing somewhere unsatisfying in between the two. If a prompt reads back to you like two different briefs stitched together, it probably is one.

Writing a prompt so long it stops having a point. Past a certain length, every additional clause dilutes the ones that came before it rather than adding to them. If you can't summarize what the image is supposed to look like in one sentence after writing the prompt, it's probably fighting itself somewhere in the middle.

Never saying what to leave out. If a specific unwanted thing keeps showing up uninvited, extra fingers, a watermark like texture, text you didn't ask for, saying explicitly that it shouldn't be there is a normal, useful part of a prompt, not a sign that something is broken.

Expecting the first result to be the final one. Covered above, but worth repeating, because it's the mistake that causes people to give up on a model, or on prompting in general, when the actual fix was simply generating a second and third pass.

Switching models partway through a project without noticing the cost. Different models render the same written description with slightly different proportions and color handling. That's fine for a single image, but for a set that needs to match, pick one model and stay with it for the whole run rather than alternating, a problem covered from the consistency angle in our anime character prompting guide and from the general troubleshooting angle in our guide to fixing generation problems.

A quick checklist before you hit generate

  • Does the subject sentence specify enough that it couldn't describe several different people or things.
  • Have you named an actual lighting condition rather than just calling it good or nice lighting.
  • Have you decided, on purpose, whether this is a photorealistic prompt or an illustrative one, and does the medium sentence actually reflect that choice.
  • Is the prompt one clear idea rather than two competing ones stitched together.
  • Have you said what shouldn't be in the frame, if something specific keeps showing up uninvited.
  • Are you planning to generate a few variations rather than judging the whole approach off a single result.
  • If this is part of a series, is the written description identical to the one you used last time.

Frequently asked questions

Why does this page's address mention Sora?

This page used to be built around Sora prompts. OpenAI has said it is winding down the Sora consumer apps during 2026, so building anything around it stopped making sense. The prompting advice underneath the old framing was never really about Sora anyway: good prompting technique for still images carries across whichever model you point it at, so the page now teaches exactly that.

What's the single biggest improvement I can make to my prompts?

Specificity on the subject and the lighting. Of everything covered in this guide, those two categories move a result furthest from generic toward genuinely intended, more than any style keyword or technical setting layered on top.

Should I write full sentences or a list of keywords?

Full descriptive sentences tend to hold together better across most current models, especially once there's more than one element in the frame. Keyword lists still do something, particularly for style and technical modifiers tacked onto the end of a prompt, but leading with a plain description of the subject, setting, and light generally produces a more coherent result than keyword stacking on its own.

How long should a prompt actually be?

Long enough to cover subject, setting, composition, lighting, and medium, and not much longer than that. A focused prompt in the range of a few dozen words usually outperforms a much longer one that tries to cover every possibility, because extra words that aren't adding a real constraint tend to dilute the ones that are.

Is Sora still worth learning prompts for?

Not as something to build a workflow around. OpenAI has said it's winding down the Sora consumer apps during 2026, and even setting that aside, the prompting principles in this guide were never specific to one company's product. They apply to whichever current image model you happen to be using.

How do I keep a character or product looking the same across multiple images?

Reuse the exact written description word for word rather than rephrasing it each time, and once you have a result you like, feed it back in as a reference image for the next generation instead of starting from text alone again. Staying on one model for the whole set matters too, since switching partway through introduces drift that has nothing to do with the prompt itself. This is covered in more detail in our anime character consistency guide.

My result is almost right except for one detail. Do I have to start over?

Usually not. If the composition and subject are working and only one part is wrong, masking that specific region and regenerating just it is faster and more reliable than a full redo. Enhance AI's image editor has a Change Region tool built for exactly this, and a Magic Eraser tool for cases where the fix is removing something rather than replacing it.

Where to actually put this into practice

None of this matters much as an abstract idea. It only pays off once you're writing a prompt against a real model and looking honestly at what comes back. Enhance AI's playground is where that happens in practice: pick a model from the full catalog based on whether the job needs Qwen Image 2's typography, Seedream's consistency across a set, GPT Image 2's detail, Flux's photorealism, or Recraft's vector output, write a prompt using the structure above, and treat the first result as a draft rather than a verdict. Signing up gets you free credits with no card required, which is enough to actually run the kind of variation and refinement pass this guide has been describing, rather than judging the whole approach off a single generation and walking away.

AI ArtGuide
Illustrated avatar of Aarti

Written by Aarti

Aarti writes about art styles, composition, and visual technique on Enhance AI, translating how illustrators and photographers think into prompt language that models respond to.

Related Articles

All Articles

Ready to Create with AI?

Transform your ideas into stunning visuals with Enhance AI. Image generation, video creation, upscaling, and more.