← Back to list

Finding Balance With AI — Creative Generation

Prologue — The Picture That Looked Like Everyone’s Picture

Sundar Krishnamurthy · 2026-06-16 13:20 · 1 claps · 12.4 min read
#ai #creative #photo-editing #video-generation
Open on Medium ↗
Wiki topics: MM · Multimodal & Generative Media AI · AI · General

Finding Balance With AI — Creative Generation

Prologue — The Picture That Looked Like Everyone’s Picture

When I built my AI recipe generator application, my favourite flourish was the feature that visualised each recipe as an AI-generated image. Type your pantry in, get a dish out, and — voilà — a glossy photograph of a meal that did not exist. My first attempts at those food images, though, taught me something equally humbling and enlightening. I’d type “a delicious bowl of pasta” and receive a picture that was technically flawless and completely anonymous. Warm light from nowhere in particular. A garnish of indeterminate herbs and stuff. That faintly airbrushed, paintlike sheen that has a subtle hint of its origins. It wasn’t bad by any meabs; however it was worse than bad — it was generic. It looked a picture that not really a picture!

That, in a nutshell, is the creative version of the problem this series is about. In the previous post, we met the proverbial “dial”: give AI too much control and you get slop; hold back too much and you get nothing worth your tokens’ spend. The dial travels into this post’s subject matter fully intact — but the failure modes here are visual, the iteration loops are faster, and the gap between slop and craft comes down to something I’ll keep returning to: the human finishing pass. Its also worth keeping in mind that in this subject the dial comes with a cost since creative generation with AI takes up a lot more of those pesky tokens than generative text.

One structural shift frames this whole post, so let me state it up front. The way good creative AI work happens has changed. It is no longer “type a magic prompt, accept the output.” It is an editing loop: generate a strong draft, keep what works, edit the weak parts conversationally and interate. The people getting remarkable results are not better prompt-whisperers. They are better editors. Industry analysis describes the field as having moved decisively to these editing-first creative loops — and once you see your role that way, everything in this chapter clicks into place.

The people getting good results aren’t writing magic prompts. They’re running a loop.

We’ll move from static to motion to synthesis: image generation, photo editing, video, animation, and finally presentation decks — the use case where text, structure, and visuals all converge, and where every lesson in the post converges with a definitve intent.

Image Generation — Specificity, References, Iteration

Generating original images — illustrations, concept art, hero images, icons — is where most people meet creative AI, and where the dial’s two extremes are easiest to spot.

Too far toward AI: “make a cool logo for my coffee brand.” You get something generic, derivative, on-trend-but-soulless — indistinguishable from everyone else’s cool coffee logo, because the model averaged across everyone else’s. Too far toward yourself: a two-hundred-word prompt micromanaging every pixel, followed by rejecting everything that isn’t the exact image in your head. The first surrenders taste; the second refuses the medium.

The sweet spot is a brief that is specific without being suffocating, references that anchor the style, and an editing loop where you steer toward the image in your head one adjustment at a time. Three important elements do most of the work here. Let me explain them:

Be concrete about the visual ingredients. Subject, composition, style, lighting, palette, mood, medium. This is the visual version of the homogenization finding from the previous post: vague briefs converge on the generic AI look everyone now recognises; specific briefs — materials, lighting, named styles — produce distinctive output. Specificity is the anti-slop, again.

Reference more, describe less. Here’s the counterintuitive bit: as the models have improved, the practical guidance has converged on prompts getting shorter — thirty to sixty words is a common working range — with reference images carrying the load for style and character consistency. References now beat description. Supplying one or two images that capture the look you want is the single highest-leverage input available.

Structure the prompt by intent. Start with what you need to accomplist most from the prompt — the core element that you want to generate as opposed to the background or minutea that you could fix iteratively. The first words of your prompt set the tone, next the crux, lastly the optionals.

Be careful about using Text. Text on generated creatives, by far, are the ones that fail the most. While the models have gotten really good at rendering text inside creative content like images or illustrations etc, the structural rendering of characters, placement in the spatial view of the creative still trip up models the most. The key rule here is to make sure to supply the text exactly rather than have the AI make it up and to precisely describe the positioning, font style, sizing and placement upfront along with the crux of the image than at the end. But be ready to iterate and fine tune along the way.

Change one thing at a time. The most repeated practitioner tip in the research, and the cheapest. Iterate on a single variable per pass — the lighting, then the composition, then the palette — so you can actually tell what moved the result. That’s how you build a mental model of the controls instead of pulling a slot machine lever.

Put together, a balanced brief looks like this:

“Hero image for an artisan coffee brand. Warm, editorial style — think a high-end café shot. A single ceramic pour-over on a wood counter, soft morning side-light, muted earth tones, shallow depth of field, minimalist. [Attach 1–2 reference images.] Generate 3 variations, then I’ll refine.”

Treat the first output as a draft, not a verdict — expecting one-shot perfection and giving up when it misses is how most people abandon the tool one round too early. And do not skip the finishing pass: the small artifacts and off details that scream “AI made this” are exactly what your eye exists to catch. One responsibility note, kept short because it matters more than it lectures: avoid style-mimicking living artists, and check rights and provenance before using outputs commercially. It’s the creative-work mirror of the previous post’s verification beat.

The model gives you a strong draft. Your taste and your finishing pass are what make it not-slop.

Photo Editing — The Restraint Is the Skill

Here is the underappreciated sibling of image generation, and arguably the more useful one for real work: editing images you already have. Background removal, inpainting — adding, removing, or replacing elements — retouching, upscaling, style transfer. The same industry analysis behind the editing-loop shift goes a step further and argues that AI image editing now matters more than generation for real creative work, because the conversational “fix the weak parts” loop is where outputs become usable. This section deserves more weight than it usually gets.

The calibration dial here is scope discipline. Too far toward AI: “make this photo better” — and you receive an over-processed, plasticky, faintly uncanny result in which the model has helpfully “improved” several things you wanted left alone. Too far toward yourself: spending an evening on fiddly manual edits AI could do in seconds. The sweet spot is targeted, well-scoped instructions, worked region by region, with a human eye on whether the untouched parts actually stayed untouched.

The crucial practice: name what to change and what to preserve. The failure mode of editing models is changing more than you asked — inpainting in particular can subtly alter adjacent areas — so the preservation clause is not politeness, it’s the instruction doing the protecting. Be instructive and ask it to enforce strictly.

“Remove the cardboard box in the background behind the product, keep the product, the table, and the lighting exactly as they are. Don’t retouch the product itself. Then I’ll review before any other changes.”

Two smaller habits worth adopting. First, work iteratively on one region at a time — the photographic cousin of “change one thing at a time.” Second, inspect at full size before using anything. Artifacts hide at thumbnail scale; the classic tells — warped hands, mangled text, melted edges — only reveal themselves when you zoom in. Trusting an edit you’ve only seen small is how those tells end up on a billboard.

And the brief, non-negotiable line on ethics: editing for craft is fine; altering documentary, journalistic, or evidentiary images in ways that mislead crosses into deception. The tool doesn’t know the difference. You do. Importantly be very cautious of the ownership of the original and respect all copyright.

Tell it exactly what to change and what to protect. The restraint is the skill.

Video and Motion — A Concepting Tool, Not a Film Studio

Let’s be fair to the technology first, because the leap has been genuinely remarkable. The blurry fifteen-second clips with melting fingers of 2024 have given way to models producing up to 4K footage with synchronized native audio and multi-shot sequences with tools now able to generate pseudo-realistic videos, ads etc fully with generative models. Remarkably better — and still bounded. Both halves of that sentence matter, and most coverage only gives you one of them.

The bounds are specific, and knowing them is what sets realistic expectations: practitioner benchmarks consistently flag hand and face artifacts in close-ups, camera logic that drifts between cuts, and soft, unconvincing motion on fast action. Which is why “make a 2-minute product video” — the too-much-control end of the dial — produces incoherent, drifting, artifact-ridden output that ignores what the medium can currently do. The opposite error is dismissing AI video entirely and missing where it is already excellent.

So where is it excellent? The honest two-column answer: strong for concepting, paid-social variants, previsualization, abstract and mood scenes, and hook testing. Weak for legally sensitive claims, exact product accuracy, regulated industries, human likeness, and final product demos. The pattern underneath: reach for it where approximate beats exact.

The working method follows from the bounds. Keep shots short and motion simple. Specify like a director — camera, lighting, mood, pacing, not just subject. Use reference frames for consistency. Generate several takes and select ruthlessly, the way a photographer shoots a roll to keep one frame.

“5-second cinematic B-roll: slow push-in on a steaming coffee cup on a café table, warm morning light, soft focus background, gentle steam motion. Generate 4 takes; I’ll pick the cleanest and we’ll refine the camera move.”

The sharpest heuristic I found in the research, and the one that cuts through the hype cycle: judge the technology by your own usable-take rate, not by cherry-picked demos. A demo reel tells you what’s possible once; your take rate tells you what’s dependable for your work.

A word on tools, deliberately light: this is the fastest-moving corner of the entire series. The landscape reshuffled hard in 2026 — OpenAI announced the discontinuation of Sora’s consumer app around late April, while the likes of Google’s Veo 3.1, Kling 3.0, and Runway Gen-4.5 were being cited as the leaders, with multi-image character consistency becoming table stakes. Treat any tool name, including those, as a dated snapshot and verify what’s current when you’re choosing. The principles — short shots, director-grade specificity, references, generate-and-select — will outlive every name on that list.

And for anything you actually publish: review for rights, disclosure, and brand policy, copyrights. Every time. The speed of generation is precisely why the review can’t be skipped.

Score your usable-take rate, not the demo reel.

Animation and Motion Graphics — Timing Is Human

A short section, honestly held — because this corner of the field is younger than its siblings. Animation and motion-graphics AI is less mature and less benchmarked than image or video generation; most credible material treats it as an extension of video tooling rather than a settled category. So I’ll claim less and frame it the way the evidence supports: as a pipeline assist, not an end-to-end animator.

First, the distinction from the previous section, because they’re easy to blur: video generation produces filmed-style footage; animation is designed motion — logo stings, explainer movement, looping graphics, kinetic type — where the timing is authored, not captured. And that distinction is exactly where the dial sits. Ask AI to “animate my brand explainer” and you get generic motion-template output with no narrative timing and no brand specificity — motion that happens to the design rather than for it. Hand-animate everything yourself and you’re doing in-betweens a machine could grind through.

The sweet spot: AI handles the labour-intensive and exploratory stages — motion drafts, in-betweening, style variations — while you direct timing, narrative, and brand coherence. Because here is the durable insight: timing, easing, and intent are precisely what current tools handle worst and humans handle best. That’s not a limitation to mourn; it’s the cleanest division of labour in the whole chapter.

“Animate this logo [attach]: have the icon draw itself in over ~1.5s with smooth ease-out, then the wordmark fades up. Clean, premium feel — no bounce. Give me a couple of timing variations to compare.”

Notice the brief defines what moves, the easing, the duration, the feel, and asks for timing variants to compare — the directorial decisions stay in the prompt, and therefore in your hands. The recurring technical gremlin is consistency: style and character drift across frames, inherited from video generation. The mitigations are the same — reference inputs and short, controlled segments. And never skip the human pass on timing. It is, quite literally, what separates designed from generated.

AI can generate motion. Timing and intent are what make motion feel deliberate — and those are yours.

Presentation Decks — The Chapter’s Stress Test

And so to the culmination — the use case where everything in this post converges. A deck is text, structure, and visuals in one artifact, which makes it the perfect stress test for the balance theme at full complexity. It is also where the reviews of AI deck tools land almost word-for-word on this series’ thesis, which I find quietly delightful.

Here’s what those experts consistently report about the current generation of tools: bespoke products, excellent at fast, well-structured first drafts and at avoiding wall-of-text slides; weak on generic language, stock-photo-quality visuals for niche topics, and the occasional redundant slide repeating a point with slight rewording. What do the experts say in this matter? Inject your specific data, examples, and your own voice. Let the AI generate, ideas, copy and graphics. In other words: the human pass is the value-add. Their words, our thesis.

The dial, then. Too far toward AI: “make a pitch deck for my startup” — a slick-looking but generic deck full of placeholder language and stock visuals; polished emptiness that reads as AI slop from the second slide. And watch for the deck-equivalent of hallucination: invented stats and claims conjured to fill a template. Too far toward yourself: building every slide by hand when AI could nail the structure and first-draft layout in minutes. The sweet spot: AI generates structure, layout, and first-draft copy; you inject the real data, the specific examples, the narrative voice, and the design polish for the audience that matters.

Three practices carry this use case. Work outline-first — lock the narrative structure before generating a single slide. (Chapter 1 readers will recognise this as the deck version of section-by-section drafting; the consensus workflow across tools says the same.) Feed it your real content and expect to replace generic copy with specifics. Match the tool to the stakes — fast AI decks are great for internal work, brainstorms, and web-shared content, and riskier for high-polish client or investor decks where every pixel is being judged.

“Build a 10-slide deck from this content [paste/upload]. Audience: angel investors. Narrative: problem → solution → traction → ask. Use my real metrics exactly as given; flag any slide where you’d normally invent a number. Clean, modern, minimal text per slide. I’ll refine copy and visuals after.”

That “flag any slide where you’d normally invent a number” clause is relevant from the previous posts’ flag-gaps-don’t-fill-them instruction wearing a suit. This is exactly the same but in a different context

One concrete, oft-overlooked gotcha the research surfaces: export fidelity. Web-first AI decks frequently lose layout and editability when exported to PowerPoint — layouts flatten, text becomes uneditable. Decide your output format before you build, not after. And as with video, the tool landscape moves quickly — one major player, Tome, exited presentations back in 2025 — so anchor on the principles and re-verify the tool list whenever you’re choosing.

The finishing pass for a deck is the whole chapter in miniature: tighten the copy (specificity), swap the weak visuals (taste and references), cut the redundant slides (editing loop), check the export (verification). The closer the audience, the heavier that pass should be.

Use AI to skip the blank page and the formatting grind — not to outsource the message.

Epilogue — Where Slop Becomes Craft

Five creative use cases, one consistent shape. Specific briefs beat vague ones. References beat description. Iteration beats one shot expectations of success. Scope what changes and protect what doesn’t. And in every single case, the final transformation — the one that turns a competent draft into something with a point of view — happens in the human finishing pass. That pass isn’t the overhead of the workflow. It is the workflow’s point. Every time something doesn’t work, iterate. Each iteration is a learning and with practice the iterations reduce without any loss of fidelity in the outcome.

I started this post with an anonymous bowl of pasta, so let me end with what fixed it: a tighter brief, two reference images, three rounds of one-change-at-a-time edits, and a closing inspection at full size. Nothing magic. Just the dial, set deliberately, plus a loop. The image that came out the other end finally looked like my picture.

If you’ve found a creative-AI habit that beats the ones here — a referencing trick, an editing-loop rhythm, a finishing-pass checklist — I’d genuinely like to hear about it in the comments. Craft is communal, and the loop gets better the more of us compare notes on it.


메타데이터
post_id
cefeb4e96afd
slug
finding-balance-with-ai-creative-generation-cefeb4e96afd
url
https://medium.com/@sundar.krishnamurthy/finding-balance-with-ai-creative-generation-cefeb4e96afd
canonical_url
https://medium.com/@sundar.krishnamurthy/finding-balance-with-ai-creative-generation-cefeb4e96afd
author_url
https://medium.com/@sundar.krishnamurthy
status
ok
fetched_at
2026-06-20 20:29:01