How I Train Artist LoRAs: Not Style Filters, but Visual Logic
How I Train Artist loras: Not style Filters, but Visual Logic
Photo by Nik on Unsplash
When I first started training artist loras, I thought the most important questions were technical ones.
Is the learning rate too high? Are the training steps not enough? Should I increase the network dimension? Is the dataset too small? Are the captions detailed enough?
Of course, all of these things matter. Especially when training Flux loras,parameters, resolution, repeats, caption quality, and dataset structure can all noticeably affect the final result. But over time, I began to realize that what truly determines the upper limit of an artist lora is often not the technical setup itself, but a more fundamental question:
WHAT EXACTLY DO YOU WANT THE MODEL TO LEARN FROM THIS ARTIST?
It sounds simple, but during training, it is very easy to overlook.
Today, whether you are training for sd1.5, sdxl, pony, or flux, lora tutorials are everywhere. YouTube, Reddit, Civitai, or any AI communities, there are plenty of people explaining how to prepare a dataset, how to write captions, how to set parameters, and how to run training with kohya or other tools.
But many tutorials are mainly about “how to train a lora”.
That’s not exactly the same as “how to train the visual system of a specific artist.”
If your goal is only to make a style filter, the process is not that complicated. You can collect a batch of images, use Wd Tagger, Joycaption, or similar tools to generate captions (or tags), adjust the learning rate, steps, and other parameters, let the GPU finish the training, and get a .safetensors file. Then you load it into WebUI, ComfyUI, or Forge and start testing.
In many cases, this will indeed give you a lora that “kind of looks like” the artist.
But that resemblance is often only surface-level.
The model may learn a certain color tendency, some repeated patterns, a type of brush texture, or recurring subject matter. For example, if you train a Klimt lora, it may start generating gold and decorative patterns everywhere. If you train a Rousseau lora, it may keep producing jungles, plants, and animals. If you train a Gérôme lora, it may constantly generate classical costume, stone columns, drapery, and Orientalist interiors.
This is not useless. As a visual filter, this type of lora can absolutely produce a strong recognizable effect.
But that is not really what I am after.
What interests me more is a different question:
If this artist were alive today, how would they handle modern subjects?
I do not want the model to simply copy what the artist already painted. I want it to preserve the artist’s visual decisions when facing subjects the artist never actually painted.
For example, what if Gérôme painted a modern laboratory, an intimacy relationship, a contemporary portrait, or a scene that does not belong to 19th century academic painting at all? If the lora only learned “classical clothing,” “stone walls,” “exotic interiors,” and “historical themes,” then it has not really learned Gérôme.
It has only learned a few easily recognizable symbols.
I have made this mistake myself.
At first, I thought that if the captions were more “advanced” and more abstract, the model would better understand the visual structure I wanted. So I wrote a lot about visual order, optical control, spatial tension, surface finish, and so on. It sounded correct, almost like art criticism.
But the training result was not good.
The model did not automatically understand those abstract concepts. Sometimes it only learned a vague classical atmosphere, or turned everything into a generic “oil painting” look. That made me realize something: for lora training, abstract art criticism is not the same thing as learnable visual information.
The model needs anchors.
It needs to know what is actually in the image: bodies, clothing, walls, floors, fabric, metal, stone, shadows, the distance between figures, the direction of their gaze, and the layers of space.
At the same time, it also needs to know how these things are organized and rendered: how the edges are controlled, how light falls across the structure, how materials are differentiated, and how figures are placed within the space.
So later, I slowly changed my captions from “style descriptions” into something closer to “visual analysis.”
This is now the core idea behind how I train artist loras.
I do not want the caption to simply tell the model:
“This is an oil painting in a certain style.”
I want the caption to tell the model how the figure stands, how the body weight is distributed, how the space is framed by walls, steps, columns, or the ground plane, how light unifies skin, fabric, and stone, and how different materials remain clear without breaking the whole image apart.
It sounds tedious, but this is exactly the difference between an artist lora and a simple style filter.
Take Gérôme as an example. If we understand him only through subject matter, it is very easy to reduce him to classical figures, Orientalist scenes, historical narratives, and smooth surfaces. But those are only the outer layer of his work. What is more valuable is how he controls the image: his spaces are often stable, the figures are precisely placed, the edges are extremely clear, materials are carefully differentiated, and skin, fabric, stone, and metal all operate within the same controlled optical system.
This cannot be solved by simply writing “in the style of Jean-Léon Gérôme.”
Likewise, Klimt is not just gold and patterns. Many Klimt loras easily become “gold leaf everywhere,” but his visual logic also includes how the figure’s silhouette is absorbed into decoration, how the body becomes a symbol inside a flattened space, and how pattern functions as both background and composition.
Rousseau is not just jungle either. What’s more interesting is his frontal composition, flattened space, almost childlike sense of order, and the strange but stable way he compresses plants, figures, animals, and background into a single visual plane.
So I think the first step in training an artist lora is not collecting images or adjusting parameters.
The first step is asking: What are the transferable features of this artist?
In other words, if you remove the original subject matter, what can still remain?
If a Gérôme lora can only generate ancient scenes, then maybe it has only learned the subject matter. But if it can handle modern figures, modern interiors, and modern objects while still preserving that precise, cold, highly finished academic visual control, then it starts to move closer to what I actually want.
Dataset selection should also follow this logic.
I do not think more images are always better. Especially when training a lora for a specific artist, what matters most is not quantity, but structure. You need to know why each image is included. Is it there for the model to learn figures? Space? Surface? Composition? Pattern? Edges? Or the staging and viewing relationship between characters?
If all the images come from the same type of subject, the model may easily confuse the artist’s style with the subject itself.
If all the captions only describe objects, the model may struggle to learn the deeper visual organization.
If the captions are too abstract, the model may not be able to find stable anchors at all.
So my usual approach is to divide the dataset into different visual layers, instead of simply throwing every image into one folder. Some images are better for learning full composition. Some are better for figure construction. Some are better for architectural space. Some are better for surface finish. Others are better for learning the staging, gaze structure, and relationship between figures.
This is not complexity for the sake of complexity.
The purpose is simply to make the training intention clearer: What do I want the model to learn from this group of images?
Large language models can help a lot in this process. They can help summarize an artist’s characteristics, analyze artworks, provide art historical context, and assist with caption writing. But I think the most important thing is that you can’t fully hand over judgment to AI.
AI can easily produce art criticism that sounds correct but is actually empty.
For example: “strong composition,” “dramatic lighting,” “rich details,” “masterful brushwork.” These phrases are not wrong, but they may not be very useful for training. They are too generic. You can apply them to almost any painter.
A useful caption should stay close to the image itself.
Do not just say “the lighting is dramatic.” Say how the light falls across the face, shoulder, fabrics, and wall. Do not just say “the composition is stable.” Say how the figures, ground, architecture, and background form that stability. Do not just say “the materials are realistic.” Say how skin, fabric, metal, and stone differ in edge, reflection, and tonal transition.
These captions may not sound beautiful, but they are more trainable.
The more I do this, the more I feel that training an artist lora is process of translating a way of seeing into something the model can learn.
You are not telling the model, “This image looks good.”
You are telling it why the image is organized this way, and how that organization might still work when the subject changes.
In the end, parameters still matter. Learning rate, steps, dim, alpha, repeats, resolution, buckets, optimizer, all of them can affect the result. But if the dataset structure and caption logic are not clear from the beginning, parameter tuning often becomes a form of damage control.
Parameters can decide whether a lora is stable.
But the dataset and captions decide what it is actually learning.
That is my current view of artist loras.
I am not that interested in training a model that merely produces screenshots that look like they came from a certain painter’s portfolio. I am more interested in training a model that can preserve the artist’s visual logic when facing new subjects.
Of course, this goal is difficult, and it can never be achieved perfectly. A lora is still a low-rank adaptation on top of a base model. It does not retrain an entire painting model from scratch. It will always be limited by the base model, dataset size, caption quality, and inference prompts.
But the limitations are also what make it interesting.
Because you are forced to keep asking:
What is truly central to this artist? What is just subject matter noise? Which symbols will cause overfitting? Which descriptions can help the model transfer the visual logic to new content?
For me, the most interesting part of training artist lora is not recreating the past.
It’s a kind of visual thought experiment:
If this artist were facing today’s world, how might they reorganize it?
That is why I keep making these loras. Not because I think AI can replace artists, but because it offers a strange and fascinating way to rethink what an artist’s visual language is actually made of.
If you also want to train a lora for a specific artist, my suggestion is: do not rush to open the training script.
Look at the works first. Look at many of them. Do not only look at subject matter. Do not only look at color.
Look at how the image is organized, how figures are placed, how space is built, how materials are separated, how edges are handled, and how light unifies the whole image.
Then decide what your dataset should contain, what your captions should say, and how your parameters should serve that goal.
Because the technical workflow can be copied.
But the way you look at an artist is much harder to copy.
If you are interested in this training approach, you can also check out some of the models I trained with a similar method:
Civitai: https://civitai.red/user/mariano_art
메타데이터
- post_id
- 5bb9449a0701
- slug
- how-i-train-artist-loras-not-style-filters-but-visual-logic-5bb9449a0701
- url
- https://medium.com/@Mariano_S/how-i-train-artist-loras-not-style-filters-but-visual-logic-5bb9449a0701
- canonical_url
- https://medium.com/@Mariano_S/how-i-train-artist-loras-not-style-filters-but-visual-logic-5bb9449a0701
- author_url
- https://medium.com/@Mariano_S
- status
- ok
- fetched_at
- 2026-06-09 15:37:30