Cheaper AI, better results
7 ways to burn fewer tokens and less money
Design + AI
Cheaper AI, better results
7 ways to burn fewer tokens and less money
Subscriptions are a blessing and a curse at the same time. On the one hand, you receive the top, always up-to-date tool and many frequent improvements. On the other hand, you need to make recurring payments, and once you pause them, you lose access to the solution.
Recently, when I was reviewing my tooling budget (related only to the design and product building category), I realized that as a creative pro, I could easily cross $500 in monthly payments for AI-related tools alone.
That amount paid on a monthly basis by a single designer is insane if it does not reflect an increase in revenue. Not so long ago, I remember designers complaining about Figma or Framer raising their subscriptions by a couple of dollars a month.
What’s more, when I say that we live in an AI bubble, one of the facts confirming that statement is that literally all major AI companies, like OpenAI or Anthropic, burn far more money than they earn on their solutions.
One day, investors will ask for their money, and those companies will need to show a profit. Imagine you paid $200 a month for your AI-assisted design or coding plan subscription, and after a year you can’t imagine working any other way. However, now the company raises the cost of this plan to $400 a month, or to generate a stable profit and reflect energy costs, they decide to set the subscription at $600 a month.
Feels unbelievable?
When you look at OpenAI and Anthropic reports from 2025, they would need to raise their plans by 50 to 70% (depending on sources) to become profitable.
All of this led me to a conclusion that we should seriously think about how to cut the costs of AI tooling, without compromising the quality of the results.
In these Note, I want to show you a bunch of techniques I already apply to reduce the costs of your AI tools a lot. In some cases, even to zero.
Let’s begin.
1. Limit the toolset
The first and the most obvious advice. I already wrote an extensive note about how and why you should limit the number of paid AI tools. You may read it here. You don’t need to use every possible combination. In most cases, as a professional, it’s completely enough to pay for two or three tools regularly.
It will not only save you money but also reduce the FOMO that you need to test the feature or model released two hours ago. No, you don’t have to test it.
Most creative professionals won’t feel a huge difference when working with Opus 4.7 and GPT 5.5, so pick one.
2. Think before acting
The one consistent thing that improves the quality of the output and reduces the cost of further prompting and iterations is preparation.
There is a known saying attributed to Abraham Lincoln:
“Give me six hours to chop down a tree and I will spend the first four sharpening the axe.”
This is valid not only for this type of hard physical work but also for collaboration with AI. Here are some real cases that should encourage you to rethink your workflow:
When you vibe code, turn on plan mode in Cursor, Codex, or Claude first. Even before that, take the time to describe the context and create a PRD. Put that context and tech stack in CLAUDE.md or AGENTS.md (depending on the tool you use). For the UI part, the minimum should be creating a DESIGN.md file, but ideally, you should set up the design system upfront.
When you ask AI about benchmarking or preparing a report from multiple documents, think of the proper structure and insights that the document should include.
Want to generate an image or a video? Don’t treat the AI tool like a casino slot machine and ask it to prepare countless series of iterations, hoping that “the next one will be good enough.” First, sketch the result you want to achieve (even with pen and paper), and think of the composition, lighting, materials, and styling. If you lack the knowledge on how to describe them to an LLM, look for real photos showing that element and ask AI to describe it for you.
All of this preparation may seem like a delay, but when you look at the wider perspective, you realize that you saved time and money, because the number of bugs and iterations was significantly reduced.
3. Take care of prompt hygiene
When interacting with AI, you have the illusion that you interact with a conscious being. However, it’s still a computer. The more precise the command you enter, the higher the quality of the output you should get.
If you haven’t done so yet, learn to prompt efficiently. The simplest and most efficient framework for better communication with an LLM is RACE:
- Role: define the role AI should take.
- Action: explain what task the AI should perform.
- Context: provide background, constraints, audience, or goals you want to achieve.
- Expectations: specify the needed output format, quality, tone of voice, constraints, or success criteria.
This means that instead of writing “Generate onboarding flow for invoicing web app.”, you should describe the task like in the following example:
“Act as a senior UX designer specializing in fintech onboarding and conversion optimization. Analyze the onboarding flow and suggest improvements that reduce drop-off and increase trust. The app helps freelancers track invoices and taxes. 15 percent of users currently abandon the flow during identity verification. The target audience is first-time freelancers aged 22 to 35. The onboarding flow has 5 screens and currently takes about 4 minutes to complete. The link showing the current flow [link].
4. Support yourself with temporary tools when vibe coding
With this technique you build correct expectations around the output, and narrow the area of exploration for AI, and finally again get better quality result with less iterations.
When building a website or an app, or trying to polish the prototype, there might be a lot of details that need to be refined or experimented with (iterations and experiments with different variants are a natural part of the design process).
Instead of prompting back and forth, which will cost you time and money, to make precise adjustments or experiment with the details, ask AI to build a temporary tool that would help you do that.
In many cases, it would look like the Tweaks panel known from Claude Design, with a few sliders, dropdowns, or a color picker.
Claude Design Tweaks panel — parametric design in action
Sometimes, you may simply ask AI for a tool to drag an image within a frame and then copy the values that should land on the final web page.
5. Simplify output communication
While the commands you give to AI should be precise, as an output, we often get friendly, casual messages. While they make AI feel more human, the truth is that we also pay for them.
In some cases, you do not care if AI says: “Great, now I have full context. I analyzed the attached documents. Here is the report created based on your instructions:”.
It would be completely fine to get a reply like: “Report done:”.
That idea was turned into a skill that you may use with top LLMs. It’s called “Caveman.” The author promises that using it will help you reduce up to 75% of output tokens.
To learn more, check the official repository of the Caveman project.
6. The right model for the right task
As I am writing this note, Opus 4.7 and GPT 5.5 are the top models. When you use them, you literally feel the difference in the results. However, they are also the most expensive ones when it comes to usage, and you may run out of your token limit quickly.
What you must realize is that for most AI tasks, lower models like Sonnet will give you exactly the same quality of result. It’s a good default for most coding, analysis, writing, and chat requests. For simpler tasks, we should even switch to Haiku, which is perfect for classification, extraction, simple Q&A, or formatting.
If you run out of your plan quickly but you didn’t build the habit of switching models depending on task complexity, now is your time to leverage that.
7. Use local models
The ease of use for paid solutions is incredible. You just enter the web URL or download an app, and that’s it. Many people think that launching AI locally is a domain reserved for software developers and computer geeks.
However, the truth is different. Actually, my 10-year-old son was able to install Ollama (a tool to run LLMs on a device) and Gemma 4 (Google’s local model) on my MacBook. Obviously, I assisted him.
Ollama lets you simply run LLMs on your machine
It is much easier than you would expect and gives you 100 percent free AI (except for electricity costs), with no token limitations.
Sure, the local models are slower, but for most tasks, it’s a matter of seconds or minutes (depending on your machine). However, many regular tasks delegated to Sonnet and Haiku level models may be a good enough solution.
If you would like to start your journey with a local LLM, simply ask GPT or Claude to guide you through the installation of Ollama and help you choose a Gemma model that will fit your machine (that varies depending on your device RAM).
Personally, I believe that as technology advances, local models will become the go-to solution for 90% of AI tasks. You just need a proper hardware setup, and Apple devices are well-tailored for that.
Next steps
If you feel like you spend too much on your AI tools, take every point I described here and follow them one after another. You should feel the difference when you see your bill next month.
When you do not treat the money cost as the main factor, I still recommend you do three things:
First, be more conscious about your prompting technique to get better results.
Then, experiment with building temporary tools when vibe coding. You save time and gain much more control over the output.
Finally, try to install a local model. If cost reduction is not a value for you, then 100% privacy may be convincing.
Let me know how it works for you!
This article was released originally as the part of **Thalion’s Notes**, the newsletter with over 8k reades, where I also share reflections deeper than simple UI design tutorials. On Saturdays I send notes on reasonable use of AI, productivity, creativity and more meaningful growth.
메타데이터
- post_id
- 35b22872773d
- slug
- cheaper-ai-better-results-35b22872773d
- url
- https://medium.com/design-bootcamp/cheaper-ai-better-results-35b22872773d
- canonical_url
- https://medium.com/design-bootcamp/cheaper-ai-better-results-35b22872773d
- author_url
- https://medium.com/@uxmisfit
- status
- ok
- fetched_at
- 2026-06-09 15:37:30