What I learned running a vision model locally on a 16GB MacBook Air
Reading stamps on scanned forms with Qwen3-VL and MLX — from memory limits and quantization to pixels, visual tokens and bandwidth math.
What I learned running a vision model locally on a 16GB MacBook Air
Reading stamps on scanned forms with Qwen3-VL and MLX — from memory limits and quantization to pixels, visual tokens and bandwidth math.
I wanted to answer a simple question: can I use my laptop to read the stamps on scanned documents — name, address, phone — using an open-source vision model, fully local?
Most guides to running LLMs locally on a Mac are limited to performant machines with 64GB of memory or more. Impressive numbers — but not much help if you’re sitting in front of the Mac most people actually own.
Mine is a MacBook Air M4 with 16GB — no fan, no Pro chip. This article follows the whole process of finding the best setup for reading stamps on it: which models actually fit, how fast they run, which image settings worked, and why. Along the way, it covers what quantization actually does to a weight, why DPI is a red herring, and where the time really goes when a model reads an image.
Why not vLLM and Docker
My first instinct was the setup I’d use on a server: vLLM in a Docker container. On a Mac, that’s a dead end. Docker runs containers inside a Linux virtual machine, and that VM has no access to the Metal GPU. Everything would run on the CPU.
So I went native, with MLX — Apple’s machine learning framework, built for unified memory — and mlx-vlm, the package that runs vision-language models on top of it.
pip install mlx-vlm
One thing worth knowing if you also use LM Studio: it ships its own MLX runtime (mlx-llm-mac-arm64-apple-metal-advsimd), completely independent from the mlx-vlm you install with pip. Two runtimes, two versions. If a model works in one and not the other, that's usually why.
What fits in 16GB
16GB sounds like a lot until you start subtracting. macOS only lets the GPU use roughly 65–75% of unified memory — on this machine, about 10–11GB. That’s the real budget for the model, its context and the image.
Here’s what I tested, all MLX builds from mlx-community:

Which raises the obvious question: what’s the difference between 8-bit and 4-bit, really?
Quantization, without the hand-waving
A model is billions of numbers. How you store each one decides how much memory it takes.
BF16 is the usual starting point. Each weight takes 16 bits: 1 for the sign, 8 for the exponent, 7 for the mantissa. Same range as a 32-bit float, less precision, half the memory.
Quantization goes further: store each weight as a small integer, plus enough information to approximately recover the original value. MLX’s 4-bit scheme works in groups of 64 weights. Each group gets a scale and a bias, and every weight becomes an integer from 0 to 15. No calibration data needed.
Here’s the idea on a toy group of 4 weights instead of 64:
weights: 0.12 -0.53 0.31 0.90
min = -0.53 → bias
max = 0.90
scale = (0.90 - (-0.53)) / 15 = 0.0953
quantized: round((w - bias) / scale)
7 0 9 15
restored: q × scale + bias
0.137 -0.530 0.328 0.900
error: 0.017 0.000 0.018 0.000
Small errors, a quarter of the memory. At 8-bit, you get 256 levels instead of 16, so the errors shrink further.
AWQ is a smarter variant you’ll see on Hugging Face: 4-bit integers with per-group scales, but it uses calibration data to find the weights that matter most — the ones connected to large activations — and protects them from rounding.
If you open a quantized MLX model’s .safetensors file, each quantized weight isn't one tensor anymore. It's three:
TensorWhat it holdsweightThe integers, packed into 32-bit values: 8 per U32 at 4-bit, 4 per U32 at 8-bitscalesOne scale per groupbiasesOne bias per group
That’s why a “4-bit” model is slightly bigger than a quarter of its BF16 size: the scales and biases come along.
From pixels to tokens: what the model actually sees
This is where most of my wrong assumptions were, so it’s worth slowing down.
An image is just a grid of numbers. Each pixel is three values — red, green, blue. That’s it. There’s no DPI inside the model, no inches, no paper size. DPI is only a conversion rate between a physical size and a number of pixels:
pixels = inches × DPI
An A4 page is about 11.7 × 8.3 inches. At 72 DPI, that’s roughly 842×596 pixels. At 100, about 1170×830. At 300, about 3500×2480. Same page, same paper — but once it’s a file, only the pixel count matters.
The model doesn’t read pixels one by one. Qwen3-VL first resizes the image so its total pixel count lands between a min_pixels and a max_pixels limit, keeping the aspect ratio and rounding each side to a multiple of 32. Then it cuts the image into 16×16 pixel patches and merges them in groups of 2×2:
image (W × H)
│ resize to fit min/max pixels, sides rounded to multiples of 32
▼
16×16 px patches
│ merge 2×2
▼
visual tokens → 1 token = 32×32 px
So the token count is simply the image area divided by 1,024. My final image, 1792×1248 pixels, is 56 × 39 blocks of 32 — exactly 2,184 visual tokens.
Tokens are the budget. Every visual token has to be processed before the model writes a single word, and each one takes memory in the context. More pixels means more tokens, a longer wait, and more memory — so resolution is a trade, not a free upgrade:

The metric that actually matters is text height in pixels, after the resize. Not DPI, not megapixels. Here’s the rule of thumb I ended up with: around 20 pixels of text height is comfortable, below 10 the model starts guessing.
You can estimate it on a napkin. A point is 1/72 of an inch, so a 10-point line of text is about 0.14 inches tall:

That’s the whole reason upscaling helped: it doesn’t add information, it gives each character more patches to be read through. And stamp text is often smaller and more degraded than printed form text, so it needs that margin more than anything else on the page.
Real scans are worse than you think
My first tests used PNGs I had rendered from PDFs myself: 842×596 pixels — an A4 page in PDF points, rendered at 72 DPI. Too small, so the obvious fix seemed to be rendering at 300 DPI.
Before doing that, I checked what was actually inside a real scanned PDF:
pdfimages -list scan.pdf
A scanned PDF is usually just a container around an image. And that image was far more modest than I expected: 1176×820 pixels, at 100 ppi, stored with a 16-color indexed palette. That changes everything about rendering. The scanner captured 1176×820 pixels of information, and nothing more. Rendering the page at 300 DPI doesn’t recover detail — it interpolates pixels that were never captured, producing a bigger file with the same information and more tokens to process.
So the right workflow is three steps:
- Inspect what’s really in the PDF with
pdfimages -list - Extract the image as-is, no re-rendering:
pdfimages -png scan.pdf page
- Resize deliberately, based on the text height you need — not on a DPI number that sounds high.
What the model got right — and wrong
A typical stamp looks something like this (invented example):
STE EXAMPLE
5, Rue Exemple - Casablanca
Tél : +212 5XX XX XX XX
N° XXX
Three findings stood out.
Cropping helps, even with the same pixels. A crop of just the stamp was read better than the full page, at the same resolution. On the full page, the stamp competes for attention with all the printed form text around it.
Where pixels run out, the model invents. My prompt explicitly said to write [illegible] when text couldn't be read. Where the resolution wasn't enough, the model didn't do that - it produced plausible local text instead. A believable street name is worse than an honest "[illegible]", because nothing flags it as wrong.
Upscaling helps — up to a point. Upscaling gives each character more patches, and reading improved. But pushing further, to a 3.9-megapixel floor, magnified the scanner’s dithering too . A “N°” became “YN”. There’s a sweet spot, and past it you’re amplifying noise instead of text.
The numbers
Final setup: Qwen3-VL-4B-Instruct at 8-bit, full page upscaled about 1.5× with a 2.2-megapixel floor — 1792×1248 pixels, 2,184 visual tokens.
One measured run:

Two things jump out from this table, and both come down to how the hardware works.
Decode speed has a ceiling you can compute. Generating each token means reading the model’s weights from memory. So the upper bound is:
max tokens/s ≈ memory bandwidth ÷ bytes read per token
The M4 has about 120 GB/s of memory bandwidth. The model is about 4.5GB. That gives a ceiling around 27 tokens/s. I measured 21.0 — roughly 94 GB/s of effective bandwidth, about 78% of the theoretical ceiling. The napkin math holds.
Images cost prefill, not decode. Prefill took about 75% of the total time. Visual tokens are processed once, in parallel, and that work is limited by compute — not bandwidth. It’s also why upscaling is expensive: a 1.52× upscale in each dimension multiplies the visual tokens about 2.3×. More pixels means a longer prefill, even if the answer is the same length.
Watching it happen
Timings only mean something if the machine isn’t fighting you in the background. For a live view I used mactop, a terminal dashboard for Apple Silicon — GPU usage, power, memory and temperatures in one screen:
sudo mactop
macOS also has good built-ins:

The one to watch is swap. When memory runs short, macOS first compresses pages in RAM, then moves them to the SSD:
RAM ──▶ compressed memory ──▶ swap on SSD
fast slower much slower
If swap grows during a run, your timings are ruined — part of the model is now being read from disk instead of memory, and the bandwidth math no longer applies. Check swap before and after every benchmark.

mactop dashboard
Takeaways
- On a Mac, go native. Docker has no GPU access; MLX does.
- Your real budget is ~65–75% of RAM. On 16GB, plan for 10–11GB.
- Quantization is a trade you can reason about.
- DPI is a red herring. Count pixels — specifically, text height after resize.
- Every 32×32 pixels is a token. Resolution is a budget, not a free upgrade.
- Extract scans natively. Rendering at 300 DPI invents nothing useful.
- Crop to what matters, and upscale moderately — too much magnifies noise.
- Don’t trust a prompt to prevent hallucinations. Where pixels run out, the model may fill the gap.
Support my work: buymeacoffee.com/red1
메타데이터
- post_id
- c19c3a62df59
- slug
- what-i-learned-running-a-vision-model-locally-on-a-16gb-macbook-air-c19c3a62df59
- url
- https://medium.com/@redouanekarzazi/what-i-learned-running-a-vision-model-locally-on-a-16gb-macbook-air-c19c3a62df59
- canonical_url
- https://medium.com/@redouanekarzazi/what-i-learned-running-a-vision-model-locally-on-a-16gb-macbook-air-c19c3a62df59
- author_url
- https://medium.com/@redouanekarzazi
- status
- ok
- fetched_at
- 2026-10-01 06:07:15