Skip to main content

Running AI image generation locally in 2026: VRAM limits, costs and honest expectations

AI article illustration for ai-jarvis.eu
Local AI image generation stopped being a hobbyist curiosity a while ago — it is now a legitimate workflow in studios, agencies and one-person shops. The question in 2026 is no longer "can I run it?" but "which card do I actually need, and what do I get for the money?" Most of the confusion comes from VRAM figures quoted for full FP16 weights that almost nobody uses in practice. Here is the honest version: model tiers, real VRAM ceilings, EU street prices, and the places where local still loses to a €10 subscription.

The three tiers of local image models

Image models are not one thing. Treat them as three families with completely different appetites.

Tier 1 — SD 1.5 and the SDXL family. Still the workhorses. SDXL at 1024×1024 has an enormous ecosystem of fine-tunes and ControlNet adapters, runs fast, and fits on modest hardware. This is where you should start, regardless of what GPU you own.

Tier 2 — FLUX.1 and SD 3.5. Diffusion transformers with far better prompt adherence, anatomy and — crucially — text rendering. FLUX.1 schnell is Apache 2.0 and runs in four steps; FLUX.1 dev is stronger but ships under a non-commercial licence, so check Black Forest Labs' current terms before you build a product on it. Stability's SD 3.5 family uses a community licence that is free below a revenue threshold.

Tier 3 — big DiT models and video. Higher-fidelity 12B+ models and any video generation (Wan, Hunyuan-style pipelines) move you into 24 GB territory or heavy quantisation with painful generation times. Real, but not the place to begin.

One thing to get out of the way: Ollama is for language models. It has nothing to do with image generation. For images you want ComfyUI, Forge, Fooocus, InvokeAI or the raw diffusers library.

The VRAM ladder: what each card class really gets you

These are practical figures, not marketing numbers. They assume quantised or FP8 weights where that is the sensible default, and include the overhead of a typical workflow with a VAE and one or two adapters.

Card classComfortablePossible with effortRealistic output
6–8 GB VRAMSD 1.5, SDXL at 768–1024 px with tilingFLUX.1 schnell GGUF Q4Good for learning and batch work; slow on FLUX
12 GB VRAMSDXL 1024 px, SD 3.5 MediumFLUX.1 dev FP8 with offloadingThe entry point for serious work
16 GB VRAMSDXL, FLUX.1 dev FP8, SD 3.5 MediumSD 3.5 Large at reduced batchThe current sweet spot for one operator
24 GB VRAMFLUX.1 dev FP16, SD 3.5 LargeShort video clips, heavy ControlNet stacksStudio-grade single-GPU setup
32 GB+ / unified memoryEverything above, larger batchesLonger video, multi-model pipelinesOverkill for stills, sane for video

On our AI Arena rig we run a 16 GB RTX 5060 Ti, and that card is the line where "local image generation" stops being a compromise. SDXL runs without offloading, and quantised FLUX variants fit without swapping weights to system RAM. Below 12 GB, you spend more time tuning flags than generating pictures.

Quantisation is why 8 GB cards still work

GGUF and FP8 quantisation shrink model weights to a fraction of their original size with a quality hit that is often invisible at normal viewing size. A Q4 quantisation of a 12B image model can run on hardware that would choke on the FP16 version. The trade-offs are real, though: quantised models generate slower per step, some adapters and ControlNets behave differently, and very aggressive quantisation does soften fine detail in faces and text.

The practical rule: quantise the model, not your expectations. If you need the last few percent of fidelity for a paid job, that is a signal to consider a 24 GB card or the cloud.

RAM, storage and power matter more than people admit

A GPU is not the whole bill. You want 32 GB of system RAM minimum — 64 GB if you plan to load and unload large checkpoints. Model files are large: a single FP16 FLUX checkpoint is in the tens of gigabytes, and a working library of checkpoints, LoRAs and upscalers fills a terabyte fast. Budget an NVMe SSD. And check your PSU: modern mid-range cards draw 180–300 W under load, with transient spikes above that.

The software stack, without the folklore

  • ComfyUI (GitHub) — node-based, ugly at first glance, and the de facto standard in 2026. Maximum control, maximum rope to hang yourself with.
  • Forge / Automatic1111 — familiar, form-based, good for SDXL and quick iteration.
  • Fooocus — the "just work" option. Minimal settings, sensible defaults, ideal for a first afternoon.
  • InvokeAI — the most polished canvas-style workflow, popular in small studios.

Models come from Hugging Face and Civitai. Both are reachable from the EU without friction, and Civitai's community fine-tunes remain one of the strongest arguments for running locally at all.

AMD, Intel and Apple Silicon: the honest state

NVIDIA is still the path of least resistance because CUDA has the best-supported kernels. AMD's ROCm works on recent Radeon cards and under Linux/WSL, but you should expect to debug. Intel's Arc cards with 12–16 GB are genuinely interesting value on paper; software maturity is the weak point. Apple Silicon runs image generation comfortably through MPS — a Mac with 32 GB or more of unified memory handles SDXL and quantised FLUX fine, just slower than a discrete card at the same price.

Local versus cloud: the actual cost math

Here is a calculation we did ourselves, because the online debate is mostly vibes.

Assume a 16 GB mid-range card at roughly €470, generating 2,000 images a month at 15 seconds each on a 250 W GPU. That is 30,000 seconds of compute — about 2.1 kWh per month. At a European household rate of roughly €0.30/kWh, your electricity bill for all that generation is around €0.63 a month. Image generation is not an energy problem; holding a gaming session is more expensive.

Now the comparison. Cloud image APIs commonly land somewhere between $0.02 and $0.05 per image depending on model and resolution — call it €0.03. Two thousand images a month is €60. The €470 card pays for itself in under eight months, and then it is free at the margin. Subscriptions like Midjourney start around €10/month but cap you on fast generations.

Flip the volume: if you generate 150 images a month, cloud or a subscription is cheaper and you keep the €470. Buy hardware for throughput and privacy, not for prestige.

The European angle: GDPR, the AI Act and your desk

This is where local generation earns its keep beyond cost.

GDPR. When you generate locally, the prompt never leaves your machine. That matters if you are working with client material, internal product shots, or anything involving identifiable people. No data processing agreement with a US provider, no transfer mechanism to justify, no sub-processor list to update. For a lot of European SMEs, that alone justifies the hardware.

AI Act. Transparency obligations under Article 50 for synthetic content apply from 2 August 2026. If you publish AI-generated or AI-manipulated images publicly in the EU, you should be ready to disclose that, and deepfake-style content carries explicit labelling duties. Running locally does not exempt you — but it does give you full control over which metadata you embed, including C2PA provenance markers. Most hosted tools write their own metadata for you, on their terms.

Availability. Every major open-weights image model is downloadable inside the EU. There is no regional gate on weights, and no subscription required. The only European-specific cost is the card itself, and EU street prices for 16 GB mid-range GPUs currently sit noticeably above US listings once VAT is included.

Where local still loses

Be honest about the gaps. Frontier hosted models still handle complex multi-subject scenes, precise typography and instruction-following better than anything you will run on a mid-range card. Cloud tools also give you instant access to the newest architecture without a download. And if you generate rarely, local is a hobby, not a strategy.

What local wins: unlimited iterations, no per-image pricing, no content filter rewriting your prompt, offline operation, and — for European professionals — a clean data story you can defend to a client.

Three concrete configurations

BudgetHardwareWhat you can do
~€900Used 12 GB card, 32 GB RAM, 1 TB NVMeSDXL at full speed, quantised FLUX slowly
~€1,400New 16 GB card (RTX 5060 Ti class), 64 GB RAM, 2 TB NVMeEverything except big DiT models and video
~€2,800+24 GB card or Apple Silicon with 64 GB unified memoryFull-precision FLUX, SD 3.5 Large, short video

If you want to see how we benchmark these sorts of workloads — tokens per second for LLMs, time-to-first-token, VRAM headroom — the methodology lives at /ai-arena, and there is more practical workflow material in our magazine archive.

Can I run modern image models without an NVIDIA GPU?

Yes, with caveats. Apple Silicon handles SDXL and quantised FLUX through MPS, just slower. AMD's ROCm works on recent Radeon cards but expects you to be comfortable debugging. Intel Arc has good VRAM per euro; software maturity is the trade-off. If you want zero friction, CUDA is still the answer.

Do I need a separate PC for this, or can I use my everyday machine?

Your everyday machine is fine as long as it has the VRAM and cooling. Image generation is bursty — heavy for twenty seconds, idle after. Just check your power supply can handle the transient spikes and that your case has airflow. A dedicated rig only makes sense if you want to leave batch jobs running.

Is generating images locally legal in the EU?

Generating is legal. Publishing is where obligations begin: from 2 August 2026, Article 50 of the AI Act requires transparency about synthetic content, and deepfake-style material carries explicit labelling duties. Also check each model's licence — some, like FLUX.1 dev, are non-commercial, while others allow commercial use below a revenue threshold.

Discussion

No comments yet — be the first to share your thoughts.
X

Don't miss out!

Subscribe for the latest news and updates.