How to Run Qwen-Image-2.1 Locally with ComfyUI (+ Prompts)
I recently started testing Qwen-Image-2.1 locally in ComfyUI, and the part that makes it interesting is not just image quality. The visual generation component is only 7B parameters, yet it can handle text-to-image generation, image editing, transparent RGBA output, multiple reference images, and native 2K generation.
More importantly, once the model is on your machine there is no per-image API charge or generation quota. Your real limits are your hardware, storage, generation time, and electricity.
In this guide I’ll show the setup I used, where the official ComfyUI model files go, how to load the Qwen workflow, and a set of prompts I used to test photorealism, text rendering, Kerala scenes, product photography, transparent assets, and difficult multi-person compositions.
What is Qwen-Image-2.1?
Qwen is Alibaba’s AI model family. Qwen-Image-2.1 was released on September 20, 2026 as a unified image generation and editing model. According to the official model card, its visual generation component has 7B parameters.
Some of the features worth testing are:
- Text-to-image generation
- Image editing in the same model
- Native transparent RGBA image generation
- Up to 10 reference images for multi-subject composition and editing
- Improved typography and fine detail
- Native 2K generation with multiple supported aspect ratios
Qwen calls the model open-source, while the published Hugging Face weights currently use the Qwen Research license. If you plan to redistribute or use the weights commercially, read the license rather than assuming an Apache/MIT-style license.
Official resources:
Qwen-Image-2.1 GitHub
Official Hugging Face model
ComfyUI-packaged model files
Why I’m using ComfyUI
ComfyUI is a node-based local inference interface. Instead of hiding the pipeline behind a single “Generate” button, it exposes the model loader, text encoder, sampler, latent size, VAE and output nodes as a graph.
That makes it useful when you want to experiment with local models rather than just consume a hosted image API.
For this test I’m running the model locally on my Mac mini. ComfyUI also supports Windows, Linux and Apple Silicon. Qwen-Image-2.1 received native ComfyUI support on release day.
Step 1: Install or update ComfyUI
Recommended: ComfyUI Desktop
If you are starting fresh, the easiest route on Windows or macOS is the official ComfyUI Desktop application. Follow the current installation instructions from the official ComfyUI documentation.
If you already use ComfyUI, update it before doing anything else. Qwen-Image-2.1 is a very new model, and an older ComfyUI build may not contain the required native nodes or workflow templates.
This is particularly important with the Desktop build because stable releases can arrive slightly later than the latest Git version.
Manual installation
If you prefer the latest source version, or the Desktop build does not yet expose the Qwen-Image-2.1 workflow, you can use the official GitHub repository:
git clone https://github.com/Comfy-Org/ComfyUI.git
cd ComfyUI
python3 -m venv venv
source venv/bin/activate
pip install -r requirements.txt
python main.py
On Windows, activate the virtual environment using the Windows equivalent before installing the requirements. GPU/backend dependencies can vary, so use the current ComfyUI installation documentation for your NVIDIA, AMD, Intel or Apple Silicon setup rather than copying an old PyTorch command from a random tutorial.
Step 2: Download the Qwen-Image-2.1 model files
For ComfyUI, I recommend using the files published by Comfy-Org rather than manually converting the original repository.
Open:
https://huggingface.co/Comfy-Org/Qwen-Image-2.1
The official ComfyUI package currently contains the following model groups.
Diffusion model
qwen_image_2.1_bf16.safetensors
qwen_image_2.1_int8_convrot.safetensors
Text encoder
qwen3vl_8b_bf16.safetensors
qwen3vl_8b_int8_convrot.safetensors
qwen3vl_8b_w4a8.safetensors
VAE
qwen_image_2.1_vae_bf16.safetensors
You do not need to download every precision variant. BF16 is the straightforward quality-first option if you have enough memory. Quantized variants exist for lower-memory setups; performance and backend compatibility can vary, so test the variant that suits your machine.
Step 3: Put the files in the correct folders
This is the part that causes most “model not found” problems.
Your ComfyUI folder should look like this:
ComfyUI/
└── models/
├── diffusion_models/
│ └── qwen_image_2.1_bf16.safetensors
│
├── text_encoders/
│ └── qwen3vl_8b_bf16.safetensors
│
└── vae/
└── qwen_image_2.1_vae_bf16.safetensors
If you choose one of the quantized diffusion models or text encoders, place it in the same corresponding directory.
After copying the files, restart ComfyUI so the loaders can discover them.
Step 4: Load the official workflow
ComfyUI provides official workflow templates for both text-to-image and image editing:
If the templates are available inside your version of ComfyUI, load the Qwen-Image-2.1 workflow directly from the workflow/template browser. Otherwise download the JSON workflow and open or drag it into ComfyUI.
Then verify that the loader nodes point to the diffusion model, text encoder and VAE you downloaded.
Step 5: Start with the default generation settings
The official Qwen pipeline uses 40 inference steps by default. The model also supports native 2K output.
Some official aspect-ratio examples are:
1:1 2048 x 2048
4:3 2400 x 1792
3:4 1792 x 2400
3:2 2528 x 1696
2:3 1696 x 2528
16:9 2752 x 1536
9:16 1536 x 2752
For vertical Reels/Shorts experiments, 9:16 is obviously useful. On a memory-constrained machine, however, I would first get the composition right at a more manageable resolution before pushing every test to the maximum size.
How I write prompts for Qwen-Image-2.1
I get better testing value from prompts that combine several requirements instead of asking for a generic “beautiful portrait.” A useful structure is:
Subject
+ environment
+ action
+ lighting
+ camera/composition
+ materials/textures
+ exact text if needed
+ constraints
For example, instead of:
Malayali man drinking tea
I would describe the location, weather, clothing, lens, skin texture and background elements. That lets me judge whether the model actually follows the prompt.
Prompt 1: Cinematic Kerala portrait
Cinematic photograph of a 28-year-old Malayali man standing at a roadside tea shop in Kozhikode during heavy monsoon rain at night. He is wearing a slightly wet dark green shirt and mundu, holding a small glass of hot chai. Rainwater reflecting warm shop lights on the road, an old Kerala private bus softly visible in the background, Malayalam shop boards, motorcycles parked nearby, realistic humid atmosphere, natural Indian skin texture, subtle imperfections, documentary photography, Kodak Portra 800 aesthetic, 35mm lens, shallow depth of field, cinematic but completely believable.
This is a good realism test because people from Kerala can immediately spot fake architecture, clothing, skin texture, buses and signboards.
Prompt 2: Premium technology advertisement
Premium minimalist studio advertisement for a futuristic transparent wireless earbud case floating above a matte black pedestal. The charging case is made from smoky transparent glass with visible precision-engineered internal components. Two earbuds float beside it. Black seamless background, controlled rim lighting, subtle reflections, extremely clean industrial design photography. At the top, clearly render the headline: "HEAR EVERYTHING." Below it in smaller typography: "VX AUDIO". Luxury technology campaign, high-end product photography, physically accurate materials and reflections, 4K commercial advertising image.
This tests typography, glass, reflections, product geometry and whether the text is actually legible.
Prompt 3: Malayalam text rendering
Design a realistic modern roadside café advertisement photographed in Kerala. A large illuminated signboard above the café must clearly display the Malayalam text:
"ചായ കുടിച്ചിട്ട് പോകാം"
Directly underneath, in smaller English letters:
"ONE MORE CHAI"
Several customers are sitting inside the café drinking tea. Warm tungsten lighting, rainy evening, reflections on the street, motorcycles parked outside, realistic Kerala architecture, candid documentary photograph. The Malayalam lettering must be correctly spelled, clearly readable and naturally integrated into the physical signboard.
This is deliberately difficult. Image models often look impressive until you zoom into non-English text.
Prompt 4: Kerala food photography
Ultra-realistic cinematic advertising photograph of Kerala beef biryani served in a traditional brass uruli, long-grain basmati rice with visible spices, tender roasted beef pieces, caramelized onions, cashews and coriander. Steam rising naturally from the food. A small bowl of raita and lime pickle beside it. Dark premium restaurant background, dramatic side lighting, subtle warm highlights, shallow depth of field, 50mm food photography, realistic oil and moisture, extremely appetizing but natural, no artificial plastic appearance.
Food is surprisingly useful for spotting overprocessed textures. Rice grains, moisture, steam and meat texture give the model plenty of opportunities to fail.
Prompt 5: Transparent PNG / RGBA asset
One of the most interesting Qwen-Image-2.1 features is native transparent image generation. Qwen recommends explicitly describing the image as RGBA and requesting an alpha channel.
This is an RGBA image with transparency. Create a premium isolated 3D icon of a futuristic artificial intelligence processor. The processor is made from brushed dark aluminium with a glowing translucent glass brain embedded in the centre. Tiny circuit traces extend toward the edges of the chip. Three-quarter perspective, realistic materials, studio lighting, crisp edges, subtle reflections. The object must be completely isolated. The image has alpha channel and the background is transparent. No floor, no external background and no text.
After generation, place the PNG over another image or video layer. That makes it immediately obvious whether you received real transparency rather than a fake checkerboard background.
Prompt 6: Miniature Kerala town
Create an extraordinarily detailed miniature diorama of a busy Kerala town junction viewed from above at a 45-degree angle. Tiny private buses, auto-rickshaws, motorcycles, pedestrians carrying umbrellas, roadside tea shops, coconut trees and colourful local buildings. It should appear to be a meticulously handcrafted physical scale model photographed with a macro lens, realistic miniature materials, tiny painted signs and road markings, subtle imperfections, shallow depth of field and tilt-shift photography. Do not make it look like a 3D render.
Prompt 7: Difficult multi-person composition
Ultra-realistic candid photograph inside a crowded Indian commuter train during golden hour. About fifteen passengers are visible at different depths of the scene: a student reading a book, an elderly man looking through the window, a mother holding a sleeping child, two office workers talking quietly and a tea vendor walking through the aisle carrying paper cups. Warm sunlight enters through the windows and creates complex natural highlights across faces and metal surfaces. Every person should have anatomically correct hands, unique natural facial features and believable expressions. Nothing is posed. Documentary street photography, 35mm Leica aesthetic, realistic Indian skin tones, natural motion blur, subtle film grain, densely composed but visually coherent.
This is the kind of prompt I like to use near the end of a model test. One portrait can hide a lot of weaknesses. Fifteen people interacting in a coherent scene cannot.
Prompt 8: Science-fiction Kerala without the generic cyberpunk look
A cinematic frame from a fictional science-fiction movie set in Kerala in the year 2045. A traditional Kerala tea shop stands beside a futuristic electric highway during a violent monsoon storm. Autonomous vehicles pass through heavy rain while three elderly men casually drink tea inside the shop as if nothing unusual is happening. Coconut trees bending in the wind, futuristic city lights faintly visible in the distance, grounded realistic production design, not cyberpunk, natural colours, atmospheric rain, anamorphic lens, subtle film grain, realistic human proportions, high-budget Indian science-fiction cinematography.
The phrase “not cyberpunk” matters here. Many image models tend to turn any request containing “future” into neon purple and blue cityscapes. Constraints are part of the test.
Image editing and multiple references
Qwen-Image-2.1 is not limited to text-to-image. The same model supports image editing and can take up to 10 reference images.
The official example is conceptually simple: load an input image and describe the change you want, such as:
Change the background to a sunset beach while preserving the person, facial identity, clothing and pose.
For multiple references, you can provide several subjects and describe how they should appear together in the final composition.
For creators, this is potentially more useful than pure text-to-image because it opens the door to product consistency, recurring characters, thumbnails, ad variations and controlled edits without sending the source files to a hosted image API.
Troubleshooting
The Qwen workflow or nodes are missing
Update ComfyUI. Qwen-Image-2.1 only landed on September 20, 2026. If the Desktop stable release has not caught up yet, use the latest supported build or current Git version.
The model does not appear in the loader
Check the folder names. The diffusion model belongs in models/diffusion_models, the Qwen3-VL encoder belongs in models/text_encoders, and the VAE belongs in models/vae. Restart ComfyUI after adding them.
You run out of memory
Try the packaged quantized model/text-encoder variants, lower your working resolution while iterating, close other memory-heavy applications, and avoid generating large batches until the basic workflow is stable.
The image looks good but ignores part of the prompt
Reduce ambiguity. Describe who is doing what, where objects should appear, what text must be written, and which attributes must remain unchanged. Complex prompts are useful for testing, but contradictory prompts are not.
Is local generation really “unlimited”?
In the API-credit sense, yes: after downloading the model, you are not paying a cloud provider for every generation and there is no hosted-service quota stopping you after a fixed number of images.
But “unlimited” does not mean free compute. Local generation still costs hardware time, storage, electricity and patience. The tradeoff is that the workload is yours to run and control.
Final thoughts
What makes Qwen-Image-2.1 interesting to me is the direction of travel. Open/local models are no longer interesting only for coding and text tasks. Image generation is moving quickly too.
A 7B visual generator that can already produce convincing portraits, product shots, typography, editing and transparent assets locally changes what a small creator or developer can build without depending entirely on paid APIs.
I’m still testing where it breaks — especially hands, crowds, Malayalam text, spatial instructions and identity consistency — because those failure cases matter more than a folder full of cherry-picked portraits.
If you are experimenting with Qwen-Image-2.1 in ComfyUI, start with the official workflow, verify the model files are in the correct folders, and then make your prompts progressively harder. That tells you much more about the model than a generic “beautiful cinematic portrait” benchmark.