

FLUX.2 Klein 9B Workflow: Local Setup Guide
Run FLUX.2 Klein 9B in ComfyUI: model files, VRAM needs, 4-step settings, text-to-image and edit workflows, error fixes, and when hosted FLUX.2 fits better.
Short answer: the fastest way to run FLUX.2 Klein 9B locally is ComfyUI with three files: the FP8 diffusion model from Black Forest Labs, the Qwen3 8B text encoder and the FLUX.2 VAE. Use the distilled checkpoint at 4 steps and CFG 1.0 for speed, or the Base checkpoint at 20 to 50 steps and CFG around 4 to 5 when you want more variety or plan to train LoRAs. The 9B weights are released under the FLUX Non-Commercial License, so check the licensing terms before you put the model itself into a paid product.
If you do not want to manage a GPU, drivers and model files, the hosted FLUX.2 Pro, Flex and Max models on APIMart cover the same text-to-image and reference-editing jobs through one API. APIMart does not host Klein itself; this guide is about running Klein on your own machine and shows where a hosted FLUX.2 model is the simpler choice.
Key Takeaways
- FLUX.2 Klein was released by Black Forest Labs on January 15, 2026, in 4B and 9B sizes, each as a 4-step distilled model and an undistilled Base model (BFL announcement).
- Klein 9B pairs a 9B flow transformer with an 8B Qwen3 text encoder and needs about 29GB of VRAM at full precision, according to the Hugging Face model card.
- The distilled 9B model runs at 4 steps with guidance 1.0; the Base 9B card example uses 50 steps with guidance 4.0, and ComfyUI's 9B text-to-image template uses 20 steps with CFG 5.
- Both 9B checkpoints use the FLUX Non-Commercial License, while both 4B checkpoints are Apache 2.0.
- FP8 weights cut VRAM use by up to 40% and NVFP4 by up to 55%, per BFL's figures for RTX GPUs.
- One Klein model handles text-to-image, single-reference editing and multi-reference editing; in ComfyUI the edit workflow adds
ReferenceLatentnodes to the same graph.
FLUX.2 Klein at a Glance
| Klein 9B (distilled) | Klein Base 9B | Klein 4B (distilled) | Klein Base 4B | |
|---|---|---|---|---|
| Developer | Black Forest Labs | Black Forest Labs | Black Forest Labs | Black Forest Labs |
| Release | January 15, 2026 | January 15, 2026 | January 15, 2026 | January 15, 2026 |
| Sampling steps | 4 | 20 to 50 (undistilled) | 4 | More than 4 (undistilled) |
| Guidance / CFG | 1.0 | 4.0 (card example) | 1.0 | Not officially confirmed |
| Text encoder | Qwen3 8B | Qwen3 8B | Qwen3 4B (ComfyUI file) | Qwen3 4B (ComfyUI file) |
| VRAM (full precision) | About 29GB | About 29GB | About 13GB | About 13GB |
| License | FLUX Non-Commercial | FLUX Non-Commercial | Apache 2.0 | Apache 2.0 |
| Best for | Fast local generation and editing | Fine-tuning, LoRA, research | Consumer GPUs, edge | Fine-tuning on smaller GPUs |
| Tasks | Text-to-image, single and multi-reference editing | Same | Same | Same |
Sources: BFL announcement, Klein 9B card, Klein Base 9B card, Klein 4B card and the ComfyUI Klein tutorial.
Which Klein Variant Should You Download?
Distilled vs Base
The distilled checkpoints are step-distilled to 4 sampling steps. That is what makes Klein fast: BFL says the family can generate or edit images in under 0.5 seconds on modern hardware, and the ComfyUI team measured about 1.2 seconds end to end for the 4B distilled model on an RTX 5090.
The Base checkpoints are not distilled. They need many more steps, but BFL describes them as having higher output diversity, and they are the ones to use for fine-tuning and LoRA training. The Base 9B model card calls it "ideal for fine-tuning, LoRA training, research, and custom pipelines where control matters more than speed."
4B vs 9B
The 9B model is BFL's flagship Klein model; BFL says it matches or exceeds models five times its size. The 4B model is the accessible option: it fits in about 13GB of VRAM, which BFL maps to an RTX 3090 or 4070 and above. If your GPU has less than 24GB, start with the 4B workflow or with FP8 9B weights and the low-VRAM tips below.
License: what you can and cannot do
This is the part most local guides skip. Klein 4B and Base 4B are Apache 2.0, which allows commercial use. Klein 9B and Base 9B use the FLUX Non-Commercial License. Under that license, the model and its derivatives may only be used for non-commercial purposes, and the license explicitly excludes commercial or production use of the model. The same license says BFL claims no ownership of outputs and that you may use outputs for any purpose, including commercial ones, except where the license prohibits it, such as using outputs to train a competing model.
In practice: experimenting, research and personal projects on 9B are fine. Running 9B inside a paid product or a production pipeline needs a commercial license from BFL (bfl.ai/licensing), or you can switch to the Apache-licensed 4B model or a hosted FLUX.2 model. This is not legal advice, so read the license text yourself.
| Your use case | Suggested option |
|---|---|
| Learning, research, personal art | Klein 9B or Base 9B locally |
| Commercial product, self-hosted | Klein 4B (Apache 2.0), or a BFL commercial license for 9B |
| Commercial product, no GPU to manage | Hosted FLUX.2 Pro or Flex through an API |
| LoRA training | Klein Base 9B or Base 4B |
Prerequisites: Hardware, Software and Access
GPU and memory
| Setup | What to expect |
|---|---|
| 9B at BF16 | About 29GB VRAM per the model card; RTX 4090 and above with offloading |
| 9B at FP8 | Up to 40% less VRAM than BF16 per BFL; exact 9B figure not officially confirmed |
| 9B at NVFP4 | Up to 55% less VRAM per BFL; needs a recent NVIDIA RTX GPU |
| 4B distilled, FP8 in ComfyUI | 8.4GB VRAM measured by ComfyUI on an RTX 5090 |
| 4B Base, FP8 in ComfyUI | 9.2GB VRAM measured by ComfyUI on an RTX 5090 |
BFL also reports that FP8 is up to 1.6x faster and NVFP4 up to 2.7x faster, benchmarked on RTX 5080 and 5090 cards at 1024x1024.
ComfyUI version
Klein uses newer FLUX.2 nodes such as Flux2Scheduler and EmptyFlux2LatentImage. The ComfyUI docs say that if the Klein templates are missing or nodes fail to load, your install is outdated, and they suggest the latest (Nightly) build. The Desktop app can lag behind the main release, so update before you troubleshoot anything else.
Hugging Face access
The 9B repositories are gated. Sign in to Hugging Face, open the Klein 9B FP8 repository, accept the license and Acceptable Use Policy, then create a read token for command-line downloads.
Step 1: Download the Model Files
The ComfyUI tutorial lists these files for the 9B workflow:
| File | Source | ComfyUI folder |
|---|---|---|
flux-2-klein-9b-fp8.safetensors (distilled) | black-forest-labs/FLUX.2-klein-9b-fp8 | models/diffusion_models/ |
flux-2-klein-base-9b-fp8.safetensors (Base) | black-forest-labs/FLUX.2-klein-base-9b-fp8 | models/diffusion_models/ |
qwen_3_8b_fp8mixed.safetensors | Comfy-Org/flux2-klein-9B | models/text_encoders/ |
flux2-vae.safetensors | Comfy-Org/flux2-dev | models/vae/ |
You only need one diffusion model to start. Pick the distilled file for speed, or the Base file if you plan to train LoRAs.
export HF_TOKEN=hf_xxx
curl -L -H "Authorization: Bearer $HF_TOKEN" \
-o ComfyUI/models/diffusion_models/flux-2-klein-9b-fp8.safetensors \
https://huggingface.co/black-forest-labs/FLUX.2-klein-9b-fp8/resolve/main/flux-2-klein-9b-fp8.safetensors
curl -L -o ComfyUI/models/text_encoders/qwen_3_8b_fp8mixed.safetensors \
https://huggingface.co/Comfy-Org/flux2-klein-9B/resolve/main/split_files/text_encoders/qwen_3_8b_fp8mixed.safetensors
curl -L -o ComfyUI/models/vae/flux2-vae.safetensors \
https://huggingface.co/Comfy-Org/flux2-dev/resolve/main/split_files/vae/flux2-vae.safetensors
The resulting layout should look like this:
ComfyUI/
└── models/
├── diffusion_models/
│ └── flux-2-klein-9b-fp8.safetensors
├── text_encoders/
│ └── qwen_3_8b_fp8mixed.safetensors
└── vae/
└── flux2-vae.safetensors
Do not mix encoders across sizes. The 4B workflow uses qwen_3_4b.safetensors, and the 9B workflow uses the Qwen3 8B encoder.
Step 2: Build the Text-to-Image Workflow
The easiest start is the built-in template. In ComfyUI, open the template browser and load Flux.2 Klein 9B Text to Image (image_flux2_text_to_image_9b). The template JSON is public in the Comfy-Org workflow_templates repository, so you can inspect it before you run it.
The nodes in the graph
| Node | Role | Key setting |
|---|---|---|
UNETLoader (Load Diffusion Model) | Loads the Klein transformer | flux-2-klein-9b-fp8.safetensors |
CLIPLoader | Loads the Qwen3 encoder | type flux2 |
VAELoader | Loads the FLUX.2 VAE | flux2-vae.safetensors |
CLIPTextEncode | Encodes the positive prompt | Your prompt |
EmptyFlux2LatentImage | Creates the empty latent | Width and height, e.g. 1024x1024 |
Flux2Scheduler | Builds the sigma schedule | Steps, width, height |
CFGGuider | Applies guidance | CFG value |
KSamplerSelect | Chooses the sampler | euler |
RandomNoise | Seeds the run | Fixed seed for reproducibility |
SamplerCustomAdvanced | Runs the sampling loop | Inputs only |
VAEDecode then SaveImage | Decodes and saves | Filename prefix |
The shipped 9B template loads the Base checkpoint with 20 steps and CFG 5. If you switch UNETLoader to the distilled file, change Flux2Scheduler to 4 steps and CFGGuider to 1.0, which matches the distilled model card and ComfyUI's distilled edit template. Leaving Base-style settings on the distilled model wastes time and can overcook the image.
Settings that work
| Setting | Distilled 9B | Base 9B |
|---|---|---|
| Steps | 4 | 20 (ComfyUI template) to 50 (model card example) |
| CFG / guidance | 1.0 | 4.0 to 5.0 |
| Sampler | euler | euler |
| Resolution | 1024x1024 to start | 1024x1024 to start |
| Seed | Fix it while tuning prompts | Fix it while tuning prompts |
Keep Flux2Scheduler width and height in sync with the latent size. The official templates wire both from the same width and height values, so copy that pattern when you build your own graph or change to a non-square size.
A prompt to test the pipeline
Klein reads natural-language descriptions well, so write prompts as full sentences with subject, setting, lighting and any text you want rendered:
A vintage motorcycle parked outside a 1950s roadside diner at dusk, neon sign reading "OPEN LATE", wet asphalt reflecting pink and teal light, 35mm photo, shallow depth of field
If the first image comes out in a few seconds and the sign text is legible, the install is working. For more prompt structure ideas that also apply to FLUX models in general, see our FLUX 3 prompting guide.
Step 3: Build the Image Editing Workflow
Klein uses the same model for editing, so you do not need a separate Kontext-style checkpoint. ComfyUI ships two 9B edit templates: image_flux2_klein_image_edit_9b_distilled and image_flux2_klein_image_edit_9b_base.
How the edit graph differs
| Added node | What it does |
|---|---|
LoadImage | Loads your source or reference image |
ImageScaleToTotalPixels | Rescales the image to about 1 megapixel |
GetImageSize | Passes the scaled size to the latent and scheduler |
VAEEncode | Turns the reference image into a latent |
ReferenceLatent | Attaches the reference latent to the conditioning |
ConditioningZeroOut | Builds the empty negative conditioning |
The reference latent is attached to both the positive prompt and the zeroed negative conditioning, and both go into CFGGuider. Because GetImageSize drives EmptyFlux2LatentImage and Flux2Scheduler, the output keeps the aspect ratio of your first reference image.
Single-reference edits
Load one image and describe the change, not the whole scene. The distilled edit template uses 4 steps and CFG 1, and the same settings are the right starting point for your own edits.
Change the jacket to deep red leather, keep the face, pose, background and lighting unchanged
Multi-reference edits
For two or more images, the template chains one ReferenceLatent per image on both the positive and negative paths. Refer to images by order in the prompt:
Place the white handbag from image 2 on the woman's shoulder in image 1, match the soft window light of image 1
BFL's documentation lists support for up to 4 reference images for Klein through its API (FLUX.2 overview); more references in a local graph also mean more memory and slower runs.
Troubleshooting and Low-VRAM Optimization
Common errors and fixes
| Symptom | Likely cause | Fix |
|---|---|---|
| Template missing or red "missing node" boxes | ComfyUI too old for FLUX.2 nodes | Update ComfyUI; use the Nightly build if needed |
| 401 or 403 when downloading the 9B file | Gated repo, license not accepted or no token | Accept the license on Hugging Face and pass a read token |
| Shape or size mismatch when loading the text encoder | 4B encoder used with 9B model, or wrong CLIP type | Use the Qwen3 8B encoder and set CLIPLoader type to flux2 |
| CUDA out of memory | BF16 weights or high resolution | Use FP8 weights, start at 1024x1024, close other GPU apps |
| Blurry or burnt images on the distilled model | Base settings (20+ steps, CFG 4 to 5) | Set 4 steps and CFG 1.0 |
| Noisy, unfinished images on Base | Too few steps | Use 20 to 50 steps |
| Edit ignores the reference | ReferenceLatent not wired to conditioning | Connect it on both positive and negative paths |
Low-VRAM checklist
- Use the FP8 checkpoints first; BFL quotes up to 40% lower VRAM. NVFP4 goes further on supported RTX cards.
- Use the
fp8mixedtext encoder rather than a full-precision one. - Start ComfyUI with
--lowvramso model parts are offloaded to system RAM when the GPU fills up. - Generate at 1024x1024 or below and upscale afterward instead of sampling at very large sizes.
- Keep batch size at 1 and reduce the number of reference images in edit graphs.
- If 9B still does not fit, the 4B distilled model ran in 8.4GB in ComfyUI's measurement.
python main.py --lowvram
Running Klein 9B with Diffusers instead
If you prefer Python over a node graph, the model card shows a Diffusers example. It needs the development version of Diffusers:
pip install git+https://github.com/huggingface/diffusers.git
import torch
from diffusers import Flux2KleinPipeline
pipe = Flux2KleinPipeline.from_pretrained(
"black-forest-labs/FLUX.2-klein-9B", torch_dtype=torch.bfloat16
)
pipe.enable_model_cpu_offload() # offloads idle parts to CPU to save VRAM
image = pipe(
prompt="A cat holding a sign that says hello world",
height=1024,
width=1024,
guidance_scale=1.0,
num_inference_steps=4,
generator=torch.Generator("cuda").manual_seed(0),
).images[0]
image.save("flux-klein.png")
For Base 9B, swap the repository to black-forest-labs/FLUX.2-klein-base-9B and use guidance_scale=4.0 with num_inference_steps=50, as in that card's example.
When Hosted FLUX.2 Pro or Flex Is the Better Choice
Local Klein is excellent for fast iteration, private experiments and LoRA work. It is less convenient when you need commercial rights on the 9B model, guaranteed uptime, or more capacity than one GPU provides.
Local Klein 9B vs hosted FLUX.2
| Need | Local Klein 9B | Hosted FLUX.2 on APIMart |
|---|---|---|
| Hardware | Your GPU, about 29GB at BF16 | None |
| Commercial use of the model | Needs a BFL commercial license | Governed by the API terms, not the open-weight license |
| Setup | ComfyUI, model files, updates | One API key |
| Reference images | BFL lists up to 4 for Klein | Up to 8 per request |
| Step and guidance control | Full | Flex only (steps 1 to 50, guidance 1.5 to 10) |
| Output size | Your GPU's limit | Up to 4MP |
| Custom LoRAs | Yes, with Base | No |
| Cost model | Your hardware and power | Per image; see the model page for current pricing |
Picking a hosted model
APIMart exposes three FLUX.2 model IDs on the /v1/images/generations endpoint: flux-2-pro for balanced everyday production, flux-2-flex when you want to tune steps and guidance (for example typography-heavy posters), and flux-2-max for the highest quality at a slower speed. Prices depend on output resolution tier and the number of reference images; check the FLUX.2 model page before you budget.
curl --request POST \
--url https://api.apimart.ai/v1/images/generations \
--header 'Authorization: Bearer <token>' \
--header 'Content-Type: application/json' \
--data '{
"model": "flux-2-flex",
"prompt": "Minimal poster with the headline SUMMER SALE and a small line reading 50% OFF",
"resolution": "2K",
"size": "3:4",
"steps": 50,
"guidance": 6.5
}'
The call is asynchronous and returns data[0].task_id. Poll GET https://api.apimart.ai/v1/tasks/{task_id} until status is completed, then read the image URL from result.images. To edit with references, pass up to 8 public URLs in image_urls. For a production pattern with retries and storage, see our guide to FLUX image workflows for developers, and for a speed and VRAM comparison against another fast open model, read Z-Image Turbo vs Flux.
A practical hybrid
Many teams prototype prompts and LoRAs on local Klein, then send final or customer-facing renders to a hosted FLUX.2 model. Prompts written in plain descriptive sentences transfer well between the two, so you rarely need to rewrite them.
FAQs
What is the difference between FLUX.2 Klein 9B and Klein Base 9B?
Klein 9B is step-distilled to run in 4 steps with guidance 1.0, which makes it fast enough for near real-time generation. Klein Base 9B is the undistilled foundation model; it needs 20 to 50 steps but gives more output diversity and is the recommended starting point for fine-tuning and LoRA training.
Can I use FLUX.2 Klein 9B commercially?
Not the model itself without a separate agreement. Both 9B checkpoints use the FLUX Non-Commercial License, which limits use of the model to non-commercial purposes; BFL offers commercial licenses at bfl.ai/licensing. The license says outputs may be used for any purpose, including commercial ones, with some exceptions, so read it in full. Klein 4B is Apache 2.0 if you need a self-hosted commercial option.
How much VRAM does FLUX.2 Klein 9B need?
The model card says about 29GB at full precision, which points to an RTX 4090 or better with CPU offloading. FP8 weights reduce memory use by up to 40% and NVFP4 by up to 55%, according to BFL; an exact VRAM figure for 9B FP8 is not officially confirmed. On cards with 16GB or less, the 4B model is the safer choice.
What steps and CFG should I use in ComfyUI?
For the distilled 9B model, use 4 steps and CFG 1.0 with the euler sampler. For Base 9B, ComfyUI's template uses 20 steps and CFG 5, and the Hugging Face example uses 50 steps and guidance 4.0. Start at 1024x1024 in both cases.
Does Klein need a separate model for image editing?
No. One Klein checkpoint handles text-to-image, single-reference editing and multi-reference editing. In ComfyUI you add LoadImage, VAEEncode and ReferenceLatent nodes to the text-to-image graph, or load the official edit template.
Is FLUX.2 Klein available on APIMart?
No. APIMart offers the hosted FLUX.2 Pro, Flex and Max models, not Klein. They are a good alternative when you would rather call an API than manage an open-weight license, need up to 8 reference images or do not have a suitable local GPU.
What to Watch Next
Klein is moving quickly in the community, with many adapters, fine-tunes and quantizations already listed on its Hugging Face page. Watch for ComfyUI template updates, new quantized weights and BFL licensing changes, since each can change what fits on your GPU and what you can ship. If your workload grows past one machine or needs clear commercial terms, test the same prompts on hosted FLUX.2 Pro or Flex before you scale.
About APIMart Team
APIMart Team writes model guides, side-by-side comparisons and pricing breakdowns for developers building with AI. APIMart gives you one API for 500+ chat, image and video models, so you can test the models covered here against each other before you ship.
Choose the model you want in the model marketplace
Try chat, image and video models in the APIMart model marketplace, and experience model capabilities quickly with one unified API.