Video summary

Clean Character Swap with Flux.2 Klein + LoRA Detailing + 4K Upscale (Full Workflow)

Main summary

Key takeaways

Technology

Main purpose

  • A full Flux.2 “clean character swap” workflow (follow-up to prior Flux.2 Klein workflows) showing how to: 1) swap a subject from one image into the pose/background of a target image, 2) optionally improve face quality with masking, 3) apply LoRA-based detailing, and 4) upscale to 4K.

Model options and sampling-step guidance

  • Uses the Klein 9B distilled model, expected to work in about 4–8 sampling steps.
    • The speaker notes this may be a mistake and suggests sampling steps may need to be higher.
  • Recommended adjustment mentioned:
    • Start small (e.g., 4) but likely increase up to around 20–25, depending on results.
  • Mentions alternative model formats:
    • FP8 precision distilled models (planned/linked similarly).
    • GGUF models loaded via a U-Net loader (GGUF placed inside a U-Net folder).

Core workflow: Subject + target scene/pose conditioning

The workflow explicitly sets required components:

  • Text encoder (CLIP) via Load-Clip (select the red-colored node).
  • VAE via the VAE node (also indicated by red-colored nodes).

LoRA usage

  • One LoRA can be selected to be used across the workflow.
  • A float/value is adjusted for image scaling consistency—especially when using larger images.
  • Later, a different LoRA may be used for the detailing pass.

Node group 1: “Processing subject”

Inputs

  • Subject image (image 1)
  • resized using an image resize / pixel-size parameter
  • prompt to change/remove background

Prompt concept

  • Uses a prompt like replacing background with white to isolate/standardize the subject.

Output

  • A generated subject image without the original background, so it can serve as the character reference while keeping the target scene.

Node group 2: “Processing target pose / target scene”

Inputs

  • Target image (image 2)
  • model + CLIP encoder + VAE connections
  • megapixel size / image scale adjustment

Prompt concept

  • Intended to remove clothes from the person in the target scene.

Notes

  • OpenPose was tried but made the workflow more complex; a simpler technique “works.”

Preview handling

  • The speaker suggests bypassing preview nodes and saving only the final output.

Node group 3: Character swap output (“subject + target pose”)

Inputs

  • Both generated images from groups 1 and 2

Prompt concept

  • A trigger-style prompt (speaker references a Civit AI example) instructing the model to place the character from image 1 into the pose/scene of image 2.

Logic described

  • Using both images + the prompt conditions so the character matches the target.

Output

  • Result -1: a clean character swap result.

Optional module: Improving face (mask-based refinement)

  • An “improving result -1 face” node group exists.
  • Condition
    • Optional; disabled unless needed.
  • How it works
    • Uses image 1 result as the base
    • Reloads the original subject image to create a face mask
    • Requires a mask—otherwise the node errors
  • Goal prompt
    • Change image 1 face to match the face from image 2
  • Outcome
    • In the speaker’s run, they didn’t see changes, but they provide it as a fix if face quality is wrong.
  • Outputs
    • Produces an “improved face” image, which can replace the earlier face reference in later steps.

LoRA detailing pass (adding realism/details)

  • After the base swap, a details pass generates “result -2” (or “result two”) using an added LoRA.
  • Speaker notes:
    • A prompt acts as a trigger word for the LoRA.
    • The LoRA connection is not always directly wired to the same graph section; diffusion may begin from the model, with another LoRA used above.
  • LoRA source
    • Downloaded from Hugging Face as a .safetensors file (labeled realistic).
  • Testing note
    • Works reasonably for a 3D character, but they recommend decreasing LoRA weight slightly.
  • Resolution note
    • After detailing, you may need to increase resolution, then upscale to 4K.

4K upscale (Seed VR2 / DIT + VAE)

  • Upscaling step to produce a 4K detailed image.
  • To make “seed VR2” work:
    • Requires a DIT model and VAE (downloaded automatically by the nodes).
  • VAE
    • A preferred VAE model is selected from available options.

Hardware considerations

  • Mentions smaller GGUF models for lower-memory GPUs.
  • Refers viewers to “Seed VR2 DIT models” and notes there are 3B and 7B sizes.
  • Guidance:
    • FP16 is good with a strong GPU; otherwise use smaller models.

Runtime result

  • Example run took about 175 seconds and produced a 4K image.

Final claimed outcome / comparison

The workflow’s final output produces a character swap that is:

  • clean
  • with details added
  • and 4K upscale results that look good in side-by-side comparison against the target/expected images.

Main speakers / sources

  • Main speaker: The video’s creator/instructor (no name provided in the subtitles).

External sources referenced

  • Civit AI (example prompt reference)
  • Hugging Face (LoRA download)
  • GitHub (Seed VR2 model details referenced)
  • GGUF models (loaded via U-Net loader; model files placed in the U-Net folder)

Original video