Video summary

5 способов создать ИИ аватар в нейросетях | Полный гайд

Main summary

Key takeaways

Technology

Summary: 5 ways to create AI avatars (full guide)

The video presents five different technical approaches to create “non-roasted” (i.e., usable/realistic) AI avatars. It includes tutorial-style workflows using different services/models, plus practical constraints and recommendations around video formats, voice/emotion quality, motion artifacts, and token/cost tradeoffs. It also covers supporting tools such as prompt bots, teleprompters, and voice extraction.


1) Avatar from video using HGEN (face + voice transfer)

Core workflow

  • Record ~2 minutes of yourself talking on camera (horizontal or vertical).
  • Upload the video to HGEN.
  • Provide consent for using your face and voice.
  • Wait about ~30 minutes for generation.

After creation

  • Pick from multiple generated avatar versions (e.g., versions tailored for horizontal vs YouTube/lessons).
  • Generate new outputs by providing either:
    • a script (written text), or
    • a recorded voice uploaded directly.

Quality limitations

  • Script-to-speech may cause stress/emotion inaccuracies—the avatar can sound wrong if intonation matters.
  • Emotions can be generated incorrectly.

Gesture/motion behavior

  • The system uses the source video for gestures, and mainly modifies:
    • lip-sync
    • facial expressions
  • Recommendations:
    • Record gestures/emotion consistent with the desired final style (otherwise mismatches can occur—e.g., monotone vs enthusiastic).
    • Avoid covering the face with hands (can lead to errors).

Use cases + tradeoffs

  • Ideal for short conversational reels (Instagram/TikTok).
  • For long videos/lessons: you may spend less time recording yourself than dealing with avatar reruns; long runs can introduce more issues.

Review/guide element

  • Mentions a teleprompter review: selecting a Pixair prompter based on reviews.
  • Highlights benefits like speed/font control and live script editing with sync to a smartphone.

2) Avatar from photo + audio (simpler, uses HGEN via Synx)

Setup

  • Upload a portrait photo and an audio track (spoken or synthesized).
  • Generate through an aggregator-like workflow (Synx), selecting the HGEN model.

Audio extraction guide

The video explains ways to convert MP4 → MP3:

  • “Google MP4 to MP3” using free converters
  • Using a local tool (mentions “ccat” as a faster option) to extract audio by disabling video and keeping audio

Token/cost detail

  • Aspect ratio selection costs a small amount (~3 tokens mentioned).

Important capture recommendations

  • Don’t use photos where the avatar is too small or far away—aim for clear visibility of the mouth.
  • Keep the background simple:
    • fewer details
    • fewer people (background people may get unintentionally animated/warped)

Emotion/face clarity

  • Better framing improves facial expression fidelity and helps avoid blur artifacts.

3) “CLLE motion control” / motion transfer (animate a character using another video)

Goal

  • Transfer movements from a source video (dance/speech gestures) onto a character/avatar.
  • Example emphasis: maintaining pose/angle/lighting alignment to reduce errors.

Tooling workflow

  • Use Nanoban to generate/prepare the avatar character image:
    • create reference images (screenshot of source + character image from Google)
    • emphasize correct shadows + lighting so the character fits the scene
  • Then use Synx for motion transfer:
    • load the source video (movement)
    • load the avatar image (character)
    • prompt example: replace the person in the video with the character from the photo (including lighting)

Prompt sensitivity

  • The video recommends keeping prompt guidance; missing details can lead to artifacts or misplacement.

Practical note

  • Presented as “become anyone” style creative content, but results depend heavily on alignment (size/lighting/pose).

4) Avatar generation with Google Veo (prompt-based or first-frame-based)

Two approaches

1) Prompt-only generation

  • Generate directly from text prompts:
    • character description
    • dialogue
    • actions
  • Uses Hix as an aggregator.
  • Model choice mentioned: “Google Veo 3.1 Fast” (described as cheaper than a more expensive alternative).
  • Cost driver: writing long VO3 prompts manually can be tedious—so bots help generate structured prompts.

2) Animate using a first frame

  • First generate a consistent avatar image/frame (often via Nanoban).
  • Then animate it using that initial image to gain consistency.

Prompt automation (guide/review element)

  • Uses GPT-chat bots to generate large prompts.
  • Can accept a voice message to create the dialogue portion.
  • Mentions template-based workflows (Claude/GPT style).
  • Downside of prompt-only: less control over exact identity/background; it can vary between runs.

Outcome

  • Enables dialogues between characters and voice acting, but consistency typically requires the first-frame method.

5) UGC-style avatar creation via HGEN (“UGC bloggers”)

What it is

A guided/templated approach: choose a predefined UGC “blogger” avatar archetype (e.g., podcaster, ASMR, product advertiser).

Process

  • In Hix:
    • go to UGC Factory
    • choose the blogger archetype
  • Pick the model:
    • Veo 3 / Veo 3.1 Fast for lower cost
    • “720” mentioned as sufficient quality for social media
  • Upload:
    • a picture of the product (e.g., cream container)
    • optionally a generated/edited image of the avatar girl
  • Provide:
    • script/dialogue text
    • tone/intonation/language
    • background ambience (or “no”)

Key recommendation

  • Generating everything randomly is less efficient because video generation costs far more than images.
  • Better method:
    • first generate the exact avatar girl in Nanoban with the correct product context
    • refine details (lighting, camera placement)
    • then use the final girl image for UGC video generation

Common artifact warning

  • Product labels with small text/symbols can “float” or distort when held.
  • Suggestion: generate the product closer to camera to reduce issues.

Voice customization

  • Notes limited voice variety in the Google VO3.1 pipeline.
  • Workaround:
    • extract voice
    • use a dedicated voice changer to match the desired voice

Voice and dubbing support tool (within method 5)

Uses voice change/dubbing

  • After extracting MP3 voice earlier (via “capcat”), use a “best voice changer” tool mentioned as:
    • Laps neural network / Laps

Workflow

  • Choose Voice Changer
  • Upload extracted voice (MP3)
  • Select a target voice
  • Dubbing generates quickly (~3–5 seconds)
  • Export audio and continue editing in Capcat (or another editor) to remove the old voice and keep the new one

Quality tip

  • Voice matching may require multiple tries, especially when switching between mismatched languages/character types (e.g., Indian voice vs Russian speaker).

Cost

  • Generations described as cheap (“pennies”) or with initial free credits.

Main speakers / sources

  • Speaker: the video creator (“neurocreators” / “guys”), a single main narrator demonstrating tools and personal workflow/review notes.
  • Main services/models mentioned:
    • HGEN (avatar generation from video/photo)
    • Synx (video/photo avatar pipeline aggregator)
    • Nanoban (character/image generation and preparation)
    • Veo (Google Veo 3.1 Fast / VO3-style avatar generation)
    • Hix (aggregator used for Veo generation and UGC Factory)
    • ElevenLabs (voice generation/voice outputs referenced)
    • Capcat (audio extraction and editing referenced)
    • Laps neural network (voice changer/dubbing referenced)
    • Pixair prompter (teleprompter product mentioned with a review)

Original video