Video summary
5 способов создать ИИ аватар в нейросетях | Полный гайд
Main summary
Key takeaways
Summary: 5 ways to create AI avatars (full guide)
The video presents five different technical approaches to create “non-roasted” (i.e., usable/realistic) AI avatars. It includes tutorial-style workflows using different services/models, plus practical constraints and recommendations around video formats, voice/emotion quality, motion artifacts, and token/cost tradeoffs. It also covers supporting tools such as prompt bots, teleprompters, and voice extraction.
1) Avatar from video using HGEN (face + voice transfer)
Core workflow
- Record ~2 minutes of yourself talking on camera (horizontal or vertical).
- Upload the video to HGEN.
- Provide consent for using your face and voice.
- Wait about ~30 minutes for generation.
After creation
- Pick from multiple generated avatar versions (e.g., versions tailored for horizontal vs YouTube/lessons).
- Generate new outputs by providing either:
- a script (written text), or
- a recorded voice uploaded directly.
Quality limitations
- Script-to-speech may cause stress/emotion inaccuracies—the avatar can sound wrong if intonation matters.
- Emotions can be generated incorrectly.
Gesture/motion behavior
- The system uses the source video for gestures, and mainly modifies:
- lip-sync
- facial expressions
- Recommendations:
- Record gestures/emotion consistent with the desired final style (otherwise mismatches can occur—e.g., monotone vs enthusiastic).
- Avoid covering the face with hands (can lead to errors).
Use cases + tradeoffs
- Ideal for short conversational reels (Instagram/TikTok).
- For long videos/lessons: you may spend less time recording yourself than dealing with avatar reruns; long runs can introduce more issues.
Review/guide element
- Mentions a teleprompter review: selecting a Pixair prompter based on reviews.
- Highlights benefits like speed/font control and live script editing with sync to a smartphone.
2) Avatar from photo + audio (simpler, uses HGEN via Synx)
Setup
- Upload a portrait photo and an audio track (spoken or synthesized).
- Generate through an aggregator-like workflow (Synx), selecting the HGEN model.
Audio extraction guide
The video explains ways to convert MP4 → MP3:
- “Google MP4 to MP3” using free converters
- Using a local tool (mentions “ccat” as a faster option) to extract audio by disabling video and keeping audio
Token/cost detail
- Aspect ratio selection costs a small amount (~3 tokens mentioned).
Important capture recommendations
- Don’t use photos where the avatar is too small or far away—aim for clear visibility of the mouth.
- Keep the background simple:
- fewer details
- fewer people (background people may get unintentionally animated/warped)
Emotion/face clarity
- Better framing improves facial expression fidelity and helps avoid blur artifacts.
3) “CLLE motion control” / motion transfer (animate a character using another video)
Goal
- Transfer movements from a source video (dance/speech gestures) onto a character/avatar.
- Example emphasis: maintaining pose/angle/lighting alignment to reduce errors.
Tooling workflow
- Use Nanoban to generate/prepare the avatar character image:
- create reference images (screenshot of source + character image from Google)
- emphasize correct shadows + lighting so the character fits the scene
- Then use Synx for motion transfer:
- load the source video (movement)
- load the avatar image (character)
- prompt example: replace the person in the video with the character from the photo (including lighting)
Prompt sensitivity
- The video recommends keeping prompt guidance; missing details can lead to artifacts or misplacement.
Practical note
- Presented as “become anyone” style creative content, but results depend heavily on alignment (size/lighting/pose).
4) Avatar generation with Google Veo (prompt-based or first-frame-based)
Two approaches
1) Prompt-only generation
- Generate directly from text prompts:
- character description
- dialogue
- actions
- Uses Hix as an aggregator.
- Model choice mentioned: “Google Veo 3.1 Fast” (described as cheaper than a more expensive alternative).
- Cost driver: writing long VO3 prompts manually can be tedious—so bots help generate structured prompts.
2) Animate using a first frame
- First generate a consistent avatar image/frame (often via Nanoban).
- Then animate it using that initial image to gain consistency.
Prompt automation (guide/review element)
- Uses GPT-chat bots to generate large prompts.
- Can accept a voice message to create the dialogue portion.
- Mentions template-based workflows (Claude/GPT style).
- Downside of prompt-only: less control over exact identity/background; it can vary between runs.
Outcome
- Enables dialogues between characters and voice acting, but consistency typically requires the first-frame method.
5) UGC-style avatar creation via HGEN (“UGC bloggers”)
What it is
A guided/templated approach: choose a predefined UGC “blogger” avatar archetype (e.g., podcaster, ASMR, product advertiser).
Process
- In Hix:
- go to UGC Factory
- choose the blogger archetype
- Pick the model:
- Veo 3 / Veo 3.1 Fast for lower cost
- “720” mentioned as sufficient quality for social media
- Upload:
- a picture of the product (e.g., cream container)
- optionally a generated/edited image of the avatar girl
- Provide:
- script/dialogue text
- tone/intonation/language
- background ambience (or “no”)
Key recommendation
- Generating everything randomly is less efficient because video generation costs far more than images.
- Better method:
- first generate the exact avatar girl in Nanoban with the correct product context
- refine details (lighting, camera placement)
- then use the final girl image for UGC video generation
Common artifact warning
- Product labels with small text/symbols can “float” or distort when held.
- Suggestion: generate the product closer to camera to reduce issues.
Voice customization
- Notes limited voice variety in the Google VO3.1 pipeline.
- Workaround:
- extract voice
- use a dedicated voice changer to match the desired voice
Voice and dubbing support tool (within method 5)
Uses voice change/dubbing
- After extracting MP3 voice earlier (via “capcat”), use a “best voice changer” tool mentioned as:
- Laps neural network / Laps
Workflow
- Choose Voice Changer
- Upload extracted voice (MP3)
- Select a target voice
- Dubbing generates quickly (~3–5 seconds)
- Export audio and continue editing in Capcat (or another editor) to remove the old voice and keep the new one
Quality tip
- Voice matching may require multiple tries, especially when switching between mismatched languages/character types (e.g., Indian voice vs Russian speaker).
Cost
- Generations described as cheap (“pennies”) or with initial free credits.
Main speakers / sources
- Speaker: the video creator (“neurocreators” / “guys”), a single main narrator demonstrating tools and personal workflow/review notes.
- Main services/models mentioned:
- HGEN (avatar generation from video/photo)
- Synx (video/photo avatar pipeline aggregator)
- Nanoban (character/image generation and preparation)
- Veo (Google Veo 3.1 Fast / VO3-style avatar generation)
- Hix (aggregator used for Veo generation and UGC Factory)
- ElevenLabs (voice generation/voice outputs referenced)
- Capcat (audio extraction and editing referenced)
- Laps neural network (voice changer/dubbing referenced)
- Pixair prompter (teleprompter product mentioned with a review)