Blog · Sep 28, 2026
Interactive avatars, also called realtime avatars, are AI characters you can video call. They show up as tutors, sales reps, support agents, game characters, and companions, anywhere a voice agent could use a face. Language models already let these characters hold a conversation. The hard part is rendering them in realtime: a face and body that react live, fast enough to keep up with the conversation.
Interactive avatars have been rendered three ways so far: inside a 3D game engine, as a deepfake, and as a two-stage model that turns audio into an intermediate representation and then into pixels.
A character world model is a fourth way of rendering an interactive avatar. The face, the body, and the scene around them are generated live, frame by frame, as the conversation happens, rather than played back from something built ahead of time. This post covers each of these approaches, and what each one can and cannot show.
3D game engine avatars
A 3D model driven by blendshapes and canned animations. 3D models are controllable, predictable and cheap to run, but every character takes weeks or months to build. Every change — like a new outfit or gesture — takes another few weeks. And this approach is limited to cartoons.
Weeks to build
Cartoon only
Deepfake avatars
A video recording of a person talking is played forwards and backwards in an infinite loop, while new lip sync is applied to the mouth region. This approach never escapes the uncanny valley since the face and body language does not match the audio being said. Cannot be very expressive otherwise the gestures feel repetitive.
Always uncanny
Photorealistic only
Multi-hour training
Two-stage avatars
A two-stage model that goes from audio to keypoints (or some intermediate representation) and then keypoints to RGB pixels. The generative model can either be zero-shot or trained per character. Techniques range from a dead puppet moving its lips to great — but this approach cannot do avatar-object interactions, moving backgrounds, or non-human characters. No provider taking this approach has done hand gestures or whole body motions, though it is theoretically possible.
No avatar-object interactions
No non-human characters
Character world models
A character world model is a world model centered on a character. It listens, talks, and reacts in real time, and it generates the whole scene, not just a face. There is no rig, no recording, and no keypoint stage. Every pixel of every frame is generated from scratch, conditioned on the audio, a prompt, and what the character is doing and feeling.
LemonSlice is the first company to build a character world model. Character World Model 1 (CWM-1) is an end-to-end video diffusion transformer, and it runs alongside an emotion engine that reads the conversation and predicts the character's next actions and emotions. Nothing in it is specific to one character, so the same model animates a photo of you, an anime character, or a picture of your dog.
Every pixel generated, end to end
Any character, any style
Hands, objects, and moving scenes
Actions and emotions you can steer
In this recording, you can see CWM-1 in action. It generates the whole world around the character in realtime, including the background, interactions with objects, the character's body language, and the character's emotions.
The full write-up, with demos, is in Introducing Character World Model-1.
Where this goes
Each of the older approaches fixes what an avatar can show before any training happens: the rig, the recording, the keypoints. A character world model moves that decision into the data. What it can render is whatever it has learned to render, and that list keeps growing.
CWM-1 is where we are today. The avatar Turing test, a live video call where you cannot tell the other side is not a person, is where we are going.
Frequently asked questions





















