Skip to main content
Clip from our home page playground demo — a LiveKit agent built with the checklist on this page. We wrote these recommendations based on what we learned shipping it.

Use a ringing sound and UI

Avatar video takes a few seconds to initialize. A phone-call metaphor works well: users understand ringing, pickup, conversation, and hang-up.

Handle avatar connect/disconnect events

Read this section if you experience any of the following issues:
  • Empty room (user sees black screen)
  • User gets no video or audio, but audio/text input are still active
  • Avatar disappears
Use the @lemonsliceai/avatar npm package on the frontend to detect when the avatar’s first video frame has rendered, then switch from ringing to your active call UI. Listen for the avatar leaving to return to an inactive state so users can rejoin. See Use a ringing sound and UI for the UX pattern.
Also listen for ParticipantDisconnected when the avatar leaves (call completion, idle timeout, errors, etc.).
On the backend, listen for these events and call agent_session.generate_reply() when the LemonSlice avatar joins to prevent idle time before the avatar speaks. See the LiveKit starter project for a complete implementation.

Handle room errors

Read this section if you experience any of the following issues:
  • Avatar fails to join a call
  • Crashed calls — avatar leaves or audio cuts out
Listen for Disconnected events to handle network errors, WebRTC failures, or join failures.

Catch pipeline errors

Read this section if you experience any of the following issues:
  • Avatar does not speak
  • Session dies unexpectedly
Subscribe to AgentSession error events on your backend. Errors with err.recoverable == False mean the pipeline is dead — end the session gracefully.

Handle startup failures

Read this section if you experience any of the following issues:
  • The call never connects
If the avatar fails to join, give users a way to exit gracefully instead of waiting indefinitely.
The LiveKit starter project monitors agent startup and can send a failure message to the room. Catch it on the frontend:

Check timeouts

Read this section if you experience any of the following issues:
  • Avatar suddenly exits a call
  • Calls end after the same number of minutes every time
Several timeouts can affect a session. Confirm each is set to the value you intend:
  • LemonSlice idle timeout (default 60 seconds) — resets when the avatar is talking
  • LemonSlice GPU timeout (default 30 minutes) — contact support@lemonslice.com if you need longer calls
  • Third-party timeouts (LiveKit, Daily, ElevenLabs, etc.)
Set idle_timeout on AvatarSession:
Setting idle_timeout to -1 disables the LemonSlice idle timeout. Ensure sessions are properly terminated to avoid stale calls and runaway billing.

Budget your latency

Read this section if your avatar feels slow to respond — including delayed replies after the user finishes speaking.
If the conversation lags, the bottleneck is usually your LLM — not LemonSlice video. Check LLM response times and tool-calling latency first.
Agent pipeline latency
Humans expect a quick back-and-forth. Measure each pipeline step (STT, LLM, TTS, avatar video) and set a latency budget upfront. LLM times often creep from sub-second to three or four seconds as you add function calls, reasoning, or multimodal inputs. If a step blows your budget, move it off the critical path — run it async, set expectations with the user, or defer it.

Use VAD to decrease perceived latency

Perceived latency matters more than absolute latency. Voice Activity Detection (VAD) lets you start reacting before a user has fully finished speaking, which effectively pulls your entire pipeline forward. Good turn detection means you can kick off generation as soon as intent is clear.

Optional: Show a transcription

Displaying transcription can improve usability, especially for catching STT errors and making the agent’s timing feel more predictable. Users can see when their speech has been “accepted,” which reduces ambiguity around when the agent will respond. However, transcription introduces its own UX risks. Many pipelines expose both fast, low-accuracy interim results and slower, higher-accuracy final transcripts (e.g., interim_results=false to suppress partials for Deepgram). When both are surfaced, users can see text rapidly change or correct itself, which feels unstable and undermines trust. We recommend disabling interim results if you choose to show transcriptions.