Real-time AI avatar

Your model,on camera,answering back.

KiKaSuite joins your video calls as a lip-synced 3D presence. It listens while you talk, thinks mid-sentence, and speaks back in under a second, on the language model, speech and voice providers you choose.

Round Trip

Under 900 ms

Providers

Any LLM

Deploy

Cloud or on premise

K
On air
KiKaSuite, EN, warm, 22 kHz

KiKaSuite
EN, warm, 22 kHz

GPT-4

/

Deepgram

/

ElevenLabs

+2 tools

Round Trip

312 ms

Gemini
OpenAI
Anthropic
Whisper
Deepgram
ElevenLabs
Cartesia
Google Cloud STT
LiveKit
MCP
VRM / GLB
Ollama
Capabilities

An open layer between your stack and the call

Not a black box with a face bolted on. Every stage of the loop with KiKaSuite is something you can inspect, replace, or run yourself.

Real-Time Voice

Streaming speech to text and text to speech run in parallel, so replies begin before the sentence ends. No awkward pauses, no push to talk.

Lifelike 3D Faces

Viseme accurate lip sync with micro expressions and gaze. Pick a ready made avatar or bring your own VRM or GLB model.

Any Language Model

Gemini, OpenAI compatible endpoints, or a local Ollama. Swap providers with a config change instead of a rewrite.

Tools Mid Sentence

Expose real capabilities over the Model Context Protocol. KiKaSuite can look something up while it is still talking to you.

Built for Teams

Organizations, credits, roles and audit trails ship in the box, the boring parts a real deployment needs.

Runs Where You Do

Hosted, in your own network, or fully offline on your own GPUs. The transport is the same either way.

The Loop

Five hops from your voice to hers

Every stage streams into the next instead of waiting for it to finish. That is the whole trick behind sub second replies.

01

You Speak

live

Mic audio streams in over the call in short frames.

02

Transcribe

about 140 ms

Partial transcripts arrive mid word from your speech to text provider.

03

Think

about 380 ms

Your model starts drafting on the first stable phrase, tools included.

04

Speak

about 180 ms

Tokens turn into audio chunks as they are generated.

05

Render

about 90 ms

Visemes drive the face and the video track goes out to the call.

Total Budget

790 ms median

P95

1.1 s

Surfaces

Built for product work, not just a demo reel

Four surfaces that turn a talking head into something you can actually ship and operate.

Live session

She sees the screen, hears the room, and answers like a participant.

Camera in, audio in, tools available. The same session works in a browser tab, a meeting bot, or an embed on your site.

Watch a Session
Preview

Avatar Editor

Shape the face, body and outfit against a live preview, or import your own VRM or GLB rig and keep your brand character.

View Details
Preview

Built-In AI Services

Memory, retrieval over your own documents, and task runners exposed as MCP tools she can call mid conversation.

View Details
Preview

Meeting Studio

Scene layouts, recording, and live transcripts for every call she joins.

View Details
Beta

On Premise

Run the whole loop inside your network on your own GPUs. Nothing leaves the building.

View Details
No account needed

Talk to her for a minute. That is the whole pitch.

Create a free KiKaSuite account, open a session, allow the mic, and say hello.