edge0

AI for Kids That Responds
Faster and Reads the Mood

On-Device ModelHybrid Edge-CloudEmotion RecognitionCommand Dispatch

Kidodo is edge0’s in-house AI hardware brand for early learning at home, maker of the AI companion robot and the Kouyubao spoken-English translator. Both products run on one hybrid edge-cloud voice stack with an on-device multimodal model: read-along and high-frequency commands complete on-device, open conversation goes to the cloud on demand, and emotion tags plus device commands are returned in the same pass.

Child interacting with the Kidodo AI companion robotEarly Learning
Two Products, One Voice CoreChild VoiceIntent & EmotionReply + Command
98.8%
Scene Command Recognition
Natural spoken-language test set on parent-child hardware
≤100ms
Local Command Response
P95, from end of speech to action callback
-75%
Cloud Token Consumption
High-frequency intents served on-device

Voice conversation in a product for kids has to be fast and stable, and it has to hear how the child feels in the moment. edge0’s hybrid edge-cloud pipeline lets the companion robot and Kouyubao answer instantly in high-frequency interaction, keep full conversations on open topics, and put emotion signals straight into how we encourage and what the device does next.

Why Hybrid Edge-Cloud?

Companionship Needs Instant Replies,
Learning Needs Full Intelligence

The on-device side guarantees speed and stability for high-frequency interaction; the cloud carries open QA, long conversations and knowledge content. One voice pipeline schedules both paths.

01

Kids Don’t Wait for the System to Think

Read-along, question-and-answer and device control happen inside continuous conversation. A pause caused by network jitter breaks a child’s attention and willingness to speak.

High-Frequency Interaction P95 ≤240ms
02

Responses Need Emotional Understanding

The same sentence can carry joy, hesitation or frustration. Recognizing words alone makes companionship and learning feedback feel flat.

Joint Semantic and Emotion Recognition
03

Two Task Types Need Two Paths

Device control wants instant stability; open dialogue and knowledge QA need cloud capability. A single all-cloud pipeline cannot serve both experience and cost.

On-Device First, Cloud on Demand
All-Cloud Path
DeviceNetworkVoice ServiceCloud ModelBack to Device
Latency follows the network
Hybrid Edge-Cloud
VoiceOn-Device UnderstandingLocal Command / Cloud DialogueUnified Reply
≤100ms
Why edge0?

One Voice Foundation for
Companionship, Learning and Translation

The AI companion robot and Kouyubao reuse the same recognition, understanding, emotion judgement and routing capabilities, then connect to their own product features through a standardized command interface.

2
Products Sharing One Voice Stack
≤100
Command Response P95 (ms)
98.5%
Scene Command Recognition
1h
First SDK Run-Through

One Entry Point, Routed to
the Right Capability

On-device handles high-frequency commands, core learning flows and emotional cues; the cloud adds knowledge, translation and open dialogue. Everything returns as one voice reply plus actions.

INPUT
“Child Voice”
Voice Router
Semantics  |  Intent  |  Emotion  |  Task Routing
Cloud
Knowledge QA  |  Translation  |  Open Dialogue
On demand
On-Device
Device Control  |  Read-Along  |  Emotion Signals
Low latency, stable
OUTPUT
Natural Reply + Emotion Tags + Device Commands
01

Tuned for Children’s Speech

Recognition is adapted to children’s pitch, speech rate, incomplete articulation and household noise, keeping real parent-child scenes stable.

02

Unified Edge-Cloud Routing

Wake-word, device control and fixed learning flows stay on-device; open QA and long conversations route to the cloud by intent.

03

Emotion Signals Enter the Dialogue

Emotion tags are emitted alongside content understanding, so the companion robot can adjust tone, encouragement and pacing.

04

Standardized Command Dispatch

Natural expressions become standard commands for playback, read-along, translation, volume and content switching — reused by both products.

Natural Responses for Kids,
Actionable Understanding for Devices

Instant High-Frequency Response
≤100ms
Child Finishes
0ms
On-Device Understanding
ASR + Intent
Voice & Command Feedback
≤240ms
High-Frequency Flows Never Wait for a Cloud Round-Trip
Emotion Recognition & Response
“I don’t really want to practice today…”
Voice Input
Content + Acoustic Cues
Joint Understanding
Emotion + Intent
Strategy Generation
Reply + Task Adjustment
Recognition Result
“A little frustrated”
Switches to Encouraging Replies, Lowers Task Difficulty

Same Voice Interaction,
a New Boundary of What’s Usable

Key Metric
Single All-Cloud Voice Stack
edge0 Hybrid Edge-Cloud
High-Frequency Commands
Every request uploaded to the cloud in full
On-device recognition and execution, P95 ≤240ms
Open Dialogue
One pipeline carries every task
Complex intents routed to cloud models on demand
Emotional Feedback
Mostly literal text understanding
Semantic, intent and emotion tags emitted together
Inference Cost
Tokens consumed on every turn
75% lower token consumption after on-device offload
Multi-Product Reuse
Robot and translator integrated separately
One SDK powers both parent-child devices

Subscribe for the latest news on Edge0 products