edge0

Glasses That See the World
and Finish the Task

Hybrid Edge-CloudOn-Device Multimodal UnderstandingMillisecond-Level LatencyFast SDK Integration

edge0 puts multimodal understanding on the glasses: recognition, reading and scene QA run from live camera frames, while voice carries conversation and complex tasks call the cloud.

Meta-Bounds smart glasses recognizing a museum exhibitSmart Glasses
High-Frequency Commands On-Device · ≤220msCamera + VoiceEVA ModelNarration / Structured Notes
95%
Device Command Recognition
High-frequency spoken commands recognized on-device
≤220ms
High-Frequency Command Latency
P95, from end of speech to device callback
-68%
Average Inference Cost
After on-device offload of frequent commands

Smart glasses live or die by response speed: people talk to what they see, and they expect an answer before the moment passes. With edge0’s EVA stack, recognition and command execution stay on the glasses, open questions reach the cloud only when needed, and everything returns as one spoken answer plus a structured UI our firmware can render directly.

Why an On-Device Model?

Glasses Must Respond Instantly
and Answer Open Questions

Device control and everyday interaction stay on the glasses; live data, open QA and complex tasks reach the cloud through the EVA Gateway on demand.

01

Voice Is the Primary Input

Smart glasses have no keyboard and no big screen. High-frequency actions must work the moment they are spoken; waiting on the cloud multiplies interaction friction.

Everyday Commands Run On-Device Instantly
02

Answers Must Drive the Interface

Scores, weather and calendars need a spoken answer and a structured UI at the same time. A plain text interface cannot carry that device experience.

Intent, Data and UI Schema in One Return
03

Complex Tasks Need the Cloud

Open QA and live data still need the network, so a single router must split high-frequency commands and complex tasks by scene.

EVA Gateway Orchestrates Edge and Cloud
All-Cloud Path
GlassesNetworkOverseas ModelTextSecond-Pass Parsing
High cost, latency varies
EVA Hybrid Edge-Cloud
VoiceIntent RoutingOn-Device / CloudVoice + UI
On Demand
Why EVA?

Teams Choose EVA for One Interface
Across Device, Cloud and UI

OEMs no longer build separate chains for command models, general LLMs, live data and UI generation; the Gateway and SDK handle unified routing and structured output.

95%
Device Command Recognition
≤220ms
High-Frequency Response P95
1h
First SDK Run-Through

Every Glance Becomes
Usable Information

The camera captures the scene and voice asks the question; the on-device model handles exhibit recognition, text understanding and key-point extraction, while complex QA and live data call the cloud through the EVA Gateway.

INPUT
Camera Frame + User Voice
EVA Router
Visual Intent Recognition  |  Task Routing
Cloud
Open QA  |  Live Info and Complex Reasoning
On demand
On-Device
Exhibit Recognition  |  Text Extraction  |  Whiteboard Notes
Low latency, stable
OUTPUT
Spoken Answer + Dynamic UI + Structured Notes
01

On-Device Multimodal Model

Image recognition, OCR, scene understanding and basic QA run on the glasses, working even on weak or broken networks.

02

Vision and Voice Together

Users ask about what is in front of them; the model combines the frame and the question to return a natural-language answer plus a dynamic UI.

03

Structured Information Extraction

Whiteboards, documents and field text become titles, key points, conclusions and to-dos, ready to save and call later.

04

Fast OEM Integration

Ships as an API Gateway and SDK inside glasses firmware, covering visual understanding, voice interaction, task orchestration and UI callbacks.

High-Frequency Commands on Device,
Complex Tasks Completed in the Cloud

The EVA SDK handles voice input, command recognition and local action callbacks on the glasses; the Gateway picks cloud services for weather, scores, calendars and open QA, then converts results into speech and UI Schema.

High-Frequency Commands On-DeviceSmart Intent RoutingComplex Tasks in the CloudDynamic UI Callbacks
On-Device Runtime
“Make the display a little brighter”
Local Speech Recognition
Local audio pre-processing
High-Frequency Command Match
Natural spoken input
Display Brightness Adjustment
Intent and parameter parsing
Status Feedback Callback
Device command callback
LOCAL ACTION
Display brightness raised
P95 ≤220ms
Weather, scores and open QA call the cloud through the EVA Gateway, billed on demand.

Everyday Actions Finish Instantly,
Complex Questions Get Full Answers

High-Frequency Commands On-Device
Metric: end of speech to device and UI callback.
≤220ms
User Finishes
0ms
On-Device Recognition
ASR + Intent
Brightness Raised
≤220ms
High-Frequency Actions Never Call the Cloud
Complex Tasks, Edge-Cloud Together
“What was the score of last night’s match?”
Intent Routing
Live data tasks
Cloud Lookup
On demand
Dynamic UI
Structured return
Result Preview
Full time 3:1 — spoken to the user in sync
Cloud Called Only for Open Questions and Live Data

Same Voice Interaction,
a New Boundary of What’s Usable

Key Metric
Generic Overseas Model API
EVA Hybrid Edge-Cloud
High-Frequency Commands
Every request goes to a generic overseas model API
On-device recognition and execution, P95 ≤220ms
Complex Tasks
Apps stitch models and data interfaces separately
EVA Gateway orchestrates cloud capabilities in one place
Interface Output
Text only, parsed again on the device
Spoken answer + dynamic UI Schema
Inference Cost
Every turn billed by cloud tokens
68% lower average cost after on-device routing
OEM Integration
Multiple vendors, multiple interfaces to adapt
One SDK for the full edge-cloud capability set

Subscribe for the latest news on Edge0 products