edge0

Sport Earbuds That Understand and Act, Without the Cloud

Beyond Cloud IntelligenceMillisecond-ReadyEasy DeploymentLow Token Cost

SHOKZ keeps high-frequency voice commands in a local loop with the edge0 on-device speech model, sending complex intents to the cloud on demand — instant, offline-ready and far fewer tokens.

Runner wearing SHOKZ open-ear sport earbudsSmart Earbuds
No Cloud Round-Trip · ≤100msVoiceOn-Device ModelDevice Command
95%
Command Recognition Accuracy
Outdoor-sport noise test set
≤100ms
On-Device Response Latency
P95, speech end to action callback
-92%
Cloud Token Consumption
High-frequency intents closed on-device

We want voice to be part of the earbuds themselves: no network dependency — the moment a user speaks, the device acts. edge0 delivered natural-language understanding, millisecond-level response and local execution inside our existing hardware resources, so we can ship a stable, consistent experience on every unit while keeping cloud cost under control at scale.

Why On-Device Models?

The Cloud Can Complete a Voice Interaction,
but Cannot Guarantee Every Command Works

Open-domain QA can go to the cloud on demand; playback control, sport data and device state are high-frequency core capabilities that must run locally.

01

The Network Can't Be a Prerequisite

Running and cycling routes cross weak-signal and fully offline zones. A voice experience that depends entirely on the cloud fails the moment the network does.

Core Commands 100% Available Locally
02

Control Must Happen Instantly

Skip, pause and volume are instant operations. A cloud round-trip adds visible waiting and breaks the feel of direct control.

P95 Response ≤100ms
03

Scale Must Not Multiply Cost

Shipped earbuds generate huge volumes of high-frequency, simple requests. Paying cloud inference per call turns into a permanent, growing cost.

Total Token Consumption -92%
Full Cloud Path
EarbudsPhoneNetworkCloud ModelBack to Earbuds
1000-2500ms
edge0 On-Device Path
EarbudsOn-Device ModelExecute
≤100ms
Why edge0?

Customers Choose edge0 Because It Puts
Near-Cloud Understanding into Existing Consumer Hardware

Traditional on-device solutions are lightweight but locked to fixed commands; cloud models understand deeply but bring network and cost constraints. edge0 connects the two with a higher “intelligence density”.

18MB
INT8 Model Package
≤300MB
Peak Runtime Memory
95%
Simple-Command Accuracy
1h
First SDK Run-Through

A Smaller Model, Carrying Denser Intelligence

Inside an 18MB package and 42MB peak memory: speech recognition, natural-language intent understanding and device command callbacks.

On-Device AvailabilityHighLowLowNatural-Language UnderstandingHigh
Fixed Keyword EnginesOn-device, but understanding is limited
Natural understanding, runs on consumer hardware
Fully Cloud Audio ModelsStrong understanding, cannot run locally

*The upper-right region represents models that understand natural phrasing and run directly on consumer hardware. Positions illustrate capability, not test scores.

01

Near-Cloud Understanding

From fixed commands to natural phrasing: “next track”, “change song” and “skip this one” all map to the same device intent.

02

Runs on Consumer Hardware

Model compression, INT8 quantization and runtime optimization fit ASR + NLU inside existing memory and compute budgets.

03

Tuned for Sport Noise

Adapted to wind, footsteps, traffic and music playback for stable outdoor command recognition.

04

SDK Wired to Device Capabilities

Standardized intents call playback and motion APIs directly — no hardware rework for the brand.

On-Device First, Cloud on Demand —
Every Interaction Spends Only the Tokens It Needs

The edge0 SDK covers the full chain from voice input, recognition and intent understanding to task routing and command callbacks, switching paths automatically. High-frequency tasks such as playback control and motion-status queries close on-device; complex intents and open-domain QA go to the cloud on demand and return through one unified result. Brands never stitch two pipelines together, cloud calls drop, and core local commands keep working on weak or missing network.

On-Device Runtime
“Play me a different song”
VAD / Denoise
Local audio preprocessing
Edge ASR
Natural speech recognition
Edge NLU
Intent & slot parsing
Action
Device command callback
Local ActionNext track playingP95 ≤180ms
Open-domain QArouted to cloud on demand

Users Feel Two Things in the End:
Instant Response, and It Works Offline

Instant Response
≤100ms
User Finishes
0ms
On-Device Processing
ASR + NLU
Volume Up
≤180ms
No Waiting for a Network Round-Trip
Works Offline
“How am I doing right now?”
User Voice
On-Device Processing
Motion Data
Results
5’24″
Pace
148
Heart Rate
176
Cadence
Still Announced Locally

Same Voice Interaction —
On-Device Changes the Usable Boundary of the Product

Key Metric
Cloud / Keyword Solutions
edge0 On-Device Model
Core Command Availability
Depends on phone and network
Works on weak or no network
Command Response
About 1000-2500ms, varies with network
P95 ≤100ms, local callback
Cloud Inference Cost
Continuous spend per recognition and parse
Zero tokens for high-frequency requests, -92% overall
Natural Language
Supported, but slow
Many phrasings map to one intent
Hardware Changes
New network and cloud link required
18MB SDK embeds into the existing design

Subscribe for the latest news on Edge0 products