
Humans are not text interfaces.
We do not experience the world one token at a time. We speak, listen, see, move, interrupt, react, hesitate, express emotion, and adapt continuously. Human intelligence is inherently real-time, multimodal, and embodied.
AI is moving in that direction, but it is still early. Today's most advanced systems can reason across audio, vision, and text, and the field is already shifting from static chat interfaces toward low-latency, multimodal interaction. GPT-4o showed the importance of real-time audio, vision, and text reasoning; Gemini 1.5 demonstrated long-context multimodal understanding across documents, video, and audio; and robotics research like RT-2 showed how vision-language models can begin translating perception and language into action. (OpenAI)
But the hard problem is not only model intelligence.
The hard problem is real-time intelligence.
A real-time AI system cannot wait for the cloud for every decision. It cannot ignore latency, bandwidth, privacy, device constraints, or the emotional nuance of human communication. It must decide what should run on-device, what should run in the cloud, and how both should collaborate continuously.
That is why we are building a full-stack frontier research lab for real-time multimodal AI.
Our thesis is simple:
The next generation of AI will not be built as a chatbot.
It will be built as a real-time perception, reasoning, and action system.
And the winning architecture will be hybrid: on-device intelligence for immediacy, privacy, and responsiveness; cloud intelligence for scale, deep reasoning, and model orchestration; and a collaborative inference layer connecting both.
This is where we believe we are uniquely positioned. We are not starting from a blank page. We come from real-time infrastructure. We understand audio, video, streaming, latency, edge conditions, developer platforms, and production-grade communication systems. Most AI teams are trying to learn real-time media from the outside. We are approaching AI from inside the real-time stack.
That gives us a structural advantage.
Because real-time multimodal AI is not just a model problem. It is an infrastructure problem, a systems problem, a latency problem, a media problem, and a product problem — all at once.
Milestone #1: Solve Speech
The first modality to solve is speech.
Speech is the most natural interface between humans and machines, but current voice AI still lacks what makes human conversation feel alive: timing, interruption handling, tone, emotion, memory, and expressive delivery.
We are building the infrastructure and models required for emotionally intelligent speech AI.
The goal is not just speech-to-text and text-to-speech. The goal is a system that can understand what was said, how it was said, why it was said that way, and how it should respond emotionally.
Our speech stack:
The technical challenge is to make this happen with extremely low latency while preserving emotional fidelity. The system must detect speech, infer intent, understand emotion, reason about response, and generate expressive voice — all in real time.
This requires hybrid inference. Some parts must happen on-device: wake word, voice activity detection, partial transcription, emotion cues, interruption detection. Other parts can run in the cloud: deeper reasoning, long-context memory, large-model orchestration, and higher-quality generation.
The research direction is clear: voice AI is moving toward realtime, controllable, emotionally aware interaction, with production systems increasingly optimizing for latency, interruption, and live conversational behavior. (OpenAI)
Milestone #2: Solve Vision and Video
Once AI can listen and speak, it must learn to see.
Vision is not just image understanding. Real-time vision means continuously interpreting the world as it changes: people, objects, movement, context, gaze, gestures, screens, environments, and intent.
Video adds another layer: time.
A single image tells you what exists. Video tells you what is happening.
We are building infrastructure and models for real-time video understanding, avatar interaction, and eventually real-time generative video experiences. The goal is a unified system that can perceive, reason, and respond across speech and vision through a shared backbone.
Our vision stack:
This is where multimodal AI becomes interactive. A model should be able to hear what a user says, see what the user sees, understand the shared context, and respond through action, voice, or avatar.
The frontier is moving toward unified multimodal models that can process long video and audio streams, reason over them, and connect perception to action. Gemini 1.5's multimodal long-context work and vision-language-action research such as RT-2 point toward this future: models that do not merely describe the world, but begin to operate within it. (arXiv)
Milestone #3: Solve the Physical World
The final frontier is the world itself.
Once AI can hear, speak, see, and reason in real time, the next step is action.
Robotics is where multimodal AI becomes embodied. A robot needs to perceive its environment, understand human instruction, reason about physical constraints, and take action safely. That requires the convergence of speech, vision, planning, and low-latency control.
This is not possible with a cloud-only architecture. Robots need on-device autonomy for reliability, safety, and immediate response. But they also need cloud intelligence for larger-scale reasoning, fleet learning, simulation, and model updates.
So the architecture must again be hybrid.
On-device models provide perception and reflexes. Cloud models provide deeper reasoning and coordination. The real-time inference layer decides how intelligence is distributed across both.
Recent robotics research is already moving in this direction, including vision-language-action models that translate visual and language understanding into robotic control, and newer work exploring VLM deployment on edge infrastructure for real-time robotic perception. (arXiv)
Why Us
The next frontier of AI will belong to teams that can combine frontier models with real-time systems engineering.
That is our advantage.
We understand the real-time layer deeply: audio, video, networking, latency, devices, edge conditions, developer experience, and scalable infrastructure. We know that milliseconds matter. We know that media quality matters. We know that reliability matters. We know that production systems behave differently from demos.
This is exactly the substrate real-time AI needs.
Most AI labs begin with models and later discover the infrastructure bottleneck. We begin with the infrastructure, and we are moving upward into models.
That makes our path different.
We are building the full stack: models, runtime, inference orchestration, media infrastructure, on-device intelligence, cloud intelligence, and developer primitives.
Our belief is that real-time multimodal AI will not be solved by a single model call. It will be solved by a full-stack system designed from first principles for human-speed interaction.
Humans are multimodal.
AI must become multimodal.
And real-time is the bridge between intelligence that answers and intelligence that participates.
