Engineering Insights

Multimodal Product UX: Vision and Voice That Earn Their Place

August 12, 2026 Volkan Karataş

Multimodal UX means the product can take more than text — a photo, a screenshot, a spoken sentence — and respond in the channel that fits the job. It is not a camera icon on every screen. It is a faster path for a task users already fail at with a keyboard.

On EduKidGames, a child pointing the camera at a worksheet is faster than typing the exercise number. On Polylingo, speaking a phrase is the product; typing is the fallback. On our admin tools, a screenshot of a broken layout is a better bug report than a paragraph. Those three sentences are the whole strategy. If we cannot name the task, we do not add the modality.

Latency is part of the interface

A voice control that answers in four seconds feels broken. A vision feature that spins while we upload a 12 MB PNG feels broken. We compress on device, stream partial transcripts, and show a skeleton of the result before the model finishes. Users forgive a slightly worse classification if the first pixels arrive in a few hundred milliseconds. They do not forgive a silent wait.

Errors have to be speakable

“Something went wrong” is worse in voice than on a form. We write error copy as if it will be read aloud: what happened, what to try, how to switch to typing. For vision, we show the crop we actually sent the model. If we guessed the wrong region of the worksheet, the parent should see that — not a confident wrong mark.

Accessibility is the design review

A voice-only flow locks out kids in a quiet classroom and users who cannot speak. A camera-only flow locks out devices without a camera permission. Every multimodal path we ship has a text equivalent, and we test it on a cheap Android the way we test a landing page on a slow 4G. Izmir school Wi-Fi is part of our QA matrix, not an afterthought.

Questions we keep getting

Should every SaaS add a voice assistant? No. Add voice where hands or literacy are the bottleneck. Admin tables are not that place.

Do you train your own vision model? Rarely. We use a hosted vision model with a tight prompt and a schema, then train only if the error is domain-specific (handwritten marks, our worksheet layout).

What about privacy? Photos of children never go to a training bucket. They hit inference, then they are deleted on the schedule in the inventory — see our AI Act notes.