Voice Mode
Talk to your assistant. Eight audio-reactive styles, from a cloudy sky and a ferrofluid blob to dithered pixels, ASCII, a CRT planet, halftone dots, a single ring and a soft aura, plus four reactive dots. A full-screen session with mute, interrupt and captions, an in-chat voice pill, a voice picker, and a browser engine that listens, waits for you to finish and replies.
Voice Mode is the visual and the controls for talking to an assistant. Pick one of eight shader styles (Sky, Dither, ASCII, CRT, Ferrofluid, Halftone, Ring or Aura) or four plain dots. Each listens, thinks and answers while moving with the sound, and monochrome styles follow your light or dark theme.
Audio levels are smoothed and measured against a learned noise floor, so a quiet room reads as quiet and the motion stays soft. Rendering pauses when the visual is off screen or the tab is hidden, and reduced motion holds every style still.
VoiceSession lays it out full screen at phone width, with captions, a status line, and mute and end within reach of a thumb. Tap the visual, press Space or just start talking to interrupt. VoiceComposer is a compact version that replaces the composer inside a chat.
useBrowserVoice runs a complete loop on the browser’s own speech recognition and speech: it waits for a pause before answering, can stay silent and show replies as captions, and ignores its own words when they come back through the speakers. For realtime voice APIs, pass the microphone and reply streams to useAudioLevels and set the state from the session’s events.
Add the Pro registry to components.json
components.json Add your token
Create a token on your account page and put it in
.env.localasHEXTAUI_PRO_TOKEN.Add the block
pnpm dlx shadcn@latest add @hextaui-pro/voice-mode
With the browser’s voice
useBrowserVoice listens with the microphone and the browser’s speech recognition, waits for a pause, calls respond with what you said, and talks back with the browser’s speech. Return your model’s reply from respond.
With a realtime voice API
For speech-to-speech models over WebRTC, drive the state from the session’s events and feed both audio streams to useAudioLevels, so the orb follows you while you talk and the model while it answers.
Inside a chat
VoiceComposer replaces the composer while a voice session runs, with a small orb, the status, a timer, mute and End. Your words and the replies can stream into the thread as text.
Anatomy
The parts you compose, from the outside in.
| Part | Description |
|---|---|
VoiceVisual | Any style by kind: sky, dither, ascii, crt, ferrofluid, halftone, ring, aura or bars. |
VoiceOrb | The shader orb. Takes the state, mute and a levels ref. |
VoiceBars | The four-dot visual, with the same props. |
VoiceShader | The shared WebGL engine. Pass your own fragment shader to build a new style with the same inputs: state, smoothed level, four bands and three colors. |
VoiceSession | The full-screen layout: actions, visual, captions, status, mute and end. |
VoiceComposer | The compact voice pill for inside a chat. |
VoicePicker | A popover of voices with tinted orbs. |
useBrowserVoice | A complete voice loop on browser APIs: listen, detect the end of a turn, respond, speak, interrupt. |
useAudioLevels | Reads a level and four bands from any MediaStream or media element into a ref, without re-rendering. |
VoiceVisual
Also accepts every div prop. Size it with a width class; it stays square.
| Prop | Type | Default |
|---|---|---|
kindWhich style to render. | "sky" | "dither" | "ascii" | "crt" | "ferrofluid" | "halftone" | "ring" | "aura" | "bars" | "sky" |
colorsThree colors for the shader. Hex, any CSS color, or currentColor. Each style ships its own. | [string, string, string] | – |
paletteSky colors, as used by VoicePicker voices. | { deep, sky, cloud } | – |
stateDrives listening, thinking and speaking motion. | VoiceState | – |
levelsLive audio. | RefObject<VoiceLevels> | – |
mutedDesaturates and stops reacting. Ring turns dashed. | boolean | false |
VoiceOrb
Also accepts every div prop. Size it with a width class; it stays square.
| Prop | Type | Default |
|---|---|---|
stateDrives churn, whirl and ripples. | "idle" | "connecting" | "listening" | "thinking" | "speaking" | "error" | "idle" |
levelsLive audio, read every frame. Usually from useAudioLevels or useBrowserVoice. | RefObject<VoiceLevels> | – |
mutedDesaturates and stops reacting. | boolean | false |
paletteThree hex colors. | { deep, sky, cloud } | blue sky |
| Prop | Type | Default |
|---|---|---|
stateThe conversation state. | VoiceState | – |
levelsPassed to the visual. | RefObject<VoiceLevels> | – |
variantWhich style to show. | VoiceVisualKind | "sky" |
colorsOverride the style’s colors. | [string, string, string] | – |
mutedMute state. | boolean | false |
onMutedChangeShows the mute button and the M shortcut. | (muted: boolean) => void | – |
onInterruptMakes the visual a button while thinking or speaking, plus Space. | () => void | – |
onEndShows the end button and Escape. | () => void | – |
captionsLatest words from each side. | { user, assistant } | – |
showCaptionsShow captions under the visual. | boolean | false |
onShowCaptionsChangeShows the captions toggle and C. | (show: boolean) => void | – |
statusReplaces the state label. | ReactNode | – |
footnoteBetween the buttons, such as a timer or minutes left. | ReactNode | – |
errorShown in place of the status. | string | null | – |
actionsTop-right controls, such as VoicePicker. | ReactNode | – |
paletteOrb colors. | VoicePalette | – |
| Prop | Type | Default |
|---|---|---|
stateThe conversation state. | VoiceState | – |
levelsPassed to the small orb. | RefObject<VoiceLevels> | – |
onEndEnds voice and brings back the composer. | () => void | – |
muted / onMutedChangeMute button. | boolean / (muted) => void | – |
onInterruptTap the orb to interrupt. | () => void | – |
startedAtStart time in ms, for the timer. | number | null | – |
useBrowserVoice(options)
Returns { state, muted, setMuted, start, end, interrupt, levels, turns, captions, error, startedAt }.
| Prop | Type | Default |
|---|---|---|
respondCalled when you finish a sentence. Return the reply; the signal aborts on interrupt. | (text, turns, signal) => Promise<string> | – |
pauseMsSilence before your turn ends. | number | 900 |
bargeInLet speaking interrupt the reply. | boolean | true |
voiceOutputSpeak replies aloud. When false, replies appear as captions with the orb still animating. | boolean | true |
voicePreferred browser voices and delivery. | { names?, pitch?, rate? } | – |
greetingSaid when the session starts. | string | – |
langRecognition and speech language. | string | "en-US" |
useAudioLevels(source, target?)
Returns a ref of { level, bands } updated every frame.
| Prop | Type | Default |
|---|---|---|
sourceMicrophone, remote WebRTC stream or an audio element. | MediaStream | HTMLMediaElement | null | – |
targetWrite into an existing ref instead of a new one. | RefObject<VoiceLevels> | – |
| Key | Action |
|---|---|
| Space | Interrupts while it’s thinking or speaking. |
| M | Mutes or unmutes your microphone. |
| C | Shows or hides captions. |
| Esc | Ends the voice session. |
| Tab | Moves through the captions toggle, voice picker, the orb when interruptible, mute and end. |
- The session is a labelled region, and the status line is a polite live region, so state changes like “Listening” and “Thinking” are announced.
- The visual is a real button labelled “Interrupt” while it can be interrupted, and disabled otherwise.
- Mute uses aria-pressed, and every control has a tooltip naming its shortcut.
- Captions give a text alternative to everything said, and replies can be shown without audio.
- Shortcuts are ignored while typing in a field, except Escape.
- With reduced motion, the orb and dots hold still and only change with the state.
- Without WebGL, every shader style falls back to a still gradient, and the dots work everywhere.
Code
9 files, added to components/blocks/voice-mode.