HextaUI

Voice Mode

Talk to your assistant. Eight audio-reactive styles, from a cloudy sky and a ferrofluid blob to dithered pixels, ASCII, a CRT planet, halftone dots, a single ring and a soft aura, plus four reactive dots. A full-screen session with mute, interrupt and captions, an in-chat voice pill, a voice picker, and a browser engine that listens, waits for you to finish and replies.

Voice Mode is the visual and the controls for talking to an assistant. Pick one of eight shader styles (Sky, Dither, ASCII, CRT, Ferrofluid, Halftone, Ring or Aura) or four plain dots. Each listens, thinks and answers while moving with the sound, and monochrome styles follow your light or dark theme.

Audio levels are smoothed and measured against a learned noise floor, so a quiet room reads as quiet and the motion stays soft. Rendering pauses when the visual is off screen or the tab is hidden, and reduced motion holds every style still.

VoiceSession lays it out full screen at phone width, with captions, a status line, and mute and end within reach of a thumb. Tap the visual, press Space or just start talking to interrupt. VoiceComposer is a compact version that replaces the composer inside a chat.

useBrowserVoice runs a complete loop on the browser’s own speech recognition and speech: it waits for a pause before answering, can stay silent and show replies as captions, and ignores its own words when they come back through the speakers. For realtime voice APIs, pass the microphone and reply streams to useAudioLevels and set the state from the session’s events.

  1. Add the Pro registry to components.json

    components.json
    {
      "registries": {
        "@hextaui-pro": {
          "url": "https://hextaui.com/r/pro/{name}.json",
          "headers": {
            "Authorization": "Bearer ${HEXTAUI_PRO_TOKEN}"
          }
        }
      }
    }
  2. Add your token

    Create a token on your account page and put it in .env.local as HEXTAUI_PRO_TOKEN.

  3. Add the block

    pnpm dlx shadcn@latest add @hextaui-pro/voice-mode

With the browser’s voice

useBrowserVoice listens with the microphone and the browser’s speech recognition, waits for a pause, calls respond with what you said, and talks back with the browser’s speech. Return your model’s reply from respond.

"use client"

import { useBrowserVoice } from "@/components/blocks/voice-mode/use-browser-voice"
import { VoiceSession } from "@/components/blocks/voice-mode/voice-session"

export function VoiceChat({ onClose }: { onClose: () => void }) {
  const voice = useBrowserVoice({
    pauseMs: 900,
    respond: async (text, turns, signal) => {
      const response = await fetch("/api/voice", {
        method: "POST",
        body: JSON.stringify({ text, turns }),
        signal,
      })
      const { reply } = await response.json()
      return reply
    },
  })

  return (
    <VoiceSession
      state={voice.state}
      levels={voice.levels}
      muted={voice.muted}
      onMutedChange={voice.setMuted}
      onInterrupt={voice.interrupt}
      onEnd={() => {
        voice.end()
        onClose()
      }}
      captions={voice.captions}
      showCaptions
      error={voice.error}
    />
  )
}

With a realtime voice API

For speech-to-speech models over WebRTC, drive the state from the session’s events and feed both audio streams to useAudioLevels, so the orb follows you while you talk and the model while it answers.

"use client"

import * as React from "react"

import { useAudioLevels } from "@/components/blocks/voice-mode/audio"
import type { VoiceState } from "@/components/blocks/voice-mode/voice-orb"
import { VoiceSession } from "@/components/blocks/voice-mode/voice-session"

export function RealtimeVoice({ token }: { token: string }) {
  const [state, setState] = React.useState<VoiceState>("connecting")
  const [mic, setMic] = React.useState<MediaStream | null>(null)
  const [reply, setReply] = React.useState<MediaStream | null>(null)
  const [muted, setMuted] = React.useState(false)
  const peer = React.useRef<RTCPeerConnection | null>(null)
  const listening = useAudioLevels(mic)
  const speaking = useAudioLevels(reply)

  React.useEffect(() => {
    const connection = new RTCPeerConnection()
    peer.current = connection
    connection.ontrack = (event) => setReply(event.streams[0])
    const events = connection.createDataChannel("oai-events")
    events.onmessage = (message) => {
      const event = JSON.parse(message.data)
      if (event.type === "input_audio_buffer.speech_started") setState("listening")
      if (event.type === "input_audio_buffer.committed") setState("thinking")
      if (event.type === "output_audio_buffer.started") setState("speaking")
      if (event.type === "output_audio_buffer.stopped") setState("listening")
    }
    navigator.mediaDevices.getUserMedia({ audio: true }).then(async (stream) => {
      setMic(stream)
      stream.getTracks().forEach((track) => connection.addTrack(track, stream))
      const offer = await connection.createOffer()
      await connection.setLocalDescription(offer)
      const answer = await fetch("https://api.openai.com/v1/realtime/calls", {
        method: "POST",
        body: offer.sdp,
        headers: { Authorization: `Bearer ${token}`, "Content-Type": "application/sdp" },
      })
      await connection.setRemoteDescription({ type: "answer", sdp: await answer.text() })
      setState("listening")
    })
    return () => connection.close()
  }, [token])

  React.useEffect(() => {
    mic?.getAudioTracks().forEach((track) => (track.enabled = !muted))
  }, [mic, muted])

  return (
    <VoiceSession
      state={state}
      levels={state === "speaking" ? speaking : listening}
      muted={muted}
      onMutedChange={setMuted}
      onEnd={() => peer.current?.close()}
    />
  )
}

Inside a chat

VoiceComposer replaces the composer while a voice session runs, with a small orb, the status, a timer, mute and End. Your words and the replies can stream into the thread as text.

"use client"

import { useBrowserVoice } from "@/components/blocks/voice-mode/use-browser-voice"
import { VoiceComposer } from "@/components/blocks/voice-mode/voice-session"

export function ChatVoice({ respond }: { respond: (text: string) => Promise<string> }) {
  const voice = useBrowserVoice({ respond: (text) => respond(text) })

  if (voice.state === "idle") {
    return <button onClick={() => void voice.start()}>Start voice</button>
  }

  return (
    <VoiceComposer
      state={voice.state}
      levels={voice.levels}
      muted={voice.muted}
      onMutedChange={voice.setMuted}
      onInterrupt={voice.interrupt}
      onEnd={voice.end}
      startedAt={voice.startedAt}
    />
  )
}

Anatomy

The parts you compose, from the outside in.

PartDescription
VoiceVisualAny style by kind: sky, dither, ascii, crt, ferrofluid, halftone, ring, aura or bars.
VoiceOrbThe shader orb. Takes the state, mute and a levels ref.
VoiceBarsThe four-dot visual, with the same props.
VoiceShaderThe shared WebGL engine. Pass your own fragment shader to build a new style with the same inputs: state, smoothed level, four bands and three colors.
VoiceSessionThe full-screen layout: actions, visual, captions, status, mute and end.
VoiceComposerThe compact voice pill for inside a chat.
VoicePickerA popover of voices with tinted orbs.
useBrowserVoiceA complete voice loop on browser APIs: listen, detect the end of a turn, respond, speak, interrupt.
useAudioLevelsReads a level and four bands from any MediaStream or media element into a ref, without re-rendering.

VoiceVisual

Also accepts every div prop. Size it with a width class; it stays square.

PropTypeDefault
kindWhich style to render.
"sky" | "dither" | "ascii" | "crt" | "ferrofluid" | "halftone" | "ring" | "aura" | "bars""sky"
colorsThree colors for the shader. Hex, any CSS color, or currentColor. Each style ships its own.
[string, string, string]–
paletteSky colors, as used by VoicePicker voices.
{ deep, sky, cloud }–
stateDrives listening, thinking and speaking motion.
VoiceState–
levelsLive audio.
RefObject<VoiceLevels>–
mutedDesaturates and stops reacting. Ring turns dashed.
booleanfalse

VoiceOrb

Also accepts every div prop. Size it with a width class; it stays square.

PropTypeDefault
stateDrives churn, whirl and ripples.
"idle" | "connecting" | "listening" | "thinking" | "speaking" | "error""idle"
levelsLive audio, read every frame. Usually from useAudioLevels or useBrowserVoice.
RefObject<VoiceLevels>–
mutedDesaturates and stops reacting.
booleanfalse
paletteThree hex colors.
{ deep, sky, cloud }blue sky
PropTypeDefault
stateThe conversation state.
VoiceState–
levelsPassed to the visual.
RefObject<VoiceLevels>–
variantWhich style to show.
VoiceVisualKind"sky"
colorsOverride the style’s colors.
[string, string, string]–
mutedMute state.
booleanfalse
onMutedChangeShows the mute button and the M shortcut.
(muted: boolean) => void–
onInterruptMakes the visual a button while thinking or speaking, plus Space.
() => void–
onEndShows the end button and Escape.
() => void–
captionsLatest words from each side.
{ user, assistant }–
showCaptionsShow captions under the visual.
booleanfalse
onShowCaptionsChangeShows the captions toggle and C.
(show: boolean) => void–
statusReplaces the state label.
ReactNode–
footnoteBetween the buttons, such as a timer or minutes left.
ReactNode–
errorShown in place of the status.
string | null–
actionsTop-right controls, such as VoicePicker.
ReactNode–
paletteOrb colors.
VoicePalette–
PropTypeDefault
stateThe conversation state.
VoiceState–
levelsPassed to the small orb.
RefObject<VoiceLevels>–
onEndEnds voice and brings back the composer.
() => void–
muted / onMutedChangeMute button.
boolean / (muted) => void–
onInterruptTap the orb to interrupt.
() => void–
startedAtStart time in ms, for the timer.
number | null–

useBrowserVoice(options)

Returns { state, muted, setMuted, start, end, interrupt, levels, turns, captions, error, startedAt }.

PropTypeDefault
respondCalled when you finish a sentence. Return the reply; the signal aborts on interrupt.
(text, turns, signal) => Promise<string>–
pauseMsSilence before your turn ends.
number900
bargeInLet speaking interrupt the reply.
booleantrue
voiceOutputSpeak replies aloud. When false, replies appear as captions with the orb still animating.
booleantrue
voicePreferred browser voices and delivery.
{ names?, pitch?, rate? }–
greetingSaid when the session starts.
string–
langRecognition and speech language.
string"en-US"

useAudioLevels(source, target?)

Returns a ref of { level, bands } updated every frame.

PropTypeDefault
sourceMicrophone, remote WebRTC stream or an audio element.
MediaStream | HTMLMediaElement | null–
targetWrite into an existing ref instead of a new one.
RefObject<VoiceLevels>–
KeyAction
SpaceInterrupts while it’s thinking or speaking.
MMutes or unmutes your microphone.
CShows or hides captions.
EscEnds the voice session.
TabMoves through the captions toggle, voice picker, the orb when interruptible, mute and end.
  • The session is a labelled region, and the status line is a polite live region, so state changes like “Listening” and “Thinking” are announced.
  • The visual is a real button labelled “Interrupt” while it can be interrupted, and disabled otherwise.
  • Mute uses aria-pressed, and every control has a tooltip naming its shortcut.
  • Captions give a text alternative to everything said, and replies can be shown without audio.
  • Shortcuts are ignored while typing in a field, except Escape.
  • With reduced motion, the orb and dots hold still and only change with the state.
  • Without WebGL, every shader style falls back to a still gradient, and the dots work everywhere.

Code

9 files, added to components/blocks/voice-mode.