Dictate into Linux applications using Doubao, Volcengine or Deepgram, with optional text polishing and configurable recording shortcuts.
What you can do with Doubao Say
Dictate text using a Doubao web account and see recognition results while speaking.
Since v1.3.1
Good to know
Audio is sent to Doubao after a recording gesture is confirmed; this unofficial backend depends on Doubao's web protocol.
Dictate with the official Volcengine Seed ASR 2.0 streaming service using your own API key.
Since v1.3.1
Good to know
Requires service access and an API key; audio is uploaded and usage is billed to your account.
Dictate US English with Deepgram Nova-3 and view interim and final transcripts.
Since v1.3.1
Good to know
Requires your own API key and paid usage; other languages and multilingual mode are not exposed.
Paste completed dictation into the focused application on Hyprland.
Since v1.3.1
Good to know
Requires `wl-copy` and a known, unchanged target window. Focus checks are not atomic; same-window caret movement is not detected. Clipboard managers may retain transcripts.
Paste completed dictation into applications and terminals on native X11.
Since v1.3.1
Good to know
Requires `xclip`, `xdotool` and usable window/PID metadata for automatic paste. Missing helpers retain text for copying. XWayland is not native X11; direct typing is unavailable.
Type dictated text into a Hyprland application without changing the clipboard.
Since v1.3.1
Good to know
Requires optional `wtype`. Newlines/tabs send Enter/Tab and may submit or execute text. Cancellation cannot undo delivered characters; failure retains text without clipboard fallback.
Start and finish dictation with tap or hold gestures using a chosen key or recorded shortcut.
Since v1.3.1
Good to know
Only one trigger is active. Fn must be reported as a Linux key. Double-tap sends Enter and may submit messages or execute commands; this gesture is configurable.
Select a PipeWire microphone and test dictation without inserting text into another application.
The three-second microphone check stays local; the voice test uses the selected online recognizer and never pastes or sends Enter.
Show all 15
Copy the most recent transcript after recognition or automatic delivery fails.
Since v1.3.1
Good to know
Only one in-memory result is retained; exiting loses it. Partial recognition can be incomplete and already delivered text cannot be undone.
Recover the previous clipboard payload after a dictated paste.
Since v1.3.1
Good to know
Best-effort restoration of one payload, only while the clipboard still contains dictation text. Rich-text formatting is lost when plain text and HTML coexist; history is not suppressed.
Control recording and send custom shortcuts with Vibekey buttons and dial controls.
Off by default; requires receiver access through the documented VID/PID-specific udev rule. No Vibekey software dependency is required.
Choose a listening waveform and reduce its animation while viewing live recognized text.
Since v1.3.1
Good to know
Classic bars is the default; reduced updates retain volume feedback. Text is inserted only after recording ends.
Use the application interface in English or Simplified Chinese, or follow the session language.
Since v1.3.1
Good to know
System mode uses Simplified Chinese for Chinese locales and English otherwise; the hosted sign-in website controls its own language.
See when a newer stable release is available and open its GitHub release page.
Since v1.3.1
Good to know
Checks contact GitHub at most once per 24 hours; updates are never downloaded or installed automatically.
Polish dictated text with editable Chinese and English prompts before inserting it.
ExperimentalSince v1.3.1
Good to know
Experimental; requires an OpenAI-compatible endpoint, model and API key. Transcripts, including provisional text, leave the device. The five-second budget or endpoint errors fall back to original text.
v1.3.1 First release with v1.0.0, v1.1.0, v1.2.0, v1.3.0
Dictate into Linux applications using configurable tap or hold shortcuts.
Details
The v1.0.0–v1.3.1 launch sequence adds native X11 paste, Deepgram US English recognition and configurable Vibekey controls. Later releases refresh microphone devices, finish Fn dictation with another key and restore plain text ahead of HTML.
Also: Native X11 paste, optional Volcengine and Deepgram recognition, experimental polishing, Vibekey controls, device recovery and clipboard restoration.
v1.3.1 app or Omarchy-plugin archives support Python 3.11–3.14; choose one route. GTK, WebKit and PipeWire are external dependencies. Git-installed plugins must be disabled before updating, checked, then re-enabled.
Added: Dictate text using a Doubao web account and see rec…Added: Dictate with the official Volcengine Seed ASR 2.0 s…Added: Dictate US English with Deepgram Nova-3 and view in…Added: Paste completed dictation into the focused applicat…Added: Paste completed dictation into applications and ter…Added: Type dictated text into a Hyprland application with…Added: Start and finish dictation with tap or hold gesture…Added: Polish dictated text with editable Chinese and Engl…Added: Select a PipeWire microphone and test dictation wit…Added: Copy the most recent transcript after recognition o…Added: Recover the previous clipboard payload after a dict…Added: Control recording and send custom shortcuts with Vi…Added: Choose a listening waveform and reduce its animatio…Added: Use the application interface in English or Simplif…Added: See when a newer stable release is available and op…Behavior changed: Start and finish dictation with tap or hold gesture…Now works: Select a PipeWire microphone and test dictation wit…Behavior changed: Polish dictated text with editable Chinese and Engl…