Local-first, hands-free voice assistant for Ollama. Talk to a local LLM and get a spoken answer — wake word → speech-to-text → local reasoning → text-to-speech, running entirely on your machine. No cloud, ever.
wake word → capture (Silero VAD end-of-turn) → faster-whisper (STT)
→ Ollama (reasoning) → Piper (TTS) → speakers ↺
The wake word (openWakeWord) and VAD (Silero) both run locally on onnxruntime.
- Python 3.10+
- Ollama running locally with a model pulled
(e.g.
ollama pull llama3.2) - A Piper voice
(
.onnx+.onnx.json) — drop it in./voices/ - A working microphone and speakers
pip install chuchoteThen grab a Piper voice
(.onnx + .onnx.json) into a voices/ folder next to where you run it (or
set piper_voice in your config), and check everything's wired up:
chuchote doctorgit clone https://github.com/Cjayy77/chuchote.git
cd chuchote
python -m venv .venv
# Windows: .venv\Scripts\activate
# macOS/Linux: source .venv/bin/activate
pip install -e .Or use the install script, which also downloads a default Piper voice and writes a config file:
./scripts/install.ps1 # Windows (PowerShell)
./scripts/install.sh # macOS/LinuxMake sure Ollama is running (ollama serve) and a voice model is in ./voices/:
chuchote start
# or override defaults:
chuchote start --model llama3.2 --wake-word hey_jarvis --voice voices/en_US-lessac-medium.onnxBy default Chuchote is always listening for the wake word (hey_jarvis).
Say it, then speak your request — Silero VAD detects when you've stopped and
the reply is transcribed, thought out, and spoken back (starting as soon as the
first sentence is ready). Press Ctrl+C to quit.
The wake-word models download automatically on first run. Other built-in words
include alexa, hey_mycroft, and hey_rhasspy (--wake-word); raise
--wake-threshold if you get false triggers. A short tone confirms the wake
word — silence it with --no-chime.
Barge-in. While Chuchote is speaking, say the wake word again to cut it off
and start a new turn (--barge-in wake, the default — echo-robust, works with
open speakers). With headphones you can use --barge-in vad so any speech
interrupts; --barge-in off lets every reply finish. In push-to-talk mode,
pressing the PTT key cuts off the reply (then hold it to speak again).
--wake-word accepts a file path as well as a built-in name, so you can
use any openWakeWord-compatible model:
chuchote start --wake-word path/to/hey_chuchote.onnxTo create one for your own phrase, use openWakeWord's automatic training
notebook — see Training New Models
in their README. It generates synthetic speech for your phrase and trains a
model in roughly an hour, no voice recordings needed; download the resulting
.onnx and point --wake-word (or wake_model in your config) at it.
Note: the training notebook runs on Google Colab (a free cloud notebook) — that's one-time, dev-machine tooling, like downloading a Piper voice. The resulting model runs fully locally; nothing about the assistant touches the network.
Prefer holding a key? Skip the wake word entirely:
chuchote start --ptt # hold space to talk, release to send
chuchote start --ptt --ptt-key ctrlNot sure everything's wired up? Run the preflight check:
chuchote doctorIt verifies Ollama is reachable and the model is pulled, a Piper voice is
present, your mic and speakers are detected, and (in wake mode) the wake-word
deps are installed — reporting each as [ ok ] / [fail] and exiting non-zero
if anything's wrong.
Chuchote isn't English-only — whisper understands ~99 languages and Piper has voices for dozens. To run it in another language, three things need to line up:
- Recognition — set
languageand use a multilingual whisper model (the plain names, not the.enones —baseon low-RAM machines,smallfor better accuracy):language = "fr" whisper_model = "base"
- Speech — download a Piper voice for that language from
VOICES.md into your
voices dir (or point
piper_voiceat it). Piper voices are one language each. - Reasoning — pick an Ollama model that's good in your language (most
modern ones are multilingual; e.g.
qwen2.5is strong for Chinese). Chuchote already asks the model to reply in whatever language you speak.
Then chuchote doctor will confirm the pieces match. Common starting points:
| Language | language |
Example Piper voice |
|---|---|---|
| French | fr |
fr_FR-siwis-medium |
| German | de |
de_DE-thorsten-medium |
| Spanish | es |
es_ES-davefx-medium |
| Italian | it |
it_IT-paola-medium |
| Portuguese (BR) | pt |
pt_BR-faber-medium |
| Dutch | nl |
nl_NL-mls-medium |
| Chinese | zh |
zh_CN-huayan-medium |
| Russian | ru |
ru_RU-dmitri-medium |
Adding any other language: find its voice on
VOICES.md, set
language to whisper's code
for it, keep a multilingual whisper_model, and you're set. The wake word stays
an English phrase (openWakeWord's built-ins are English) — or use --ptt.
Chuchote remembers the conversation. Each exchange is stored in a SQLite
database (memory.db in your per-user data dir) and the most recent turns are
fed back into the model's context every turn, so it stays coherent across turns
and restarts.
chuchote start --forget # start a session with a clean slate
chuchote forget # erase all saved memoryTune how much history is injected via history_messages in
chuchote/config.py.
Persist your settings in a config file instead of passing flags every time:
chuchote init # writes a commented config.toml to your config dirEdit the file (its path is printed by init; typically %APPDATA%\chuchote\
on Windows or ~/.config/chuchote/ elsewhere), uncomment what you want to
change, and it's picked up on the next chuchote start. Point at a specific
file with --config PATH. Precedence is defaults < config file < flags.
Flags (over sensible defaults; see chuchote/config.py for the full list and VAD tunables):
| Flag | Default | Meaning |
|---|---|---|
--model |
llama3.2 |
Ollama model |
--whisper-model |
base.en |
faster-whisper model (small.en = more accurate, needs ~1 GB free RAM) |
--language |
auto |
recognition language (en, fr, de, zh, …); needs a multilingual model |
--voice |
first .onnx in ./voices |
Piper voice model |
--wake-word |
hey_jarvis |
wake word model (alexa, hey_mycroft, hey_rhasspy) |
--wake-threshold |
0.5 |
wake sensitivity 0..1 (higher = fewer false triggers) |
--no-chime |
off | disable the wake-word acknowledgement tone |
--barge-in |
wake |
interrupt a reply: wake / vad (headphones) / off |
--ptt |
off | use push-to-talk instead of the wake word |
--ptt-key |
space |
push-to-talk key to hold (with --ptt) |
--forget |
off | clear conversation memory before starting |
--no-banner |
off | don't print the startup banner |
--config |
per-user config dir | path to a config file |
Run the test suite (covers the pure logic — sentence chunking, memory, config precedence, banner styling — with no audio/model deps needed):
pip install -e ".[dev]"
pytestTests also run in CI (GitHub Actions) on Linux and Windows, Python 3.10–3.13.
Publishing a GitHub release triggers the publish.yml workflow, which builds
and uploads to PyPI via trusted publishing
— no API token. One-time setup on pypi.org: add a trusted publisher for this
repo (owner Cjayy77, repo chuchote, workflow publish.yml, environment
pypi), and create a matching pypi environment in the repo's GitHub settings.
MIT.
