Skip to content
Media 1 for listing Listen.Think.Speak - Local Speech to Speech AI

Description

Listen.Think.Speak is a fully offline voice chat asset for Windows:

๐ŸŽ™๏ธ speech-to-text โ†’ ๐Ÿค– LLM โ†’ ๐Ÿ—ฃ๏ธ text-to-speech


Why itโ€™s great

  • โšก Super fast: typical end-to-end replies in 1,500โ€“4,000 ms (STT โ†’ LLM โ†’ TTS).

  • ๐Ÿ”’ 100% local: zero network calls; ideal for offline games and strict privacy.

  • ๐Ÿง  Context aware: built-in conversation history for multi-turn dialog.

  • ๐Ÿ—ฃ๏ธ Voices: 28 plug-and-play voice models. [English only for now.]

  • ๐Ÿงฉ Drop-in demo: press Load LLM โ†’ Record โ†’ Speak to test in minutes.


Technical details

  • ๐Ÿ–ฅ๏ธ Platform: Windows-only (x64) for now.

  • ๐Ÿงฎ Inference: CPU-optimized pipeline for TTS and STT. GPU recommended for LLM.

  • ๐Ÿง  Core model: Llama 2 7B (quantized); conversation memory retained per session.

  • ๐Ÿ”Š TTS: ONNX voice models with default 22,050 Hz sample rate.

  • ๐Ÿ—‚๏ธ Model storage: single .bin LLM model file in StreamingAssets; ONNX voices from 60mb.

  • ๐ŸŽš๏ธ Latency breakdown (typical): STT ~300โ€“800 ms โ†’ LLM ~700โ€“2,200 ms โ†’ TTS ~300โ€“1,200 ms โ†’ Total 1.5โ€“4.0 s. (Tested on intel i7, i9 | Nvidia 2080 Super, 3070-Ti Laptop, 4070-Ti)

Setup in 3 steps

  1. Tools โ†’ Listen.Think.Speak โ†’ Prerequisites โ†’ install components + download the Llama 7B core.

  2. Download one or more ONNX voices, then add them in Speak Manager โ†’ Available Models.

  3. Open the demo, click Load LLM, then Record โ†’ Speak. Done!

Perfect for

  • ๐Ÿง‘โ€๐Ÿš€ Immersive NPC conversations

  • ๐Ÿ•ต๏ธ Voice-controlled puzzle/stealth mechanics

  • ๐Ÿฐ Narrative games with dynamic banter

  • ๐Ÿ›ก๏ธ Offline environments

Current support: ๐ŸชŸ Windows only (macOS/Linux planned).



If you want fast, private, and professional in-game voice interactions without touching the cloud, Listen.Think.Speak is your plug-and-play solution.

Included formats