Balanced · Free

ZYR Voice

Natural TTS · powered by Z.ai

Parameters
GLM 5 + Z.ai TTS (tongtong voice)
Speed
Balanced
Context
32k tokens
Pricing
Free

Powered By

Z.ai GLM 5 + Z.ai TTS (24 kHz WAV)

How ZYR Voice Works

ZYR Voice uses ZYR 2's balanced pipeline to generate the text response, then passes it through the Z.ai TTS API (voice: tongtong, format: WAV, 24 kHz). Markdown formatting is stripped before synthesis (code fences become '(code block)', bold/italic markers removed, links reduced to text). Long text is split into 1024-char chunks and synthesized sequentially with exponential backoff on rate limits. WAV chunks are concatenated with patched headers into a single audio file. The voice button appears on every AI message across all ZYR models, not just ZYR Voice.

Pipeline

  1. 1ZYR 2's 3-stage pipeline (Plan → Execute → Evaluate)
  2. 2Markdown stripped for natural speech
  3. 3Z.ai TTS synthesis (tongtong voice, 24 kHz WAV)

Best For

  • Audio playback of responses
  • Accessibility (visually impaired users)
  • Hands-free consumption
  • Podcast-style content
  • Voice-first interfaces

Capabilities

  • Natural-sounding Z.ai TTS voice (tongtong)
  • 24 kHz WAV output at 16-bit mono PCM
  • Markdown stripping (no 'backtick backtick python' in audio)
  • Stop button during playback
  • Voice button on every AI message
  • Long-text chunking with rate-limit retry

Limitations

  • TTS adds 1-3s latency on first play
  • Voice quality limited to Z.ai's available voices
  • Code blocks read as '(code block)' instead of full content
  • Rate limits on Z.ai TTS API (sequential chunking with backoff)

ZYR Voice — FAQ

What is ZYR Voice?

ZYR Voice uses ZYR 2's balanced pipeline to generate the text response, then passes it through the Z.ai TTS API (voice: tongtong, format: WAV, 24 kHz). Markdown formatting is stripped before synthesis (code fences become '(code block)', bold/italic markers removed, links reduced to text). Long text is split into 1024-char chunks and synthesized sequentially with exponential backoff on rate limits. WAV chunks are concatenated with patched headers into a single audio file. The voice button appears on every AI message across all ZYR models, not just ZYR Voice.

How many parameters does ZYR Voice have?

ZYR Voice has GLM 5 + Z.ai TTS (tongtong voice). It is powered by Z.ai GLM 5 + Z.ai TTS (24 kHz WAV).

Is ZYR Voice free?

Yes. ZYR Voice is free. There are no per-token costs, no subscription, and no usage limits beyond the underlying API rate limits.

What is ZYR Voice best for?

ZYR Voice is best for: Audio playback of responses, Accessibility (visually impaired users), Hands-free consumption, Podcast-style content, Voice-first interfaces.

What is the context window of ZYR Voice?

ZYR Voice has a context window of 32k tokens.

Other ZYR Models

Try ZYR Voice

Free. No signup required.

Open ZYR Chat →