speech-build

Original：🇺🇸 English

Translated

Generate and transcribe speech using Google's Gemini-TTS and Chirp 3 models. Supports Text-to-Speech (Single/Multi-speaker), Instant Custom Voice, and Speech-to-Text (Transcription/Diarization).

7installs

Sourcecnemri/google-genai-skills

Added on2026-02-07

NPX Install

npx skill4agent add cnemri/google-genai-skills speech-build

SKILL.md Content

View Translation Comparison →

Speech Skill (TTS & STT)

Use this skill to implement audio generation and transcription workflows using the

google-genai

and

google-cloud-speech

SDKs.

Quick Start Setup

python

from google import genai
from google.genai import types
# For STT: from google.cloud import speech_v2

client = genai.Client()

Reference Materials

Text-to-Speech (TTS): Gemini-TTS, Chirp 3 HD, Instant Custom Voice.
Speech-to-Text (STT): Chirp 3 Transcription, Diarization, Streaming.
Voices & Locales: Available voices (
```
Aoede
```
,
```
Puck
```
...) and languages.
Prompting Guide: How to control style, accent, and pacing in Gemini-TTS.
Source Code: Deep inspection of SDK internals.

Common Workflows

1. Generate Speech (Gemini-TTS)

python

response = client.models.generate_content(
    model="gemini-2.5-flash-preview-tts",
    contents="Hello, world!",
    config=types.GenerateContentConfig(
        response_modalities=["AUDIO"],
        speech_config=types.SpeechConfig(
            voice_config=types.VoiceConfig(
                prebuilt_voice_config=types.PrebuiltVoiceConfig(voice_name='Kore')
            )
        )
    )
)

2. Transcribe Audio (Chirp 3)

python

# Requires google-cloud-speech
from google.cloud import speech_v2
# ... (See stt.md for full setup)
response = speech_client.recognize(...)

speech-build

NPX Install

Tags

SKILL.md Content

Speech Skill (TTS & STT)

Quick Start Setup

Reference Materials

Common Workflows

1. Generate Speech (Gemini-TTS)

2. Transcribe Audio (Chirp 3)