Speech AI - Pronunciation, STT & TTS

Pronunciation scoring, speech-to-text, and text-to-speech for language learning

¿Debería usar esto?

Calidad y seguridad

A
Calidad de la descripción
100%
Integridad del esquema
56%
Calidad de los nombres
94%
Riesgo de envenenamiento
100%
Coincidencia de permisos
100%
Cumplimiento del protocolo
100%

Basado en el análisis automatizado de las definiciones de herramientas y el cumplimiento del protocolo.

Costo de contexto

~2,594Tokens (definiciones de herramientas)
~623 BTamaño de respuesta típico
Impacto significativo en la atención (2.03% del contexto de 128k)

Este es el número aproximado de tokens que se consumen cada vez que las herramientas del servidor se cargan en el contexto de un modelo. Los recuentos más altos reducen la atención disponible para otras tareas.

Instalar

Instalación con un clic

Agrega esto a tu archivo `claude_desktop_config.json`:

{
  "mcpServers": {
    "speech-ai": {
      "url": "https://pronunciation-mcp.thankfulfield-a7857897.eastus.azurecontainerapps.io/mcp"
    }
  }
}

Puntos de conexión remotos

https://pronunciation-mcp.thankfulfield-a7857897.eastus.azurecontainerapps.io/mcpstreamable-http
https://apim-ai-apis.azure-api.net/mcp/pronunciation/mcpstreamable-http

Qué puede hacer

Inventario de herramientas

Herramientas (10)

🟢 Solo lectura🟡 Escritura🔴 Eliminación⚪ Desconocido
🟢assess_pronunciation(audio_base64, text, audio_format)

Assess English pronunciation quality from audio. Scores pronunciation at four levels: overall, sentence, word, and phoneme. Each score is 0-100. Phonemes are returned in both IPA and ARPAbet notation. Sub-300ms inference latency. Args: audio_base64: Base64-encoded audio data. Supports WAV, MP3, OGG, and WebM formats. text: The reference English text that the speaker was expected to read aloud. audio_format: Audio format hint — one of 'wav', 'mp3', 'ogg', 'webm'. Defaults to 'wav'. Returns: dict with keys: - overallScore (int 0-100): Overall pronunciation quality - sentenceScore (int 0-100): Sentence-level fluency and accuracy - words (list): Per-word scores, each containing: - word (str): The word - score (int 0-100): Word pronunciation score - phonemes (list): Per-phoneme scores with IPA/ARPAbet notation - decodedTranscript (str): What the model heard (ASR transcript) - transcript (str): Reference text - confidence (float 0-1): Scoring confidence - warnings (list[str]): Quality warnings if any - audioQuality (dict): Audio metrics (SNR, peak/RMS dB, etc.)

Esquema de entrada

{
  "type": "object",
  "properties": {
    "audio_base64": {
      "description": "Base64-encoded audio data. Supports WAV, MP3, OGG, and WebM formats.",
      "maxLength": 20000000,
      "type": "string"
    },
    "text": {
      "description": "The reference English text that the speaker was expected to read aloud.",
      "maxLength": 10000,
      "type": "string"
    },
    "audio_format": {
      "default": "wav",
      "description": "Audio format hint — one of 'wav', 'mp3', 'ogg', 'webm'.",
      "type": "string"
    }
  },
  "required": [
    "audio_base64",
    "text"
  ]
}
🟢check_pronunciation_service

Check if the pronunciation assessment service is healthy and ready. Returns: dict with keys: - status (str): 'healthy' or error state - modelLoaded (bool): Whether the scoring model is loaded - version (str): API version

Esquema de entrada

{
  "type": "object",
  "properties": {}
}
🟢get_phoneme_inventory

Get the full phoneme inventory supported by the pronunciation scorer. Returns a list of all English phonemes the engine can assess, including ARPAbet symbol, IPA equivalent, example word, and phoneme category (vowel, consonant, diphthong). Returns: list of dicts, each with keys: - arpabet (str): ARPAbet symbol (e.g. 'AA', 'TH') - ipa (str): IPA notation - example (str): Example word containing the phoneme - category (str): vowel, consonant, or diphthong

Esquema de entrada

{
  "type": "object",
  "properties": {}
}
🟢transcribe_audio(audio_base64, audio_format, include_timestamps)

Transcribe audio to text with word-level timestamps. Converts spoken English audio into text with optional word-level timestamps and per-word confidence scores. Args: audio_base64: Base64-encoded audio data (WAV, MP3, OGG, FLAC, WebM). audio_format: Audio format hint. Auto-detected from magic bytes if omitted. include_timestamps: Whether to include word-level timing (default: true). Returns: dict with keys: - text (str): Full decoded transcript - words (list): Per-word results with timestamps, each containing: - word (str): The transcribed word - start (float): Start time in seconds - end (float): End time in seconds - confidence (float 0-1): Word-level confidence - audioDurationMs (int): Audio duration in milliseconds - metadata (dict): Processing time, audio length, model version - audioQuality (dict): Audio metrics (SNR, peak/RMS dB, etc.)

Esquema de entrada

{
  "type": "object",
  "properties": {
    "audio_base64": {
      "description": "Base64-encoded audio data. Supports WAV, MP3, OGG, FLAC, and WebM formats.",
      "maxLength": 20000000,
      "type": "string"
    },
    "audio_format": {
      "anyOf": [
        {
          "type": "string"
        },
        {
          "type": "null"
        }
      ],
      "default": null,
      "description": "Audio format hint — 'wav', 'mp3', 'ogg', 'flac', 'webm'. Auto-detected if omitted."
    },
    "include_timestamps": {
      "default": true,
      "description": "If true, include word-level start/end times and confidence.",
      "type": "boolean"
    }
  },
  "required": [
    "audio_base64"
  ]
}
🟢check_stt_service

Check if the speech-to-text service is healthy and ready. Returns: dict with keys: - status (str): 'healthy' or error state - modelLoaded (bool): Whether the STT model is loaded - version (str): API version

Esquema de entrada

{
  "type": "object",
  "properties": {}
}
🟢synthesize_speech(text, voice, speed)

Generate natural speech audio from English text. Produces high-quality speech with 12 English voices. Returns base64-encoded WAV audio (16-bit PCM, 24kHz mono) along with metadata. Available voices: - af_heart (default), af_bella, af_nicole, af_sarah, af_sky (American female) - am_adam, am_michael (American male) - bf_emma, bf_isabella (British female) - bm_george, bm_lewis, bm_daniel (British male) Args: text: English text to synthesize (1-5000 characters). voice: Voice ID. See list above. Defaults to 'af_heart'. speed: Speed multiplier from 0.5 to 2.0 (default: 1.0). Returns: dict with keys: - audio_base64 (str): Base64-encoded WAV audio (16-bit PCM, 24kHz) - duration_ms (str): Audio duration in milliseconds - voice (str): Voice ID used - text_length (str): Input text character count - processing_ms (str): Synthesis time in milliseconds

Esquema de entrada

{
  "type": "object",
  "properties": {
    "text": {
      "description": "English text to convert to speech. Max 5000 characters.",
      "maxLength": 5000,
      "type": "string"
    },
    "voice": {
      "anyOf": [
        {
          "type": "string"
        },
        {
          "type": "null"
        }
      ],
      "default": null,
      "description": "Voice ID (e.g. 'af_heart', 'am_adam'). Uses default if omitted."
    },
    "speed": {
      "default": 1,
      "description": "Speech speed multiplier (0.5 = half speed, 2.0 = double).",
      "type": "number"
    }
  },
  "required": [
    "text"
  ]
}
🟢list_tts_voices

List all available text-to-speech voices with metadata. Returns: dict with keys: - voices (list): Available voices, each with id, name, gender, accent, grade - defaultVoice (str): Default voice ID

Esquema de entrada

{
  "type": "object",
  "properties": {}
}
🟢check_tts_service

Check if the text-to-speech service is healthy and ready. Returns: dict with keys: - status (str): 'healthy' or error state - modelLoaded (bool): Whether the TTS model is loaded - version (str): API version

Esquema de entrada

{
  "type": "object",
  "properties": {}
}
🟢transcribe_audio_pro(audio_base64, language, diarize)

Transcribe audio with Whisper Large V3 Turbo — multilingual STT. Supports 99 languages with automatic language detection, word-level timestamps, per-word confidence scores, and optional speaker diarization (identifies who spoke each word). Best-in-class WER (~2%). Args: audio_base64: Base64-encoded audio (WAV, MP3, OGG, FLAC, WebM). language: Language code. Auto-detected if omitted. Supports 99 languages. diarize: Enable speaker diarization (default: false). When true, each word includes a speaker label (e.g. SPEAKER_00, SPEAKER_01). Returns: dict with keys: - text (str): Full decoded transcript - words (list): Per-word results with timestamps, each containing: - word (str), start (float), end (float), confidence (float 0-1) - speaker (str|null): Speaker label when diarize=true - speakers (dict|null): Speaker info with count and labels - audioDurationMs (int): Audio duration in milliseconds - metadata (dict): Processing time, language, languageProbability - audioQuality (dict): Audio metrics (SNR, peak/RMS dB, etc.)

Esquema de entrada

{
  "type": "object",
  "properties": {
    "audio_base64": {
      "description": "Base64-encoded audio data. Supports WAV, MP3, OGG, FLAC, and WebM formats.",
      "maxLength": 20000000,
      "type": "string"
    },
    "language": {
      "anyOf": [
        {
          "type": "string"
        },
        {
          "type": "null"
        }
      ],
      "default": null,
      "description": "Language code (e.g. 'en', 'es', 'zh'). Auto-detected when omitted."
    },
    "diarize": {
      "default": false,
      "description": "Enable speaker diarization to identify who spoke each word.",
      "type": "boolean"
    }
  },
  "required": [
    "audio_base64"
  ]
}
🟢check_whisper_service

Check if the Whisper STT Pro service is healthy and ready. Returns: dict with keys: - status (str): 'healthy' or error state - modelLoaded (bool): Whether the Whisper model is loaded - diarizeLoaded (bool): Whether the diarization pipeline is loaded - version (str): API version - modelName (str): Whisper model name (e.g. 'large-v3-turbo')

Esquema de entrada

{
  "type": "object",
  "properties": {}
}

Comunidad

Califica este servidor

Evidencia

Observaciones recientes

verificadoversión no registrada10 herramientas
verificadoversión no registrada10 herramientas
verificadoversión no registrada10 herramientas
verificadoversión no registrada10 herramientas
verificadoversión no registrada10 herramientas
verificadoversión no registrada10 herramientas