Speech AI - Pronunciation, STT & TTS

Pronunciation scoring, speech-to-text, and text-to-speech for language learning

使うべきか

品質と安全性

A
説明の品質
100%
スキーマの完全性
56%
命名の品質
94%
ポイズニングのリスク
100%
権限の一致
100%
プロトコルへの準拠
100%

ツール定義とプロトコルへの準拠に関する自動分析に基づいています。

コンテキストコスト

~2,594トークン数(ツール定義)
~623 B一般的なレスポンスサイズ
注意への影響は大きい(128k コンテキストの 2.03%)

これは、サーバーのツールがモデルのコンテキストに読み込まれるたびに消費されるおおよそのトークン数です。数が多いほど、ほかのタスクに使える注意が減ります。

インストール

ワンクリックインストール

これを `claude_desktop_config.json` ファイルに追加してください:

{
  "mcpServers": {
    "speech-ai": {
      "url": "https://pronunciation-mcp.thankfulfield-a7857897.eastus.azurecontainerapps.io/mcp"
    }
  }
}

リモートエンドポイント

https://pronunciation-mcp.thankfulfield-a7857897.eastus.azurecontainerapps.io/mcpstreamable-http
https://apim-ai-apis.azure-api.net/mcp/pronunciation/mcpstreamable-http

できること

ツール一覧

ツール(10)

🟢 読み取り専用🟡 書き込み🔴 削除⚪ 不明
🟢assess_pronunciation(audio_base64, text, audio_format)

Assess English pronunciation quality from audio. Scores pronunciation at four levels: overall, sentence, word, and phoneme. Each score is 0-100. Phonemes are returned in both IPA and ARPAbet notation. Sub-300ms inference latency. Args: audio_base64: Base64-encoded audio data. Supports WAV, MP3, OGG, and WebM formats. text: The reference English text that the speaker was expected to read aloud. audio_format: Audio format hint — one of 'wav', 'mp3', 'ogg', 'webm'. Defaults to 'wav'. Returns: dict with keys: - overallScore (int 0-100): Overall pronunciation quality - sentenceScore (int 0-100): Sentence-level fluency and accuracy - words (list): Per-word scores, each containing: - word (str): The word - score (int 0-100): Word pronunciation score - phonemes (list): Per-phoneme scores with IPA/ARPAbet notation - decodedTranscript (str): What the model heard (ASR transcript) - transcript (str): Reference text - confidence (float 0-1): Scoring confidence - warnings (list[str]): Quality warnings if any - audioQuality (dict): Audio metrics (SNR, peak/RMS dB, etc.)

入力スキーマ

{
  "type": "object",
  "properties": {
    "audio_base64": {
      "description": "Base64-encoded audio data. Supports WAV, MP3, OGG, and WebM formats.",
      "maxLength": 20000000,
      "type": "string"
    },
    "text": {
      "description": "The reference English text that the speaker was expected to read aloud.",
      "maxLength": 10000,
      "type": "string"
    },
    "audio_format": {
      "default": "wav",
      "description": "Audio format hint — one of 'wav', 'mp3', 'ogg', 'webm'.",
      "type": "string"
    }
  },
  "required": [
    "audio_base64",
    "text"
  ]
}
🟢check_pronunciation_service

Check if the pronunciation assessment service is healthy and ready. Returns: dict with keys: - status (str): 'healthy' or error state - modelLoaded (bool): Whether the scoring model is loaded - version (str): API version

入力スキーマ

{
  "type": "object",
  "properties": {}
}
🟢get_phoneme_inventory

Get the full phoneme inventory supported by the pronunciation scorer. Returns a list of all English phonemes the engine can assess, including ARPAbet symbol, IPA equivalent, example word, and phoneme category (vowel, consonant, diphthong). Returns: list of dicts, each with keys: - arpabet (str): ARPAbet symbol (e.g. 'AA', 'TH') - ipa (str): IPA notation - example (str): Example word containing the phoneme - category (str): vowel, consonant, or diphthong

入力スキーマ

{
  "type": "object",
  "properties": {}
}
🟢transcribe_audio(audio_base64, audio_format, include_timestamps)

Transcribe audio to text with word-level timestamps. Converts spoken English audio into text with optional word-level timestamps and per-word confidence scores. Args: audio_base64: Base64-encoded audio data (WAV, MP3, OGG, FLAC, WebM). audio_format: Audio format hint. Auto-detected from magic bytes if omitted. include_timestamps: Whether to include word-level timing (default: true). Returns: dict with keys: - text (str): Full decoded transcript - words (list): Per-word results with timestamps, each containing: - word (str): The transcribed word - start (float): Start time in seconds - end (float): End time in seconds - confidence (float 0-1): Word-level confidence - audioDurationMs (int): Audio duration in milliseconds - metadata (dict): Processing time, audio length, model version - audioQuality (dict): Audio metrics (SNR, peak/RMS dB, etc.)

入力スキーマ

{
  "type": "object",
  "properties": {
    "audio_base64": {
      "description": "Base64-encoded audio data. Supports WAV, MP3, OGG, FLAC, and WebM formats.",
      "maxLength": 20000000,
      "type": "string"
    },
    "audio_format": {
      "anyOf": [
        {
          "type": "string"
        },
        {
          "type": "null"
        }
      ],
      "default": null,
      "description": "Audio format hint — 'wav', 'mp3', 'ogg', 'flac', 'webm'. Auto-detected if omitted."
    },
    "include_timestamps": {
      "default": true,
      "description": "If true, include word-level start/end times and confidence.",
      "type": "boolean"
    }
  },
  "required": [
    "audio_base64"
  ]
}
🟢check_stt_service

Check if the speech-to-text service is healthy and ready. Returns: dict with keys: - status (str): 'healthy' or error state - modelLoaded (bool): Whether the STT model is loaded - version (str): API version

入力スキーマ

{
  "type": "object",
  "properties": {}
}
🟢synthesize_speech(text, voice, speed)

Generate natural speech audio from English text. Produces high-quality speech with 12 English voices. Returns base64-encoded WAV audio (16-bit PCM, 24kHz mono) along with metadata. Available voices: - af_heart (default), af_bella, af_nicole, af_sarah, af_sky (American female) - am_adam, am_michael (American male) - bf_emma, bf_isabella (British female) - bm_george, bm_lewis, bm_daniel (British male) Args: text: English text to synthesize (1-5000 characters). voice: Voice ID. See list above. Defaults to 'af_heart'. speed: Speed multiplier from 0.5 to 2.0 (default: 1.0). Returns: dict with keys: - audio_base64 (str): Base64-encoded WAV audio (16-bit PCM, 24kHz) - duration_ms (str): Audio duration in milliseconds - voice (str): Voice ID used - text_length (str): Input text character count - processing_ms (str): Synthesis time in milliseconds

入力スキーマ

{
  "type": "object",
  "properties": {
    "text": {
      "description": "English text to convert to speech. Max 5000 characters.",
      "maxLength": 5000,
      "type": "string"
    },
    "voice": {
      "anyOf": [
        {
          "type": "string"
        },
        {
          "type": "null"
        }
      ],
      "default": null,
      "description": "Voice ID (e.g. 'af_heart', 'am_adam'). Uses default if omitted."
    },
    "speed": {
      "default": 1,
      "description": "Speech speed multiplier (0.5 = half speed, 2.0 = double).",
      "type": "number"
    }
  },
  "required": [
    "text"
  ]
}
🟢list_tts_voices

List all available text-to-speech voices with metadata. Returns: dict with keys: - voices (list): Available voices, each with id, name, gender, accent, grade - defaultVoice (str): Default voice ID

入力スキーマ

{
  "type": "object",
  "properties": {}
}
🟢check_tts_service

Check if the text-to-speech service is healthy and ready. Returns: dict with keys: - status (str): 'healthy' or error state - modelLoaded (bool): Whether the TTS model is loaded - version (str): API version

入力スキーマ

{
  "type": "object",
  "properties": {}
}
🟢transcribe_audio_pro(audio_base64, language, diarize)

Transcribe audio with Whisper Large V3 Turbo — multilingual STT. Supports 99 languages with automatic language detection, word-level timestamps, per-word confidence scores, and optional speaker diarization (identifies who spoke each word). Best-in-class WER (~2%). Args: audio_base64: Base64-encoded audio (WAV, MP3, OGG, FLAC, WebM). language: Language code. Auto-detected if omitted. Supports 99 languages. diarize: Enable speaker diarization (default: false). When true, each word includes a speaker label (e.g. SPEAKER_00, SPEAKER_01). Returns: dict with keys: - text (str): Full decoded transcript - words (list): Per-word results with timestamps, each containing: - word (str), start (float), end (float), confidence (float 0-1) - speaker (str|null): Speaker label when diarize=true - speakers (dict|null): Speaker info with count and labels - audioDurationMs (int): Audio duration in milliseconds - metadata (dict): Processing time, language, languageProbability - audioQuality (dict): Audio metrics (SNR, peak/RMS dB, etc.)

入力スキーマ

{
  "type": "object",
  "properties": {
    "audio_base64": {
      "description": "Base64-encoded audio data. Supports WAV, MP3, OGG, FLAC, and WebM formats.",
      "maxLength": 20000000,
      "type": "string"
    },
    "language": {
      "anyOf": [
        {
          "type": "string"
        },
        {
          "type": "null"
        }
      ],
      "default": null,
      "description": "Language code (e.g. 'en', 'es', 'zh'). Auto-detected when omitted."
    },
    "diarize": {
      "default": false,
      "description": "Enable speaker diarization to identify who spoke each word.",
      "type": "boolean"
    }
  },
  "required": [
    "audio_base64"
  ]
}
🟢check_whisper_service

Check if the Whisper STT Pro service is healthy and ready. Returns: dict with keys: - status (str): 'healthy' or error state - modelLoaded (bool): Whether the Whisper model is loaded - diarizeLoaded (bool): Whether the diarization pipeline is loaded - version (str): API version - modelName (str): Whisper model name (e.g. 'large-v3-turbo')

入力スキーマ

{
  "type": "object",
  "properties": {}
}

コミュニティ

このサーバーを評価する

エビデンス

最近の観測

検証済みバージョンは記録されていませんツール 10 件
検証済みバージョンは記録されていませんツール 10 件
検証済みバージョンは記録されていませんツール 10 件
検証済みバージョンは記録されていませんツール 10 件