oruk Speech

Hosted speech-to-text + speech emotion/tone analysis for agents. No install; trial keys built in.

我该使用它吗

质量与安全性

B
描述质量
100%
模式完整度
64%
命名质量
80%
投毒风险
100%
权限匹配度
100%
协议合规性
100%

基于对工具定义和协议合规性的自动分析。

上下文开销

~2,311token 数(工具定义)
~2.4 KB典型响应大小
对注意力有中等影响(占 128k 上下文窗口的 1.81%)

这是每次将服务器的工具加载到模型上下文窗口时所消耗的大致 token 数。数值越高,可用于其他任务的注意力就越少。

安装

一键安装

将以下内容添加到你的 `claude_desktop_config.json` 文件中:

{
  "mcpServers": {
    "speech": {
      "url": "https://oruk.ai/mcp"
    }
  }
}

远程端点

https://oruk.ai/mcpstreamable-http

它能做什么

工具清单

工具(7)

🟢 只读🟡 写入🔴 删除⚪ 未知
🟢oruk_analyze_speech(audio_url, audio_base64, filename, model, detail, ...)

Transcribe English audio AND score how it was said in one call: transcript, tagged transcript, selected scores from 15 emotion and 16 speaking-style labels, and time-local segments. Use this when the user cares about both the words and the delivery — meetings, support calls, interviews, voice notes. Accepts wav/flac/mp3/m4a/ogg/webm. Up to 30 MB via audio_url or 8 MiB decoded via audio_base64; up to 60 minutes of English speech. Returns compact summaries by default. For words only use oruk_transcribe_audio; for tone only use oruk_analyze_tone.

输入模式

{
  "type": "object",
  "properties": {
    "audio_url": {
      "description": "Publicly fetchable audio file URL (wav, flac, mp3, m4a, ogg, webm; up to 30 MB / 60 minutes of English speech).",
      "type": "string",
      "maxLength": 2000,
      "format": "uri"
    },
    "audio_base64": {
      "description": "Base64-encoded audio bytes for local files (up to 8 MiB decoded). Prefer audio_url for anything larger.",
      "type": "string",
      "maxLength": 11500000
    },
    "filename": {
      "description": "Original filename including extension (e.g. call.wav). Helps decoding when audio_base64 is used.",
      "type": "string",
      "maxLength": 160
    },
    "model": {
      "description": "oruk-resonance (full local pipeline: transcription, emotion, style, affect, analysis) or oruk-fourier (parallel transcript and native 15-label emotion, with the shared 16-label style model). Resonance is the default.",
      "type": "string",
      "enum": [
        "oruk-resonance",
        "oruk-fourier"
      ]
    },
    "detail": {
      "description": "compact (default) returns top label scores and condensed segments; full preserves all returned labels, segments, and word-level timings, subject to response-size limits. It does not expose unreturned label scores.",
      "type": "string",
      "enum": [
        "compact",
        "full"
      ]
    },
    "diarize": {
      "description": "Label speakers (oruk-resonance only; the model is switched to oruk-resonance automatically). Speaker diarization locates turns, then Resonance scores each turn with its own text, emotions, and styles. Use for calls, meetings, and interviews. Included in subscription plan minutes. Processing details: https://oruk.ai/security#processing.",
      "type": "boolean"
    },
    "api_key": {
      "description": "Only for temporary keys from oruk_create_trial_key. Permanent keys belong in your MCP client config as an \"Authorization: Bearer <key>\" header, never in tool arguments.",
      "type": "string",
      "maxLength": 200
    }
  },
  "additionalProperties": false
}
🟢oruk_transcribe_audio(audio_url, audio_base64, filename, model, detail, ...)

Transcribe prerecorded English audio to text with time-ordered segments and word timings. Use this when only the words matter. Accepts wav/flac/mp3/m4a/ogg/webm. Up to 30 MB via audio_url or 8 MiB decoded via audio_base64; up to 60 minutes of English speech. Does not score emotion or tone — use oruk_analyze_speech for transcript + tone together, or oruk_analyze_tone for tone alone.

输入模式

{
  "type": "object",
  "properties": {
    "audio_url": {
      "description": "Publicly fetchable audio file URL (wav, flac, mp3, m4a, ogg, webm; up to 30 MB / 60 minutes of English speech).",
      "type": "string",
      "maxLength": 2000,
      "format": "uri"
    },
    "audio_base64": {
      "description": "Base64-encoded audio bytes for local files (up to 8 MiB decoded). Prefer audio_url for anything larger.",
      "type": "string",
      "maxLength": 11500000
    },
    "filename": {
      "description": "Original filename including extension (e.g. call.wav). Helps decoding when audio_base64 is used.",
      "type": "string",
      "maxLength": 160
    },
    "model": {
      "description": "oruk-resonance (full local pipeline: transcription, emotion, style, affect, analysis) or oruk-fourier (parallel transcript and native 15-label emotion, with the shared 16-label style model). Resonance is the default.",
      "type": "string",
      "enum": [
        "oruk-resonance",
        "oruk-fourier"
      ]
    },
    "detail": {
      "description": "compact (default) returns top label scores and condensed segments; full preserves all returned labels, segments, and word-level timings, subject to response-size limits. It does not expose unreturned label scores.",
      "type": "string",
      "enum": [
        "compact",
        "full"
      ]
    },
    "diarize": {
      "description": "Label speakers (oruk-resonance only; the model is switched to oruk-resonance automatically). Speaker diarization locates turns, then Resonance scores each turn with its own text, emotions, and styles. Use for calls, meetings, and interviews. Included in subscription plan minutes. Processing details: https://oruk.ai/security#processing.",
      "type": "boolean"
    },
    "api_key": {
      "description": "Only for temporary keys from oruk_create_trial_key. Permanent keys belong in your MCP client config as an \"Authorization: Bearer <key>\" header, never in tool arguments.",
      "type": "string",
      "maxLength": 200
    }
  },
  "additionalProperties": false
}
🟢oruk_analyze_tone(audio_url, audio_base64, filename, model, detail, ...)

Score how speech sounds without transcribing it: selected emotion (happy, frustrated, worried, …) and speaking-style (sarcastic, confident, hesitant, warm, …) scores per acoustic segment. Runs the Resonance encoder and affect head only — the transcription decoder is never invoked, so nothing is transcribed and it consumes the same subscription audio minutes as unified analysis. Use this when the user asks about mood, delivery, sentiment, sarcasm, or emotional dynamics in audio. Up to 30 MB via audio_url or 8 MiB decoded via audio_base64; up to 60 minutes of English speech. Labels use model-specific thresholds; the highest-scoring emotion is returned if none passes, and styles can be empty. Outputs describe delivery, not probabilities of inner state. Need the words too? Use oruk_analyze_speech.

输入模式

{
  "type": "object",
  "properties": {
    "audio_url": {
      "description": "Publicly fetchable audio file URL (wav, flac, mp3, m4a, ogg, webm; up to 30 MB / 60 minutes of English speech).",
      "type": "string",
      "maxLength": 2000,
      "format": "uri"
    },
    "audio_base64": {
      "description": "Base64-encoded audio bytes for local files (up to 8 MiB decoded). Prefer audio_url for anything larger.",
      "type": "string",
      "maxLength": 11500000
    },
    "filename": {
      "description": "Original filename including extension (e.g. call.wav). Helps decoding when audio_base64 is used.",
      "type": "string",
      "maxLength": 160
    },
    "model": {
      "description": "oruk-resonance (full local pipeline: transcription, emotion, style, affect, analysis) or oruk-fourier (parallel transcript and native 15-label emotion, with the shared 16-label style model). Resonance is the default.",
      "type": "string",
      "enum": [
        "oruk-resonance",
        "oruk-fourier"
      ]
    },
    "detail": {
      "description": "compact (default) returns top label scores and condensed segments; full preserves all returned labels, segments, and word-level timings, subject to response-size limits. It does not expose unreturned label scores.",
      "type": "string",
      "enum": [
        "compact",
        "full"
      ]
    },
    "diarize": {
      "description": "Label speakers (oruk-resonance only; the model is switched to oruk-resonance automatically). Speaker diarization locates turns, then Resonance scores each turn with its own text, emotions, and styles. Use for calls, meetings, and interviews. Included in subscription plan minutes. Processing details: https://oruk.ai/security#processing.",
      "type": "boolean"
    },
    "api_key": {
      "description": "Only for temporary keys from oruk_create_trial_key. Permanent keys belong in your MCP client config as an \"Authorization: Bearer <key>\" header, never in tool arguments.",
      "type": "string",
      "maxLength": 200
    }
  },
  "additionalProperties": false
}
🟢oruk_check_usage(api_key)

Verify that an Oruk API key works and report the subscription, remaining audio minutes, and recent API usage. Use this after setup or to diagnose access and usage limits. Requires the Authorization header from your MCP config or a temporary api_key.

输入模式

{
  "type": "object",
  "properties": {
    "api_key": {
      "description": "Only for temporary keys from oruk_create_trial_key. Permanent keys belong in the Authorization header of your MCP client config.",
      "type": "string",
      "maxLength": 200
    }
  },
  "additionalProperties": false
}
🟡oruk_create_trial_key

Mint a real, temporary oruk API key with no account required: 3 requests, expires in 30 minutes, spends from a capped shared budget. Use this when no Authorization header is configured and the user wants to try transcription or tone analysis right now. Pass the returned key as the api_key argument of the audio tools. Share the signup link with the user so they can keep using oruk afterwards (7-day free trial on self-serve plans).

输入模式

{
  "type": "object",
  "properties": {},
  "additionalProperties": false
}
🟢oruk_list_models

List oruk’s speech models with lifecycle, current subscription plans, and explicitly labeled legacy reference rates, the five API tasks, the 15 emotion and 16 speaking-style labels, and audio limits. No API key required. Use this to choose a model, estimate cost before analyzing long audio, or see which labels exist.

输入模式

{
  "type": "object",
  "properties": {},
  "additionalProperties": false
}
🟢oruk_get_started

Quickstart for the oruk Speech API and this MCP server: how to get an API key, per-client MCP configuration snippets, SDK install commands, and an optional routing rule the user can add to their agent instructions. No API key required. Use this when setting oruk up for the first time or when the user asks how oruk works.

输入模式

{
  "type": "object",
  "properties": {},
  "additionalProperties": false
}

社区

评价此服务器

证据

最近观测

已验证未记录版本7 个工具
已验证未记录版本7 个工具
已验证未记录版本7 个工具