together-ai

Run Together AI chat, embeddings and images; manage fine-tunes, batches and endpoints.

사용해야 할까요

품질 및 안전성

A
설명 품질
100%
스키마 완전성
85%
이름 품질
80%
오염 위험
100%
권한 일치
100%
프로토콜 준수
100%

도구 정의와 프로토콜 준수에 대한 자동 분석을 기반으로 합니다.

컨텍스트 비용

~4,270토큰 (도구 정의)
~1.4 KB일반적인 응답 크기
상당한 주의 영향 (128k 컨텍스트의 3.34%)

이는 서버의 도구가 모델의 컨텍스트에 로드될 때마다 소비되는 대략적인 토큰 수입니다. 수치가 높을수록 다른 작업에 사용할 수 있는 주의가 줄어듭니다.

설치

원클릭 설치

`claude_desktop_config.json` 파일에 다음을 추가하세요:

{
  "mcpServers": {
    "together-ai": {
      "url": "https://together-ai.usefulapi.io/mcp"
    }
  }
}

원격 엔드포인트

https://together-ai.usefulapi.io/mcpstreamable-http

할 수 있는 일

도구 목록

도구 (23)

🟢 읽기 전용🟡 쓰기🔴 삭제⚪ 알 수 없음
🟢together_whoami

Identify the API key: its organization, project and project slug. The project slug forms the `<project_slug>/<endpoint_slug>` model name for dedicated-endpoint inference. A cheap way to confirm the key works. Together: GET /whoami.

입력 스키마

{
  "type": "object",
  "properties": {},
  "$schema": "https://json-schema.org/draft/2020-12/schema"
}
🟢together_list_models(dedicated)

List Together's models with type (chat, language, code, image, embedding, moderation, rerank), context length, organization, license and per-token pricing. Together: GET /models.

입력 스키마

{
  "type": "object",
  "properties": {
    "dedicated": {
      "description": "Only return models that can run on dedicated endpoints.",
      "type": "boolean"
    }
  },
  "$schema": "https://json-schema.org/draft/2020-12/schema"
}
🟢together_list_files

List uploaded data files (fine-tune, eval and batch-api inputs, plus job outputs) with size, type, purpose and validation status. Together: GET /files.

입력 스키마

{
  "type": "object",
  "properties": {},
  "$schema": "https://json-schema.org/draft/2020-12/schema"
}
🟢together_get_file(file_id)

Fetch one file's metadata, including its processing_status and validation_report (why a fine-tune training file was rejected). Together: GET /files/{id}.

입력 스키마

{
  "type": "object",
  "properties": {
    "file_id": {
      "type": "string",
      "minLength": 1,
      "description": "The file id, e.g. file-abc123."
    }
  },
  "required": [
    "file_id"
  ],
  "$schema": "https://json-schema.org/draft/2020-12/schema"
}
🟢together_list_fine_tunes

List fine-tuning jobs with status, base model, output model name and training settings. Together: GET /fine-tunes.

입력 스키마

{
  "type": "object",
  "properties": {},
  "$schema": "https://json-schema.org/draft/2020-12/schema"
}
🟢together_get_fine_tune(fine_tune_id)

Fetch one fine-tuning job: status, progress, hyperparameters, token counts, cost and the output model name. Together: GET /fine-tunes/{id}.

입력 스키마

{
  "type": "object",
  "properties": {
    "fine_tune_id": {
      "type": "string",
      "minLength": 1,
      "description": "The job id, e.g. ft-abc123."
    }
  },
  "required": [
    "fine_tune_id"
  ],
  "$schema": "https://json-schema.org/draft/2020-12/schema"
}
🟢together_list_fine_tune_events(fine_tune_id)

List the event log of one fine-tuning job (queued, started, checkpoint saved, epoch completed, errors). The first place to look when a job failed. Together: GET /fine-tunes/{id}/events.

입력 스키마

{
  "type": "object",
  "properties": {
    "fine_tune_id": {
      "type": "string",
      "minLength": 1,
      "description": "The job id, e.g. ft-abc123."
    }
  },
  "required": [
    "fine_tune_id"
  ],
  "$schema": "https://json-schema.org/draft/2020-12/schema"
}
🟢together_list_batches

List batch inference jobs with status, progress, model and input/output/error file ids. Together: GET /batches.

입력 스키마

{
  "type": "object",
  "properties": {},
  "$schema": "https://json-schema.org/draft/2020-12/schema"
}
🟢together_get_batch(batch_id)

Fetch one batch job: status (VALIDATING, IN_PROGRESS, COMPLETED, FAILED, EXPIRED, CANCELLED), progress, and the output_file_id / error_file_id once done. Together: GET /batches/{id}.

입력 스키마

{
  "type": "object",
  "properties": {
    "batch_id": {
      "type": "string",
      "minLength": 1,
      "description": "The batch job id."
    }
  },
  "required": [
    "batch_id"
  ],
  "$schema": "https://json-schema.org/draft/2020-12/schema"
}
🟢together_list_endpoints(type, usage_type, mine)

List endpoints with model, owner and state (PENDING, STARTING, STARTED, STOPPING, STOPPED, ERROR). Use mine=true and type=dedicated to see what is running on your account and billing by the minute. Together: GET /endpoints.

입력 스키마

{
  "type": "object",
  "properties": {
    "type": {
      "description": "Filter by endpoint type.",
      "type": "string",
      "enum": [
        "dedicated",
        "serverless"
      ]
    },
    "usage_type": {
      "description": "Filter by usage type.",
      "type": "string",
      "enum": [
        "on-demand",
        "reserved"
      ]
    },
    "mine": {
      "description": "Only endpoints owned by the caller.",
      "type": "boolean"
    }
  },
  "$schema": "https://json-schema.org/draft/2020-12/schema"
}
🟢together_get_endpoint(endpoint_id)

Fetch one dedicated endpoint: state, model, hardware, autoscaling bounds and display name. Together: GET /endpoints/{endpointId}.

입력 스키마

{
  "type": "object",
  "properties": {
    "endpoint_id": {
      "type": "string",
      "minLength": 1,
      "description": "The endpoint id, e.g. endpoint-d23901de-...."
    }
  },
  "required": [
    "endpoint_id"
  ],
  "$schema": "https://json-schema.org/draft/2020-12/schema"
}
🟢together_list_hardware(model)

List hardware configurations for dedicated endpoints with GPU type/count/memory and price in cents per minute. Pass a model to get only compatible configurations with live availability. Together: GET /hardware.

입력 스키마

{
  "type": "object",
  "properties": {
    "model": {
      "description": "Only hardware compatible with this model, with availability.",
      "type": "string"
    }
  },
  "$schema": "https://json-schema.org/draft/2020-12/schema"
}
🟢together_list_evaluations(status, limit)

List LLM-as-a-judge evaluation jobs (classify, score, compare) with status, parameters and results once completed. Together: GET /evaluation.

입력 스키마

{
  "type": "object",
  "properties": {
    "status": {
      "description": "Filter by status: pending, queued, running, completed, error, user_error.",
      "type": "string"
    },
    "limit": {
      "description": "Maximum number of jobs to return.",
      "type": "integer",
      "minimum": 1,
      "maximum": 1000
    }
  },
  "$schema": "https://json-schema.org/draft/2020-12/schema"
}
🔴together_chat_completion(model, messages, max_tokens, temperature, top_p, ...)

Run a chat completion on a Together model (billed per token). Non-streaming. For a dedicated endpoint pass its `<project_slug>/<endpoint_slug>` as the model. Together: POST /chat/completions.

입력 스키마

{
  "type": "object",
  "properties": {
    "model": {
      "type": "string",
      "minLength": 1,
      "description": "Model name, e.g. meta-llama/Llama-3.3-70B-Instruct-Turbo."
    },
    "messages": {
      "minItems": 1,
      "type": "array",
      "items": {
        "type": "object",
        "properties": {
          "role": {
            "type": "string",
            "enum": [
              "system",
              "user",
              "assistant",
              "tool"
            ]
          },
          "content": {
            "type": "string"
          }
        },
        "required": [
          "role",
          "content"
        ]
      },
      "description": "The conversation so far."
    },
    "max_tokens": {
      "description": "Maximum tokens to generate.",
      "type": "integer",
      "minimum": 1,
      "maximum": 9007199254740991
    },
    "temperature": {
      "type": "number",
      "minimum": 0,
      "maximum": 2
    },
    "top_p": {
      "type": "number",
      "minimum": 0,
      "maximum": 1
    },
    "top_k": {
      "type": "integer",
      "minimum": 0,
      "maximum": 9007199254740991
    },
    "repetition_penalty": {
      "type": "number"
    },
    "stop": {
      "description": "Stop sequences.",
      "type": "array",
      "items": {
        "type": "string"
      }
    },
    "seed": {
      "description": "Seed for reproducible sampling.",
      "type": "integer",
      "minimum": -9007199254740991,
      "maximum": 9007199254740991
    },
    "reasoning_effort": {
      "description": "Reasoning effort for reasoning models that support it.",
      "type": "string",
      "enum": [
        "low",
        "medium",
        "high"
      ]
    }
  },
  "required": [
    "model",
    "messages"
  ],
  "$schema": "https://json-schema.org/draft/2020-12/schema"
}
🔴together_create_embeddings(model, input)

Generate vector embeddings for one or more texts (billed per token). Together: POST /embeddings.

입력 스키마

{
  "type": "object",
  "properties": {
    "model": {
      "type": "string",
      "minLength": 1,
      "description": "Embedding model, e.g. BAAI/bge-large-en-v1.5."
    },
    "input": {
      "anyOf": [
        {
          "type": "string"
        },
        {
          "minItems": 1,
          "type": "array",
          "items": {
            "type": "string"
          }
        }
      ],
      "description": "A text, or a list of texts, to embed."
    }
  },
  "required": [
    "model",
    "input"
  ],
  "$schema": "https://json-schema.org/draft/2020-12/schema"
}
🔴together_generate_image(model, prompt, negative_prompt, width, height, ...)

Generate images from a prompt (billed per image/megapixel). Returns image URLs by default rather than base64, to keep responses small. Together: POST /images/generations.

입력 스키마

{
  "type": "object",
  "properties": {
    "model": {
      "type": "string",
      "minLength": 1,
      "description": "Image model, e.g. black-forest-labs/FLUX.1-schnell."
    },
    "prompt": {
      "type": "string",
      "minLength": 1,
      "description": "What to draw."
    },
    "negative_prompt": {
      "description": "What to steer away from.",
      "type": "string"
    },
    "width": {
      "description": "Width in pixels.",
      "type": "integer",
      "minimum": 64,
      "maximum": 9007199254740991
    },
    "height": {
      "description": "Height in pixels.",
      "type": "integer",
      "minimum": 64,
      "maximum": 9007199254740991
    },
    "n": {
      "description": "Number of images.",
      "type": "integer",
      "minimum": 1,
      "maximum": 4
    },
    "steps": {
      "description": "Number of generation steps.",
      "type": "integer",
      "minimum": 1,
      "maximum": 9007199254740991
    },
    "seed": {
      "type": "integer",
      "minimum": -9007199254740991,
      "maximum": 9007199254740991
    },
    "guidance_scale": {
      "description": "Prompt adherence; higher is more literal.",
      "type": "number"
    },
    "image_url": {
      "description": "Input image URL, for models that support editing.",
      "type": "string"
    },
    "output_format": {
      "type": "string",
      "enum": [
        "jpeg",
        "png"
      ]
    },
    "response_format": {
      "description": "url (default here) or base64. base64 can be very large.",
      "type": "string",
      "enum": [
        "url",
        "base64"
      ]
    }
  },
  "required": [
    "model",
    "prompt"
  ],
  "$schema": "https://json-schema.org/draft/2020-12/schema"
}
🔴together_create_fine_tune(model, training_file, validation_file, suffix, n_epochs, ...)

Start a fine-tuning job on an uploaded training file (billed per token processed). Stop it with together_cancel_fine_tune. Together: POST /fine-tunes.

입력 스키마

{
  "type": "object",
  "properties": {
    "model": {
      "type": "string",
      "minLength": 1,
      "description": "Base model to fine-tune."
    },
    "training_file": {
      "type": "string",
      "minLength": 1,
      "description": "File id of an uploaded training file (purpose fine-tune)."
    },
    "validation_file": {
      "description": "File id of an uploaded validation file.",
      "type": "string"
    },
    "suffix": {
      "description": "Suffix for the fine-tuned model's name (max 64 chars).",
      "type": "string",
      "maxLength": 64
    },
    "n_epochs": {
      "description": "Passes over the training data.",
      "type": "integer",
      "minimum": 1,
      "maximum": 9007199254740991
    },
    "n_checkpoints": {
      "description": "Intermediate checkpoints to save.",
      "type": "integer",
      "minimum": 1,
      "maximum": 9007199254740991
    },
    "n_evals": {
      "description": "Evaluations on the validation set during training.",
      "type": "integer",
      "minimum": 0,
      "maximum": 9007199254740991
    },
    "batch_size": {
      "description": "Batch size, or 'max' (the default).",
      "anyOf": [
        {
          "type": "integer",
          "minimum": 1,
          "maximum": 9007199254740991
        },
        {
          "type": "string",
          "const": "max"
        }
      ]
    },
    "learning_rate": {
      "type": "number",
      "exclusiveMinimum": 0
    },
    "warmup_ratio": {
      "type": "number",
      "minimum": 0,
      "maximum": 1
    },
    "max_seq_length": {
      "type": "integer",
      "minimum": 1,
      "maximum": 9007199254740991
    },
    "from_checkpoint": {
      "description": "Continue from a previous job: <job_id>, <output_model_name>, optionally with :<step>.",
      "type": "string"
    },
    "training_type": {
      "description": "Full fine-tune or LoRA. Together defaults to LoRA when omitted.",
      "oneOf": [
        {
          "type": "object",
          "properties": {
            "type": {
              "type": "string",
              "const": "Full"
            }
          },
          "required": [
            "type"
          ]
        },
        {
          "type": "object",
          "properties": {
            "type": {
              "type": "string",
              "const": "Lora"
            },
            "lora_r": {
              "type": "integer",
              "minimum": 1,
              "maximum": 9007199254740991,
              "description": "Rank of the LoRA adapter matrices."
            },
            "lora_alpha": {
              "type": "number",
              "description": "Scaling factor applied to the LoRA adapter weights."
            },
            "lora_dropout": {
              "description": "Dropout on LoRA adapter inputs.",
              "type": "number",
              "minimum": 0,
              "maximum": 1
            },
            "lora_trainable_modules": {
              "description": "Comma-separated target modules, or all-linear for the model defaults.",
              "type": "string"
            }
          },
          "required": [
            "type",
            "lora_r",
            "lora_alpha"
          ]
        }
      ]
    },
    "training_method": {
      "description": "Supervised fine-tuning (sft, the default) or preference tuning (dpo).",
      "oneOf": [
        {
          "type": "object",
          "properties": {
            "method": {
              "type": "string",
              "const": "sft"
            },
            "train_on_inputs": {
              "anyOf": [
                {
                  "type": "boolean"
                },
                {
                  "type": "string",
                  "const": "auto"
                }
              ],
              "description": "Whether prompt/user tokens contribute to the loss; 'auto' lets Together decide."
            }
          },
          "required": [
            "method",
            "train_on_inputs"
          ]
        },
        {
          "type": "object",
          "properties": {
            "method": {
              "type": "string",
              "const": "dpo"
            },
            "dpo_beta": {
              "type": "number"
            },
            "dpo_normalize_logratios_by_length": {
              "type": "boolean"
            },
            "dpo_reference_free": {
              "type": "boolean"
            },
            "rpo_alpha": {
              "type": "number"
            },
            "simpo_gamma": {
              "type": "number"
            }
          },
          "required": [
            "method"
          ]
        }
      ]
    }
  },
  "required": [
    "model",
    "training_file"
  ],
  "$schema": "https://json-schema.org/draft/2020-12/schema"
}
🔴together_cancel_fine_tune(fine_tune_id)

Cancel a running fine-tuning job. Cannot be resumed, but a new job can continue from its last checkpoint via from_checkpoint. Together: POST /fine-tunes/{id}/cancel.

입력 스키마

{
  "type": "object",
  "properties": {
    "fine_tune_id": {
      "type": "string",
      "minLength": 1,
      "description": "The job id, e.g. ft-abc123."
    }
  },
  "required": [
    "fine_tune_id"
  ],
  "$schema": "https://json-schema.org/draft/2020-12/schema"
}
🔴together_create_batch(input_file_id, endpoint, model_id, completion_window, priority)

Start an asynchronous batch job over an uploaded JSONL input file (purpose batch-api), at a discount to real-time inference. Together: POST /batches.

입력 스키마

{
  "type": "object",
  "properties": {
    "input_file_id": {
      "type": "string",
      "minLength": 1,
      "description": "File id of the uploaded JSONL request file."
    },
    "endpoint": {
      "type": "string",
      "enum": [
        "/v1/chat/completions",
        "/v1/audio/transcriptions",
        "/v1/audio/translations"
      ],
      "description": "The API each line of the input file is sent to."
    },
    "model_id": {
      "description": "Model to process the requests with.",
      "type": "string"
    },
    "completion_window": {
      "description": "Time window for completion, e.g. 24h.",
      "type": "string"
    },
    "priority": {
      "description": "Processing priority.",
      "type": "integer",
      "minimum": -9007199254740991,
      "maximum": 9007199254740991
    }
  },
  "required": [
    "input_file_id",
    "endpoint"
  ],
  "$schema": "https://json-schema.org/draft/2020-12/schema"
}
🔴together_cancel_batch(batch_id)

Cancel a batch job that has not finished. Together: POST /batches/{id}/cancel.

입력 스키마

{
  "type": "object",
  "properties": {
    "batch_id": {
      "type": "string",
      "minLength": 1,
      "description": "The batch job id."
    }
  },
  "required": [
    "batch_id"
  ],
  "$schema": "https://json-schema.org/draft/2020-12/schema"
}
🔴together_create_endpoint(model, hardware, autoscaling, display_name, availability_zone, ...)

Deploy a model on dedicated GPUs. The endpoint STARTS AUTOMATICALLY and bills per minute of uptime until stopped — set inactive_timeout to auto-stop it, and use together_list_hardware for valid hardware ids. Together: POST /endpoints.

입력 스키마

{
  "type": "object",
  "properties": {
    "model": {
      "type": "string",
      "minLength": 1,
      "description": "The model to deploy."
    },
    "hardware": {
      "type": "string",
      "minLength": 1,
      "description": "Hardware id, e.g. 1x_nvidia_a100_80gb_sxm."
    },
    "autoscaling": {
      "type": "object",
      "properties": {
        "min_replicas": {
          "type": "integer",
          "minimum": 0,
          "maximum": 9007199254740991,
          "description": "Replicas kept running even with no load."
        },
        "max_replicas": {
          "type": "integer",
          "minimum": 1,
          "maximum": 9007199254740991,
          "description": "Maximum replicas to scale up to under load."
        }
      },
      "required": [
        "min_replicas",
        "max_replicas"
      ],
      "description": "Replica bounds for autoscaling."
    },
    "display_name": {
      "description": "Human-readable name.",
      "type": "string"
    },
    "availability_zone": {
      "description": "Availability zone, e.g. us-central-4b.",
      "type": "string"
    },
    "inactive_timeout": {
      "description": "Minutes of inactivity before auto-stop; 0 disables it.",
      "type": "integer",
      "minimum": 0,
      "maximum": 9007199254740991
    },
    "disable_speculative_decoding": {
      "type": "boolean"
    },
    "state": {
      "description": "Initial state. Pass STOPPED to create without starting (and without billing).",
      "type": "string",
      "enum": [
        "STARTED",
        "STOPPED"
      ]
    }
  },
  "required": [
    "model",
    "hardware",
    "autoscaling"
  ],
  "$schema": "https://json-schema.org/draft/2020-12/schema"
}
🔴together_start_endpoint(endpoint_id)

Start a stopped dedicated endpoint. It bills per minute of uptime until stopped. Reversible with together_stop_endpoint. Together: PATCH /endpoints/{endpointId} with state=STARTED.

입력 스키마

{
  "type": "object",
  "properties": {
    "endpoint_id": {
      "type": "string",
      "minLength": 1,
      "description": "The endpoint id."
    }
  },
  "required": [
    "endpoint_id"
  ],
  "$schema": "https://json-schema.org/draft/2020-12/schema"
}
🔴together_stop_endpoint(endpoint_id)

Stop a running dedicated endpoint, which stops its per-minute billing. Requests to it fail until it is started again. Together: PATCH /endpoints/{endpointId} with state=STOPPED.

입력 스키마

{
  "type": "object",
  "properties": {
    "endpoint_id": {
      "type": "string",
      "minLength": 1,
      "description": "The endpoint id."
    }
  },
  "required": [
    "endpoint_id"
  ],
  "$schema": "https://json-schema.org/draft/2020-12/schema"
}

커뮤니티

이 서버 평가하기

증거

최근 관측

검증됨버전이 기록되지 않음도구 23개