together-ai

Run Together AI chat, embeddings and images; manage fine-tunes, batches and endpoints.

Sollte ich dies verwenden

Qualität und Sicherheit

A
Qualität der Beschreibung
100%
Vollständigkeit des Schemas
85%
Qualität der Benennung
80%
Risiko der Vergiftung
100%
Übereinstimmung der Berechtigungen
100%
Einhaltung des Protokolls
100%

Basierend auf einer automatisierten Analyse der Tool-Definitionen und der Einhaltung des Protokolls.

Kontextkosten

~4,270Tokens (Tool-Definitionen)
~1.4 KBTypische Antwortgröße
Erhebliche Auswirkung auf die Aufmerksamkeit (3.34% von 128k Kontext)

Dies ist die ungefähre Anzahl der Tokens, die jedes Mal verbraucht werden, wenn die Tools des Servers in den Kontext eines Modells geladen werden. Höhere Werte verringern die Aufmerksamkeit, die für andere Aufgaben verfügbar ist.

Installieren

Installation mit einem Klick

Fügen Sie dies Ihrer Datei `claude_desktop_config.json` hinzu:

{
  "mcpServers": {
    "together-ai": {
      "url": "https://together-ai.usefulapi.io/mcp"
    }
  }
}

Remote-Endpunkte

https://together-ai.usefulapi.io/mcpstreamable-http

Was es kann

Tool-Inventar

Tools (23)

🟢 Nur lesen🟡 Schreiben🔴 Löschen⚪ Unbekannt
🟢together_whoami

Identify the API key: its organization, project and project slug. The project slug forms the `<project_slug>/<endpoint_slug>` model name for dedicated-endpoint inference. A cheap way to confirm the key works. Together: GET /whoami.

Eingabe-Schema

{
  "type": "object",
  "properties": {},
  "$schema": "https://json-schema.org/draft/2020-12/schema"
}
🟢together_list_models(dedicated)

List Together's models with type (chat, language, code, image, embedding, moderation, rerank), context length, organization, license and per-token pricing. Together: GET /models.

Eingabe-Schema

{
  "type": "object",
  "properties": {
    "dedicated": {
      "description": "Only return models that can run on dedicated endpoints.",
      "type": "boolean"
    }
  },
  "$schema": "https://json-schema.org/draft/2020-12/schema"
}
🟢together_list_files

List uploaded data files (fine-tune, eval and batch-api inputs, plus job outputs) with size, type, purpose and validation status. Together: GET /files.

Eingabe-Schema

{
  "type": "object",
  "properties": {},
  "$schema": "https://json-schema.org/draft/2020-12/schema"
}
🟢together_get_file(file_id)

Fetch one file's metadata, including its processing_status and validation_report (why a fine-tune training file was rejected). Together: GET /files/{id}.

Eingabe-Schema

{
  "type": "object",
  "properties": {
    "file_id": {
      "type": "string",
      "minLength": 1,
      "description": "The file id, e.g. file-abc123."
    }
  },
  "required": [
    "file_id"
  ],
  "$schema": "https://json-schema.org/draft/2020-12/schema"
}
🟢together_list_fine_tunes

List fine-tuning jobs with status, base model, output model name and training settings. Together: GET /fine-tunes.

Eingabe-Schema

{
  "type": "object",
  "properties": {},
  "$schema": "https://json-schema.org/draft/2020-12/schema"
}
🟢together_get_fine_tune(fine_tune_id)

Fetch one fine-tuning job: status, progress, hyperparameters, token counts, cost and the output model name. Together: GET /fine-tunes/{id}.

Eingabe-Schema

{
  "type": "object",
  "properties": {
    "fine_tune_id": {
      "type": "string",
      "minLength": 1,
      "description": "The job id, e.g. ft-abc123."
    }
  },
  "required": [
    "fine_tune_id"
  ],
  "$schema": "https://json-schema.org/draft/2020-12/schema"
}
🟢together_list_fine_tune_events(fine_tune_id)

List the event log of one fine-tuning job (queued, started, checkpoint saved, epoch completed, errors). The first place to look when a job failed. Together: GET /fine-tunes/{id}/events.

Eingabe-Schema

{
  "type": "object",
  "properties": {
    "fine_tune_id": {
      "type": "string",
      "minLength": 1,
      "description": "The job id, e.g. ft-abc123."
    }
  },
  "required": [
    "fine_tune_id"
  ],
  "$schema": "https://json-schema.org/draft/2020-12/schema"
}
🟢together_list_batches

List batch inference jobs with status, progress, model and input/output/error file ids. Together: GET /batches.

Eingabe-Schema

{
  "type": "object",
  "properties": {},
  "$schema": "https://json-schema.org/draft/2020-12/schema"
}
🟢together_get_batch(batch_id)

Fetch one batch job: status (VALIDATING, IN_PROGRESS, COMPLETED, FAILED, EXPIRED, CANCELLED), progress, and the output_file_id / error_file_id once done. Together: GET /batches/{id}.

Eingabe-Schema

{
  "type": "object",
  "properties": {
    "batch_id": {
      "type": "string",
      "minLength": 1,
      "description": "The batch job id."
    }
  },
  "required": [
    "batch_id"
  ],
  "$schema": "https://json-schema.org/draft/2020-12/schema"
}
🟢together_list_endpoints(type, usage_type, mine)

List endpoints with model, owner and state (PENDING, STARTING, STARTED, STOPPING, STOPPED, ERROR). Use mine=true and type=dedicated to see what is running on your account and billing by the minute. Together: GET /endpoints.

Eingabe-Schema

{
  "type": "object",
  "properties": {
    "type": {
      "description": "Filter by endpoint type.",
      "type": "string",
      "enum": [
        "dedicated",
        "serverless"
      ]
    },
    "usage_type": {
      "description": "Filter by usage type.",
      "type": "string",
      "enum": [
        "on-demand",
        "reserved"
      ]
    },
    "mine": {
      "description": "Only endpoints owned by the caller.",
      "type": "boolean"
    }
  },
  "$schema": "https://json-schema.org/draft/2020-12/schema"
}
🟢together_get_endpoint(endpoint_id)

Fetch one dedicated endpoint: state, model, hardware, autoscaling bounds and display name. Together: GET /endpoints/{endpointId}.

Eingabe-Schema

{
  "type": "object",
  "properties": {
    "endpoint_id": {
      "type": "string",
      "minLength": 1,
      "description": "The endpoint id, e.g. endpoint-d23901de-...."
    }
  },
  "required": [
    "endpoint_id"
  ],
  "$schema": "https://json-schema.org/draft/2020-12/schema"
}
🟢together_list_hardware(model)

List hardware configurations for dedicated endpoints with GPU type/count/memory and price in cents per minute. Pass a model to get only compatible configurations with live availability. Together: GET /hardware.

Eingabe-Schema

{
  "type": "object",
  "properties": {
    "model": {
      "description": "Only hardware compatible with this model, with availability.",
      "type": "string"
    }
  },
  "$schema": "https://json-schema.org/draft/2020-12/schema"
}
🟢together_list_evaluations(status, limit)

List LLM-as-a-judge evaluation jobs (classify, score, compare) with status, parameters and results once completed. Together: GET /evaluation.

Eingabe-Schema

{
  "type": "object",
  "properties": {
    "status": {
      "description": "Filter by status: pending, queued, running, completed, error, user_error.",
      "type": "string"
    },
    "limit": {
      "description": "Maximum number of jobs to return.",
      "type": "integer",
      "minimum": 1,
      "maximum": 1000
    }
  },
  "$schema": "https://json-schema.org/draft/2020-12/schema"
}
🔴together_chat_completion(model, messages, max_tokens, temperature, top_p, ...)

Run a chat completion on a Together model (billed per token). Non-streaming. For a dedicated endpoint pass its `<project_slug>/<endpoint_slug>` as the model. Together: POST /chat/completions.

Eingabe-Schema

{
  "type": "object",
  "properties": {
    "model": {
      "type": "string",
      "minLength": 1,
      "description": "Model name, e.g. meta-llama/Llama-3.3-70B-Instruct-Turbo."
    },
    "messages": {
      "minItems": 1,
      "type": "array",
      "items": {
        "type": "object",
        "properties": {
          "role": {
            "type": "string",
            "enum": [
              "system",
              "user",
              "assistant",
              "tool"
            ]
          },
          "content": {
            "type": "string"
          }
        },
        "required": [
          "role",
          "content"
        ]
      },
      "description": "The conversation so far."
    },
    "max_tokens": {
      "description": "Maximum tokens to generate.",
      "type": "integer",
      "minimum": 1,
      "maximum": 9007199254740991
    },
    "temperature": {
      "type": "number",
      "minimum": 0,
      "maximum": 2
    },
    "top_p": {
      "type": "number",
      "minimum": 0,
      "maximum": 1
    },
    "top_k": {
      "type": "integer",
      "minimum": 0,
      "maximum": 9007199254740991
    },
    "repetition_penalty": {
      "type": "number"
    },
    "stop": {
      "description": "Stop sequences.",
      "type": "array",
      "items": {
        "type": "string"
      }
    },
    "seed": {
      "description": "Seed for reproducible sampling.",
      "type": "integer",
      "minimum": -9007199254740991,
      "maximum": 9007199254740991
    },
    "reasoning_effort": {
      "description": "Reasoning effort for reasoning models that support it.",
      "type": "string",
      "enum": [
        "low",
        "medium",
        "high"
      ]
    }
  },
  "required": [
    "model",
    "messages"
  ],
  "$schema": "https://json-schema.org/draft/2020-12/schema"
}
🔴together_create_embeddings(model, input)

Generate vector embeddings for one or more texts (billed per token). Together: POST /embeddings.

Eingabe-Schema

{
  "type": "object",
  "properties": {
    "model": {
      "type": "string",
      "minLength": 1,
      "description": "Embedding model, e.g. BAAI/bge-large-en-v1.5."
    },
    "input": {
      "anyOf": [
        {
          "type": "string"
        },
        {
          "minItems": 1,
          "type": "array",
          "items": {
            "type": "string"
          }
        }
      ],
      "description": "A text, or a list of texts, to embed."
    }
  },
  "required": [
    "model",
    "input"
  ],
  "$schema": "https://json-schema.org/draft/2020-12/schema"
}
🔴together_generate_image(model, prompt, negative_prompt, width, height, ...)

Generate images from a prompt (billed per image/megapixel). Returns image URLs by default rather than base64, to keep responses small. Together: POST /images/generations.

Eingabe-Schema

{
  "type": "object",
  "properties": {
    "model": {
      "type": "string",
      "minLength": 1,
      "description": "Image model, e.g. black-forest-labs/FLUX.1-schnell."
    },
    "prompt": {
      "type": "string",
      "minLength": 1,
      "description": "What to draw."
    },
    "negative_prompt": {
      "description": "What to steer away from.",
      "type": "string"
    },
    "width": {
      "description": "Width in pixels.",
      "type": "integer",
      "minimum": 64,
      "maximum": 9007199254740991
    },
    "height": {
      "description": "Height in pixels.",
      "type": "integer",
      "minimum": 64,
      "maximum": 9007199254740991
    },
    "n": {
      "description": "Number of images.",
      "type": "integer",
      "minimum": 1,
      "maximum": 4
    },
    "steps": {
      "description": "Number of generation steps.",
      "type": "integer",
      "minimum": 1,
      "maximum": 9007199254740991
    },
    "seed": {
      "type": "integer",
      "minimum": -9007199254740991,
      "maximum": 9007199254740991
    },
    "guidance_scale": {
      "description": "Prompt adherence; higher is more literal.",
      "type": "number"
    },
    "image_url": {
      "description": "Input image URL, for models that support editing.",
      "type": "string"
    },
    "output_format": {
      "type": "string",
      "enum": [
        "jpeg",
        "png"
      ]
    },
    "response_format": {
      "description": "url (default here) or base64. base64 can be very large.",
      "type": "string",
      "enum": [
        "url",
        "base64"
      ]
    }
  },
  "required": [
    "model",
    "prompt"
  ],
  "$schema": "https://json-schema.org/draft/2020-12/schema"
}
🔴together_create_fine_tune(model, training_file, validation_file, suffix, n_epochs, ...)

Start a fine-tuning job on an uploaded training file (billed per token processed). Stop it with together_cancel_fine_tune. Together: POST /fine-tunes.

Eingabe-Schema

{
  "type": "object",
  "properties": {
    "model": {
      "type": "string",
      "minLength": 1,
      "description": "Base model to fine-tune."
    },
    "training_file": {
      "type": "string",
      "minLength": 1,
      "description": "File id of an uploaded training file (purpose fine-tune)."
    },
    "validation_file": {
      "description": "File id of an uploaded validation file.",
      "type": "string"
    },
    "suffix": {
      "description": "Suffix for the fine-tuned model's name (max 64 chars).",
      "type": "string",
      "maxLength": 64
    },
    "n_epochs": {
      "description": "Passes over the training data.",
      "type": "integer",
      "minimum": 1,
      "maximum": 9007199254740991
    },
    "n_checkpoints": {
      "description": "Intermediate checkpoints to save.",
      "type": "integer",
      "minimum": 1,
      "maximum": 9007199254740991
    },
    "n_evals": {
      "description": "Evaluations on the validation set during training.",
      "type": "integer",
      "minimum": 0,
      "maximum": 9007199254740991
    },
    "batch_size": {
      "description": "Batch size, or 'max' (the default).",
      "anyOf": [
        {
          "type": "integer",
          "minimum": 1,
          "maximum": 9007199254740991
        },
        {
          "type": "string",
          "const": "max"
        }
      ]
    },
    "learning_rate": {
      "type": "number",
      "exclusiveMinimum": 0
    },
    "warmup_ratio": {
      "type": "number",
      "minimum": 0,
      "maximum": 1
    },
    "max_seq_length": {
      "type": "integer",
      "minimum": 1,
      "maximum": 9007199254740991
    },
    "from_checkpoint": {
      "description": "Continue from a previous job: <job_id>, <output_model_name>, optionally with :<step>.",
      "type": "string"
    },
    "training_type": {
      "description": "Full fine-tune or LoRA. Together defaults to LoRA when omitted.",
      "oneOf": [
        {
          "type": "object",
          "properties": {
            "type": {
              "type": "string",
              "const": "Full"
            }
          },
          "required": [
            "type"
          ]
        },
        {
          "type": "object",
          "properties": {
            "type": {
              "type": "string",
              "const": "Lora"
            },
            "lora_r": {
              "type": "integer",
              "minimum": 1,
              "maximum": 9007199254740991,
              "description": "Rank of the LoRA adapter matrices."
            },
            "lora_alpha": {
              "type": "number",
              "description": "Scaling factor applied to the LoRA adapter weights."
            },
            "lora_dropout": {
              "description": "Dropout on LoRA adapter inputs.",
              "type": "number",
              "minimum": 0,
              "maximum": 1
            },
            "lora_trainable_modules": {
              "description": "Comma-separated target modules, or all-linear for the model defaults.",
              "type": "string"
            }
          },
          "required": [
            "type",
            "lora_r",
            "lora_alpha"
          ]
        }
      ]
    },
    "training_method": {
      "description": "Supervised fine-tuning (sft, the default) or preference tuning (dpo).",
      "oneOf": [
        {
          "type": "object",
          "properties": {
            "method": {
              "type": "string",
              "const": "sft"
            },
            "train_on_inputs": {
              "anyOf": [
                {
                  "type": "boolean"
                },
                {
                  "type": "string",
                  "const": "auto"
                }
              ],
              "description": "Whether prompt/user tokens contribute to the loss; 'auto' lets Together decide."
            }
          },
          "required": [
            "method",
            "train_on_inputs"
          ]
        },
        {
          "type": "object",
          "properties": {
            "method": {
              "type": "string",
              "const": "dpo"
            },
            "dpo_beta": {
              "type": "number"
            },
            "dpo_normalize_logratios_by_length": {
              "type": "boolean"
            },
            "dpo_reference_free": {
              "type": "boolean"
            },
            "rpo_alpha": {
              "type": "number"
            },
            "simpo_gamma": {
              "type": "number"
            }
          },
          "required": [
            "method"
          ]
        }
      ]
    }
  },
  "required": [
    "model",
    "training_file"
  ],
  "$schema": "https://json-schema.org/draft/2020-12/schema"
}
🔴together_cancel_fine_tune(fine_tune_id)

Cancel a running fine-tuning job. Cannot be resumed, but a new job can continue from its last checkpoint via from_checkpoint. Together: POST /fine-tunes/{id}/cancel.

Eingabe-Schema

{
  "type": "object",
  "properties": {
    "fine_tune_id": {
      "type": "string",
      "minLength": 1,
      "description": "The job id, e.g. ft-abc123."
    }
  },
  "required": [
    "fine_tune_id"
  ],
  "$schema": "https://json-schema.org/draft/2020-12/schema"
}
🔴together_create_batch(input_file_id, endpoint, model_id, completion_window, priority)

Start an asynchronous batch job over an uploaded JSONL input file (purpose batch-api), at a discount to real-time inference. Together: POST /batches.

Eingabe-Schema

{
  "type": "object",
  "properties": {
    "input_file_id": {
      "type": "string",
      "minLength": 1,
      "description": "File id of the uploaded JSONL request file."
    },
    "endpoint": {
      "type": "string",
      "enum": [
        "/v1/chat/completions",
        "/v1/audio/transcriptions",
        "/v1/audio/translations"
      ],
      "description": "The API each line of the input file is sent to."
    },
    "model_id": {
      "description": "Model to process the requests with.",
      "type": "string"
    },
    "completion_window": {
      "description": "Time window for completion, e.g. 24h.",
      "type": "string"
    },
    "priority": {
      "description": "Processing priority.",
      "type": "integer",
      "minimum": -9007199254740991,
      "maximum": 9007199254740991
    }
  },
  "required": [
    "input_file_id",
    "endpoint"
  ],
  "$schema": "https://json-schema.org/draft/2020-12/schema"
}
🔴together_cancel_batch(batch_id)

Cancel a batch job that has not finished. Together: POST /batches/{id}/cancel.

Eingabe-Schema

{
  "type": "object",
  "properties": {
    "batch_id": {
      "type": "string",
      "minLength": 1,
      "description": "The batch job id."
    }
  },
  "required": [
    "batch_id"
  ],
  "$schema": "https://json-schema.org/draft/2020-12/schema"
}
🔴together_create_endpoint(model, hardware, autoscaling, display_name, availability_zone, ...)

Deploy a model on dedicated GPUs. The endpoint STARTS AUTOMATICALLY and bills per minute of uptime until stopped — set inactive_timeout to auto-stop it, and use together_list_hardware for valid hardware ids. Together: POST /endpoints.

Eingabe-Schema

{
  "type": "object",
  "properties": {
    "model": {
      "type": "string",
      "minLength": 1,
      "description": "The model to deploy."
    },
    "hardware": {
      "type": "string",
      "minLength": 1,
      "description": "Hardware id, e.g. 1x_nvidia_a100_80gb_sxm."
    },
    "autoscaling": {
      "type": "object",
      "properties": {
        "min_replicas": {
          "type": "integer",
          "minimum": 0,
          "maximum": 9007199254740991,
          "description": "Replicas kept running even with no load."
        },
        "max_replicas": {
          "type": "integer",
          "minimum": 1,
          "maximum": 9007199254740991,
          "description": "Maximum replicas to scale up to under load."
        }
      },
      "required": [
        "min_replicas",
        "max_replicas"
      ],
      "description": "Replica bounds for autoscaling."
    },
    "display_name": {
      "description": "Human-readable name.",
      "type": "string"
    },
    "availability_zone": {
      "description": "Availability zone, e.g. us-central-4b.",
      "type": "string"
    },
    "inactive_timeout": {
      "description": "Minutes of inactivity before auto-stop; 0 disables it.",
      "type": "integer",
      "minimum": 0,
      "maximum": 9007199254740991
    },
    "disable_speculative_decoding": {
      "type": "boolean"
    },
    "state": {
      "description": "Initial state. Pass STOPPED to create without starting (and without billing).",
      "type": "string",
      "enum": [
        "STARTED",
        "STOPPED"
      ]
    }
  },
  "required": [
    "model",
    "hardware",
    "autoscaling"
  ],
  "$schema": "https://json-schema.org/draft/2020-12/schema"
}
🔴together_start_endpoint(endpoint_id)

Start a stopped dedicated endpoint. It bills per minute of uptime until stopped. Reversible with together_stop_endpoint. Together: PATCH /endpoints/{endpointId} with state=STARTED.

Eingabe-Schema

{
  "type": "object",
  "properties": {
    "endpoint_id": {
      "type": "string",
      "minLength": 1,
      "description": "The endpoint id."
    }
  },
  "required": [
    "endpoint_id"
  ],
  "$schema": "https://json-schema.org/draft/2020-12/schema"
}
🔴together_stop_endpoint(endpoint_id)

Stop a running dedicated endpoint, which stops its per-minute billing. Requests to it fail until it is started again. Together: PATCH /endpoints/{endpointId} with state=STOPPED.

Eingabe-Schema

{
  "type": "object",
  "properties": {
    "endpoint_id": {
      "type": "string",
      "minLength": 1,
      "description": "The endpoint id."
    }
  },
  "required": [
    "endpoint_id"
  ],
  "$schema": "https://json-schema.org/draft/2020-12/schema"
}

Community

Diesen Server bewerten

Nachweis

Aktuelle Beobachtungen

verifiziertVersion nicht aufgezeichnet23 Tools