Agentic RL: Credit Assignment and CLI Agents

Filter agent RL methods by supervision, critic and task setting; retrieve source links and BibTeX.

사용해야 할까요

품질 및 안전성

B
설명 품질
95%
스키마 완전성
60%
이름 품질
50%
오염 위험
100%
권한 일치
100%
프로토콜 준수
100%

발견 사항 (8)

  • LOWTool 'Agentic_RL_list_sources' doesn't follow camelCase/snake_caseAgentic_RL_list_sources에서
  • LOWTool 'Agentic_RL_search_evidence' doesn't follow camelCase/snake_caseAgentic_RL_search_evidence에서
  • LOWTool 'Agentic_RL_fetch_evidence' doesn't follow camelCase/snake_caseAgentic_RL_fetch_evidence에서
  • LOWTool 'Agentic_RL_dataset_overview' doesn't follow camelCase/snake_caseAgentic_RL_dataset_overview에서
  • LOWTool 'Agentic_RL_search_tasks' doesn't follow camelCase/snake_caseAgentic_RL_search_tasks에서
  • LOWTool 'Agentic_RL_get_task' doesn't follow camelCase/snake_caseAgentic_RL_get_task에서
  • LOWTool 'Agentic_RL_list_method_facets' doesn't follow camelCase/snake_caseAgentic_RL_list_method_facets에서
  • LOWTool 'Agentic_RL_filter_methods' doesn't follow camelCase/snake_caseAgentic_RL_filter_methods에서

도구 정의와 프로토콜 준수에 대한 자동 분석을 기반으로 합니다.

컨텍스트 비용

~787토큰 (도구 정의)
~412 B일반적인 응답 크기
중간 정도의 주의 영향 (128k 컨텍스트의 0.61%)

이는 서버의 도구가 모델의 컨텍스트에 로드될 때마다 소비되는 대략적인 토큰 수입니다. 수치가 높을수록 다른 작업에 사용할 수 있는 주의가 줄어듭니다.

설치

원클릭 설치

`claude_desktop_config.json` 파일에 다음을 추가하세요:

{
  "mcpServers": {
    "agentic-rl": {
      "url": "https://hoyant-su-agentic-rl.hf.space/gradio_api/mcp/"
    }
  }
}

원격 엔드포인트

https://hoyant-su-agentic-rl.hf.space/gradio_api/mcp/streamable-http

할 수 있는 일

도구 목록

도구 (8)

🟢 읽기 전용🟡 쓰기🔴 삭제⚪ 알 수 없음
🟢Agentic_RL_list_sources

List original papers and retrieval coverage. Discover source-linked comparisons of credit assignment, agent memory, selective observation and terminal benchmarks, with JSON, CSV and BibTeX links.

입력 스키마

{
  "type": "object",
  "properties": {}
}
🟢Agentic_RL_search_evidence(query, limit)

Search original papers on agentic reinforcement learning, credit assignment and CLI agents. Use English keywords (AND), OR and quoted phrases. Return relevant passages, source citations, equations and table cells.

입력 스키마

{
  "type": "object",
  "properties": {
    "query": {
      "type": "string",
      "description": ""
    },
    "limit": {
      "type": "integer",
      "description": "",
      "default": 5
    }
  },
  "required": [
    "query"
  ]
}
🟢Agentic_RL_fetch_evidence(evidence_id)

Fetch a complete original evidence block by the evidence_id returned from search_evidence, including section anchor, version, equations, table cells, links, and attribution.

입력 스키마

{
  "type": "object",
  "properties": {
    "evidence_id": {
      "type": "string",
      "description": ""
    }
  },
  "required": [
    "evidence_id"
  ]
}
🟢Agentic_RL_dataset_overview

Inspect ShellOps and ShellOps-Pro task counts, train/test splits, task types, published schemas, source files, license and citation.

입력 스키마

{
  "type": "object",
  "properties": {}
}
🟢Agentic_RL_search_tasks(query, partition, split, limit, offset)

Find real ShellOps CLI benchmark tasks by case-insensitive literal substring in the complete instruction, task ID or published task type. Empty query lists all tasks. Select partition 'all', 'shellops' or 'shellops_pro'; select published split 'all', 'train_src', 'train' or 'test'. Results are ordered by partition then task ID, with explicit pagination and no relevance scoring. The train subset is not double-counted.

입력 스키마

{
  "type": "object",
  "properties": {
    "query": {
      "type": "string",
      "description": ""
    },
    "partition": {
      "type": "string",
      "description": "",
      "default": "all"
    },
    "split": {
      "type": "string",
      "description": "",
      "default": "all"
    },
    "limit": {
      "type": "integer",
      "description": "",
      "default": 10
    },
    "offset": {
      "type": "integer",
      "description": ""
    }
  },
  "required": [
    "query"
  ]
}
🟢Agentic_RL_get_task(task_id, partition)

Inspect one published ShellOps or ShellOps-Pro task by its exact task_id and partition ('shellops' or 'shellops_pro'). Returns the complete instruction, actual reward specification, published reference answer/command, file-entry metadata, pinned parquet rows and workspace asset links. File content is available at the source links. No shell execution or solution verification is performed.

입력 스키마

{
  "type": "object",
  "properties": {
    "task_id": {
      "type": "string",
      "description": ""
    },
    "partition": {
      "type": "string",
      "description": ""
    }
  },
  "required": [
    "task_id",
    "partition"
  ]
}
🟢Agentic_RL_list_method_facets

List exact filter values for agent RL credit granularity, supervision, value critics and evaluation settings. Each value reports its source-supported method count.

입력 스키마

{
  "type": "object",
  "properties": {}
}
🟢Agentic_RL_filter_methods(credit_granularity, required_supervision, learned_value_critic, evaluation_setting)

Filter agent RL credit-assignment methods by research conditions and return original section evidence and BibTeX. Discover accepted values with list_method_facets. Filters combine with AND; empty strings leave a facet unrestricted. Unknown critic status never matches no. Results use publication order without a relevance or quality ranking.

입력 스키마

{
  "type": "object",
  "properties": {
    "credit_granularity": {
      "type": "string",
      "description": ""
    },
    "required_supervision": {
      "type": "string",
      "description": ""
    },
    "learned_value_critic": {
      "type": "string",
      "description": ""
    },
    "evaluation_setting": {
      "type": "string",
      "description": ""
    }
  }
}

커뮤니티

이 서버 평가하기

증거

최근 관측

검증됨버전이 기록되지 않음도구 8개
검증됨버전이 기록되지 않음도구 8개
검증됨버전이 기록되지 않음도구 8개