Skip to content

模型评估

评估、基准测试与对比 AI 模型表现

全部场景·AI 与 Agent·按 Star 数展示前 30 个匹配项目

ragflowinfiniflow
安全

领先的开源 RAG 引擎,融合 Agent 能力打造更强上下文层

87.1k63.3
career-opssantifer
安全

开源 AI 求职:扫描职位、结构化评估打分、定制简历并追踪投递

68.5k62065.3
详情
agentmemoryrohitg00
安全

基于真实基准的编码 Agent 持久记忆方案

27.3k141.3k61.9
详情
gorillaShishirPatil
安全

Gorilla: Training and Evaluating LLMs for Function Calls (Tool Calls)

13.0k65.5
MemOSMemTensor
安全

Self-evolving memory OS for LLM & AI Agents: ultra-persistent memory, hybrid-retrieval, and cross-task skill reuse, with 35.24% token saving

11.0k61.5
安全

ConardLi's open-source Skills collection, featuring web design, knowledge retrieval, image generation, and more.

10.7k61.2
pyodyzhao062
安全

A Python library for anomaly detection across tabular, time series, graph, text, image, and audio data. 60+ detectors, benchmark-backed ADEn

10.0k52.8
deeplakeactiveloopai
安全

Agent 的 AI 数据运行时:无服务器 Postgres + 多模态数据湖

9.2k61.4
详情
安全

非线智能 NoneLinear - ReLE评测:中文AI大模型能力评测(持续更新):目前已囊括374个大模型,覆盖chatgpt、gpt-5.4、谷歌gemini-3.1-pro、Claude-4.6、文心ERNIE-X1.1、ERNIE-5.0、qwen3.6-max、qwe

6.4k51
agentopsAgentOps-AI
安全

Python SDK for AI agent monitoring, LLM cost tracking, benchmarking, and more. Integrates with most LLMs and agent frameworks including Crew

5.8k63.4
darwin-skillalchaincyf
安全

达尔文.skill —— 一个让你的Skill无限进化的系统:评估→改进→测试→保留或回滚 | Autoresearch-inspired autonomous skill optimization for Claude Code. Evaluate, improve, test, keep or revert.

5.7k10.1k60.9
详情
安全

The CLI and skills that turn any coding assistant into an expert at creating, evaluating, and deploying AI agents on Google Cloud.

5.7k66.1
KilnKiln-AI
安全

Build, Evaluate, and Optimize AI Systems. Includes evals, RAG, agents, fine-tuning, synthetic data generation, dataset management, MCP, and

5.0k58.3
安全

Enhanced LanceDB memory plugin for OpenClaw — Hybrid Retrieval (Vector + BM25), Cross-Encoder Rerank, Multi-Scope Isolation, Management CLI

4.5k65.6
安全

Versioned Codex instruction deployment with preview, ownership manifests, hook isolation, scenario evaluation, and recovery.

4.0k64.9
fast-agentevalstate
安全

Code, Build and Evaluate agents - excellent Model and Skills/MCP/ACP/A2A Support

3.9k65.3
MemMachineMemMachine
安全

Universal memory layer for AI Agents. It provides scalable, extensible, and interoperable memory storage and retrieval to streamline AI agen

3.2k67.3