领先的开源 RAG 引擎,融合 Agent 能力打造更强上下文层
模型评估
评估、基准测试与对比 AI 模型表现
全部场景·AI 与 Agent·按 Star 数展示前 30 个匹配项目
开源 AI 求职:扫描职位、结构化评估打分、定制简历并追踪投递
基准测试成绩领先的免费开源 AI 记忆系统
基准测试成绩领先的免费 AI 记忆系统
简单快速的检索增强生成框架(EMNLP2025)
跑在本机的 AI 求职框架:评估职位、定制简历、备战面试
编码 MCP 工具箱:语义检索与编辑,Agent 的 IDE
基于真实基准的编码 Agent 持久记忆方案
跟 SQL 数据库聊天:Agentic 检索实现精准 Text-to-SQL
单文件无服务记忆层,替代复杂 RAG 管道
Agent 可观测与评估 SDK:链路追踪与多 Agent 调试
LLM 微调、RAG 与评估数据集制作利器
Gorilla: Training and Evaluating LLMs for Function Calls (Tool Calls)
AI Observability & Evaluation
Self-evolving memory OS for LLM & AI Agents: ultra-persistent memory, hybrid-retrieval, and cross-task skill reuse, with 35.24% token saving
ConardLi's open-source Skills collection, featuring web design, knowledge retrieval, image generation, and more.
A Python library for anomaly detection across tabular, time series, graph, text, image, and audio data. 60+ detectors, benchmark-backed ADEn
Agent 的 AI 数据运行时:无服务器 Postgres + 多模态数据湖
Open-source context retrieval layer for AI agents
非线智能 NoneLinear - ReLE评测:中文AI大模型能力评测(持续更新):目前已囊括374个大模型,覆盖chatgpt、gpt-5.4、谷歌gemini-3.1-pro、Claude-4.6、文心ERNIE-X1.1、ERNIE-5.0、qwen3.6-max、qwe
Python SDK for AI agent monitoring, LLM cost tracking, benchmarking, and more. Integrates with most LLMs and agent frameworks including Crew
达尔文.skill —— 一个让你的Skill无限进化的系统:评估→改进→测试→保留或回滚 | Autoresearch-inspired autonomous skill optimization for Claude Code. Evaluate, improve, test, keep or revert.
The CLI and skills that turn any coding assistant into an expert at creating, evaluating, and deploying AI agents on Google Cloud.
Build, Evaluate, and Optimize AI Systems. Includes evals, RAG, agents, fine-tuning, synthetic data generation, dataset management, MCP, and
Enhanced LanceDB memory plugin for OpenClaw — Hybrid Retrieval (Vector + BM25), Cross-Encoder Rerank, Multi-Scope Isolation, Management CLI
Versioned Codex instruction deployment with preview, ownership manifests, hook isolation, scenario evaluation, and recovery.
Code, Build and Evaluate agents - excellent Model and Skills/MCP/ACP/A2A Support
Awesome list for AI agent harness engineering: tools, patterns, evals, memory, MCP, permissions, observability, and orchestration.
A Comprehensive Benchmark to Evaluate LLMs as Agents (ICLR'24)
Universal memory layer for AI Agents. It provides scalable, extensible, and interoperable memory storage and retrieval to streamline AI agen
