Arize-ai/phoenix
Phoenix is an open-source AI observability platform for experimentation, evaluation, and troubleshooting of large language model applications, supporting multiple frameworks and LLM providers with flexible deployment options.
Awesome MCP › Testing & Debugging
MCPBench is an open-source evaluation framework designed specifically for benchmarking Model Context Protocol (MCP) servers. It supports the evaluation of three main types of MCP servers: Web Search, Database Query, and GAIA tasks. The framework is compatible with both local and remote MCP servers, allowing for flexible deployment and testing scenarios. MCPBench measures key performance metrics such as task completion accuracy, latency, and token consumption, all under consistent settings of large language models (LLMs) and agents. This ensures a fair and standardized comparison across different MCP servers, including popular ones like Brave Search and DuckDuckGo. The framework provides a comprehensive setup process, including launching MCP servers either locally or remotely, and running evaluations on various tasks. It supports configuration through JSON files that specify server details and run commands, facilitating easy integration and automation. MCPBench also includes datasets for benchmarking, such as a WebSearch dataset with 600 QA pairs from diverse domains like Frames, news, and technology, and a Database Query dataset. Users can add their own datasets in a specified JSON format to extend the evaluation capabilities. MCPBench is inspired by the LangProBe benchmark and aims to provide a robust and transparent evaluation environment for MCP servers. The project is actively maintained and open-sourced, with documentation and experimental results available to the community. It requires Python 3.11 or higher, along with nodejs and jq for installation and operation. Overall, MCPBench serves as a critical tool for researchers and developers working with MCP servers, enabling them to assess and compare server performance comprehensively and consistently. It contributes to advancing the development and deployment of MCP technologies by providing standardized benchmarks and evaluation protocols.
https://github.com/modelscope/MCPBench
Phoenix is an open-source AI observability platform for experimentation, evaluation, and troubleshooting of large language model applications, supporting multiple frameworks and LLM providers with flexible deployment options.
MCP Inspector is a developer tool that provides a visual UI and CLI for testing, debugging, and interacting with Model Context Protocol (MCP) servers, enhancing MCP server development workflows.
MCPJam Inspector is a developer tool for testing and debugging MCP servers, supporting multiple protocols and LLM interaction, designed to streamline MCP development workflows.
OpenInference is an open-source extension of OpenTelemetry that provides comprehensive tracing and observability for AI applications, including support for the Model Context Protocol (MCP).
OpenMCP is an all-in-one VSCode plugin that integrates development, testing, and management tools for Model Context Protocol (MCP) server debugging and large model interaction.
MCP Node.js Debugger is an MCP server that enables AI coding assistants like Cursor and Claude Code to debug Node.js applications at runtime by setting breakpoints and inspecting runtime state.
Open source MCP analytics and authentication platform that deploys an OAuth gateway in front of MCP servers and aggregates prompt analytics, generated setup instructions and real time debug logs.
Swift app for macOS, iOS and visionOS that connects to local and remote MCP servers to browse and exercise their prompts, resources and tools while testing and debugging.