Chinese-Llm-Benchmark
Visit Toolchinese-llm-benchmark is an Open Source & Models tool that provides a comprehensive evaluation platform for Chinese AI large language models. It includes rankings and a large defect library for over 370 models.
chinese-llm-benchmark is an Open Source & Models tool that provides a comprehensive evaluation platform for Chinese AI large language models. It includes rankings and a large defect library for over 370 models.
About
chinese-llm-benchmark, also known as ReLE (Really Reliable Live Evaluation for LLM), is a continuously updated platform for evaluating Chinese AI large language models. It currently covers 375 models, including commercial options like ChatGPT, Google Gemini, Claude, and Ernie, as well as open-source models such as Llama, GLM, and Mistral. The benchmark offers multi-dimensional capability assessments across 7 domains, including education, healthcare, finance, law, reasoning, language, and agent/tool calling, with approximately 300 detailed sub-dimensions. Beyond providing rankings, it features a defect library containing over 2 million entries, facilitating research and improvement of large models. The platform also offers free evaluation services for private large models.
Capabilities
Pricing & Plans
Open Source
Free
FAQs