PyPI - evalscope - Versions diffs - 0.13.2__tar.gz → 0.14.0__tar.gz - Mend

evalscope 0.13.2tar.gz → 0.14.0tar.gz

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.

Potentially problematic release.

This version of evalscope might be problematic. Click here for more details.

Files changed (362) hide show

{evalscope-0.13.2/evalscope.egg-info → evalscope-0.14.0}/PKG-INFO RENAMED Viewed

@@ -1,6 +1,6 @@
 Metadata-Version: 2.1
 Name: evalscope
-Version: 0.13.2
+Version: 0.14.0
 Summary: EvalScope: Lightweight LLMs Evaluation Framework
 Home-page: https://github.com/modelscope/evalscope
 Author: ModelScope team
@@ -47,12 +47,12 @@ Requires-Dist: ms-opencompass>=0.1.4; extra == "opencompass"
 Provides-Extra: vlmeval
 Requires-Dist: ms-vlmeval>=0.0.9; extra == "vlmeval"
 Provides-Extra: rag
-Requires-Dist: langchain<0.3.0; extra == "rag"
-Requires-Dist: langchain-community<0.3.0; extra == "rag"
-Requires-Dist: langchain-core<0.3.0; extra == "rag"
-Requires-Dist: langchain-openai<0.3.0; extra == "rag"
+Requires-Dist: langchain<0.4.0,>=0.3.0; extra == "rag"
+Requires-Dist: langchain-community<0.4.0,>=0.3.0; extra == "rag"
+Requires-Dist: langchain-core<0.4.0,>=0.3.0; extra == "rag"
+Requires-Dist: langchain-openai<0.4.0,>=0.3.0; extra == "rag"
 Requires-Dist: mteb==1.19.4; extra == "rag"
-Requires-Dist: ragas==0.2.9; extra == "rag"
+Requires-Dist: ragas==0.2.14; extra == "rag"
 Requires-Dist: webdataset>0.2.0; extra == "rag"
 Provides-Extra: perf
 Requires-Dist: aiohttp; extra == "perf"
@@ -93,12 +93,12 @@ Requires-Dist: transformers>=4.33; extra == "all"
 Requires-Dist: word2number; extra == "all"
 Requires-Dist: ms-opencompass>=0.1.4; extra == "all"
 Requires-Dist: ms-vlmeval>=0.0.9; extra == "all"
-Requires-Dist: langchain<0.3.0; extra == "all"
-Requires-Dist: langchain-community<0.3.0; extra == "all"
-Requires-Dist: langchain-core<0.3.0; extra == "all"
-Requires-Dist: langchain-openai<0.3.0; extra == "all"
+Requires-Dist: langchain<0.4.0,>=0.3.0; extra == "all"
+Requires-Dist: langchain-community<0.4.0,>=0.3.0; extra == "all"
+Requires-Dist: langchain-core<0.4.0,>=0.3.0; extra == "all"
+Requires-Dist: langchain-openai<0.4.0,>=0.3.0; extra == "all"
 Requires-Dist: mteb==1.19.4; extra == "all"
-Requires-Dist: ragas==0.2.9; extra == "all"
+Requires-Dist: ragas==0.2.14; extra == "all"
 Requires-Dist: webdataset>0.2.0; extra == "all"
 Requires-Dist: aiohttp; extra == "all"
 Requires-Dist: fastapi; extra == "all"
@@ -121,7 +121,7 @@ Requires-Dist: plotly<6.0.0,>=5.23.0; extra == "all"
 </p>
 <p align="center">
-<img src="https://img.shields.io/badge/python-%E2%89%A53.8-5be.svg">
+<img src="https://img.shields.io/badge/python-%E2%89%A53.9-5be.svg">
 <a href="https://badge.fury.io/py/evalscope"><img src="https://badge.fury.io/py/evalscope.svg" alt="PyPI version" height="18"></a>
 <a href="https://pypi.org/project/evalscope"><img alt="PyPI - Downloads" src="https://static.pepy.tech/badge/evalscope"></a>
 <a href="https://github.com/modelscope/evalscope/pulls"><img src="https://img.shields.io/badge/PR-welcome-55EB99.svg"></a>
@@ -199,6 +199,8 @@ Please scan the QR code below to join our community groups:
 ## 🎉 News
+- 🔥 **[2025.04.10]** Model service stress testing tool now supports the `/v1/completions` endpoint (the default endpoint for vLLM benchmarking)
+- 🔥 **[2025.04.08]** Support for evaluating embedding model services compatible with the OpenAI API has been added. For more details, check the [user guide](https://evalscope.readthedocs.io/en/latest/user_guides/backend/rageval_backend/mteb.html#configure-evaluation-parameters).
 - 🔥 **[2025.03.27]** Added support for [AlpacaEval](https://www.modelscope.cn/datasets/AI-ModelScope/alpaca_eval/dataPeview) and [ArenaHard](https://modelscope.cn/datasets/AI-ModelScope/arena-hard-auto-v0.1/summary) evaluation benchmarks. For usage notes, please refer to the [documentation](https://evalscope.readthedocs.io/en/latest/get_started/supported_dataset.html)
 - 🔥 **[2025.03.20]** The model inference service stress testing now supports generating prompts of specified length using random values. Refer to the [user guide](https://evalscope.readthedocs.io/en/latest/user_guides/stress_test/examples.html#using-the-random-dataset) for more details.
 - 🔥 **[2025.03.13]** Added support for the [LiveCodeBench](https://www.modelscope.cn/datasets/AI-ModelScope/code_generation_lite/summary) code evaluation benchmark, which can be used by specifying `live_code_bench`. Supports evaluating QwQ-32B on LiveCodeBench, refer to the [best practices](https://evalscope.readthedocs.io/en/latest/best_practice/eval_qwq.html).
@@ -212,15 +214,14 @@ Please scan the QR code below to join our community groups:
 - 🔥 **[2025.02.13]** Added support for evaluating DeepSeek distilled models, including AIME24, MATH-500, and GPQA-Diamond datasets，refer to [best practice](https://evalscope.readthedocs.io/en/latest/best_practice/deepseek_r1_distill.html); Added support for specifying the `eval_batch_size` parameter to accelerate model evaluation.
 - 🔥 **[2025.01.20]** Support for visualizing evaluation results, including single model evaluation results and multi-model comparison, refer to the [📖 Visualizing Evaluation Results](https://evalscope.readthedocs.io/en/latest/get_started/visualization.html) for more details; Added [`iquiz`](https://modelscope.cn/datasets/AI-ModelScope/IQuiz/summary) evaluation example, evaluating the IQ and EQ of the model.
 - 🔥 **[2025.01.07]** Native backend: Support for model API evaluation is now available. Refer to the [📖 Model API Evaluation Guide](https://evalscope.readthedocs.io/en/latest/get_started/basic_usage.html#api) for more details. Additionally, support for the `ifeval` evaluation benchmark has been added.
+<details><summary>More</summary>
 - 🔥🔥 **[2024.12.31]** Support for adding benchmark evaluations, refer to the [📖 Benchmark Evaluation Addition Guide](https://evalscope.readthedocs.io/en/latest/advanced_guides/add_benchmark.html); support for custom mixed dataset evaluations, allowing for more comprehensive model evaluations with less data, refer to the [📖 Mixed Dataset Evaluation Guide](https://evalscope.readthedocs.io/en/latest/advanced_guides/collection/index.html).
 - 🔥 **[2024.12.13]** Model evaluation optimization: no need to pass the `--template-type` parameter anymore; supports starting evaluation with `evalscope eval --args`. Refer to the [📖 User Guide](https://evalscope.readthedocs.io/en/latest/get_started/basic_usage.html) for more details.
 - 🔥 **[2024.11.26]** The model inference service performance evaluator has been completely refactored: it now supports local inference service startup and Speed Benchmark; asynchronous call error handling has been optimized. For more details, refer to the [📖 User Guide](https://evalscope.readthedocs.io/en/latest/user_guides/stress_test/index.html).
 - 🔥 **[2024.10.31]** The best practice for evaluating Multimodal-RAG has been updated, please check the [📖 Blog](https://evalscope.readthedocs.io/zh-cn/latest/blog/RAG/multimodal_RAG.html#multimodal-rag) for more details.
 - 🔥 **[2024.10.23]** Supports multimodal RAG evaluation, including the assessment of image-text retrieval using [CLIP_Benchmark](https://evalscope.readthedocs.io/en/latest/user_guides/backend/rageval_backend/clip_benchmark.html), and extends [RAGAS](https://evalscope.readthedocs.io/en/latest/user_guides/backend/rageval_backend/ragas.html) to support end-to-end multimodal metrics evaluation.
 - 🔥 **[2024.10.8]** Support for RAG evaluation, including independent evaluation of embedding models and rerankers using [MTEB/CMTEB](https://evalscope.readthedocs.io/en/latest/user_guides/backend/rageval_backend/mteb.html), as well as end-to-end evaluation using [RAGAS](https://evalscope.readthedocs.io/en/latest/user_guides/backend/rageval_backend/ragas.html).
-<details><summary>More</summary>
 - 🔥 **[2024.09.18]** Our documentation has been updated to include a blog module, featuring some technical research and discussions related to evaluations. We invite you to [📖 read it](https://evalscope.readthedocs.io/en/refact_readme/blog/index.html).
 - 🔥 **[2024.09.12]** Support for LongWriter evaluation, which supports 10,000+ word generation. You can use the benchmark [LongBench-Write](evalscope/third_party/longbench_write/README.md) to measure the long output quality as well as the output length.
 - 🔥 **[2024.08.30]** Support for custom dataset evaluations, including text datasets and multimodal image-text datasets.
@@ -503,6 +504,10 @@ Reference: Performance Testing [📖 User Guide](https://evalscope.readthedocs.i
 ![wandb sample](https://modelscope.oss-cn-beijing.aliyuncs.com/resource/wandb_sample.png)
+**Supports swanlab for recording results**
+![swanlab sample](https://sail-moe.oss-cn-hangzhou.aliyuncs.com/yunlin/images/evalscope/swanlab.png)
 **Supports Speed Benchmark**
 It supports speed testing and provides speed benchmarks similar to those found in the [official Qwen](https://qwen.readthedocs.io/en/latest/benchmark/speed_benchmark.html) reports:

{evalscope-0.13.2 → evalscope-0.14.0}/README.md RENAMED Viewed

@@ -10,7 +10,7 @@
 </p>
 <p align="center">
-<img src="https://img.shields.io/badge/python-%E2%89%A53.8-5be.svg">
+<img src="https://img.shields.io/badge/python-%E2%89%A53.9-5be.svg">
 <a href="https://badge.fury.io/py/evalscope"><img src="https://badge.fury.io/py/evalscope.svg" alt="PyPI version" height="18"></a>
 <a href="https://pypi.org/project/evalscope"><img alt="PyPI - Downloads" src="https://static.pepy.tech/badge/evalscope"></a>
 <a href="https://github.com/modelscope/evalscope/pulls"><img src="https://img.shields.io/badge/PR-welcome-55EB99.svg"></a>
@@ -88,6 +88,8 @@ Please scan the QR code below to join our community groups:
 ## 🎉 News
+- 🔥 **[2025.04.10]** Model service stress testing tool now supports the `/v1/completions` endpoint (the default endpoint for vLLM benchmarking)
+- 🔥 **[2025.04.08]** Support for evaluating embedding model services compatible with the OpenAI API has been added. For more details, check the [user guide](https://evalscope.readthedocs.io/en/latest/user_guides/backend/rageval_backend/mteb.html#configure-evaluation-parameters).
 - 🔥 **[2025.03.27]** Added support for [AlpacaEval](https://www.modelscope.cn/datasets/AI-ModelScope/alpaca_eval/dataPeview) and [ArenaHard](https://modelscope.cn/datasets/AI-ModelScope/arena-hard-auto-v0.1/summary) evaluation benchmarks. For usage notes, please refer to the [documentation](https://evalscope.readthedocs.io/en/latest/get_started/supported_dataset.html)
 - 🔥 **[2025.03.20]** The model inference service stress testing now supports generating prompts of specified length using random values. Refer to the [user guide](https://evalscope.readthedocs.io/en/latest/user_guides/stress_test/examples.html#using-the-random-dataset) for more details.
 - 🔥 **[2025.03.13]** Added support for the [LiveCodeBench](https://www.modelscope.cn/datasets/AI-ModelScope/code_generation_lite/summary) code evaluation benchmark, which can be used by specifying `live_code_bench`. Supports evaluating QwQ-32B on LiveCodeBench, refer to the [best practices](https://evalscope.readthedocs.io/en/latest/best_practice/eval_qwq.html).
@@ -101,15 +103,14 @@ Please scan the QR code below to join our community groups:
 - 🔥 **[2025.02.13]** Added support for evaluating DeepSeek distilled models, including AIME24, MATH-500, and GPQA-Diamond datasets，refer to [best practice](https://evalscope.readthedocs.io/en/latest/best_practice/deepseek_r1_distill.html); Added support for specifying the `eval_batch_size` parameter to accelerate model evaluation.
 - 🔥 **[2025.01.20]** Support for visualizing evaluation results, including single model evaluation results and multi-model comparison, refer to the [📖 Visualizing Evaluation Results](https://evalscope.readthedocs.io/en/latest/get_started/visualization.html) for more details; Added [`iquiz`](https://modelscope.cn/datasets/AI-ModelScope/IQuiz/summary) evaluation example, evaluating the IQ and EQ of the model.
 - 🔥 **[2025.01.07]** Native backend: Support for model API evaluation is now available. Refer to the [📖 Model API Evaluation Guide](https://evalscope.readthedocs.io/en/latest/get_started/basic_usage.html#api) for more details. Additionally, support for the `ifeval` evaluation benchmark has been added.
+<details><summary>More</summary>
 - 🔥🔥 **[2024.12.31]** Support for adding benchmark evaluations, refer to the [📖 Benchmark Evaluation Addition Guide](https://evalscope.readthedocs.io/en/latest/advanced_guides/add_benchmark.html); support for custom mixed dataset evaluations, allowing for more comprehensive model evaluations with less data, refer to the [📖 Mixed Dataset Evaluation Guide](https://evalscope.readthedocs.io/en/latest/advanced_guides/collection/index.html).
 - 🔥 **[2024.12.13]** Model evaluation optimization: no need to pass the `--template-type` parameter anymore; supports starting evaluation with `evalscope eval --args`. Refer to the [📖 User Guide](https://evalscope.readthedocs.io/en/latest/get_started/basic_usage.html) for more details.
 - 🔥 **[2024.11.26]** The model inference service performance evaluator has been completely refactored: it now supports local inference service startup and Speed Benchmark; asynchronous call error handling has been optimized. For more details, refer to the [📖 User Guide](https://evalscope.readthedocs.io/en/latest/user_guides/stress_test/index.html).
 - 🔥 **[2024.10.31]** The best practice for evaluating Multimodal-RAG has been updated, please check the [📖 Blog](https://evalscope.readthedocs.io/zh-cn/latest/blog/RAG/multimodal_RAG.html#multimodal-rag) for more details.
 - 🔥 **[2024.10.23]** Supports multimodal RAG evaluation, including the assessment of image-text retrieval using [CLIP_Benchmark](https://evalscope.readthedocs.io/en/latest/user_guides/backend/rageval_backend/clip_benchmark.html), and extends [RAGAS](https://evalscope.readthedocs.io/en/latest/user_guides/backend/rageval_backend/ragas.html) to support end-to-end multimodal metrics evaluation.
 - 🔥 **[2024.10.8]** Support for RAG evaluation, including independent evaluation of embedding models and rerankers using [MTEB/CMTEB](https://evalscope.readthedocs.io/en/latest/user_guides/backend/rageval_backend/mteb.html), as well as end-to-end evaluation using [RAGAS](https://evalscope.readthedocs.io/en/latest/user_guides/backend/rageval_backend/ragas.html).
-<details><summary>More</summary>
 - 🔥 **[2024.09.18]** Our documentation has been updated to include a blog module, featuring some technical research and discussions related to evaluations. We invite you to [📖 read it](https://evalscope.readthedocs.io/en/refact_readme/blog/index.html).
 - 🔥 **[2024.09.12]** Support for LongWriter evaluation, which supports 10,000+ word generation. You can use the benchmark [LongBench-Write](evalscope/third_party/longbench_write/README.md) to measure the long output quality as well as the output length.
 - 🔥 **[2024.08.30]** Support for custom dataset evaluations, including text datasets and multimodal image-text datasets.
@@ -392,6 +393,10 @@ Reference: Performance Testing [📖 User Guide](https://evalscope.readthedocs.i
 ![wandb sample](https://modelscope.oss-cn-beijing.aliyuncs.com/resource/wandb_sample.png)
+**Supports swanlab for recording results**
+![swanlab sample](https://sail-moe.oss-cn-hangzhou.aliyuncs.com/yunlin/images/evalscope/swanlab.png)
 **Supports Speed Benchmark**
 It supports speed testing and provides speed benchmarks similar to those found in the [official Qwen](https://qwen.readthedocs.io/en/latest/benchmark/speed_benchmark.html) reports:

{evalscope-0.13.2 → evalscope-0.14.0}/evalscope/backend/rag_eval/__init__.py RENAMED Viewed

@@ -1,4 +1,4 @@
-from evalscope.backend.rag_eval.backend_manager import RAGEvalBackendManager
+from evalscope.backend.rag_eval.backend_manager import RAGEvalBackendManager, Tools
 from evalscope.backend.rag_eval.utils.clip import VisionModel
 from evalscope.backend.rag_eval.utils.embedding import EmbeddingModel
 from evalscope.backend.rag_eval.utils.llm import LLM, ChatOpenAI, LocalLLM

{evalscope-0.13.2 → evalscope-0.14.0}/evalscope/backend/rag_eval/backend_manager.py RENAMED Viewed

@@ -8,6 +8,12 @@ from evalscope.utils.logger import get_logger
 logger = get_logger()
+class Tools:
+    MTEB = 'mteb'
+    RAGAS = 'ragas'
+    CLIP_BENCHMARK = 'clip_benchmark'
 class RAGEvalBackendManager(BackendManager):
     def __init__(self, config: Union[str, dict], **kwargs):
@@ -47,9 +53,19 @@ class RAGEvalBackendManager(BackendManager):
         from evalscope.backend.rag_eval.ragas.tasks import generate_testset
         if testset_args is not None:
-            generate_testset(TestsetGenerationArguments(**testset_args))
+            if isinstance(testset_args, dict):
+                generate_testset(TestsetGenerationArguments(**testset_args))
+            elif isinstance(testset_args, TestsetGenerationArguments):
+                generate_testset(testset_args)
+            else:
+                raise ValueError('Please provide the testset generation arguments.')
         if eval_args is not None:
-            rag_eval(EvaluationArguments(**eval_args))
+            if isinstance(eval_args, dict):
+                rag_eval(EvaluationArguments(**eval_args))
+            elif isinstance(eval_args, EvaluationArguments):
+                rag_eval(eval_args)
+            else:
+                raise ValueError('Please provide the evaluation arguments.')
     @staticmethod
     def run_clip_benchmark(args):
@@ -59,17 +75,17 @@ class RAGEvalBackendManager(BackendManager):
     def run(self, *args, **kwargs):
         tool = self.config_d.pop('tool')
-        if tool.lower() == 'mteb':
+        if tool.lower() == Tools.MTEB:
             self._check_env('mteb')
             model_args = self.config_d['model']
             eval_args = self.config_d['eval']
             self.run_mteb(model_args, eval_args)
-        elif tool.lower() == 'ragas':
+        elif tool.lower() == Tools.RAGAS:
             self._check_env('ragas')
             testset_args = self.config_d.get('testset_generation', None)
             eval_args = self.config_d.get('eval', None)
             self.run_ragas(testset_args, eval_args)
-        elif tool.lower() == 'clip_benchmark':
+        elif tool.lower() == Tools.CLIP_BENCHMARK:
             self._check_env('webdataset')
             self.run_clip_benchmark(self.config_d['eval'])
         else:

{evalscope-0.13.2 → evalscope-0.14.0}/evalscope/backend/rag_eval/cmteb/arguments.py RENAMED Viewed

@@ -20,6 +20,12 @@ class ModelArguments:
     encode_kwargs: dict = field(default_factory=lambda: {'show_progress_bar': True, 'batch_size': 32})
     hub: str = 'modelscope'  # modelscope or huggingface
+    # for API embedding model
+    model_name: Optional[str] = None
+    api_base: Optional[str] = None
+    api_key: Optional[str] = None
+    dimensions: Optional[int] = None
     def to_dict(self) -> Dict[str, Any]:
         return {
             'model_name_or_path': self.model_name_or_path,
@@ -31,6 +37,10 @@ class ModelArguments:
             'config_kwargs': self.config_kwargs,
             'encode_kwargs': self.encode_kwargs,
             'hub': self.hub,
+            'model_name': self.model_name,
+            'api_base': self.api_base,
+            'api_key': self.api_key,
+            'dimensions': self.dimensions,
         }

{evalscope-0.13.2 → evalscope-0.14.0}/evalscope/backend/rag_eval/ragas/arguments.py RENAMED Viewed

@@ -21,7 +21,6 @@ class TestsetGenerationArguments:
     """
     generator_llm: Dict = field(default_factory=dict)
     embeddings: Dict = field(default_factory=dict)
-    distribution: str = field(default_factory=lambda: {'simple': 0.5, 'multi_context': 0.4, 'reasoning': 0.1})
     # For LLM based evaluation
     # available: ['english', 'hindi', 'marathi', 'chinese', 'spanish', 'amharic', 'arabic',
     # 'armenian', 'bulgarian', 'urdu', 'russian', 'polish', 'persian', 'dutch', 'danish',

{evalscope-0.13.2 → evalscope-0.14.0}/evalscope/backend/rag_eval/ragas/tasks/testset_generation.py RENAMED Viewed

@@ -67,9 +67,14 @@ def get_persona(llm, kg, language):
 def load_data(file_path):
-    from langchain_community.document_loaders import UnstructuredFileLoader
+    import nltk
+    from langchain_unstructured import UnstructuredLoader
-    loader = UnstructuredFileLoader(file_path, mode='single')
+    if nltk.data.find('taggers/averaged_perceptron_tagger_eng') is False:
+        # need to download nltk data for the first time
+        nltk.download('averaged_perceptron_tagger_eng')
+    loader = UnstructuredLoader(file_path)
     data = loader.load()
     return data

{evalscope-0.13.2 → evalscope-0.14.0}/evalscope/backend/rag_eval/ragas/tasks/translate_prompt.py RENAMED Viewed

@@ -2,7 +2,6 @@ import asyncio
 import os
 from ragas.llms import BaseRagasLLM
 from ragas.prompt import PromptMixin, PydanticPrompt
-from ragas.utils import RAGAS_SUPPORTED_LANGUAGE_CODES
 from typing import List
 from evalscope.utils.logger import get_logger
@@ -16,10 +15,6 @@ async def translate_prompt(
     llm: BaseRagasLLM,
     adapt_instruction: bool = False,
 ):
-    if target_lang not in RAGAS_SUPPORTED_LANGUAGE_CODES:
-        logger.warning(f'{target_lang} is not in supported language: {list(RAGAS_SUPPORTED_LANGUAGE_CODES)}')
-        return
     if not issubclass(type(prompt_user), PromptMixin):
         logger.info(f"{prompt_user} is not a PromptMixin, don't translate it")
         return

{evalscope-0.13.2 → evalscope-0.14.0}/evalscope/backend/rag_eval/utils/embedding.py RENAMED Viewed

@@ -1,10 +1,12 @@
 import os
 import torch
 from langchain_core.embeddings import Embeddings
+from langchain_openai.embeddings import OpenAIEmbeddings
 from sentence_transformers import models
 from sentence_transformers.cross_encoder import CrossEncoder
 from sentence_transformers.SentenceTransformer import SentenceTransformer
 from torch import Tensor
+from tqdm import tqdm
 from typing import Dict, List, Optional, Union
 from evalscope.backend.rag_eval.utils.tools import download_model
@@ -18,10 +20,10 @@ class BaseModel(Embeddings):
     def __init__(
         self,
-        model_name_or_path: str,
+        model_name_or_path: str = '',
         max_seq_length: int = 512,
         prompt: str = '',
-        revision: Optional[str] = None,
+        revision: Optional[str] = 'master',
         **kwargs,
     ):
         self.model_name_or_path = model_name_or_path
@@ -139,7 +141,7 @@ class CrossEncoderModel(BaseModel):
             max_length=self.max_seq_length,
         )
-    def predict(self, sentences: List[List[str]], **kwargs) -> List[List[float]]:
+    def predict(self, sentences: List[List[str]], **kwargs) -> Tensor:
         self.encode_kwargs.update(kwargs)
         if len(sentences[0]) == 3:  # Note: For mteb retrieval task
@@ -154,6 +156,46 @@ class CrossEncoderModel(BaseModel):
         return embeddings
+class APIEmbeddingModel(BaseModel):
+    def __init__(self, **kwargs):
+        self.model_name = kwargs.get('model_name')
+        self.openai_api_base = kwargs.get('api_base')
+        self.openai_api_key = kwargs.get('api_key')
+        self.dimensions = kwargs.get('dimensions')
+        self.model = OpenAIEmbeddings(
+            model=self.model_name,
+            openai_api_base=self.openai_api_base,
+            openai_api_key=self.openai_api_key,
+            dimensions=self.dimensions,
+            check_embedding_ctx_length=False)
+        super().__init__(model_name_or_path=self.model_name, **kwargs)
+        self.batch_size = self.encode_kwargs.get('batch_size', 10)
+    def encode(self, texts: Union[str, List[str]], **kwargs) -> Tensor:
+        if isinstance(texts, str):
+            texts = [texts]
+        embeddings: List[List[float]] = []
+        for i in tqdm(range(0, len(texts), self.batch_size)):
+            response = self.model.embed_documents(texts[i:i + self.batch_size], chunk_size=self.batch_size)
+            embeddings.extend(response)
+        return torch.tensor(embeddings)
+    def encode_queries(self, queries, **kwargs):
+        return self.encode(queries, **kwargs)
+    def encode_corpus(self, corpus, **kwargs):
+        if isinstance(corpus[0], dict):
+            input_texts = ['{} {}'.format(doc.get('title', ''), doc['text']).strip() for doc in corpus]
+        else:
+            input_texts = corpus
+        return self.encode(input_texts, **kwargs)
 class EmbeddingModel:
     """Custom embeddings"""
@@ -165,6 +207,10 @@ class EmbeddingModel:
         revision: Optional[str] = 'master',
         **kwargs,
     ):
+        if kwargs.get('model_name'):
+            # If model_name is provided, use OpenAIEmbeddings
+            return APIEmbeddingModel(**kwargs)
         # If model path does not exist and hub is 'modelscope', download the model
         if not os.path.exists(model_name_or_path) and hub == HubType.MODELSCOPE:
             model_name_or_path = download_model(model_name_or_path, revision)

{evalscope-0.13.2 → evalscope-0.14.0}/evalscope/backend/rag_eval/utils/llm.py RENAMED Viewed

@@ -2,7 +2,7 @@ import os
 from langchain_core.callbacks.manager import CallbackManagerForLLMRun
 from langchain_core.language_models.llms import LLM as BaseLLM
 from langchain_openai import ChatOpenAI
-from modelscope.utils.hf_util import GenerationConfig
+from transformers.generation.configuration_utils import GenerationConfig
 from typing import Any, Dict, Iterator, List, Mapping, Optional
 from evalscope.constants import DEFAULT_MODEL_REVISION
@@ -16,9 +16,9 @@ class LLM:
         api_base = kw.get('api_base', None)
         if api_base:
             return ChatOpenAI(
-                model_name=kw.get('model_name', ''),
-                openai_api_base=api_base,
-                openai_api_key=kw.get('api_key', 'EMPTY'),
+                model=kw.get('model_name', ''),
+                base_url=api_base,
+                api_key=kw.get('api_key', 'EMPTY'),
             )
         else:
             return LocalLLM(**kw)

{evalscope-0.13.2 → evalscope-0.14.0}/evalscope/backend/vlm_eval_kit/backend_manager.py RENAMED Viewed

@@ -1,4 +1,5 @@
 import copy
+import os
 import subprocess
 from functools import partial
 from typing import Optional, Union
@@ -66,8 +67,9 @@ class VLMEvalKitBackendManager(BackendManager):
                     del remain_cfg['name']  # remove not used args
                     del remain_cfg['type']  # remove not used args
-                    self.valid_models.update({model_type: partial(model_class, model=model_type, **remain_cfg)})
-                    new_model_names.append(model_type)
+                    norm_model_type = os.path.basename(model_type).replace(':', '-').replace('.', '_')
+                    self.valid_models.update({norm_model_type: partial(model_class, model=model_type, **remain_cfg)})
+                    new_model_names.append(norm_model_type)
                 else:
                     remain_cfg = copy.deepcopy(model_cfg)
                     del remain_cfg['name']  # remove not used args

{evalscope-0.13.2 → evalscope-0.14.0}/evalscope/benchmarks/arc/arc_adapter.py RENAMED Viewed

@@ -134,7 +134,7 @@ class ARCAdapter(DataAdapter):
         if self.model_adapter == OutputType.MULTIPLE_CHOICE:
             return result
         else:
-            return ResponseParser.parse_first_option(text=result)
+            return ResponseParser.parse_first_option(text=result, options=self.choices)
     def match(self, gold: str, pred: str) -> float:
         return exact_match(gold=gold, pred=pred)

{evalscope-0.13.2 → evalscope-0.14.0}/evalscope/benchmarks/data_adapter.py RENAMED Viewed

@@ -314,11 +314,15 @@ class DataAdapter(ABC):
         kwargs['metric_list'] = self.metric_list
         return ReportGenerator.gen_report(subset_score_map, report_name, **kwargs)
-    def gen_prompt_data(self, prompt: str, system_prompt: Optional[str] = None, **kwargs) -> dict:
+    def gen_prompt_data(self,
+                        prompt: str,
+                        system_prompt: Optional[str] = None,
+                        choices: Optional[List[str]] = None,
+                        **kwargs) -> dict:
         if not isinstance(prompt, list):
             prompt = [prompt]
         prompt_data = PromptData(
-            data=prompt, multi_choices=self.choices, system_prompt=system_prompt or self.system_prompt)
+            data=prompt, multi_choices=choices or self.choices, system_prompt=system_prompt or self.system_prompt)
         return prompt_data.to_dict()
     def gen_prompt(self, input_d: dict, subset_name: str, few_shot_list: list, **kwargs) -> Any:

{evalscope-0.13.2 → evalscope-0.14.0}/evalscope/benchmarks/general_qa/general_qa_adapter.py RENAMED Viewed

@@ -40,7 +40,7 @@ class GeneralQAAdapter(DataAdapter):
             for subset_name in subset_list:
                 data_file_dict[subset_name] = os.path.join(dataset_name_or_path, f'{subset_name}.jsonl')
         elif os.path.isfile(dataset_name_or_path):
-            cur_subset_name = os.path.basename(dataset_name_or_path).split('.')[0]
+            cur_subset_name = os.path.splitext(os.path.basename(dataset_name_or_path))[0]
             data_file_dict[cur_subset_name] = dataset_name_or_path
         else:
             raise ValueError(f'Invalid dataset path: {dataset_name_or_path}')

{evalscope-0.13.2 → evalscope-0.14.0}/evalscope/benchmarks/hellaswag/hellaswag_adapter.py RENAMED Viewed

@@ -108,7 +108,7 @@ class HellaSwagAdapter(DataAdapter):
         if self.model_adapter == OutputType.MULTIPLE_CHOICE:
             return result
         else:
-            return ResponseParser.parse_first_option(result)
+            return ResponseParser.parse_first_option(result, options=self.choices)
     def match(self, gold: str, pred: str) -> float:
         return exact_match(gold=str(gold), pred=str(pred))

{evalscope-0.13.2 → evalscope-0.14.0}/evalscope/benchmarks/live_code_bench/live_code_bench_adapter.py RENAMED Viewed

@@ -18,7 +18,6 @@ logger = get_logger()
     extra_params={
         'start_date': None,
         'end_date': None,
-        'num_process_evaluate': 1,
         'timeout': 6
     },
     system_prompt=
@@ -33,7 +32,6 @@ class LiveCodeBenchAdapter(DataAdapter):
         extra_params = kwargs.get('extra_params', {})
-        self.num_process_evaluate = extra_params.get('num_process_evaluate', 1)
         self.timeout = extra_params.get('timeout', 6)
         self.start_date = extra_params.get('start_date')
         self.end_date = extra_params.get('end_date')
@@ -84,7 +82,7 @@ class LiveCodeBenchAdapter(DataAdapter):
             references,
             predictions,
             k_list=[1],
-            num_process_evaluate=self.num_process_evaluate,
+            num_process_evaluate=1,
             timeout=self.timeout,
         )
         return metrics['pass@1'] / 100  # convert to point scale

evalscope 0.13.2__tar.gz → 0.14.0__tar.gz

Potentially problematic release.

evalscope 0.13.2tar.gz → 0.14.0tar.gz