README: Evals provide a framework for evaluating large language models (LLMs) or systems built using LLMs. We offer an existing registry of evals to test different dimensions of OpenAI mo…
GitHub 项目简介: Evals is a framework for evaluating LLMs and LLM systems, and an open-source registry of benchmarks.
README: Without evals, it can be very difficult and time intensive to understand how different model versions might affect your use case.
README: If you are going to be creating evals, we suggest cloning this repo directly from GitHub and installing the requirements using the following command: ```sh pip install -e . ```
README: If you don't want to contribute new evals, but simply want to run them locally, you can install the evals package via pip: ```sh pip install evals ```